Wednesday, March 12, 2014

Dialogue

Information sharing through trusted, vetted channels is an integral part of a successful security operations program. For the purpose of this blog posting, let's assume that an organization already has in place the ability to leverage their host and network forensics infrastructure to both identify information worth sharing and capitalize upon information they receive through trusted, vetted channels. Even with this in place, it can still be difficult for an organization to share information. What could be limiting the sharing? There may be many factors, but one such factor I've seen repeatedly is not a technical limitation, but rather, an organizational limitation.

Legal and privacy professionals have an obligation to protect the organizations and data they represent. Most legal and privacy professionals come from rigorous legal and/or regulatory backgrounds, but they are not necessarily technical, and they don't usually have an operational background in security. Thus, when security professionals within an organization try to gain approval for an information sharing program, a game of telephone often ensues. Allow me to explain:

As security professionals, we might say "we would like to share lists of domain names we have observed engaged in malicious activity". Legal and privacy professionals might hear "they want to share lists that may include our clients' or partners' domain names". Or, we might say "we would like to share lists of email addresses we have observed sending phishing emails into the enterprise". Legal and privacy professionals might hear "they want to share lists of internal email addresses and potentially contents of email".

And so on -- there is no shortage of examples that I could bring here. As you can see, each party comes from their respective angle, and each party has difficulty understanding where the other party is coming from. This can easily lead to impasse, frustration, and deadlock within an organization, to the detriment of security operations. What can be done to remedy this? As security professionals, it is our duty to engage legal and privacy professionals in a dialogue. Will we have to educate them? Yes, absolutely. Will we have to be educated on certain issues ourselves and possibly change some of our policies and procedures? Of course. Will we reach a mutual understanding in the end that leads to better security operations and reduced risk for the enterprise? I truly believe so, and in fact, I have seen this with my own eyes. Because of this, it is incumbent upon us as security professionals to engage legal and privacy professionals in a dialogue. It may not come as naturally to us as other aspects of our jobs, but the stakes are too high for us not to.

Tuesday, March 11, 2014

100 a Day

One of the goals of an incident response team should be to handle no more than 100 alerts a day. At first, this may sound like a ridiculous assertion. However, I think that if we examine this more closely, you will agree that it makes sense. Let's take an analytical approach and go to the numbers.

As previously discussed on this blog and elsewhere, one hour from detection to containment should be the goal in incident response. Put another way, one hour should be the time allotted to work an alert, perform all required analysis, forensics, and investigation, and take any necessary containment actions. Let's say we have each of our analysts working an eight hour shift. Assuming 100% productivity for each analyst, that allows each analyst to work approximately eight incidents per day. Let's assume that we want to work 96 alerts properly each day (since 100 is not divisible by eight). That works out to a requirement to have 12 analysts on shift (or spread across multiple shifts) to give proper attention to each alert. What happens if analyst cycles are taken away from incident response and lent to other tasks? The numbers look worse. What happens if the necessary analysis, forensics, and investigation take more than an hour (due to technology, process, or other limitations)? The numbers look even worse yet.

So, if you're the type of enterprise that has 500 analysts sitting in your SOC or Incident Response Center, you can probably stop reading this blog post and get back to your daily routine. What's that you say? The analyst is the scarcest resource, and you don't have enough of them? Yes, of course. I know.

Let's face it -- the numbers are sobering. Even a large enterprise with a large incident response team can realistically handle no more than 100-200 alerts in a given day. Sometimes I meet people who tell me that "we handle 5,000 incidents per day". I don't believe that for a second (putting aside, for now, the fact that incidents, events, and alerts are not the same thing). Either that organization is not paying each alert the attention it deserves, or the alerts are of such low value to security operations that it wouldn't make much difference whether they fired or not. One need only look to the recent Nieman Marcus intrusion to see the devastating effects of having too large a volume of noisy, low fidelity, false-positive prone alerts that drown out any activity of true concern (http://www.businessweek.com/articles/2014-02-21/neiman-marcus-hackers-set-off-60-000-alerts-while-bagging-credit-card-data).

Clearly, the challenge becomes populating the alerting queue with reliable, high fidelity, actionable alerts for analysts to review in priority order (priority will be the subject of an upcoming blog post). This process is sometimes referred to as content development and can be outlined at a high level as follows:
  • Collect the data of highest value and relevance to security operations and incident response. As previously discussed on this blog, fewer data sources providing higher value at lower volume/size, while still maintaining the required visibility are desired.
  • Identify goals and priorities for detection and alerting in line with business needs, security needs, management/executive priorities, risk/exposure, and the threat landscape. Use cases can be particularly helpful here.
  • Craft human language logic designed to extract only the events relevant to the goals and priorities identified in the previous step.
  • Convert the human language logic into precise, incisive, targeted queries designed to surgically extract reliable, high fidelity, actionable alerts with few to no false positives
  • Continually iterate through this process, identifying new goals and priorities, developing new content, and adjusting existing content based on feedback obtained through the incident response process.
Resources are limited. Every alert counts. Make every alert worth the analyst's attention.

Monday, March 10, 2014

Buyer Beware

A couple of weeks ago, I attended the RSA conference in San Francisco.
I always enjoy attending RSA, as it provides a unique opportunity to
engage many different aspects of the larger security community at the
same time. The conference is attended by vendors, practitioners/enterprises, researchers, industry analysts, journalists, investors, and others. I was fortunate enough to take part in several interesting and engaging discussions during the week.  I would like to discuss one observation I made during the conference in this posting.

I took some time during the week to walk the vendor expo two or three
times. What I saw there inspired this blog, though it didn't necessarily surprise me. Not every vendor on the floor was guilty of this, but many, many vendors proffered a technology or solution for "big data", "security analytics", and/or "big data security analytics". In other words, many (though not all) vendors said they provided a solution for the same "space". Since I spent over a decade
on the enterprise/operational side, I can sympathize with the confusion this can bring to the enterprise audience. Leaders in the enterprise have many responsibilities, and it is difficult for them to keep track of the large number of vendors and what each vendor's specialty is.

Marketing is unlikely to change in the near future, and as such, it appears that the words "buyer beware" are important words for the enterprise. Many enterprises want to be doing "big data" and "security analytics", and thus, it's not particularly surprising that many vendors are offering "big data" and "security analytics" solutions. But what does it actually mean to do "big data" and "security analytics"? I think it's helpful to take a step back and think a level deeper about this in order to better understand it.

At a high level, "big data" and "security analytics" are about the two very different, but equally important concepts of collection and analysis. Allow me to explain. Before it is possible to run analytics, one needs the right data upon which to run those analytics. Before "big data" emerged as a buzzword, this was called "collection" or "instrumentation of the network". Further, in order to run analytics, one also needs a high performance platform upon which to issue the precise, targeted, incisive queries required by analytics. Before "security analytics" emerged as a buzzword, this was sometimes called analysis or forensics, among other terms. Collection and analysis, at enterprise speeds, are both equally important. If you think about it, you can't really have one without the other. Or, to put it another way, what good does the greatest collection capability provide without a way to analyze that data in a timely and accurate manner? Similarly, what good does the greatest analytical capability provide without the underlying data to support it?

As I walked around the expo floor, two families of "big data security analytics" products jumped out at me:

1) Analysis platforms that struggle with collection/consumption of data
2) Collection platforms that struggle with the analysis component (either because of performance, analytical capability, or both)

So, what about a platform that can do both collection and analysis at enterprise speeds? That's what I call a real "big data security analytics" platform -- one that lives up to the intent and spirit of the marketing buzzwords. Think about the ramifications of a single platform that provides excellent collection and excellent analysis. That's a great way to bring "big data security analytics" to your organization with reduced complexity and at a reduced cost.

If you're going to do "big data", it's worth thinking about how to do it right.

Friday, March 7, 2014

Measuring Security Intelligence Value

There was recently a discussion around measuring the value of a security intelligence program in the Twitterverse. Several well-known security experts took part in the discussion, and it was quite interesting to see everyone's thoughts on the subject. Further, this is a discussion that I am hearing more and more in the security operations space, and rightfully so. Security intelligence is a complex topic that requires more elaboration than Twitter's character limit allows for, and in fact, it requires more elaboration than I can realistically put into a blog posting. That being said, I will give it a shot. I have a large amount of operational experience in this area, so if you'll indulge me, I'll provide some thoughts on the topic here in this blog posting.

Generating security intelligence data (i.e., functioning as a source of intelligence) is an interesting topic, but not one that I will discuss in this post. Threat assessment is another fascinating topic, but also not one that I will discuss in this post. Additionally, intelligence sharing is also a hot topic, but again, not one that I will discuss in this post. Instead, I will focus on security intelligence as it relates to defending a large network in the context of a broader security operations program (i.e., functioning as a consumer of intelligence). In this context, at a high level, security intelligence involves consuming a piece of information (e.g., domain name, URL pattern, file name, MD5 hash, etc.), along with some context (e.g., exploit site, callback domain, drop site, malicious attachment MD5, etc.), and subsequently leveraging that information, in the right context, against host data and/or network traffic data.

Over the course of my career, I have seen security intelligence programs that work well, along with those that do not work as well. Before I get into a discussion of metrics around security intelligence programs, here are a few observations relating to challenges that organizations often encounter when implementing a security intelligence program:
  • Information lacks context: The best information in the world is useless unless we know in which context to use it -- context is key. Remember, only information with the proper context can qualify as intelligence.
  • Confusion of quantity with quality: If we have 5,000 "malicious" domain names, but 4,995 of them generate almost entirely false positives, that provides far less value than 10 reliable, high fidelity malicious domain names. Not only does the first example detect less true positives, but the volume of false positives overwhelm the work queue to the detriment of security operations.
  • Lack of indicator reliability and fidelity: It is important to vet indicators and sources before they are introduced into the alerting queue and workflow. Failure to do this properly can result in an overwhelming volume of false positives that dominate the work queue and squander valuable analyst cycles.
  • Lack of appropriate data: The best intelligence in the world is useless if we can't search for it over large quantities of host and network data over long periods of time rapidly.
  • Improper tracking of intelligence and sources: This leads well into the metrics discussion -- it is extremely important to warehouse and track intelligence and its sources with enough granularity to enable metrics and measurement.
  • Lack of integration with the workflow: If leveraging security intelligence is a pain, analysts will do it less, or they won't do it at all. This is a workflow/efficiency issue that can come at a great cost to an organization's overall security posture.
Measuring the value of a security intelligence program can be a difficult task. I'm sure there are many ways to approach the challenge. As mentioned above, proper warehousing and tracking of intelligence and its sources is a necessary precursor. Assuming that is in place, here are a few measurement approaches that I have found helpful over the course of my career:
  • Overall number of incidents/percentage of incidents identified via security intelligence vs. identified via other means.
  • Percentage of false positives per intelligence source.
  • Percentage of overall false positives resulting from security intelligence.
  • Percentage of incidents identified through security intelligence per intelligence source (also known as percentage of true positives per intelligence source).
  • Percentage of overall true positives resulting from security intelligence.
  • Mean time to detection (should decrease as security intelligence program matures).
  • Number of long-time (i.e., long undetected) intrusions uncovered via security intelligence.
This is not an exhaustive list, but it does provide a few metrics that I have found helpful for measuring and showing the value of a security intelligence program. Although it's implied here, it's perhaps worth stating explicitly that proper instrumentation of the network, for both collection and analysis is critical here. If an organization does not have total visibility into the network traffic and host data, along with the ability to incisively query that data rapidly, that organization will not be particularly successful in building a security intelligence program nor in measuring it.

A security intelligence program is a great thing, and every large organization should have one. It pays to consider how to make your security intelligence program the best one it can be.

Thursday, March 6, 2014

It's All About the Workflow


In a previous blog post entitled "The Scarcest Resource", I discussed how, of all the resources necessary for security operations and incident response, human analyst cycles are the most scarce. Recently, HP echoed the same sentiment in a report entitled "State of Security Operations" (https://ssl.www8.hp.com/ww/en/secure/pdf/4aa5-0501enw.pdf). The following quote from that report is particularly poignant:

"In SOCs, this results in minimal investment in the most expensive CPU in the room: the analyst."

The issue is clear, but what can an organization do to address it? There are many possible approaches one could take here, but I would like to discuss one of my favorites: workflow. Workflow is a concept that, in my experience, has the greatest return on investment for security operations when implemented correctly. With the volume, velocity, and variety of data coming at an analyst these days, it's more important than ever to focus the analyst via a single, unified work queue containing actionable, high fidelity items. Further, it's crucial that the analyst be able to perform all necessary analysis, investigation, and pivots and work each item to resolution from within the workflow. Let's have a look at what this workflow might look like and how each step of it corresponds to the incident response process:
  • On a continual basis, intelligent alerting content is developed across all sensing and instrumentation platforms using incisive, precise, targeted, finely-tuned queries designed to extract reliable, actionable, high fidelity events from the vast quantity of data. These events are the items that populate the work queue. This corresponds to the detection stage of the incident response process.
  • Working through the items in the work queue, analysts investigate each one, pivoting into and out of relevant platforms as appropriate to support the investigation. All investigation is documented within the work queue, and once analysis is complete, the analyst draws a conclusions about what has occurred. This corresponds to the analysis stage of the incident response process.
  • The analyst then proceeds through the containment, remediation, and recovery stages of the incident response process, pivoting into and out of relevant supporting systems as necessary. The stages are guided by the conclusions drawn during the analysis stage.
  • Lessons learned are gathered and documented, and detection techniques are improved accordingly. This completes the incident response process and provides a virtuous feedback loop as an added bonus.
It's interesting to note that this workflow is incredibly reliant on the population of the work queue with a sensible volume of reliable, actionable, high fidelity events. This requires sensing and instrumentation platforms designed to support incisive, precise, targeted, finely-tuned queries to extract the most relevant events, while minimizing false positives. I can't emphasize enough how critical this is to the operational workflow as a whole.

It's not easy to master your organization's workflow, but in my experience, it is the single greatest return on investment one can gain organizationally. How do you workflow?

Wednesday, March 5, 2014

The Forgotten Servers

In the enterprise, there is often a separation between network segments containing endpoint/workstation systems (e.g., laptops), network segments containing internal-facing servers (e.g., Exchange), and network segments containing external-facing servers (e.g., web servers). Fundamentally, this makes a lot of sense. Each of these segments serves a very different purpose, and as such, we would expect the traffic transiting each segment to behave differently. Further, each segment should have its own controls that permit traffic necessary for business operations, while denying traffic not befitting of that particular segment.

Although enterprises separate the various types of assets reasonably well, detection and alerting are predominantly focused on network segments containing endpoint/workstation systems. This is for several reasons, but primary among them are:
  • Identifying compromised/infected endpoint/workstation systems is relatively well understood and fairly mature, while identifying compromised/infected server systems is less well understood and not particularly mature.
  • Network segments containing servers, and particularly external-facing servers, are generally less well instrumented than network segments containing endpoint/workstation systems.
It is true that server compromises happen less often than endpoint/workstation compromises. But, it is also true that when server compromises do happen, they are often far more serious and consume far more incident response resources than endpoint/workstation compromises. Server compromises have the potential to lead to additional intrusions, data loss, theft of intellectual property, fraudulent activity, and other malicious activity. Furthermore, server compromises tend to go undetected for long periods of time, mainly because of the two reasons I outlined above.

So, given the risk, and the continued evidence that server compromises lead to bad things, it's a wonder enterprises don't study their server network data more closely. I would recommend two initial steps here, based on my own experience monitoring server networks:
  • Ensure the server network segments are properly instrumented, as it is difficult to monitor network segments for which data collection is incomplete/inadequate.
  • Dedicate some well-trained, highly-skilled analyst cycles to study the traffic on the server network segments.  When reliable, high fidelity approaches are discovered, they can be automated as appropriate.
On server network segments, the stakes are high. So why is it that enterprises almost never pay them the attention they are due?

Tuesday, March 4, 2014

Security as a Line Item

The world of security operations and incident response has traditionally been the bailiwick of governments and large enterprises. The reasons for this are fairly straightforward. Security operations and incident response are relatively resource-intensive undertakings, and large organizations have the ability to bring the necessary people, process, and technology to the table. Many small and medium-sized businesses understand the threat and see the need to perform security operations and incident response, but they do not have the necessary resources available to do so.

As we all know, attackers do not limit themselves to governments and large enterprises. While it may be true that the most prized targets are located within large organizations, small and medium-sized businesses also offer a lucrative bounty for the attacker. But how can small and medium-sized businesses practice security operations and incident response given their resource limitations? I believe that the move to the cloud plays a critical role in the solution.

Small and medium-sized businesses often outsource HR, benefits, IT, and other critical business functions to benefit from the economies of scale afforded by outsourcing. Those same organizations can also outsource security operations and incident response to leverage the same economies of scale. In other words, for certain organizations, security can be thought of as a line item on the menu of services they purchase from the cloud. Small and medium-sized businesses cannot dedicate their own people, process, and technology to security functions, but they can purchase access to a cloud provider's people, process, and technology to meet their business needs and security goals. In fact, this is already starting to happen, and the model seems to be a good one.

For cloud providers looking to sell their people, process, and technology, it is important to think about how you will differentiate yourselves and persuade your customers to choose you over another provider. Are your people adequately trained, do they have the necessary skills, and are they trustworthy? Is your process organized, well-documented, timely, accurate, and does it follow industry best-practices and guidance? Does your technology support your operational workflow, does it scale to modern speeds and data volumes, and does it enable you to exploit the value of the data you possess?

For small and medium-sized business looking to improve security via a line item, it is important to understand what you are buying. Ask to meet the people who will be reviewing your data. Ask them questions based on your priorities and business needs to understand how they think and what their world view is. Ask to review the provider's processes and understand how they will respond when an incident hits. Ask the provider what technology they use, how it scales under load and volume, and what unique capabilities that technology brings them over their competitors. Be a tough customer -- after all, it is important to remember that you can manage risk, but you cannot eliminate it.

Security as a line item is coming, and in fact, it is already here. Those that understand the value of the cloud to small and medium-sized businesses will be able to capitalize on this, while at the same time, protecting a segment of the market that has traditionally been under-served. Likewise, small and medium-sized businesses that are choosy about to where they outsource will do better than those that are not.

Do you see the clouds forming on the horizon?