Friday, March 21, 2014

Answers at the Speed of Business

Imagine yourself as the lead incident responder during a breach response. If you've been in this position, as I have, you know that it can feel a bit like being in the hot seat. During the breach response, key stakeholders will have important, time-sensitive questions they want answered. Those questions will be aimed directly at you, and you will be expected to provide answers quickly -- answers at the speed of business. The stakeholders don't just need answers -- they need them now -- or better yet, make that yesterday. These stakeholders may include executives, legal, privacy, public relations, clients, partners, and others. The questions they will ask are designed to quickly assess damage and risk to the organization, as well as what follow-on actions need to be taken from a legal, privacy, and/or public relations standpoint.

There are many questions these stakeholders might pose, but a few of the more common ones are:
  • How did this happen?
  • When did this begin?
  • Is this activity still occurring?
  • How many systems/brands/products have been affected?
  • What sensitive, proprietary, and/or confidential/private data has been taken?
  • What can be done to stop this activity/prevent it from happening again?
Performing network forensics allows us to query, interrogate, and study the data to obtain accurate answers to important stakeholder questions. As you can imagine, every moment is critical during this process. Given this, it always frustrated me that I seemed to spend a majority of my time waiting for queries to return or "munging" data (due to tool limitations), rather than actually doing analysis. I could never understand why a) vendors sold technology that didn't meet the needs of incident responders, b) the organizations I was supporting bought that technology, and c) I was expected to use something that was not properly designed for the purposes I was being forced to use it for. It always seemed like the technologies I was using were fighting me, rather than enabling and empowering me.

I've been in the hot seat enough times to know that enough is enough. The time has come for network forensics technology that meets the needs of incident responders. Anything less simply fails them. With the stakes as high as they are today, failure is not an option.

Thursday, March 20, 2014

Ask a Stupid Question....

As the saying goes, ask a stupid question, get a stupid answer. Security professionals know that in order to properly run security operations and perform incident response, we need to be able to ask intelligent questions of our data. We need to be able to issue precise, targeted, incisive queries to hone in on the most relevant data, while minimizing or eliminating time spent with data that is irrelevant. With the velocity, volume, and variety of data confronting us, this concept is more central than ever to effective security operations and incident response. Given this, I am often surprised at how few technologies truly empower the analyst to ask those intelligent questions. If your technologies only allow you to ask stupid questions, what kind of answers do you think you'll get?

Wednesday, March 19, 2014

Uber Data Source: Holy Grail or Final Fantasy?

In August 2011, I gave a talk at the GFIRST conference entitled "Uber Data Source: Holy Grail or Final Fantasy?". In this talk, I proposed that given the volume and complexity that a larger number of highly specialized data sources brings to security operations, it makes sense to think about moving towards a smaller number of more generalized data sources. One could also imagine taking this concept further, ultimately resulting in an "uber data source". I would like to discuss this concept in more detail in this blog post. For the purposes of this post, I am working within the context of network traffic data sources. I consider host (e.g., AV logs), system (e.g., Linux syslogs), and/or application level (e.g., web server logs) data sources beyond the context of this blog posting.

Let's begin by first looking at the current state of security operations in most organizations, specifically as it relates to network traffic data log collection. In most organizations, a large number of highly specialized network traffic data sources are collected. This creates a complex ecosystem of logs that clouds the operational workflow. In my experience, the first question asked by an analyst when developing new alerting content or performing incident response is "To which data source or data sources do I go to find the data I need?". I would suggest, based on my experience, that this wastes precious resources and time. Rather, the analyst's first question should be "What questions do I need to ask of the data in order to accomplish what I have set out to do?". This necessitates a "go to" data source -- the "uber data source".

Additionally, it is helpful here to highlight the difference between data value and data volume. Each data source that an organization collects will have a certain value, relevance, and usefulness to security operations. Similarly, each data source will also produce a certain volume of data when collected and warehoused. Data value and data volume do not necessarily correlate. For example, firewall logs often consume 80% of an organization's log storage resources, but actually prove quite difficult to work with when developing alerting content or performing incident response. Conversely, as an illustrative example, DHCP logs provide valuable insight to security operations, but are relatively low volume.

There is also another angle to the data value vs. data volume point. As you can imagine, collecting a large volume of less valuable logs creates two issues, among others:
  • Storage is consumed more quickly, thus reducing the retention period (this can have a detrimental effect on security operations when performing incident response, particularly around intrusions that have been present on the network for quite some time)
  • Queries return more slowly due to the larger volume of data (this can have a detrimental effect on security operations when performing incident response, since answers to important questions come more slowly)
Those who disagree with me will argue: "I can't articulate what it is, but I know that when it comes time to perform incident response, I will need something from those other data sources." To those people, I would ask this question: If you're collecting so much data, irrespective of its value to security operations, that your retention period is cut to less than 30 days and your queries take hours or days to run, are you really able to use that data you've collected for incident response? I would think not.

If we take a step back, we see that the current "give me everything" approach to log collection involves collecting a large number of highly specialized data sources. This is for a variety of reasons, but history and lack of understanding regarding each data source's value to security operations are among them. If we think about what these data sources are conceptually, we see that they are essentially meta-data from layer 4 of the OSI model (the transport layer) enriched with specific data from layer 7 of the OSI model (the application later) suiting the purpose of that particular data source. For example, DNS logs are essentially meta-data from layer 4 of the OSI model enriched with additional contextual information regarding DNS queries and responses found in layer 7 of the OSI model. I would assert that there is a better way to operate without adversely affecting network visibility.

The question I asked back in 2011 was "Why not generalize this?". For example, why collect DNS logs as a specialized data source when the same visibility can be provided as part of a more generalized data source of higher value to security operations? In fact, this has been happening steadily over the last few years. It is now possible to architect network instrumentation to collect fewer data sources of higher value to security operations. This has several benefits:
  • Less redundancy and wastefulness across data sources
  • Less confusion surrounding where to go to get the required data
  • Reduced storage cost or increased retention period at the same storage cost
  • Improved query performance
The "uber data source" is a concept that I believe the security world is coming around to and moving towards. It may be a little uncomfortable to move away from the "give me everything" approach to log collection, but if you think about it, that's really the only way forward in the era of big data. Uber me, baby.

Tuesday, March 18, 2014

The Question

When I speak at conferences or in private meetings, I inevitably get "the question" immediately after presenting:

"How do you understand our pain so well?"

The answer is simple -- I lived that pain for over a decade on the operational side before moving over to the vendor side. I've seen what enables, empowers, and facilitates security operations and incident response and what doesn't. I've seen how vendors struggle with fitting their technology into the operational workflow, rather than forcing the operational workflow to fit their technology. I've also seen where vendors typically fall short of the needs of the analysts and incident responders.

All of that pain and experience influence my professional world view, which in turn, results in a better, more operationally useful product. The best vendors I worked with while on the operational side were those that came from an operational background. Those were the vendors that best understood operational issues, gaps, and needs and sought to address them.

If you are working with vendors that don't approach your challenges from the perspective of an operational background, how can you be certain that they will truly understand your pain and deliver solutions that meet your operational needs? I'd suggest that this is something important to think about as you evaluate different technologies. I'm sure you'd prefer that your vendors were educated previously on somebody else's dime, rather than your own.

Monday, March 17, 2014

Signal-to-Noise Ratio

Recent media reports discussing the Target and Nieman Marcus breaches have indicated that, in both cases, numerous alerts fired as a result of the intrusion activity. In both cases, the alerts were not properly handled, causing the breaches to remain undetected. I'm sure there are many angles in which these reports can be dissected. Rather than play the blame game, I would like to discuss a subject that remains a challenge for our profession as a whole: the signal-to-noise ratio.

Wikipedia defines the signal-to-noise ratio as "a measure used in science and engineering that compares the level of a desired signal to the level of background noise." In other words, the more you have of what you want, and the less you have of what you don't want, the easier it is to measure something. Let's illustrate this concept by imagining a conversation between two people in a noisy cafe. If I record that conversation from the next table, upon playback, it will be very difficult for me to truly understand what was discussed. Conversely, if I record that conversation in a quiet room, it will be much easier to understand what was discussed upon playback. The signal-to-noise ratio in the second scenario is much higher than in the first scenario.

The same concept applies to security operations and incident response. In security operations, true positives are the signal, and false positives are the noise. Consider the case of two different Security Operations Centers (SOCs), SOC A and SOC B. In SOC A, the daily work queue contains approximately 100 reliable, high fidelity, actionable alerts. Each alert is reviewed by an analyst. If incident response is necessary for a given alert, it is performed. In SOC B, the daily work queue contains approximately 100,000 alerts, almost all of which are false positives. Analysts attempt to review the alerts of the highest priority. Because of the large volume of even the highest priority alerts, analysts are not able to successfully review all of the highest priority alerts. Additionally, because of the large number of false positives, SOC B's analysts become desensitized to alerts and do not take them particularly seriously.

One day, 10 additional alerts relating to payment card stealing malware fire within a few minutes of each other.

In SOC A, where every alert is reviewed by an analyst, where the signal-to-noise ratio is high, and where 10 additional alerts seems like a lot, analysts successfully identify the breach less than 24 hours after it occurs. SOC A's team is able to perform analysis, containment, and remediation within the first 24 hours of the breach. The team is able to stop the bleeding before any payment card data is exfiltrated. Although there has been some damage, it can be controlled. The organization can assess the damage, respond appropriately, and return to normal business operations.

In SOC B, where an extremely small percentage of the alerts are reviewed by an analyst, where the signal-to-noise ratio is low, and where 10 additional alerts doesn't even raise an eyebrow, the breach remains undetected. Months later, SOC B will learn of the breach from a third party. The damage will be extensive, and it will take the organization months or years to fully recover.

Unfortunately, in my experience, there are a lot more SOC B's out there than there are SOC A's. It is relatively straightforward to turn a SOC B into a SOC A, but it does require experienced professionals, organizational will, and focus. How do I know? I've turned SOC B's into SOC A's several times during my career.

We are fortunate to have some great technology choices these days that we can leverage to improve our security operations and incident response functions. These technology choices can enable us to learn of and respond to breaches soon after they occur. Before purchasing any technology intended to produce alerts destined for the work queue, we should ensure that it supports the ability to issue very precise, targeted, incisive questions of the data. This enables us to hone in on the activity we want to identify (the true positives/the signal), while minimizing the activity we do not want to identify (the false positives/the noise). As always, these technologies are tools that need to be properly leveraged as part of a larger people, process, and technology picture.

What is your signal-to-noise ratio? Is it high enough to detect the next breach, or could it stand to be strengthened? I would posit that the ratio of true positives to false positives (the signal-to-noise ratio) is an important metric that all organizations should review. Not doing so could have dire consequences.

Friday, March 14, 2014

Year of the Data Breach or Year of the Cloud?

Some people have been calling 2014 the year of the data breach. It's not difficult to understand why -- it seems that there is another breach in the news weekly, if not more often than that. People often ask me why there are so many breaches in the news of late. I can't say for sure, but I suspect it is some combination of these factors, among others:
  • Crime does pay (attackers profit by compromising organizations)
  • Difficulty in tracking down and prosecuting the attackers (for a variety of reasons)
  • Better detection techniques
  • Better information sharing
  • Decrease in stigma for owning up to a compromise
  • Greater security awareness among business leaders and executives
My thought is that 2014 will actually be remembered as the year of the cloud. Time will tell for sure, but I am already seeing a few indications that this may be the case:
  • Small and medium-sized businesses are becoming more acutely concerned by the risks and threat landscape, causing them to seek economically viable security solutions for the SMB market (reference earlier "Security as a Line Item" blog posting).
  • Tightening budgets inside enterprises and governments, causing those organizations to seek economies of scale for security solutions
  • Shortage of qualified analytical talent, causing organizations to consider de facto analyst "time-sharing" arrangements
  • Movement towards a "SOC Center of Excellence" model, allowing organizations to focus on their primary business (which is most often not security)
  • Vastly increased interest in publications and blogs discussing the cloud
If anything, I would argue that the recent press on breaches has helped to accelerate the move to the cloud that was already underway. Each new breach that comes to light likely causes several organizations to move from the thought stage to the action stage. Perhaps the year of the cloud is upon us?

Thursday, March 13, 2014

New TLDs

Recently, ICANN has delegated 100 new top level domains (TLDs). For example, it is now possible to register and use domains ending in .best, . fish, .vacations, and many others. Additional TLDs are on the way in the near future as well. The complete list of domains that have been delegated, and to whom they have been delegated can be found here: http://newgtlds.icann.org/en/program-status/delegated-strings.

There are many reasons why the list of TLDs was expanded. Instead of discussing the reasons behind TLD expansion, I would like to discuss the implications of this TLD expansion to security operations.

For starters, TLD expansion means that it is now even easier than it already was for attackers to register and use malicious domains to carry out attacks against organizations. For example, there are now an even greater number of options for registering exploit, payload delivery, callback, update, and drop site domains. Previously, we had seen attackers leverage the "user-friendly" .cc and .ms TLDs (among others) extensively because of this. I'm sure that the list of "user-friendly" domains has now been expanded considerably.

So what can an organization do to try and stay ahead of, or at least current with, the threat? Fortunately, network traffic data can be used to provide us an analytical approach to tackling this challenge. Let's take a look at some steps we might be able to take proactively to assess what TLDs are required for business operations versus for which TLDs we can consider putting controls in place:
  • Begin by running an aggregate query over several weeks or one month of network traffic data and aggregating by TLD with count. The idea here is to cover a large enough period of time so as to get as complete a picture as possible regarding normal business operations.
  • Note all TLDs that do not appear in the query results but do appear in the TLD expansion list referenced above (i.e., there is no network traffic data to those TLDs). For example, we might not see .best, .fish, or .vacations in the query results. Because it does not appear that these TLDs are necessary for business operations, controls can be put in place to block/deny traffic to and from these TLDs.
  • Note all TLDs that do appear in the list and have a high count (a large amount of traffic) to them (e.g., .com, .org, .net, etc.). A large amount of traffic indicates that the TLDs are important for business operations and should be left untouched. Note that I am only talking about controls at the TLD level here -- specific known malicious domains can and should still be blocked.
  • Note all TLDs that appear in the list and have a low count (a small amount of traffic) to them. Drill down into this traffic and analyze it more deeply. Determine whether the traffic is legitimate (i.e., necessary for business operations), recreational, suspicious, or malicious. If the traffic is not required for business operations, consider putting controls in place to block/deny traffic to and from these domains.
The threat landscape is continuously evolving. As security professionals, we continually seek opportunities to proactively protect the enterprise. In the case of the new TLDs, we can use the network traffic data and our analytical skills to allow the data to guide us towards better controls that protect our organizations without negatively impacting business operations. The data is your friend. Use it.