Real-Time Cyber Threat Intelligence Platform with Databricks
Summary
- Centripetal processes 1.9 trillion threat data points annually on Databricks, yet most organizations deploy only 5% of available cyber threat intelligence despite over 99% of breaches being preventable.
- The platform uses semantic transformation at the collection point to enable constant-time threat enrichment, eliminating costly post-processing joins and meeting the 27-second attacker breakout window constraint.
- AI classifiers accelerate analyst productivity within a deterministic foundation, ensuring detection accuracy without introducing hallucination risks into mission-critical threat identification workflows.
Real-Time Cyber Threat Intelligence Platform with Databricks

In 2025, enterprises face a critical challenge processing and operationalizing global cyber threat intelligence at scale. Over 1.9 trillion threat data points stream annually across thousands of feeds, yet organizations deploy only 5% of available intelligence despite breaches being preventable 99% of the time. Cybersecurity is fundamentally an intersection problem: massive telemetry must correlate with comprehensive threat intelligence in near-real-time to be actionable.
Learn how to build a real-time threat intelligence platform using Databricks and Kafka to stream global CTI into production. Discover semantic transformation techniques that push detection logic to the collection point, enabling constant-time threat enrichment rather than costly post-processing joins. this video covers pipeline architecture, deterministic foundation principles, and how AI classifiers enhance analyst productivity without hallucination risks.
🤝
Chapters
00:00Introduction to Proactive Cyber Defense00:24Global Threat Intelligence Scale and Challenges02:29Why Breaches Happen: The Intelligence Gap04:38Telemetry, Detection, and Intelligence08:27The Coverage Gap: 90% Intelligence, 5% Coverage10:54Scale: 1.9 Trillion Threat Intelligence Points12:35Real-Time Constraints: The 27-Second Breakout14:13Cybersecurity as an Intersection Problem16:02CTI Standards and Complexity Challenges22:29Pipeline Architecture: Semantic Transformation30:46AI in the Threat Intelligence Pipeline34:02Real-Time Detection Engine and Flattened Data35:24Deterministic Foundation and Human Analysts37:50AI Analyst Interface and Operational Results
FAQs
What is a real-time cyber threat intelligence platform?
A real-time CTI platform correlates streaming threat data with telemetry in near-real-time to enable proactive cyber defense. Centripetal's platform ingests global CTI feeds using Databricks and Kafka, processing 1.9 trillion threat data points annually to detect and enrich threats as they occur rather than through expensive post-processing.
Why do most organizations only operationalize 5% of available threat intelligence?
The coverage gap exists because traditional approaches require costly post-processing joins that slow detection pipelines, making full-scale intelligence operationalization impractical. Centripetal addresses this with semantic transformation at the data collection point, enabling constant-time enrichment without complex downstream processing so far more intelligence can be applied.
How does semantic transformation improve cyber threat detection?
Semantic transformation pushes detection logic to the point of data collection, so threat enrichment happens as data streams in rather than after the fact. This approach allows constant-time threat lookups, making it feasible to operationalize a much larger share of available intelligence within the tight latency window that modern cyber attackers exploit.
How does AI fit into a deterministic threat intelligence pipeline?
AI classifiers are layered on top of a deterministic foundation to enhance analyst productivity and surface patterns at scale. This design keeps mission-critical detection logic rules-based to avoid hallucination risks, while AI accelerates analyst response by prioritizing and contextualizing the most significant threats.
Full transcript
[00:08] So, thank everybody for being here. I'm honored to present. This is the last session of the day, I believe. So, you guys are really troopers. I guess a couple of slides that I have to put on on behalf of Databricks. Forward-looking statement. And this talk is about
[00:24] proactive cyber defense with global cyber threat intelligence, scaled analytics, and AI acceleration. So, my name is Devan. I head up the intelligence services at Centripetal, which is the company who I work for. Centripetal is one of the largest consumers of threat intelligence in the world. And we do quite a lot of that
[00:40] data processing on Databricks. And we're going to get into a little bit of how we do it. Some of the really high-level challenges that we encountered in the process of building an intelligence analytics platform. And and the work is the little in the little corner right
[00:55] there. There's a number of people who have contributed a lot to the work that we've done over the years. So, we'll begin with this. How many of you have seen this movie? Gone in 60 Seconds. Good movie, maybe, maybe not. The premise of the title, the movie has
[01:12] actually nothing to do with this talk, by the way. But the premise of the title is that you can if you have a thief who is well experienced, then you can put any car in front of them and within 60 seconds they can take it. And
[01:27] the parallel is that the title has something to do with the challenges that we face in cyber security today. That is, thanks to ChatGPT, a fake movie starring some very famous, well-known threat actors across China, Iran, and
[01:42] other places. And you can think about how one might be breached in another 30 seconds. And the reason why this is important is is that Did I just miss something? Okay. So, we're going to view the
[01:57] cybersecurity lens a problem through a different lens. Most of cybersecurity right now is focused primarily on getting the telemetry in, using the intelligence or the detection engines to kind of detect what you want to find based on an interrupt-driven process,
[02:14] right? So, there have been a lot of talks actually throughout yesterday and today regarding ingestion of large amounts of data, especially telemetry, uh but probably a little bit of uh maybe fewer sessions or fewer attention towards the flip side of
[02:29] that problem, which is on the intelligence side. So, I'm going to start out with a statistic. So, did you know that in every breach that has occurred in the history of computers, over 99% of the time there was information that was known beforehand
[02:45] that could have been used to stop the breach. Uh this is a well-known research. Um it's based on breach reports from, you know, Verizon DBIR, IBM CISA. So, that's a fact. What is cyber threat intelligence? I know most There are many of you who don't, you know, who aren't kind of new or maybe not as familiar with the
[03:01] cybersecurity space. So, cyber threat intelligence is the work product of a cyber analyst. What they do is they do research. They look at telemetry, they look at anomalies, they'll do malware um you know, detection malware detection engineering or reverse engineering, uh maybe actor tracking. Um they can
[03:18] even do financial crime tracking. So, they they do all this research. There's probably 10,000 or more if not 100,000, maybe not 100,000, but tens of thousands of individuals who work in this space. And the work product that they produce is threat intelligence. It's information about the actor, why they might be attacking somebody,
[03:34] what their motivations might be, where they're coming from. Maybe they're using infrastructure in certain countries or certain bulletproof hosting um entities. Um what kind of tools and techniques do they use? Are they looking to scam somebody out of $500, do some phishing
[03:49] emails that get spread out to a million email addresses or are they targeting your organization? Right? Are they going after the health health care sector or are they going after government? Are they trying to make a make a statement about leveraging, you know, cyber threats or attacks to support certain political or
[04:06] other kind of goals? The impact of a breach is, according to IBM data breach report, is around four four and a half million dollars. That's what it costs when the breach doesn't get detected early enough
[04:21] and it doesn't get stopped enough early enough. So, the question becomes if 99% of intelligence already exists, why do the breaches happen? So, we'll delve a little bit into that. So, this is a a Actually, I saw this
[04:38] graphic from from another presentation, very similar to this, and this describes what most of what a security operations team does in terms of looking at data that they bring in to try to detect and analyze and report on on threats and
[04:55] findings. On the left side, you see that there's a big green bar, telemetry, and the whole cyber threat analysis process can be boiled down to really three things. One is you've got telemetry that describes everything that is happening within your network. So, it could be
[05:12] your network perimeter, like your firewall, it could be your switches, it could be your endpoint, it could be your applications, it could be your database. All these logs are pulled in, and for those of you in the cybersecurity space, you know that they typically get sent to a SIM or maybe if you're sophisticated or advanced enough, you have a security data lake,
[05:28] some products like Lake Watch. And this problem in terms of ingestion has is is already well known. There are lots of products out there that does this, and lots of detection engines and capabilities. Um there are now efforts like open cybersecurity schema framework, OCSF,
[05:45] where they're trying to standardize this um so that it's easier to use. Uh you can properly put in you can put in a semantic layer on top to kind of interpret, let's say if you've got logs from two different types of firewalls, you can kind of view them in the same way. Um OpenTelemetry is another kind of
[06:01] standard that allows you to bring in various types of operating system or application level telemetry um that is uh that can help you aid in investigations. You know, for example, if somebody were to breach your laptop and get control of it and they're remotely executing
[06:17] commands on it, then all that telemetry regarding the processes and the users that are logged in, the programs that are running, the memory that's being used and so forth, all that can be can come back and that can help investigators kind of decide and assess what those risks are, what has happened, and how to mitigate.
[06:35] There's not a lot of work that's been done in terms of the focus on threat intelligence, but that's really the other half. The threat intelligence is used everywhere. In fact, pretty much all, actually every detection-oriented product today, modern product, uses
[06:50] threat intelligence whether you know that or not. Think about your antivirus. Right? It's got definitions. It looks for malicious files. Uh think about an EDR or if you even use Google, you know, um browser and you get the safe browser warning saying, "Hey, you're going to website that looks malicious. Do you
[07:07] really want to continue?" So, there's a lot of intelligence that's embedded in the products, but not much of it is actually available uh to purchase, to operationalize, and intersect with the telemetry. And the biggest part in the middle is the detection part. All right, so this is where all of the security operations
[07:23] analysts really focus their energy. So, there's an entire field that's relatively new called detection engineering, where you've got data scientists, you know, cybersecurity experts, uh maybe engineers, and they work together to kind of ascertain, well, what do you see
[07:38] in the telemetry and what do you see in the intelligence, and using that combination, you basically write detection rules. You know, if if a piece of malware, for example, gets downloaded from a website, but the way that you download that is first clicking on a link that goes to
[07:53] some known good website to verify that your internet connectivity is good, you download it, maybe the malware will then get executed, and it'll persist, and it'll carry on certain actions. So, all of these different tactics, techniques, and TTPs can be quantified, can be labeled,
[08:10] and can be described. And the detection engines are really built around, well, how do you take that information and try to find instances of those indications of breach. So, let's look at
[08:27] the breakdown. As I mentioned earlier, telemetry the the telemetry problem has mostly been solved. In other words, everybody who's working in the cybersecurity space already has, generally, a SIM. They already have the capability to bring it in. Yes. And according to various analyses, um it looks like most advanced
[08:44] or sophisticated enterprises have about 90% coverage of their telemetry. So, that's pretty good. You know, you've got some shadow IT, you've got some applications, legacy systems that you can't get telemetry out of, but that's pretty good coverage, you know, over maybe 80-90%.
[09:00] On the flip side, most enterprises have about 5% coverage of intelligence. So, the intelligence is driving the detections, and you've got all the telemetry, but if you only have 5% of what's known bad, what is the probability of you actually
[09:15] detecting an arbitrary threat? In fact, we've done a bit of uh academic research on this over the last few years, and we know, based on evidence, that the maximum probability of success of any enterprise having the best tools, the best people, the best capabilities to
[09:32] detect detect to detect any arbitrary threat is somewhere about 4 and 1/2 to 5%. So, there's something broken here, right? And how do we attack this problem? And if we think about um that 99% number, if the intelligence
[09:48] already exists, could you hypothetically gain access to all of it? And if you had access to all of it, could it change the posture? Could you actually detect and stop all these threats? So, the breaches actually happen in those
[10:04] error points. It's the 90% or 95% of intelligence that you don't possess on average. And it's probably in that 10% where you don't have visibility into your own environment and your telemetry. And what's unknown is how good are the detection engines? It's hard to
[10:21] ascertain that because all the detection engines are right of today are written based on the intelligence that they possess or the behavior activities that you see in the telemetry. So, it's very very hard to determine whether the detection engines are giving
[10:36] you complete coverage, but you have to start with at least maybe 90% coverage in telemetry and maybe 90% coverage in threat intelligence, then you can opine about whether the detection engines are actually doing their jobs. So, let's look at that. Let's look at the totality of intelligence that is produced in the world today.
[10:54] 1.9 trillion threat contexts produced in 2025. Prior to that, it was a lot lower. For this year, it's projected to be around 4 trillion, and I think we'll actually exceed that by quite a bit. To give you just a little clarification
[11:11] on what this intelligence context means, it's finished intelligence that are sold by intelligence producers out there in the world today. So, this is not the telemetry that's producing the intelligence. Each of these those intelligence indicators the context is probably based on millions and billions
[11:27] of telemetry that those producers are consuming and analyzing. This is the work product of the tens of thousands of of of of analysts. So, this represents about 100 megabytes per second of threat
[11:44] intelligence data per second. Megabytes of data that's streaming in that in order to consume that intelligence, you have to process. And this is where Databricks and the analytics comes in. This 1.9 trillion also represents
[12:00] the totality of all organizational cyber threat risk that is known to man. Pretty much. And that represents about 80 to 90% by volume. So, it probably sees about, you know, maybe 2 trillion. But, that's the number. If If you're thinking about assessing risk of an
[12:15] enterprise against cyber threats, well, that's it. You've got 9 trillion data points active today. And it's changing rapidly. That translates, by the way, to about 3.2 petabytes per year.
[12:35] All right. So, this next part is from the recent Cycraft Strike uh cyber threat report where they observe that the fastest time to uh breakout time for an e-crime is 27 seconds. So, the first dimension that we were looking at was how much totality of
[12:51] volume of data that we have we have to consume. But, now we're looking at the constraints placed on the detection, right? So, we have to consume the telemetry fast enough. We have to consume the intelligence fast enough. And then the detection engine has to be fast enough and has to be faster than 27 seconds. Now, on average, it's higher.
[13:08] It's, you know, maybe 30 minutes or an hour. But, that trend has been diminishing quite a bit. And in fact, the fastest known privately discussed um breakout time is under 20 seconds, so around 12 or 13 seconds. And this is driven
[13:23] quite a lot through our automation. Also driven a lot by AI. Malicious actors are using it. You've heard a lot of you know media articles about mythos and so forth. But the trend this is that this is going further and further down and it's it may take us maybe another year
[13:39] to get to 10 seconds or on average or or or five. It's certainly not going to get to zero. But our ability to consume data and operationalize that intelligence to actually make a difference in detection has to be measured in seconds. So it's a real time system that we have
[13:55] to we have to build. And in fact there's a lot of discussion or claims around well what is machine speed? Machine speed is probably less than 1 second end to end. So that's what we're striving at. So let's look at what that means here.
[14:13] Um I think at the cyber security keynote um Dom from Apple mentioned that he had a slide up that said cyber security is a data problem. Well we all know it's a data problem. But what kind of data problem is it?
[14:29] I would put put forth that it is actually an intersection problem. Right? Because you have massive amounts of telemetry and this is these are actually real numbers. Right? So you can have an enterprise or a multi-enterprise like an MSP that's consuming 86 billion events
[14:45] that you have to search through let's say over um 24 15 to 24 40 out of 1 week period. So that's the pool of telemetry that you're looking at. And you're while you're looking at that you're probably ingesting a million events per second over a 24-hour period.
[15:01] Some enterprises are much larger than this. Many are smaller. On the flip side of that you've got out of the 4 trillion. Right? So you've got about 11 billion pool of indicators that may be relevant at this very moment at any given second. And then we have data that's coming in
[15:17] at 127 context per second. So fundamentally you have to build a real-time system from the beginning, end to end, from consumption to processing to analytics to the data lake.
[15:32] And everything that you do in the data lake has to be real-time. You can't do batch-oriented approaches anymore because any latency that you add will push you towards exceeding that 30-second goal. If we can
[15:47] process all of this, and we solve this problem, then you can operationalize the totality of intelligence, and you can de-risk the organization from all of the known threats. That then allows you to bubble up zero-days, the things that are very,
[16:02] very difficult to find. So, I'm going to dig a little bit into cyber threat intelligence. I know, you know, some of you here uh are not familiar with this. Uh I mentioned, you know, in terms of, well, yeah, it's a threat actor, you know, context. It's the um you know, the TTPs, the tactics,
[16:17] techniques, and procedures. Uh it's the endpoints, like the URLs where you're downloading it from, the hashes of the files, um you know, the the series of procedure steps that, you know, a threat actors take that typically tend to be uh representative of their their approaches.
[16:33] But, in threat intelligence, the standard is a little bit more difficult, right? In telemetry, I mentioned, you know, standards like OCSF and OpenTel, OpenTelemetry, where most solutions already provide some level of support or adoption, such as Databricks. In cyber threat intelligence, it's
[16:49] almost a free-for-all. Okay? There is a standard called STIX, but if you look at the complexity of this, right? This is essentially a graph. It's essentially a graph with a lot of relations between objects. You have observables, uh you have, you know, actors, you have campaigns, you have all
[17:05] these things, and then different producers of intelligence will represent them in different ways. On the right side is a JSON snippet. And very often, what you have are graph-like concepts that are delivered as disjoint
[17:23] objects. In other words, you've got 10, 20, 30, 100,000, 10,000 objects that you have to combine together, join together to figure out which actor is doing what. Or if you're looking at a particular malware, what kind of observables or indicators
[17:38] represent that kind of malware activity, or whatever it is, right? So, you the pivot points require a lot of joins. And going back to the need for a real-time analysis, if the logic that you have in the middle
[17:54] to apply that intelligence requires you to pull together 30, 40 different tables and go over multiple columns just to extract that what to find, you've already lost because the amount of scale that you're talking about in in decomposing
[18:10] and processing these nested JSON becomes too extreme. If you're only looking at maybe, you know, 100,000 events per second, you could probably do it. But not at the millions. So, a lot of the techniques So, in the process of what as we were building this platform,
[18:26] one of the earliest lessons that we learned was that we would prototype. We would prototype solutions for detections, for transformations, for semantic layers, and things like that. And we would find that, oh, it works great with 1,000, 10,000, 100,000, a million records. But then we would push to 10 million,
[18:42] 100 million, a billion records, and suddenly the performance would actually increase or degrade exponentially or logarithmically. And furthermore, we would run into the bounds of of the the the limitations of the platform itself. For example, Spark drivers.
[18:57] I can't tell you how many times we crashed those things, and how many times we had to up-scale them and fit them into bigger VMs or bigger nodes, only to have to rescale them up again because the intelligence volume keeps changing. And some of the big challenges around cyber threat intelligence is is the
[19:12] delivery, right? So, this is just one ex- example, but if you look at let's say um so, in the ecosystem there are about 400 threat intelligence providers commercial that sell their their intelligence today. You can go out there and buy them if you have the money. Probably 80% of them
[19:30] have their own APIs. Many of them do support STIX, but it's partial. So, they actually give richer, better information if you use their direct APIs. Most of them uh give you bulk exports, which means that you go download a
[19:47] million records, then you have to process all of that, right? Others, what they'll do is they'll actually give you three different files and say, "Well, you've got to join them together based on some key." Others yet will actually give you a differential API. What that means is I can say, "Give me everything that you
[20:03] know today, and then tomorrow I can say, 'Give me everything that you've learned that's different from today.'" Really fantastic, but very few intel providers support that. The challenge here is that when you're trying to bring in all this data into a data lake, then you've got a lot of custom transformations that has to be done.
[20:19] And in that process lies a lot of challenges. I'm going to go into how we attack this problem. Part of it is in Databricks, and part of it is outside. And this is very very different from a traditional ETL type of approach because we're actually pushing
[20:35] the semantic layer all the way to the left. So, what I mean by that is threat intelligence from a practitioner's perspective is really just data, right? It's coming in in a very different structured way, but you have to interpret it, and you
[20:50] have to understand how to apply it in order for that intelligence to actually have an impact. That's what investigators do, or if you try to pull that intelligence onto the wire or into a detection engine, that's what the detection engines do. The issue the challenge is is that when
[21:05] we try to create a single funnel of all this intelligence through a single pipe, it didn't work very well because every provider has a different definition for the same thing, and they have a different score for the same thing. So, one provider's risk score, let's say,
[21:20] um of a 99 might actually be not as good as another provider's 87. Uh one provider might classify, you know, a a threat actor as APT 123. Um The other one might say it's crazy bear.
[21:37] So, there's a lot of semantic standardizations and normalization and mapping that needs to occur. And one of the biggest challenges is in, well, out of this complexity, if you're going to build a detection engine for this one threat, how much of
[21:53] that information do you need? It's rarely all. It's actually a small percentage. But, in the industry, we're focused on, let's bring all the data in in its raw form, then let's do the transformations and analysis and the
[22:09] semantic analysis in the data lake. It works. We tried it. But, it doesn't scale. So, this is what we ended up building. So, this is the very beginning of the cyber threat intelligence pipeline.
[22:29] What we have are data collectors. There's many, many of these. Okay. Each data collector is specific to a provider and specific to the semantics, the meaning, the the underlying data that's coming from each provider. All of this is running on a Kafka-based
[22:45] pipeline. Uh and the reason why we do that is because we have to take the data out of Kafka and run analytics against it before we put it right back in. So, it's a multi-stage pipeline, but the pipeline is actually a network. Okay. It's a pipeline network. The very first thing that we We used to
[23:01] do the semantic analysis much later. Like we would uh ingest everything to a bronze table, maybe elevate to a silver by doing some semantic analysis. Maybe we'll customize and create a materialized view for gold. Very typical approach.
[23:16] It's difficult to scale that when the definition of what is different changes. Okay? So, if you think about that intelligence, when you have a data collector and think about CDC, like CDC is a great example of a capability of Dataworks. Well, you can't say well, it's very difficult to say for this particular
[23:32] input the definition of a material changes that you've got to ignore these 10 different fields only account for these other 10 different fields, and only allow variations of a certain range in these other fields. That only works for that one provider. It doesn't work for everybody else.
[23:48] So, the ability to create the define the semantics of what means in order to detect a threat based on the intelligence you have to define that up front, and it's actually very much tied to the original data that comes from the source. And by the time you push this all the way into a data lake you lose
[24:05] that that context because everybody kind of brings it together. So, we ended up actually putting a semantic transformer way way up front. So, at the time of collection, we collect everything, we dump the data into raw, and what that does is we can always replay it. We don't even put it into a structured
[24:21] bronze table. We dump the raw so that we can replay the original download, then we can actually determine, well, was the semantic mapping incorrect? Uh was it insufficient? And you'll see why that's important later. The semantic transformer, what it does is
[24:37] it actually looks at the intelligence that's being provided and it's a combination of just the data as well as a function. So, to apply intelligence, you have to define a matching function. In other words, if you have an IP address right? And you're looking it for it in
[24:52] in in a log or in traffic well, where does that IP address have to show up in order for this to be a match? That's critically important. That is the matching function. The matching function could be extremely sophisticated. It could be like well, go offset into the
[25:07] payload and maybe 200 bytes and look for these extract out and maybe a section of it and then hash it and then see if there's a match. So, the semantic transformer essentially allows us to push the detection logic
[25:24] that normally is done downstream all the way up to the front and you make it part of the set. And within Kafka we use Avro. What that means is we have built-in versioning of the schema. So, when we ingest both the data and the matching functions and the
[25:39] criteria, that means that if the matching function changes, if the provider decides one day that domain name of badguy.com should not be an exact match, it should actually be an entire subdomain match. The indicator didn't change, but the matching function changed.
[25:55] That constitutes a relevant change. Now, if you have data that comes in and says um so, that we do a a semantic transformer and then after that we do a semantic CDC, change data capture. What that means is is that once we apply the semantic
[26:11] layer on the data and the transformation, we can then ascertain will the did the effect of that actually change downstream? Is it really a new threat context? And did it affect the change in the detection engine or the detection rule? If it didn't, then
[26:27] there's no change. It's almost like saying um you know, you've got an IP address, you've got some threat context and then maybe you you have some comment in there saying it's uh uploaded or provided by some some additional descriptive information. If that descriptive information is not
[26:43] relevant for detection it's considered to be ignorable during the CDC process. So, the semantic CDC is not a data CDC, It's a semantic layer CDC. Once we do that, then we immediately go into and this is again a pipeline, a real-time pipeline. We go into a
[26:59] preemptive join and detection scoped views. We construct that. We're transforming the data into a view that becomes very easy to look up. Right? So, if you look back at this
[27:14] the schema here, right? If we're trying to join the data against a threat actor or an incident or a particular TTP, we want to pivot and structure the data in flight so that down the road you don't have to go through a complex query
[27:32] or complex merge. What you want to do is you want to join on a single key if you if possible, two keys if possible, and so you pivot. You pivot and rechange the schema of the data containing only the semantic information, and that allows you to enable much
[27:47] faster queries down the road. So, this approach is generally uh it's a data parallel approach, right? So, you're replicating the data multiple times, but we're doing this in flight, which means that the only limitation that you have is the network bandwidth. And since we're dealing with single sets
[28:03] of data, right? So, you're collecting intelligence right now, we're not talking about gigabytes of data. We're talking about maybe megabytes of data. So, you can hold that information and do the transformations in flight one by one, and it's extremely fast.
[28:21] Once we do that, we actually capture that context that is pivoted on the lookup keys, and we store that into a gold table. This gets stored into Databricks. What this allows us to do is we can do extremely quick in Databricks joins and merges, and we can also have uh an
[28:38] in-flight um joining. From there, now we continue on down the path. What we we do a detection CDC. What the detection CDC means is that in the first part, all we're doing is we're looking at what the data has
[28:53] changed and whether meaning of that data has changed. If the data has not changed, then the next thing that we need to look at is, well, did the function change, right? Did the detection technique change? So, we then look at computing,
[29:09] well, should the application of that threat context be different than what it was before? If it isn't, there's no need to update downstream. Right? So, every every step of this way, really, every time you have a CDC, you have a reduction of data. So, even though you're ingesting up front 4
[29:25] trillion data points in a given year, you're reducing in the real time that that that that volume. Now, you're losing some of that through the the data parallelization where you're like if you have five different pipelines in a forgiven provider, then
[29:40] you have five x replication, but that's just temporary, right? So, it actually ends up being faster. The goal here is to make it fast, not necessarily make it efficient in terms of storage. Once we have the CDC, then we generate, essentially, or classify them into what are detection rules. So, most threat
[29:57] context that you're trying to detect will fall into a certain classification. If you're looking for an IP address, it's only going to be present in certain parts of the packets or certain parts of the logs. If you're looking for a domain name or URL, they're only going to be present in certain places. So, in most cases, you can classify a lot of these
[30:13] uh these these intelligence indicators or contexts into a bucket of detection rules, but sometimes those detection rules have variations. For example, that domain name example, you could have a regular expression. You might match a whole bunch of them, or you might have a suffix, prefix. Could
[30:29] be uh a a partial. You're looking at a relative URL of a uh path of a URL. So, all that generates placement into detection rules. Let's move on. Um where does AI fall into this category?
[30:46] Or or into this pipeline. The very first part is is that when we do the joins of all the threat intelligence that comes in, what happens is that at that stage all the different data that's coming in from the different providers are essentially being merged into one. So
[31:01] now you have a complete context of all threat information that exists for that pivot point for let's say for that IP address. So at that point we can use an AI classifier to say, all right, what information that is present in the data did we not classify
[31:19] correctly? I can assess that. And then we can train it to essentially create a new transformer or a change the transformer as code. So the key here is is that we're not using the classifier to actually process the data. We're
[31:34] actually using it to produce code. The code is deterministic. You can look at it and you can tell whether it's broken. Now we're moving on. At the tail end of that pipeline is the detection rules. Those detection rules go into a rules transformer. The rules transformer again gets down uh
[31:52] pushed into a snapshot tables um and we split them into current snapshot, what's active right now, as well as historical. And again we do preemptive uh add, merge, divide, uh and remove detection rules, and we push that up to the preemptive joins of detection all the
[32:07] way up front. So what we're doing is anytime we have a change in the semantic analysis that is triggered downstream, we use that information to actually trigger a change upstream to say, well, we're going to change the semantic approach. So it's all about pushing the analysis
[32:25] up front to the beginning of the pipeline. The human analysts and the AI analysts, they're simply working with that context to produce the placement of those rules. So what happens after this? All the detection rules go into rule transformer, they drop into Databricks,
[32:41] they get into the snapshot and historical gold tables. Then we'll generate deployment policy policies. In other words, what part of that intelligence is relevant for a particular detection engine and push them out. Each of those detection engines could be per customer, it could
[32:56] be per environment, it could be specific to the risk profile. So, all this is done in a pipeline. The only dependencies that we have on on Databricks is just the reason rights of the data source. It's extremely quick.
[33:14] So, what are we really doing here? I mean, that's a lot of graphics, charts. We're shifting the semantic analysis and the detection rule creation all the way to the beginning. It's the collection point. Once we receive it, most of that is actually driven by how the producer of that intelligence tells us to
[33:30] actually action their information. They tell us exactly what to do with it. We want to identify relevant threat data for detection, make those pivots, and then CDC the semantic differentials, and then push all of that into detection
[33:46] rules that get deployed into detection engines. Those detection engines are what actually map to uh uh logs. So, now let's look at telemetry. This is very very simple. This problem's already been solved, right? So, raw logs coming into a semantic uh transformer,
[34:02] and then you have semantic logs with CTI. Very very simple. The detection engine is quite interesting because it's actually a simple concept because what we've done is we've taken what would normally take a detection engine a very long time to
[34:18] process through the structured nested data, and instead we've turned that data, we've dereferenced it, we've flattened it so that every comparison that you have to make can be done in constant time or log n time, which means it's extremely fast. So, the joins
[34:35] the latency that we're able to achieve with this kind of pipeline is we have a latency The latency is a delay. It's not an It's not an insufficient capability to process. It's just that it takes us about 1 second, less than 1 second, to actually process and enrich all the telemetry as it's coming in before it even hits Databricks
[34:52] at the golden platinum layers. The latency of processing that intelligence is about under 30 seconds. We're striving to make it under 15, but it's pretty good given that it used to be a lot longer than that. Used to be 15 minutes. The latency of doing the intersections
[35:07] through these transformations less than 50 microseconds. We can do 200 million lookups per second to essentially enrich and classify telemetry that's coming in. This is real time.
[35:24] So, I'm going to close out with what is the final impact of the on on on the consumer of this, right? It's the human analysts, the people who are sitting in front of a system, a SIM, a data lake, and trying to interact with this data. So, one of the key aspects of this is that in order to leverage intelligence
[35:39] at scale, you've got to use a deterministic foundation, right? So, it means that telemetry and intelligence are facts. You don't want to make them subject to AI-based errors. You don't want to have hallucination around, "Well, did it really say this?" So, this is why there isn't very much AI in the pipelines.
[35:57] So, that means that whenever you have data, evidence that suggests a threat activity it's based on deterministic data. It's based on the truth as you know it at that time. So, there cannot be a misinterpretation of the data in that case. The second part is the
[36:13] determination of presence of risk is also based on the intersection of that, right? So, the algorithms to do that intersection is computer science. It's not AI. It's not making a guess. What that means is if that threat is observed, it was there. It was there. The only question is
[36:28] was it significant? Was it material? So, that's the part that actually humans do, right? We as humans and analysts will go and say, "Here are the symptoms. I believe, based on my investigation, that this is a threat actor with legitimate threat activity." So, that's a guess. That's an opinion,
[36:44] but it's based on the facts that are presented to you. So, we've modeled it a very similar way. So, the assessment of the impact of the risk is opined by the AI analyst that we are leveraging. And so, the only questionable part that the AI analyst can get wrong
[37:01] is that it's misinterpreting or going to the wrong conclusion based on the facts that are presented. That limits the scope of the impact of a bad AI decision. That's really, really important. This process has allowed us to to deploy the 2 trillion threat
[37:17] indicators last year. Our false positivity rate was under 350 last year. Previous year was 333. So, I don't know how many zeros are in front of that. It's quite a lot. But this is unheard of, to be able to operationalize that much data, attach that context to it, make
[37:33] decisions, and have basically one intervention per day across the entire customer base. That's astounding. So, I'm going to finish off with a couple of slides of how our analysts are using AI to interface with this. We have spent
[37:50] probably 95% of our engineering and data science time building the data foundation and the pipelines. That's where the investment is. Then we built a harness. The harness, I think we built in maybe a month or two. And we're using an off-the-shelf LLM that's sitting
[38:05] outside. Very few token uses. And what we're doing is we're driving the access to this data that is deterministic, that is semantically formatted. And we're simply asking the AI to do the things that a human analyst does.
[38:21] So, this is a couple of snapshots of of that interaction. Our analyst can go in there and say, "All right, what should I do at work today? I just walked in." It's going to go in there and give you the priority hunt of the top unwritten findings that it discovered based on the intelligence and the intersection of
[38:37] that with telemetry over the last 24 hours or whatever that period is. And you can see examples of this customer has three separate high-scoring malicious domains active simultaneously. Right? Endpoint isolation is warranted. There's evidence that's being collected based on the intelligence, the presence of it on the logs, and it's making
[38:53] determinations and recommendations on what should be done. Here's another example. You can then say, "Well, let's dig into that particular set of domains for that." Just ask it a question. And what it's doing is that it's pulling that intelligence context together that is already there and making assessments
[39:10] and a report essentially, writing up the findings on what it thinks it is. And our analysts are interacting and saying, "Yeah, that looks right." Or no, it doesn't look right. You can see these descriptions of, you know, things like DNS layer blocking. It's not there. There are, you know, it could be buried by high volume, etc.,
[39:25] etc. And it'll go in there and orchestrate, query additional flows to figure out whether there's evidence to support that finding. And they'll generate summary assessments. All this off of two prompts. Third prompt. Give it an instruction that, "Well, this looks real." The analyst makes that
[39:41] decision and says, "We're going to draft the finding. Please do it for me." It'll summarize everything, make it succinct. And you can ask it to generate a completely written report that brings in all the threat context, all the reasons why the detections actually happened, and all the
[39:57] underlying evidence. Create a narrative and hand this on a silver platter to the analyst to make a decision. So, this is the outcome of investing in real-time data analytics and semantic analysis to position the
[40:14] data for actual consumption by both humans as well as AI. Thank you.
Learn more about the Databricks Data and AI platform.
The information provided herein is for general informational purposes only and may not reflect the most current product capabilities or configurations.