Skip to main content

Agentic Security: Scaling Threat Detection with LLM Agents and Lakehouse Architecture

Summary

  • Adobe built Caspian, a security lakehouse on the Databricks Data and AI platform, that normalizes multi-source security logs into a three-tier catalog structure—landing, curated, and team-specific—giving LLM agents a consistent, queryable foundation for threat detection across the entire enterprise.
  • LLM agents in Caspian extract indicators of compromise and TTPs from threat intelligence feeds, automatically write hunting queries against normalized security data, and orchestrate parallel investigations across email, identity, cloud, and endpoint systems simultaneously.
  • Delta-based workflows replace rigid job dependencies, severity scoring automatically routes confirmed threats to the appropriate response teams, and SQL warehouses enable detections to run faster than attackers can move through an environment.

Agentic Security: Scaling Threat Detection with LLM Agents and Lakehouse Architecture

Watch: Agentic Security: Scaling Threat Detection with LLM Agents and Lakehouse Architecture
Scattered security data hides threat patterns that don't exist in isolation. When endpoint logs, identity events, and network traffic sit in separate silos, security teams miss attack chains that span systems, even when all the signals are there. This talk demonstrates how to build an agentic security platform using LLM agents and a lakehouse architecture to correlate security data and orchestrate threat hunts at enterprise scale.
You'll learn how Adobe built Caspian on Databricks to normalize multi-source security logs into a unified data schema organized into landing, curated, and team-specific catalogs. See how LLM agents extract indicators and TTPs from threat feeds, write custom hunting queries against normalized data, and orchestrate parallel workflows across email, identity, cloud, and endpoint systems. Discover how Delta-based workflows replace rigid job dependencies, how severity scoring routes threats to response teams, and how SQL warehouses enable detection faster than attackers move.
🤝

Chapters

FAQs

What is Caspian and what security problem does it solve?

Caspian is Adobe's security data platform built on the Databricks Data and AI platform that solves the problem of security signals being scattered across isolated systems. It normalizes logs from endpoints, identity systems, email, and cloud services into a unified schema, allowing LLM agents to detect attack chains that span multiple systems and would otherwise remain invisible.

How do LLM agents perform threat hunting in Caspian?

LLM agents in Caspian ingest threat intelligence feeds, extract indicators of compromise and TTPs (tactics, techniques, and procedures), then automatically generate and execute SQL hunting queries against the normalized security data. Multiple specialized agents run in parallel across email, identity, cloud, and endpoint datasets to detect coordinated attacks.

What is the Pyramid of Pain and why does it shape Adobe's security approach?

The Pyramid of Pain ranks threat indicators by how much disrupting them hurts attackers—hashes and IPs at the bottom are easily changed, while behaviors and TTPs at the top are much harder for attackers to modify. Caspian's architecture focuses on detecting behavioral patterns at the top of the pyramid, making defenses more resilient against adversaries who can trivially rotate low-level indicators.

How does Caspian route threats to the right response teams?

After agents complete a threat hunt, Caspian applies severity scoring to classify findings by risk level. High-severity findings are automatically routed to the relevant security response team based on the affected system domain—email, identity, cloud, or endpoint—enabling faster response without requiring manual triage of every alert.

Full transcript

[00:07] Today's session is architecting agentic security. Honestly, it is about how we use LLM agents and lakehouse architecture to scale threat detection and also fleet sweep orchestration. Fleet sweep orchestration means it's about hunting
[00:23] bad actors in our environment. That is like fleet sweep. How we orchestrated I'm just going to walk you through. Everything you see today is like a blueprint of what we implemented in production and some of the things that we're still experimenting with agents and all. We are still experimenting. So,
[00:39] some of the things are like more in experimental fashion. So, I'm going to provide some blueprints as well. Before we jump into the session, a brief overview about myself. I am Saikiran Uppu. I'm currently working as a senior security researcher and AI researcher at
[00:55] Adobe. I primarily work in cyber threat research and intelligence team where we kind of focus on adversarial tracking, infrastructure tracking. We also look at APTs, nation threat actors, etc. We use various techniques of research and
[01:10] automation AI to hunt dark web forums, breach forums. And also we have premium feeds, we have open source intel, we get intel from various industry partners. We get uh intelligence from government agencies
[01:26] like FBI, etc. to about latest cyber threat actors, criminals, etc. We use that signals and process them to identify those attackers are doing some activity in our industry or we also host a lot of customer
[01:42] information. So, customer run managed infrastructure, we also identify if those attackers are having any malicious activity in those environment. Apart from that, our team also do a lot of threat profiling basically to identify threat actors and model them so that in
[02:00] any future attacks we can model to the previous threat personas. That is primarily our team's responsibility. Apart from the work, I'm quite active in the industry. I'm speaker at B-Sides down initiative other conferences as well.
[02:15] Jumping to the straight to the problem. Uh On the left-hand side you can see this call it pyramid of pain. This is the ranking order of the indicators in the increasing order of the brain pain that is caused to the attackers. Attackers are any bad entity that kind
[02:31] of attacks our organization. So, if you see it is all put it in the increasing order of the pain that is caused to the attackers. If we remove that particular indicator of the out of the attack chain. For example, if you see on the bottom half, those are hashes, domains, and IPs. So,
[02:47] if you kind of block an IP address, attacker will simply pivot to another IP address. If you block a domain, attacker will simply create a new IP new domain. It's like a so cheap. It's like a $12 per domain. So, they're going to pivot so fast. For attackers, it's easy to
[03:02] create like create these uh indicators in the lower half of this pyramid, but it's for the defenders it's very hard because we try to block one IP address, they come up with a new IP address. It's very hard. But, the when we shift to the top off of the pyramid,
[03:18] that is where it becomes harder for the attackers. They can't really quickly pivot. Those are network host artifacts, tools, and TTPs. TTPs means techniques, tactics, and procedures. This is not like which IP they are attacking from.
[03:34] It is like how they are attacking. What is their behavior? Do they always follow certain methodology? Do they If they do certain payloads, do they do first step after second step? What hops do they do? Every time the attackers kind of follow the similar methodology
[03:49] or behavior. They don't change that often. They kind of change on infrastructure from AWS to Cloudflare. uh These are all proxies. They kind of change quickly, but the TTPs they kind of like it's like a their behavior. So, they can't change it fast. So, most of
[04:04] the traditional security tools and detection rules kind of sit in the lower half. We kind of have this knuckles security rules and also detection engineering logic entirely built on the lower half. We hardly have anything built on the top because it's very hard
[04:19] to map those and write it to the detection rules. So, if we miss those, the intrusions or even you are recently seeing all the supply chain attacks, all the GitHub packages, all the with everything with AI related threats, those bypass
[04:35] the lower half and completely focus on the TTP part. So, one takeaway from this slide is like you can't really hash a phone call. You can hash a binary, but you can't really hash a phone call. So, the detection leaves in the bottom half. We're going
[04:50] to talk about some of the problems in the next slide. So, imagine this is like a real threat I'm going to talking about. These are some of the real examples like in the past couple of months. Imagine like
[05:06] imagine a phone call right at like 2:00 a.m. to help desk. The employer suddenly calls saying that I have an urgent meeting. This is my manager. This is my building number. They get this information from social engineering. They are using LinkedIn these days to gather all this information and they
[05:23] call it to the help desk saying saying that simply can you reset my MFA device. Help desk simply confirms couple of couple of information related to the LDAP username and manager information and they simply add that MFA device
[05:39] because the help desk is generally being trust trustworthy and being want to help the employee, they simply add that device. Because this MFA adding MFA devices completely sit in the top off of the pyramid that we discussed in the previous slide, it completely bypasses
[05:55] the detection rules that generally we set up for any malicious IPs, malicious C2 domains, or any payloads. Here, if you observe, there is no attachment, no domain, no hash, nothing. So, it's all about the trust because help desk generally trust people with information. They
[06:12] simply verify that information and add a new MFA. Once they add a new device, the you get entire SSO, entire everything you get through that new device. So, that here the trust is the TTB. So, once the attacker bypasses the trust, he
[06:28] don't need to bypass any MFA. That's the take away take away from this slide. I'm going to simply say like once once this process is done, the attacker is in our system. The signals do exist, but they are like scattered across various places. Because we don't have a
[06:44] correlation correlation among these kind of TTB, we kind of don't aware like this person is like adding a new MFA and through which they are able to get access to various our infra. We don't have any correlation and we completely missed that. If you look
[07:01] at the slide, right? There are three separate entities that the signals live in. The first entity is where the all the IDP logs like all entries sit. So, you can see how new factor got enrolled and new device got added. These all
[07:19] when you look at in isolation, those are all correct. Those are all a traditional employee device logs sit in. Those are all correct. In the second second logs, those you can see service desk desk tickets. It shows like how incident is open, how
[07:35] the help desk person confirmed with the manager's name, building name, etc. And how the incident is resolved. In the third bucket, you see all endpoint related logs. This is like how the new device is added, how the new VPN connection is made, and there are no detections, basically. This is like an
[07:51] isolated silo. If you look at it, it is all looks good. You might be thinking that this is like a more like an identity problem or like a help desk training problem, but it is much beyond. Recently, if you might be aware of Team PCP or Shy Hulud, they kind of target
[08:08] these kind of attacks. Even Shiny Hunters and Scattered Spider, they kind of trick people into gathering information from LinkedIn's social engineering and do these kind of attacks because these are very hard to catch unless you have pretty good training with the help desk. Even
[08:25] employees needs to be well aware what they're clicking, how they're communicating with threat actors. Next example is This is This is like a pretty recent example May 26th. It is about the AntV a supply
[08:40] chain package. If you look at the left-hand side, right? There's like a CI build log. You might have seen this like tens of thousands of times. You can see how the normal steps are installed. Like you can see like NPM packages installed. Everything is like normal.
[08:57] It It says like some pre-install and post-install and it had added 840 packages. Everything is normal. But this is like a mini Shy Hulud campaign that targeted the AntV NPM package developer. They kind of
[09:12] compromise them. They injected this backdoor into the package. And the backdoor What the backdoor does is while the in the forefront it looks all clean, but in the background what it's doing is it's installing a bun runtime. And after that it's dropping like a 500 kilobytes
[09:29] of package. So it is all in the background. And once that is installed, it is exfiltrating all the secrets. It's enumerating all the cloud cloud enumeration endpoints. This is like I IMDS and also it is trying to
[09:45] probe every secrets within that git CI build and it is extracting all the secrets from the memory. And once that NPM install succeed at the last step, it exfiltrates
[10:01] all this information to a C2 server at 443 port. This is all hidden. There is no fishing line, nothing. No fishing line, no malicious domain to block. Here that TTPs are like depend trust on the supply chain dependency. We
[10:16] kind of trust these kind of packages in the dependency graph. That's where the TTPs they exploit. And like you may ask like how these this kind of information we missed. So
[10:32] every detection layer you reach had an answer for this. That's how the mini sha hollowed package this particular package. They had the hash and signature scanning. So instead of packaging one binary, they had like 640 NPM packages. They kind of changed slightly the
[10:48] obfuscation method and the how they decrypt the package and they produce 640 packages. If you would have just strictly gone with the atomic route where you block indicators, you would have missed this completely because you have to block every single artifact even the signature like if you
[11:05] have written error signature or some other signature, it won't block them because these are typically created to bypass those methods. And even the GitHub secret masking, right? Instead of directly looking at environment variables, they are directly
[11:22] looking at the memory within that host. So instead of instead of going through the environment route, they are directly going through the memory and capturing all the secrets. That's how traditional endpoint security devices or any other solutions present even hit
[11:37] those kind of solutions. I highlighted few other things like how the attacker is crafting this malicious packages. Here the key point is like atoms kind of change 640 times, but the behavior didn't because irrespective of watch
[11:52] supply chain NPM package they published, the entire behavior or their patterns remain the same. So, here is how we started like it's like a 2-3 years back. We have this problem of segregated data sets. Signals
[12:09] do exist. We have monitoring for all the endpoint devices. We have like identity logs etc. But the problem is everything sits in their own console. For example, identity log sits in their auth console. Endpoint device sits everything in the
[12:26] same index. There is no correlation. Every network logs sits in their own their vendor console. Everything is like segregated. The problem is like the analyst tags, right? Like if there is an intrusion or a breach or a signal, we
[12:41] need to open like hundreds of tabs to investigate and there is no correlation. There is no single column correlation between any log. So, that's the real problem. The The problem is like signal signal was never missing. Signal is always there, but the problem
[12:57] is there is no proper correlation or SE single tag that we can associate from the entire pipeline. So, the idea is like what if we can introduce certain tags so that we can track from the entire all the application layers, all the network layer, every layer possible.
[13:18] So, so we build this platform called Caspian. So, it is our security lakehouse on Databricks. So, every source that you've seen in the previous slide, which is all scattered, we kind of normalize that into single nomenclature, single data schema. So, every
[13:35] every data set, we brought them into this funnel kind of solution where everything sits into one Caspian data lake. We identified common common columns on every log source
[13:50] so that even date, even source, where the log is coming from any correlation columns with other data sets we have identified and mapped it into this particular solution called Caspian. So, here it's like one single source and one
[14:06] single schema for all the data sets. So, that's how we designed this and we have even normalized the column names so that everything is in the same nomenclature. We have published an article like I put this in the slide. We have published a detailed
[14:21] blog on how to normalize these data sets and all. We have have two workspaces to test it like a stage and prod where you can test all the other data schemas and how the data is ingestioning in the stage and the prod is where we kind of utilize it
[14:38] for our data threat hunting solutions. Coming to the next slide, we have built this into three catalogs, classic medallion of in security terms. Landing is where all the raw logs come in. Here we ingest all the sources from
[14:54] EDR, firewall, all the you know, pan traffic, all the OS traffic. Every kind of traffic you first come into the landing landing landing view and then we develop these views that is like more curated. I think all the data experts know this. So, it
[15:11] is like a purpose-built and these are clean and normalized. For example, one example that I want to point out is you have a lot of endpoint detection logs. These are huge petabytes of scale. So, we have split it into various views. For example, all the process logs come
[15:27] into one view, all the network logs come into one view. So, it's easier to search and all the indexing will happen on these logs. So, we get faster results. We are almost have able to achieve like 15 minutes worth of time so that we can search at
[15:42] least 90 days of fleet sweep in our environment. And finally, we even developed the team specific catalogs where specific teams and enriched with the CMDB data sits in these particular catalogs where a lot of our teams
[15:59] research and put our data lakes in this particular catalog. So, the now the big picture is like how we are Now, we proceed to the big picture.
[16:14] So, as as the new emerging threats come as I discussed in the first slide either through various public blog posts or premium feeds or even our dark web forums or breach. How do we even digest that, right? Like we created the system
[16:31] internally. It kind of scans this and identifies these particular indicators. Indicators are what I mentioned in the first slide. It can be atomic indicators or TTPs. It extract both of them. And a typical extraction looks like a needle and the
[16:47] haystack is like almost like a 10 petabytes of data that we collect on a daily basis. How do we search? That's a big problem. We on the right we have the Caspian haystack. It's almost like 10 petabytes of EDR data alone. Additionally, we have
[17:02] all the like network logs worth of 200 TB, identity of 400 45 TB, etc. So, here is the shape of the challenge, right? Like 4 KB of TTP extraction. How do we search in this big haystack? So, the
[17:19] architecture is the bridge. We have used this architecture of agentic orchestration where agents can help us hunt in those big haystack. I'll show you how this works in the next slide. So, the first step is like we created
[17:35] this different agents for purpose-built for threat collection, threat digestion, etc. The first agent is what reads on the internet. So, every day a lot of well-known sources publish a lot of articles with the detailed
[17:51] indicators, detailed TTPs. Even you may know Mandiant Unit 42. These publish detailed articles. So, we kind of hook these agents on the internet every few minutes like 20 minutes. They go on internet, fetch any new article which is relevant to the
[18:07] threats that we see. So, once they identify the threats, we are using this pre-filtering logic because these massive hunts are quite expensive. Because to scan for the last 90 days or even 365 days in an environment, it's
[18:22] quite computationally expensive. So, we are using this multi-gated approach where we using some LLM logic to identify if there are really indicators in the in the context. And also is the indicators and TTPs even relevant to us. A lot of
[18:39] times what we see is this threat is only applicable to certain industries. For example, a lot of threats are applicable to only banking. So, we kind of like put it a lower priority. If it is applicable to technology and even other industries, we kind of rate it high.
[18:55] Also, we have a large footprint on Acrobat and Photoshop across the billions of devices in the internet. So, we do have like like touch points on various devices. So, we do consider a lot of other industries as one of the
[19:10] core threat areas. So, uh once the LL agents extract the articles, it kind of gives the verdict like should we process it or should we skip it based upon various filters that we use. So, I just highlighted two examples where it can skip based upon
[19:27] the confidence. Uh we also having a lot of evaluation for this. We using LLM judges to even uh evaluate the work of these agents. So, uh just in a few examples, right? In the first article, it was like a basic
[19:42] uh like a security awareness article. It kind of skips through that article. It doesn't process. And the second article is more like a threat actors abusing OAuth tokens. That's what we discussed in the first uh slide as well. So, uh agents read it. It understands. And it also uh we provided a knowledge base of
[19:59] our past uh investigations, incidents, and everything. We provided the knowledge base. So, it can able to associate this particular threat actor with any of the previous uh threat actors that we've seen in our environment. So, it gives it a high priority and it gets uh fastest
[20:15] processing. So, the next one is like as I mentioned, the agents kind of read and understand and extract the TTPs. So, this is the agent that turns the pros into something you can hunt. Uh no human uh is involved
[20:31] in this. Everything is like done by agents. So, agents kind of read. They kind of write these queries. These queries are custom written. Uh They kind of know the entire data schema as I mentioned in the previous uh slide as how we correlated various sources and brought like a common data schema. So,
[20:48] we are giving this data schema as a input definition for these agents to write those custom queries. So, most of the queries uh like it it was able to write very good. And we are using uh LLM evaluators to uh even make the queries better.
[21:06] Uh And you can see on the left and right. This is like a very basic bullet like a blueprint. You can customize this and how we build that actions and also how you make the each each uh
[21:23] behavioral analysis step and how to write queries better for each of the step. Uh I just give you one example. And also at the last we are also using something called grounding. So, a lot of times LLMs kind of hallucinate. So, what we are doing is every time we write the
[21:38] query, we are re-evaluating the query in back in the article or back in the threat digest so that it is like correct query and not like a like a some hallucinated query. Because these queries are computationally expensive, we are pre-checking in the back in the source
[21:53] and then we are executing this in the background. Uh coming to the next slide, this is how the entire orchestration workflows. I just give a few examples. It's not like
[22:09] a entire suite. This is like a five or six five or six nodes. What we start is like we keep these agents on a loop where every few minutes they hunt internet, extract indicator, extract TTPs and use our internal data
[22:26] schema and it knows context about all the log sources in our environment. And it writes custom queries and launch this custom mini agents that what you call. It It launches email agent, auth agent, cloud agent, etc. And it hunts for
[22:43] It starts from simple first. Do we have any email hits? Like suppose if a threat article mentions about a email address from which these particular phishing campaigns are getting launched, it kind of searches there first. If there are any hits, it proceeds to any URL based
[22:59] hits. So, did any of our employees click any of that URL? Similarly, it will proceed to authentication hits. Like it decides on its own. We don't need to put them in an order. So, it decides on its own which to launch, which not to launch. So, that we computationally save a lot of
[23:16] compute on our end. Similarly, if the attack is coming from any cloud infrastructure or it's targeting any IM users or any secrets in the cloud, it particularly launches the cloud. So, we do have footprint on three clouds. So, it knows context about all the three
[23:32] clouds and all the queries how to hunt them. And uh one of the interesting use cases we do have the endpoint systems on our entire employee devices and even our cloud footprint and also data center. So, it have this custom curated views of all
[23:49] the data sets. So, it launches these custom queries because as I mentioned the data set of the EDR is quite in the terabytes. This is one of the challenging hunts that we have to perform. And finally, once the agents kind of
[24:06] talk among each other and collude to one particular verdict, what is like actually happening? Is the threat real or the or the threat like not real or the mostly the hits that we see in our environment are mostly false positive. Should we need to escalate it further?
[24:22] Once this agent decide the home human in loop comes in. So, like lot of our team kind of investigate the final output and they kind of either need to escalate it to hunting team or detection engineering team or even to see sort response for immediate escalation. So, that is how we
[24:38] are performing all the like a fleet sweep orchestration. This is everything is present in the data bricks. We are using the jobs and auto launch feature where you can custom build this like a plug-and-play architecture where you can add more data sets based upon
[24:55] your environment. Uh coming to the next slide, right? How do the writer and five agents agree what they're done? Uh so, this is like a one delta row we are writing. We initially started
[25:10] experimenting with the jobs and tasks in Databricks that you know. Then we kind of encountered issues where a lot of jobs are waiting for indicators. A lot of times the blogs are like dark web forums are down. So, we kind of took this very in the naive
[25:25] approach, right? Wired the Wired the all the jobs in a like a sequential manner. So, every job is waiting for the previous one to complete. Then then it took us a long time because even a lot of times the indicators are stale.
[25:42] You know, these IP address and domains keep changing. So, we have to keep the blogs updated with the information. So, that's how they we took this approach where we started listening for this poll-based poll-based jobs where they kind of poll
[25:59] back to the domains or something like that and extract the indicators and proceed to the next task. Uh we kind of divided our entire orchestration into three workloads.
[26:15] These are more like a three runtimes. One build that like a makes sense. We the writer node. This runs once per article and once per article. These are heavy and rare. So, it kind of fetches fetches the article extract the indicator extract
[26:31] the IPs and it calls all the other other agents to do the job of hunting. Actually, this is like more like a continuous loop where we look out for all the dark web forums as well and breach forums extra. So, we get very fast intel before any of the other
[26:48] vendor tools comes into picture. And the next one is like hunts. This is all the SQL warehouse. You might see in in the data bricks you have an option to use the SQL warehouse. This is like a five jobs that I've seen in the previous slide, but you can expand it to cover
[27:04] your entire orchestration fleet. So, one of the advantages by using this is like the speed and also the parallel processing. It is quite efficient with all the read heavy right light kind of loads. And
[27:19] it also it also has like caching feature. So, it is like good if you have the same queries running again and again. So, that's one of the hunts workspace that we use for SQL in warehouse. And finally, it is the listeners. It is more like an ad hoc
[27:36] ad hoc base workflows. This is what we are using for the all the polar's. All the polar's and agents how agents communicate with each other like a beaconing. So, one agent initiates a communication with another agent using this beaconing. So, it's very ad hoc
[27:52] traffic. So, we are relying mostly on the serverless serverless workloads that you see in the data bricks. Coming to the next one. Once the all the agents agree on the particular verdict,
[28:09] we kind of post it internally in our either through various communication channels like a messaging or like email. It kind of immediately escalates. And also the one of the feature that we implemented is the priority because we almost like get like 200 different
[28:25] blog posts like breach forum alerts. How do we even prioritize, right? So, then we come up this come up with this order prioritizing metrics where if a lot of times the attacker simply probe our environment, there won't be any real
[28:41] attack. So, they simply hit our firewalls simply hit our DMZ zone that we are treating as lower priority. If we are seeing any traffic that is originating inside our environment and going to the external entity more like a C2 domain that we are
[28:57] treating as a high high severity. Also, as I mentioned previously, we are also tracking email email hits. If there are any clicks on the email that we are treating as high priority. If there are no hits, simply it's like a fishing or scam that we are
[29:14] treating as low priority. Based upon the correlation between various data sets, we are increasing the priority or decreasing the priority. Once the priority is set to critical, it kind of escalates it to the C-CERT for immediate response.
[29:33] Also, about what agents are doing today. So, as I mentioned, we are doing like three agents. Those are already in production and we have fourth experimental agent. So, the first one is like triage agent. This is This is the agent that goes on internet, checks for every article, and every indicators it can extract and
[29:49] bring it to the environment. This decides whether it's worth the fleet. It can search the title. It can search the content. And once it decides on the verdict of skip or submit, it goes to the next agent.
[30:04] Agent 2 is more like an extraction. It understand deep into our environment. It understands our past investigations, incidents, and it can correlate whether this new threat is even worth investigation worth investigating. And the third one is the agent that
[30:20] escalates. So, once the fleet's entire fleet sweep is done, it kind of escalates to various other teams. So, the fourth agent that I mentioned is like more experimental. It's already running in production. It kind of auto hunts. Basically, they
[30:36] read the work curated data schema and it auto hunts instead of we manually writing queries, we going to going to the different dashboards, we're writing different queries, it kind of generate these kind of queries and auto run them and extract the results and then post it
[30:53] to the like like other agents for the critical escalation. Here the LLM's kind of route and extract and the SQL warehouses is where the entire fleet sweep is done. Right now we
[31:08] are using this like a hierarchical approach where any threat that is critical, we are running it for the past 30 days for immediate escalation. If it is more like a historical scanning, we are approaching it for like a 90 days to see if we have any historical matches in our
[31:24] environment. Finally, this is the slide just just wanted to take away from the slide uh One is like shape computes the work based typically based upon the workflow
[31:40] and the workload. You can decide which compute environment to use. Three workloads, more like a three runtimes. Heavy and rare run on job clusters, read heavy and parallel on work houses, and idle and bursty on serverless. So, if you're like
[31:56] serverless are much costlier. So, if you have idle and bursty workloads, you can use on the serverless. Two is like decoupled through shared state, not job dependencies. Instead of approaching through job dependencies where a lot of the jobs kind of stuck,
[32:12] you can use this one delta row kind of approach where jobs write to a SQL database. So, based upon the other agents can read these rows and launch their respective jobs. Three is like let the LLM route the work
[32:28] and let the SQL do the job. So, let the agents think and write the queries and SQL warehouses actually run run and process the output.
[32:46] Finally, these are the takeaways final takeaways like kill one one job dependencies and make it more like a polling approach. That's what is working better for us and second one is like a give the most boring work to the agents. We have seen seen some tremendous
[33:01] progress. We have able to identify some of the threats faster than any of the vendors that publish indicators through our threat intel platform. A lot of indicators that come to the threat intel platform are much like delayed because our agents are hunting faster on
[33:17] internet or we also have some personas on telegram channels or dark web. We were able to know about a threat faster than any of our threat intel platforms provide us.
[33:35] Yeah, I think that's all I have for today. If you have any questions, you can always reach out. This is more like a blueprint. We have a detailed blog that published on our website and also about the Caspian platform. You can refer to that and if you have any questions, feel free to reach out to me.
[33:52] Thank you.

Learn more about the Databricks Data and AI platform.

The information provided herein is for general informational purposes only and may not reflect the most current product capabilities or configurations.