Phishing Triage at Scale: Agentic AI with Knowledge Graphs and Databricks
Summary
- CVS Health built LangFish, an agentic phishing triage system on Databricks, reframing the problem from classification to investigation—combining knowledge graphs for long-term memory with multi-agent orchestration to detect coordinated phishing campaigns across 200,000 employees.
- The system deploys specialized sub-agents for analyzing email headers, sender reputation, URLs, and content in parallel, with MLflow tracing providing explainable decisions and audit trails that meet enterprise security and compliance requirements.
- Knowledge graphs connect related emails across campaigns through label propagation, enabling the system to recognize patterns that span multiple messages and continuously improve through feedback loops as analysts validate triage decisions.
Phishing Triage at Scale: Agentic AI with Knowledge Graphs and Databricks

Phishing attacks exploit human trust with increasing sophistication, yet traditional classification approaches struggle to adapt. this video reveals why phishing triage is an investigation problem, not classification, and demonstrates how CVS Health automated their phishing response using agentic AI, knowledge graphs, and Databricks to reduce analyst workload at enterprise scale.
Learn how to build long-term memory with knowledge graphs that connect related emails across campaigns, design multi-agent systems with specialized sub-agents for headers, senders, URLs, and content analysis, and use MLflow tracing for explainable decisions. Discover patterns for consistent reasoning, continuous improvement feedback loops, and real-time streaming pipelines that scale to 200,000 employees.
🤝
Chapters
00:00Introduction and Problem Statement00:39Phishing vs Classification: Investigation Framework02:16Phishing Definition and Evolution05:45Phishing as Entry Vector and Security Risk06:35Why Agents are Needed: Limitations of Classifiers and LLMs08:58Four Challenges: Scale, Consistency, Memory, and Explanation11:05System Architecture and Data Pipeline13:43Knowledge Graphs: Building Memory from Email Data16:35Graph Conversion and Label Propagation19:51Evaluation and Threat Posture Framework23:33LangFish Implementation: Real-Time Streaming Architecture27:37Multi-Agent Triage System and Sub-Agents28:31Consistency, MLflow Observability, and Continuous Improvement32:20Example: Detecting Credential Phishing with LangFish34:45Conclusions: Investigation-Based Security at Scale
FAQs
Why is phishing triage an investigation problem rather than a classification problem?
Phishing attacks are coordinated campaigns where individual emails may appear borderline but reveal a clear pattern when connected to related messages. Classification assigns a label to a single email in isolation, while an investigation-based approach uses memory through knowledge graphs and a structured process to assess emails in the context of known campaigns and attacker behaviors.
How does LangFish use knowledge graphs for phishing memory?
LangFish builds a knowledge graph from email metadata, sender domains, URLs, and content patterns, connecting emails that share infrastructure or content with previously seen phishing campaigns. Label propagation spreads threat classifications across the graph, so a newly observed email connected to a known malicious campaign is flagged even if it appears clean in isolation.
What are the specialized sub-agents in CVS Health's LangFish system?
LangFish uses a multi-agent architecture where specialized sub-agents independently analyze different aspects of each email: headers for routing and authentication anomalies, sender reputation, embedded URLs for malicious indicators, and content for social engineering signals. A supervisor agent aggregates their findings into a final triage decision with an explanation.
How does MLflow tracing support explainability in phishing triage?
MLflow tracing records the inputs, reasoning steps, and outputs of each agent in the triage pipeline, creating an auditable decision log for every phishing verdict. This enables security analysts to review exactly why the system made a decision, supports continuous improvement by identifying where agents make errors, and provides audit trails for compliance purposes.
Full transcript
[00:06] So we're going to go ahead and do some introductions. Um, Virendra? Uh, good morning everyone. So I'm Virendra Dhiman. I'm lead data scientist at CVS Health and I'm going to present the session here with Andrew. Hey there. Uh, my name is Andrew Henson. I am one of the distinguished engineers
[00:22] at CVS Health and I focus on data science more recently agentic AI as it pertains to security. All right. So what we'll do today, our talk, as you know, is automating phishing triage with agentic AI and knowledge graphs memory on Databricks.
[00:39] Uh, and so what we're going to do is kind of jump in, but one of the first things I want to kind of like level set is uh, triaging phishing isn't exactly a classification problem. So I think we tend to think of phishing or spam and we kind of maybe confuse the two.
[00:54] Uh, but one thing I the primarily would like for you to walk away with is that really when you do phishing it's more of an investigation, not a clearly a triage or classification problem. Uh, and so what we will hope to convey today is that if you can combine memory with a structured process, and of course
[01:11] with agents, then you can effectively do uh, triage at scale with agents. But from the perspective of an investigation. All right. So I'm going to spend probably more time than necessary, but I'm going to spend a little bit of time just talking about what phishing is.
[01:28] Uh, and I want to kind of hopefully motivate uh, what the problem is. Uh, and so oh goodness. Okay. Technical problems. Uh, but so I'll just start with kind of a joke. I'm not good at jokes, but hopefully this works. Uh,
[01:44] what did the hacker, oh sorry, hacker's out of office message say? Any any guesses? Okay. Gone fishing, right? Fishing talk. Terrible joke, so I apologize.
[01:59] Okay. So, but imagine you you're at your desk and you receive this message. Uh and today, you know, FIFA Cup and you want to click on this link. Uh and interestingly, if you were to click on this link, uh it could compromise your system and set off a chain of events that could be really bad for you, your
[02:16] personal finances, your company's health, uh many things. Uh and so, then, you know, the question is what what exactly is phishing? And then later, we'll talk about what differences with with spam. Phishing in this case is really just a social engineering attack with the aim to get your credentials uh
[02:32] to do something malicious. And so, it's basically someone sending you uh something to entice you to uh click on a link or go to a download an attachment or anything of that nature. So, like phishing, I'm sure we've all heard of it, but again, just to level set for this audience, this is this is the idea
[02:48] that we'll run with. So, one thing I want to do is draw the distinction between phishing and spam. And so, they might on the surface seem simple or similar, but you'll see hopefully that the there's a big difference. Spam is in general innocuous. We we don't really There's no
[03:04] malicious intent other than just to annoy you. Uh and then if you kind of, you know, read spam, you really just waste your time. You don't really compromise your your financial health or your company's health. Uh it's just just an annoyance. And then there is a plethora of things that exist to kind of help you deal with spam. Maybe not the
[03:21] perfect solution, but many solutions exist. Uh but one thing that's different is that phishing is is more polymorphic in that it the the attacks vary. They change maybe on the course of a day, months. Uh they're pretty sophisticated uh and they're more difficult to carry out effectively. Uh and then more
[03:38] importantly though, the intent of phishing is really like I mentioned earlier, is to steal your credentials to do something uh malicious. And then if you if you make a mistake with phishing, instead of wasting your time, you could potentially uh open your company up to a breach or your financial uh, health and
[03:54] things of that nature. So, now I want to kind of walk through a little bit of history uh, just to kind of like how did we get here and and where have we where have we been? So, phishing kind of started in the early 90s and it was kind of a AOL thing with chats. Uh, so people were going to the
[04:10] chat and they kind of uh, put messages to get you to give them your password. So, maybe pretend to be a friend or whatever other scenario. And then later in the 2000s we started to see phishing kind of increase to just sending out a bunch of messages just like you fish in real life and hope for
[04:26] the best. So, the spray and pray kind of idea. And so, you send out a bunch of fake messages. Hopefully someone clicks. The remediation was simple. You can kind of create spam block lists and things like that. Uh, but then in 2004 we started to see more targeted exercises where attackers would understand uh, maybe some
[04:42] information about the uh, victim uh, and then they could maybe cater the message specifically for that victim and then that's where we started to see things like spear phishing, uh, really targeted research, deliberate efforts. Uh, and then in the 2000s we started to see that evolve to more business email
[04:57] compromise, invoice fraud, credential harvesting. Uh, and then as you've seen more recently supply chain attacks. Uh, and then now today with LLMs we're seeing really sophisticated attacks uh, where grammar is is perfect. It
[05:13] sounds like a human. It knows things about you. Uh, it's rules, standard methods don't necessarily apply because these things are polymorphic. They change uh, as time goes. So, like as you can see phishing has evolved alongside us uh, and it's only really getting worse. And at scale
[05:28] of agents it's it's only going to get worse. Uh, and then so really then why is it that phishing works? It's the simplest possible attack. It just needs you us to do something malicious. Uh, and then really it's just this needle in the haystack. If I can send if I'm an attack if I can send you enough information or
[05:45] emails or messages and get you to make one mistake then then they win and the attacker wins in that case. So, it's just the one time you let your guard down, you're attacked. Uh so, and then why is this so important,
[06:01] right? Because and again, I do apologize for spending so much time, but I I really want to impress upon you how important phishing is. Uh so, phishing is interesting in that it's the initial frame for the attack. So, some attacks start at different vectors, but phishing is really kind of the entry, the gateway. You can kind of see
[06:18] here like from a minor perspective, attachments, links, and then just if you click one of these, it can open up all of this just from clicking a bad link and then ultimately down to a ransomware type of attack. And so, just the the one link that you thought was okay, it the attacker, they're in the
[06:35] system and then they're able to move around just from that front door. All right. So, what what are we presenting here? So, in the case of Agari triage, where do we sit, right? Because we're not exactly uh up front trying to block in sort of a
[06:50] spam sort of filter. What we're actually trying to do is investigate after phishing has been reported. And so, there's a pretty big difference. And so, we're not sort of proposing a scalable way to do spam detection, which would be very difficult with Agari because you have millions and millions
[07:05] uh depending on the size of your enterprise, but what we're saying is more, okay, if I can use Agari Agari for phishing at the scale of a report. So, you have your standard front protection, uh users report phishing, what if we automate that part? That's really the goal of what we're presenting here. So,
[07:21] after your standard defenses have kind of either failed or let things through, this is where phishing what we're proposing kind of sits. Uh and so, like again, that's different from sort of a gateway type of filtering and then more towards when we report,
[07:38] then we actually do the triage after it's been reported. Today that actually happens, you have socks and things of that nature that go in and do reported fishing. Um the big thing here is that it's not it's not filtering, it's investigation. So that's that's a key
[07:53] point. All right. So in more of just, you know, some what if what can we do? Uh so can we take classifiers? Can we take uh LLMs? Classifiers sort of work. So you can train a machine learning model. Naive Bayes is really common one. You have different approaches.
[08:08] Uh but imagine that attackers are changing over some period of time. There's a behavioral component. Uh there's components that are specific to uh a a campaign uh that will have different signatures. And so if you train or attempt to train a larger model to do that, it gets really difficult. Uh and then the volume
[08:25] can change. And so attacks can be set up over the course of weeks or months or days and then completely different a week, month, or day later. So then maybe we can throw LLMs at these attacks and LLMs could potentially work. But the problem with LLM, at least out of the box, uh of course there's no
[08:41] memory. And so you would go into every single investigation not exactly knowing what you did in a previous investigation. So not a great thing. Uh and so really then that means that really if what we can do, if we can use an agent, we can expose it memory and give it some structure. Cater it to
[08:58] whatever it is your organization does. So if your organization have a specific set of uh standard operating procedures, you build that into your process and then you structure your memory on top of it. And here in a minute I'll actually dive into really the core of what the memory looks like for this system. There's there's many types, but really
[09:13] how do we think about it from a knowledge graph perspective? All right. And so the big there is four takeaways here. So the problem is scale. So right, we we mentioned earlier that we can't investigate every possible alert. Uh very costly tokens and things
[09:28] like that. So what we really want to do is investigate things that are most malicious. And then at the same time we don't want humans to have to investigate every possible report as well. So, uh there's a finite capacity humans each individual human that reports a a phishing message can report thousands in
[09:45] a day or a month. You only have one human can investigate maybe two, three, four, five per day. So, that's the scale is is really a problem. Consistency is another problem. What we see from triage perspective from analysts, even expert analysts, is that there's differences in
[10:01] experience and previous investigations and decisions that the yield could be totally different. So, we need a way to ensure that the agents can repeatedly produce the same result. And so, that's a really key kind of concern. The next part is like how do we do that? And so, we would do that with memory.
[10:17] So, if we can get the agent better memory, better context, then maybe we can have more consistent results. More consistent results means we can explain them better in the future. And then and the final one is explanation. So, then we really want to be able to say, "Okay, if we have an agent go through and do this, how do we explain the result?" So, we really don't
[10:33] want a black box. So, an agent can go away and make a decision, well, how do we figure out if it made a wrong decision? Or if it continues to make a wrong decision. So, we have to have a way to kind of explain it. So, these are challenges that we have to kind of solve with whatever agent system that we put forth.
[10:49] And again, triage is is a isn't a classification. So, we we have to figure out a way, how do we do these investigations? All right. So, now what I'm going to do is kind of talk about the memory. So, probably the the fun part. So, I think we all are fairly good phishing is really bad and then we have to figure
[11:05] out it's different from spam, but we got to figure out ways to do it. So, what what we're going to do is just our system in general. This is kind of a high level overview. So, we have some data ingest. So, our data comes from wherever the source is, so whether that be
[11:20] say the graph API where maybe messages were reported, whatever your source platform is. You ingest messages in some vector. So, users are going to report messages. Those messages are ingested from some source and then arrive within your your system. Um and then what you really really want
[11:36] really want to do is stream that. So, as those messages are being reported, you really want to stream that. That's where Spark Structured Streaming comes into play. Of course, that's where Databricks comes into play. Uh and then as we bring streaming streaming those states, then we can actually process them. Then the next thing is like the knowledge graph. So, then we have this input. So, then
[11:52] what do we do with it? Like how is it useful? So, this is where we can decompose these down into an actual graph. But the fun thing is how do we how do we even structure the graph, right? So, we can do memory. We know memory's great. We know there's knowledge graphs. But how do we do that?
[12:08] And so, then after we have the knowledge graph, then we have to figure out a way to use that information usefully. And so, we what we really want to do is augment the agent. So, we want to expose the memory, but we also want to augment that memory with some type of classification, machine learning. So, um
[12:24] LLMs aren't perfect at predictions, but machine learning does a really great job. We can leverage methods from that. So, that's where label propagation, semi-supervised types of approaches come into play. And then from there, we can send all of this information as context to an actual uh multi-agent system to
[12:39] then have that the agent investigate. So, what we're ultimately trying to do through all of this is prepare the state in a in a way that's repeatable for the agent to then process. So, we need the agent to have a repeatable set of information, context that's optimized for the problem, for your organization,
[12:57] and then have that carry out the investigation part of it, which is different because each investigation may have a different trajectory, some standard components. And then finally, we have to align it. And so, we have to have some way to do evaluations and observability. We have to know what decisions are being made. If you have
[13:12] audit kind of strings, uh there has to be a way to do the evals. But today, I'm going to focus on really just two things uh because of time. So, one thing I'm going to focus on is the the knowledge graph. And then we're going to talk a little bit Well, knowledge graph label propagation. And then we're going to talk about the
[13:27] evaluations, which how we design for this particular task. And so, there's tons of evaluation methods, but we created a couple specifically for this task to help align the agents and help us troubleshoot mistakes. So, one of the first things so memory
[13:43] from an analyst graph perspective is, well, how how do you build a memory? So, we have a email that's reported. And so, the question then becomes, well, how do we extract relationships? Um and so, typically, if you use kind of an off-the-shelf product and maybe you're doing a Wikipedia
[13:59] mining, and you can kind of run that through standard graph database, it'll do the schema mapping for you. It kind of infers that. In our case, that doesn't necessarily exist, or at least not to our knowledge. And so, what we had to do first was define a way that we can actually
[14:14] structure or relate or build the ontology for the data itself. And so, what I'm showing here is kind of how we take every single message and decompose it down into nodes and edges that we then use within our graph. And then we scale this across. And so, what you see here at the center is the
[14:30] message. So, the message as it's reported, it comes in and we kind of just assign a unique identifier to the message. And then from that, we can extract information. So, what we can extract is who reported the message, who sent the message, the body of the message, the
[14:46] URL, and so on and so forth. And then each one of these, we have a particular type of relationship that we define. Reported by, has subject, posted and you can see here, we can start to do these secondary types of relationships, which are things like similarity-based.
[15:01] And so, if we take and deconstruct or decompose from the message the body, then we can actually build a similar message body from other messages. So, the idea here is that if we take and use this model, can we relate other messages? Do they connect to each other
[15:17] from this model? And that's the really critical thing, because if I report another message and there's never really a reporter that matches or there's never really subject uh similarities or anything like that then this node will always be by itself. And we have to connect this with other nodes. And so in
[15:32] this next figure what I'll show is that this is kind of a simulation of how that happens. And so suppose we have the original message and then later we have another message. And the way that these get connected in our graph, our knowledge graph, is through things like this. So where maybe the subjects that came in between the two messages that
[15:48] were reported were similar. We can do this using semantic sort of similarity methods or we can also do this with the message body. We can actually say okay well these two message bodies, while they may differ in name or subject or or or intent, the similarity, the semantics of
[16:04] the bodies were actually the same. And so but the other thing here is that each one of these we can kind of assign different weights and strengths because ultimately we need to connect the different messages together. So we have to then define how do we we measure the strength because perhaps the subject similarity is noisy over time. And maybe
[16:20] we find that many messages have similar semantic similarity from the subject. And so we have to actually build weights and things of that nature. Um but then the question is okay after we do all of that then what do we what do we use it for? So we have this graph data this knowledge graph. So the next
[16:35] thing is then okay we we can put data we can research it we can use standard graph methods. But then if we want to do label propagation we have to figure out a way to get from a heterogeneous graph which is a knowledge graph down to a homogeneous graph which is something that we can use standard
[16:52] graph theory type networking methods. And so if we're going to do label propagation we want to do label propagation on a homogeneous graph. And so in order to do that we need to be able to take the relationships that we infer from the bigger knowledge graph and then compress those down to the actual homogeneous graph. And so what we
[17:09] can do is define motifs. And so motifs are basically just similarities that we can map between the graph. And I'll kind of just walk through just a simple one. Uh you can have direct message to message. So, basically the same message that's reported um or you can have really interesting ones where you have a message to a
[17:25] component. Uh so, in this case, these two messages are related by the fact that they share a common host. Um or they maybe have a common sender or common reporter. So, this is a sort of a straight uh relationship that we can look for. So, we understand this relationship, we can build directed
[17:41] graphs, and then we can actually uh kind of search for this in our bigger graph. We can and we actually do this with graph frames within Databricks. And so, you can use knowledge graphs uh within Databricks or you can have a graph database. Uh the second component here is a uh slightly more complicated, which
[17:57] is what we call the message to component to component to message. And so, what this is is more you'd have a message uh that maybe shares a sender uh and then a host, uh and then these two are related through some other secondary mechanism, uh and then that host is related to this message. So, in this way, these two
[18:13] messages are related by this more complicated uh relationship. So, there's two hops uh from the original message to this other message, uh and we can look for those. And so, basically what we're doing is we're going to search the graph with these motifs in mind, and then we're going to build those motifs
[18:30] uh to then get down from this knowledge heterogeneous knowledge graph to a simple graph. So, then the question is, well, what do we do with the simple graph? So, once we have that, then we can actually help the agent by then propagating the information onto it. And so, in this case, what we're doing is we've gone
[18:45] from all of these uh interesting edges and relationships to just simply that these two nodes uh are related. So, we take a new message that comes in, we infer from the knowledge graph the motifs, and then we say, "Okay, this message is now related to these other messages, which then have
[19:02] these historical distributions." So, that's to say that this message is related to a previous message that was associated with uh credential phishing, and yet another one that was associated with credential phishing, and so on and so forth. And so what we want to do is to take the fact and say, "Okay, well, this message that's new with common
[19:18] elements to these previous messages and say that if it was associated with credential fishing recently, well, then there's a good chance that this is probably the same thing. Not perfect, but it's a good chance." And so what we'll do is we'll pass this information to the agent. And so as it considers all
[19:34] of the context as well as this information, it could then start to make a good decision. So that's one component of the memory. And Verender will talk a moment later about the other components. You'll see the investigation state where we actually then go into the graph and pull other information in as well. So then
[19:51] the next part is a So hopefully kind of convince you that we we have some form of memory, not maybe your traditional memory, but it is a form of memory within the graph. And then what I want to do now is to just talk about how do we think about this from a cyber defense or security perspective when we evaluate
[20:07] these. And so there you know, your standard evaluators things like hallucination detection, groundedness, coverage, and things of that nature. But what we need to do in our in our case is actually extend those slightly to make sense in our case. And what I mean by that is that let's suppose that there are what we call a
[20:24] threat posture. And so what we're doing is we can take some set of ground truth. And the ground truth can map to one of these three categories. And so the ground truth could either be malicious or safe or suspicious. And then we can take that as the the baseline. And then what we can do is
[20:39] measure our agent against whatever this ground truth is. And what we're specifically looking for is if the ground truth has been reduced. And so if we say that there's a malicious classification, then we don't want to ever have the agent say something is
[20:56] malicious and it be fault I'm sorry, something that is truly malicious and it turns it out to be a false negative. And so that's really, really bad. So what we would do is we want to weight that in our sort of score. We want to say, "Okay, it's okay to go up in the sense that, okay, well, you took a safe
[21:11] message and you said it was uh fishing for or investigate further. That's okay generally." Uh I mean, obviously, you want to tune it. But, it's never okay to go backwards. It's never to say that if we take and and run this on ground truth and it takes a previous message that was
[21:27] reported as fishing and it says it's safe. That's what we don't really want to happen. Of course, there's a trade-off where we need to optimize those. But, in our case, we're really specific on we can't go backwards. And so, that's a really uh key concern. So, in in order to do that, so this is why our standard like F1, obviously,
[21:43] accuracy, if we have a 99% um uh reported message that's usually spam. Uh so, we we can't really use our standard metrics. And so, what we can do is look at it in over the course of a few of these uh components. And so, ultimately, the panel of uh LLMs
[21:58] judges looks something like this. Uh I just kind of showed you what this uh PPI, which is what we were just looking at. Uh and then we have things like coverage. And so, what this is is if we take the ground truth, of course, you you have to have some ground truth, and we can just say we can look at a summary
[22:13] that's produced by our agent and then compare that summary to a summary that was produced by a human. And so, then the constraint becomes, "Okay, if I take this agent, which should be performing at the level of the human, well, does the agent actually produce a report that covers everything the human would?" The
[22:28] human being the expert in this case. Uh and we can actually use an LLM as judge to take the two reports and then tell us how well the actual agent that produced this uh triage actually covered the human analyst's review. Uh then, what we also have is like
[22:43] factual accuracy. So, is it reasonable for the agent to then infer its conclusion given the data? And so, we we're trying to say, "Okay, well, does the agent actually do uh is it reasonable to infer uh this conclusion given the data?" So, if you have 1 + 1
[22:59] uh and the agent said it was three, is that reasonable? I know because one plus one should be two in that scenario. And then this is the threat posture where we look at what we were just showing, how many times does the agent go backwards, how many times does it go forward, and then we want to optimize that trade-off.
[23:16] And so from here, I'm going to pass it over to Varinder and he's going to walk you through how we actually implemented this. So now we know the problem of the phishing and the overall design how we're going to be solving it. So I want you to introduce to LangFish. Why do we call it LangFish?
[23:33] What it is? So it has two words, Lang and Fish. So because it's based on large language models, so we say it Lang and then Fish is because of solving the phishing attack problem. And so essentially it's an autonomous agent AI system that helps us automating the
[23:50] phishing triage at large scale. So in nutshell, when a user report a phishing email, it it get processed by the LangFish. And that essentially has knowledge graph and triage agent as components. And it can produce the
[24:07] decision whether it the email look like suspicious or not. If it looks suspicious, it can be escalated to the human agent or human analyst. But if it does not, then uh it can directly recommend to close that case so to reduce the backlog. And
[24:24] in this world like 99% of the cases are benign, but only 1% are phishing. But still each and every case need investigation. So because of the limited human resources, the backlog arises. So that is the problem we are trying to
[24:40] solve here at large scale. So if we if this works as expected, then 99% of cases can directly go to the close. So we will essentially you know, fasten the processing of the phishing triage.
[24:59] So to do that and because user can report fishing across the enterprise, they're like 200,000 employees in the CVS and anyone can report fishing email and those uh those reports are coming continuously. So, we need a real-time solution streaming uh system for this.
[25:15] For that uh we we use the AWS ETL and then once we have the data stream through data loader of from the data bricks, we we put the data into the bronze layer. So, we this is the medallion architecture that we're using and then we process that data. Uh
[25:33] there's a concept of cases and signals. So, you can say that when user report the fishing, that is essentially a fishing triage signal, but there could be like similar-looking emails. So, we group them into the cases. So, one case can have like 50 uh signals together.
[25:48] They they're all like looking similar. So, we process them as case level so to fasten the processing. And uh to from making from bronze to silver layer like uh it needs the cases and signal correlation plus some dedupes and another feature engineering.
[26:03] Uh with that uh once we have the silver layer data, then triage agent can pick up from there and we have a concept of uh invest uh the state. So, I'm going to introduce that to next slide, but once uh triage agent process all the data from the
[26:20] silver uh silver table, then we produce the report the final investigation summary in the gold uh state and from there then sort of sort of tool can do the final update for closing the case or triaging it at final step.
[26:35] So, let's talk about uh the investigation state. So, we have the data in the like we have the raw data and then uh we ingest it uh into our system and then we're and you talk about the knowledge graph. So, we we we put it into the
[26:50] memory and then we have we we generate a score from based on the historical records. Okay, what is the likelihood of this product this current email to be phishing or non-phishing? So, we assign a score from there, and then we enrich it from other
[27:05] data sources like threat intelligence or threat stream data. So, our threat intelligence as well. So, and then finally we build a final data for the agent to be consumed, and that is called we are calling it investigations state. And this is at
[27:21] each phase it is getting updated automatically for every new case or every new email that it reads. And now we talk about the agentic architecture. So, we are utilizing the
[27:37] multi multi-agent architecture. So, when we have the intake as a case, so we have multiple sub-agents, and each sub-agent is expert in its own domain. So, so header agent is going to look the header only, and then sender agent, body
[27:54] agent, body HTML, and the URL, and so and so forth. So, each sub-agent is going to produce its own summary of the investigation what what they think about it, and we are utilizing the cheap model with temperature zero like llama model 17 70 billion parameters. With that, and
[28:13] then each of these investigation is passed to the synthesis agent, and that's a little bit heavy model that Cloud Sonnet or Opus or it could be any the bigger model that can understand all the findings from the each sub-agent, and then produce the final verdict. And
[28:31] the verdict report, so because as you see, right, these agentic systems are generative AI systems. These are non-deterministic. Every time you ask the same question, it's going to produce a different answer. So, to control the end because we need a consistent output, we We consistency in our verdict reports
[28:47] and all that across the emails. No matter what investigation is done, it should produce exactly the same uh format uh output. So, for that we use the templates. Uh we we enforce that uh using the Pydantic and uh the HTML format
[29:04] reports. So, it can stick to uh that format. And then we finally report that uh send the final report to the SOAR tool, and SOAR tool can update uh uh the status of the the fishing uh email. And so So, this was one triage agent
[29:21] architecture, but how do we do that at scale? So, to do that in scale, that essentially a lot of components are uh stitching together. First is the streaming real-time cases. We need a uh capability to stream those real-time, and we are using Databricks Delta Loader
[29:37] for that. And then we we should be able to integrate uh information from various other systems. And that we are using tools for that. We are giving And then what for multi-agent collaboration, we are using LangGraph. And then uh as you see these sub-agents
[29:53] can process each investigation in parallel. And for that, we are using LangGraph batch. And as as I had write, the final report should be consistent across the investigations. So, for the for that, we're going to be uh we are using uh verdict agent.
[30:10] So, with that, now observability. That's another important crucial aspects in the in the company. So, when when we say that uh agent have arrived to such a decision, we want to know why. Okay,
[30:26] what what it led to that decision. So, we're using MLflow in Databricks for uh detailed traces uh of the triage agent. As you see on the right-hand side, so in the LangGraph, you can say that it's it went to the intake agent, and then uh it made a LLM call, and then went to the
[30:43] behavior agent and so and so forth. So, all that agentic decisions are being logged in the MLflow traces and we also maintain the audit audit logs and tables in Databricks. And any decision that is created by
[31:00] triage agent could be overwritten by human analyst if they think that it is not doing what it's supposed to do or it should be overwritten. So, we have that facility as well and those are also tracked in the SOAR tool and we have a continuous improvement feedback loop so
[31:16] that as Andrew mentioned in the knowledge graph, we do have the resolved cases and from those resolved cases we arrive on the patterns that we feed to the triage agent. So, as I was talking about investigation states, as you see on the left is a raw data and then we have the information
[31:33] from the knowledge graph based on the previous history of the records. So, if you see it has like pattern recognition, it has 30 days sender trend, similarity, subject samples, so things like that. And uh So, first two are the input to the
[31:49] triage agent and second two are produced by the triage agent of the from the lang graph. So, analysis load it says so these are the writeable portion that agent can write like it can take the intake summary, header analysis, sender analysis. So, these analysis are
[32:04] produced by the each sub-agent. And then the final verdict is it can it can recommend okay, what is the recommended next step, investigate further or close this case or what is the triage decision, what is the confidence level look like. So, this uh
[32:20] these fields will be produced in a consistent format to be utilized by the SOAR tool. And now if we take example like how I mean what a phishing email look like. So, as you see on the left hand side, so there is an so subject is a your Microsoft 365
[32:37] password expire in 2 hours. So, there's a So, there's a urgency is being created to for a user to enter enter their credentials. So, into So, it's expiring in 2 hours. So, they they want to do the credential harvesting. So, if we pass this to Language, then it's going to
[32:53] analyze different part of the email and it produce the report like this. So, it can say that it says that it's a credential phishing campaign and it has seen similar cases as it was saying that these campaigns are remembered in knowledge graph. So, it it's it produced
[33:10] that in the final wordings that it's related to this campaign and it has seen 14 prior cases like that. So, there's there's enough context for a human to review and it's not just making one decision that oh, it's a phishing, but it's giving enough context for a human to analyze and to make further decision.
[33:26] In And this would be as recommended as investigate further because this is a phishing. This we cannot close this case. It's It's not a benign case. And so, what's the deliverable look like? So, it it will tell you okay, what it look like. It's a credential
[33:41] phishing. It's a risk is high. What is the recommended next step is the investigate further. It can give give you the spam likelihood, phishing likelihood, and malicious intent. And on the right-hand side, if you see these are the brief summary from each agent perspective, sub agents. How do What's
[33:57] the header look like? What the sender says? And so and so forth. And what's the final verdict? So, for closing. Yeah. So, essentially we're saying that there are five systems, but there's like five components with one compound system. So,
[34:13] we have scale. The scale is one problem. So, for that we're using streaming. And the memory is being done by knowledge graph. And then the process is the multi triage multi agent triage system. There's a multi sub agents that we are utilizing. And then
[34:29] we need the explainability. That's the traces. So, why the decision is made in certain way. That is being tracked tracked using MLflow traces and the evals as Andrew explained the evals earlier. So, those are to prevent hallucination.
[34:45] So, with that I can pass to Andrew for conclusion. Thanks. Yeah, so just tying all that back. So, what we kind of talked about was fishing is a bad thing. Fishing isn't exactly classification, it's investigation. Uh and that we have
[35:01] to deal with really these four components. The scale, the consistency, memory. We don't want this to be memoryless process. Uh and we have to have this explainable. We have to be able to justify what decisions the agent made, why it made those decisions. And so, the way that we kind of accounted for those as Varun mentioned as well,
[35:17] streaming is how we will handle the scale. Uh so, we we basically structure our state, make it highly optimized. We basically then have preprocessing flows that then input and ingest and then we can do whatever we want from that perspective or going forward. Uh consistency. So, we we build a graph.
[35:34] We take our process that we use in whatever your organization would use. We map that onto our runtime and then we execute that with an agent sort of a fashion. Uh we handle the memory. We use knowledge graphs and label propagation. So, the memory is searchable. Uh you can
[35:51] build agents to actually search the memory directly. You can build queries up front that then queries the memory and then you can use that memory to then make some inferences about what the future looks like. Uh that will be used alongside because then you might say, well, why not just use that? But, there's other components
[36:08] that the agent needs to consider and the human in general would not make it off of their their prior memory as well. Uh and then finally to do the explanation, we'll use MLflow. You could use whatever really standard tracing tool you'd like, but the point there
[36:23] being is that of course you need some way to trace and then you need some way to kind of ingest and learn from those traces. Uh and then you need to do evals to make sure that everything is aligned. Uh so if anything uh triage is definitely uh not not a classification problem,
[36:39] it's an investigation problem and you really should have memory uh and some process that you model within your agent security style at scale. And so with that we're we're just do a some quick acknowledgements. So uh thanks to our our Databricks team uh Nick uh Jerry Rohan and Matt, they're really
[36:56] awesome. Uh they help us with any of these kind of challenges and from a Databricks perspective uh getting the streaming pipelines up and running. And then all of our colleagues within our security engineering team as well as our SOC partners who really helped us uh kind of make sense of the data we're data scientists.
Learn more about the Databricks Data and AI platform.
The information provided herein is for general informational purposes only and may not reflect the most current product capabilities or configurations.