Building Production AI Agents: 700 Trillion Signals to Action
Summary
- MiQ Digital built a production AI planning agent using the ReAct pattern with LangGraph and the Databricks SDK to process 700 trillion programmatic advertising signals and automate media planning workflows with human-in-the-loop oversight.
- The agent follows four design principles: transparent reasoning with human-in-the-loop control, problem decomposition into intent levels, AI-ready data optimization, and continuous output critique and evaluation to maintain quality.
- Integrating Databricks Genie into the architecture enables the agent to query enterprise data reliably, while observability at every node and CI/CD validation transform non-deterministic LLM behavior into a production-grade system.
Building Production AI Agents: 700 Trillion Signals to Action

Building production-grade AI agents that actually work requires orchestrating complex data systems with human oversight. Many organizations struggle to bridge the inference gap between fragmented data sources and actionable intelligence, especially when data spans hundreds of vendors with no common taxonomies. Databricks' unified platform, with Genie workspaces and Unity Catalog, enables teams to make that data accessible to agentic reasoning systems.
Hear from MiQ Digital on their production approach to AI agents for media planning using the ReAct pattern with LangGraph and Databricks SDK. Learn four-pillar agent design: transparent thinking and human-in-the-loop control, problem decomposition into intent levels, AI-ready data optimization, and continuous critique and evaluation. Discover how observability at every node, CI/CD validation, and context persistence transform non-deterministic LLMs into enterprise-grade systems handling 700 trillion signals in real time.
🤝
Chapters
00:00Introduction and Challenges07:29Planning Agent Solution12:20Live Demo: Planning Agent in Action22:39Technical Architecture: ReAct and Tool Integration25:49Design Principles: Transparency, Reflection, Data Readiness28:30Problem Decomposition and Intent Mapping31:56Genie Integration and Data Optimization35:08Output Critique, Evaluation, and Validation37:01Observability, CI/CD, and Key Learnings
FAQs
Why did MiQ Digital build a custom AI planning agent rather than using existing tools?
MiQ Digital needed an agent specifically designed for programmatic media planning that could reason over their proprietary 700 trillion signal dataset while integrating enterprise data from Databricks. Existing planning tools lacked the ability to connect to their data ecosystem at the required depth, and off-the-shelf agents could not meet MiQ's requirements for transparency, human-in-the-loop control, and enterprise governance.
What is the ReAct pattern and how does MiQ use it in their planning agent?
The ReAct pattern is an AI agent design that alternates between reasoning steps and action steps, allowing the agent to think through a problem before calling tools and to reflect on results before proceeding. MiQ implemented ReAct using LangGraph to give their planning agent a transparent chain of thought that human operators can inspect and intervene in when needed.
How did MiQ integrate Databricks Genie into their AI agent?
MiQ integrated Genie as a data access tool within the agent, optimizing their data assets for AI readiness so Genie can respond accurately to structured and analytical queries from the agent. This involved configuring Genie workspaces with business context and ensuring enterprise data in Unity Catalog was organized in a way the agent could effectively query.
What does observability mean for a production AI agent and why does MiQ prioritize it?
Observability means capturing detailed logs, traces, and metrics at every node in the agent's workflow so teams can diagnose why the agent made a particular decision or took a specific action. MiQ prioritizes observability because non-deterministic LLMs can behave unexpectedly, and full visibility into reasoning steps and tool calls is essential for debugging, CI/CD validation, and maintaining production reliability at scale.
Full transcript
[00:08] Okay. Welcome everyone. Thank you for joining in. I know there's a lot of sessions going around, so thank you for making time for this. We're going to try and make like the next 35 minutes work for all of you, so everyone take something back from the room. Cool.
[00:23] So, here we have the session the title again fairly provocative from 700 trillion data signals to action. We've kept it that way cuz one, it's a data engineering flex, absolutely. Sounds braggy, it is.
[00:40] But yeah, now the real magic happens when we go outside the ingestion process and after we've gotten the signals into the room. What do we do with with that then? And that is where like the magic happens. And that is what we're going to be talking about through today.
[01:09] Yeah. Like you have me. I'm Abhishek. I'm the global product lead for MIQ. And we also have Anish who's leading the product and data science side of things at MIQ. Again, we are we've been working with on this like for the last 6 months, so we
[01:24] have been through all frustrations, like escalations, and then every kind of scenario that we can possibly think so we'll try and share some of the learnings that we've had over this period. Again, in terms of how we break broken up this session, we're going to set some context in terms of why we had to build
[01:41] another agent in a world that's already filled with agents, a lot of them having planning agents of their own. Why did we have to build a new one? The next one, we're going to be talking about the solution scope and how did we arrive at the scope of what we're calling currently calling planning agent.
[01:56] And this is where the fun starts. We're going to do like a little live demo. Uh, we will start like showcasing something. This is the first time that we're bringing this out like none of our clients have seen it, none of the agencies have seen it. So, you guys will be the first ones to essentially see this. We're going to be demoing this at
[02:12] Cannes this later in the week as well. Uh, and again, uh, we're going to then deep dive into like a lot of the technical considerations that we've made, uh, some of the best practices and how we brought Genie to a level where it can actually start working with the enterprise data
[02:28] that we have in our ecosystem. Uh, good. So, without further ado, uh, just a little bit about like the organization MiQ. So, MiQ is the foremost uh, global managed service media company that's out there. We've
[02:44] been doing programmatic for the last 14 years. So, MiQ is almost as old as programmatic in itself. Uh, working across like uh, 30 different markets with like thousands of brands like running tens of thousands of campaigns every month. Uh,
[02:59] media planning is not something that we would want to do. It is a mandate for us in most scenarios. So, it's becoming like a scale challenge, which is where like the current solution really helps.
[03:15] Okay. So, as as like programmatic media service partners, we run media across like the omni channel landscape. So, starting from uh, the watching side of things where you have like the wider video coverage there, uh, to the final touch points around retail media, MiQ does it all.
[03:30] Multi-format, multi-channel, and multi-DSP is some somewhat of a tagline tagline for us. And because we are sitting at the intersection of this omni channel landscape, we are often hit with all kinds of like data elements, uh, from like different parts of like a consumer
[03:47] journey that helps us build a more consolidated view of what a consumer is thinking, how they're considering stuff, and eventually how do they end up buying. So, this not only helps in terms of like creating that broader view, but also in terms of actionability of what
[04:02] the next action is going to be from a consumer perspective. Uh and the scaling is where like Databricks essentially has really helped us like evolve over the last uh I would say 8 years now. So, we started as like the first customer for
[04:17] Databricks out of EMEA. Uh Renald, who we heard every yesterday, he was actually the first solution architect that worked with MIQ. Now, he's moved on to like being the CTO and he's moved on to better things, but he still is like uh very much in loop uh with MIQ.
[04:33] Uh again, uh the strategic positioning is also in terms of uh like how we built on top of Databricks. So, most of our like clients can essentially work with us and integrate with us in as more seamless manner as well. Uh and again, like influencing and
[04:48] engagement is something that we've uh really been hitting the nail on the head as well with Databricks cuz we've been across their product and tech councils, and we are helping them in terms of driving the next phase of what comes in the Databricks roadmaps.
[05:06] Cool. Now that that's initial context done, I'm going to quickly skip into the key challenges of why we had to build a media planner. Uh to start off with, media planning has always been inefficient, right? Uh but the last decade or so is where like we started to see like things go a lot
[05:23] quicker and a lot in a messier direction. The idea of programmatic was to bring simplicity to like digital advertising. While it did it for the first few say years, after the platform started to like essentially proliferate, and you had like a lot of data coming
[05:39] from like all touchpoints, uh the idea of simplicity has gone out of the window with programmatic, and planning media campaigns in an omni-channel environment is more complex than ever at this point. Uh and this is where Planning Agent essentially comes in in terms of simplifying that process. But to be able
[05:56] to understand what problems like it is solving for within our current internal ecosystem and like the agency ecosystems, we try to do like a brief overview of like what are the challenges our teams currently face. So there's significant knowledge gaps uh like
[06:12] based on when people have joined the organization and their comfort with programmatic, the kind of plans that we are able to create is significantly different across like different people. There's information fragmentation because like there are so many sandboxes that we're operating across, so many
[06:27] different data vendors with no common taxonomies. There is always like difficulties in terms of how do we make this data work together? And there's always a lack of actionability. Great insights, that's always been like a USP for MiQ, but the next question is what next? How do I use
[06:43] this insight to some build something more meaningful? And that actionability element has always been a challenge for us. Uh multiply this with the factors that are there at the agency level on the media planning side of things. It's time-consuming. A basic pitch can take anywhere between 6 to 8 weeks for them
[06:59] to develop. It's expensive cuz like data licenses aren't cheap. And again, uh most media planners have their uh way of working still like embedded in spreadsheets to be very honest. So again, it's fairly optimized. And all of
[07:14] this essentially feeds into the un- unstandardized way of like responding to a brief. Your brief is only as good as your planner in those scenarios, and that essentially is like a major opportunity cost that like media plan- media agencies are currently bearing.
[07:29] Uh and this is where uh Planning Agent essentially comes in. We designed Planning Agent with the very simple objective. We wanted to help agencies win. Uh if they win more, we win alongside them. So, if they get a big bigger chunk
[07:44] of the pie, we win alongside that. Uh, the idea was also to connect the different silos that currently exist within our organization as well as within the agency ecosystems. So, we can build like a layer of connected intelligence that can help people plan better. And furthermore, like it just improves
[08:01] standardization. Like you should have like a standard way in which you are essentially mimicking your best sellers or best planners response in most things that you do.
[08:16] Uh, cool. So, how we define the problem for planning agent is a three-step process. Uh, media plan would essentially encompass like you pitching against opportunity, you responding to a particular brief, and once you've won that, getting people to spend more with you. So, identifying those upsell opportunities where you could be doing
[08:32] like more with the client. So, to start off with, we started focusing on the first element. That's the pitch element. That's where the highest amount of unoptimization currently lies. You can always ask like a chat GPT to create a media plan, but that's always going to be generic. That's never going
[08:49] to have like variability. So, if you asked it to create a plan now, and you asked it to create a plan six months down the road, there's going to be no coherence across the two as well. So, this is where like having a data backing against like the media planning is something that really helps.
[09:05] Uh, so again, against each of the pillars that we identified, we listed out some key jobs to be done. Uh, across strategy and comms planning, we are trying to get a lay of the land, understanding what problems to solve. Once you've identified the right problems to solve,
[09:21] the tactical media planning framework essentially breaks it down into how you're going to solve for those. And once you've done that, you can essentially move on to the upsell where you start to bring in like the campaign performance of a live campaign data it like the plan plan that you've suggested to identify what's working and what
[09:38] needs to change. Now, this is where we also identify like the different stakeholders that are part of the general process. So, on the exact side, like these are the people that are going to be the first people in the door in terms of influencing the brand, in terms of identifying the right
[09:55] You have the hands-on planner, which are going to be the ones that are going to be creating the briefs, responding to briefs. And then the MIQ commercial teams, again, we're building this so that we can make money off it as well. So, while we're trying to solve for other people, we're also trying to solve it for ourselves. And this is where how do we
[10:11] make our sellers more intelligent in terms of responding to these pitches? Uh but again, like the factor I said earlier, like everyone is building an agent right now. There are like as many agents as there are people in the agency ecosystem. So,
[10:27] again, there is nothing new. What are what are we bringing to the table at that point in time? Why is MIQ's planning agent the best planning agent that the agency can get get their hands on? Our unique value has always been in terms of the partnerships that we've unlocked. Over the last 15 years, we've
[10:45] essentially partnered with 300 data partners, more than 300 data partner. I'm I'm sure this is outdated now. We're signing like five things like every month, so this is definitely outdated. But yeah, over 300 data partners. Uh you see the likes of like Samba, you see the likes of Vizio. These are like your TV
[11:02] data partners, which are giving you like glass level ACR behavior. This. So, you exactly know what people are watching, what ads they consuming at what point in time, across what interfaces. So, that helps you build a comprehensive view in terms of uh how the watching side of things are
[11:17] happening. You have the browsing signals that give you like decent amount of uh input into the consideration phase of a user of how do they go about researching particular stuff. What goes in terms of like driving their actions to a particular purchase or an intent behavior? And we also have like transactional
[11:33] level details from the likes of Numeratior, Ibotta, Circana, which again helps us build that like end touch point of what they bought, where they bought it, and again at what price did they buy it most importantly. So all of these combined help us create like the persona of a user.
[11:49] And and essentially help help us build like a graph against what they uh how did they start their journey, how did they come about the consideration phase, and what did they did they end up buying. And this is like our unique value proposition. So the agent that we've developed
[12:04] is like uh consuming a lot of these signals and building intelligence on top of it rather than reciting from like historical responses or institutional learnings as as like most agents do.
[12:20] Okay. So now for the fun part, uh I'm quickly going to switch to a live demo. Uh And again, guys, uh this is the first time that we're being going to be showing this externally. Uh so would love to see some reaction on the audience as well. Okay. Perfect. So this is uh for people that
[12:37] have been here at the summit before, last year we essentially demoed Sigma, which is our unified platform that helps our teams plan, execute, and measure campaigns from the same interface. Now again, these are not like uh
[12:53] like a single point solution partners that you would have essentially worked across. This is like an omni channel suite that is there where you can do like TV, YouTube, uh digital, uh retail media, all campaign setups from a singular platform, activate them through
[13:09] a multi-modal and multi-DSP framework, and essentially do like comprehensive research against like the outcomes that you're driving from the same platform itself. Now within this, the latest agentic integration that we've made is in terms of uh MIQ planning agent.
[13:25] Uh so again, when I go in here, it's a fairly simple uh brief uh I I say fairly simple experience that is there. There is a single prompt window that exist. So, I can just go ahead and type my prompt in here. All right, so I've been working with it
[13:40] travel aggregation partner and I'm creating a brief for them. The first thing that I'm trying to do is get a lay of the land. Who is the travel audience profile that I'm going to be targeting? Again, if that there might be indications where like people already know the target audience, so I can fill that in as like a brief as well. So, the
[13:58] agent essentially gets a little bit more direction in terms of where it wants to plan. Uh, once I like execute this, uh now this is interesting. I'll show you how this works. Okay. Okay. So, now that like I've just uh
[14:15] started executing this particular thing. Mhm. Let me try and refresh this.
[14:31] So, once the agent starts to pick up pick up that problem statement, it starts to break the problem statement down. It first identifies what the intent against the problem is. What are we trying to do out of it? It then starts to look at the tool descriptions. Now, again like while our data is definitely there, we've also
[14:48] integrated it with other like trusted sources that our agency partners use. We've also integrated it some institution knowledge also. So, it's looking at all the tools that it has access to and then trying to identify what what best resources do I have in terms of answering for the particular
[15:04] problem statement that is in front of me. Once it's done that like again, it's making these individual calls, it is doing the intent mapping and again like going back back and forth in terms of reflecting and refining of what it can essentially answer and with what level of confidence. Once it's done that, it
[15:21] we we like moderated like a human in the loop interaction here where the agent will essentially give you like a plan. Like this is like how I am approaching the media planning side of things. Now these are the particular steps that I'm going to be taking in terms of breaking
[15:36] this problem down. These are the insights that I'm going to discover and based on this these are the outcomes that I'm going to be essentially establishing. Now once we have essentially created a basic framework, it is to the user can essentially then start to feed in their understanding of things as well. If
[15:52] there are steps that I feel are missing, if there are pieces of information that I want to be added as part of the initial response, always use those as an input here to further augment the plan that I've currently created. So I can just do something like this. Uh
[16:17] retailer split against travel purchases. Uh so again, when I do that, what it starts to do is it will start to again re-initiate that plan and add another
[16:34] step with this as like a key constraint in itself it as well. So once I'm happy with the plan, once I've curated this enough, I think like I have a good baseline line to start with, I can essentially just kick off from that point on. Uh once again,
[16:50] uh like because it is not something where it is recited from already existing learnings, it is actually building things out on the fly. Uh uh so again, it will take some time in terms of coming with things out. Now here again, it's it's asked it's
[17:07] given me a clarification. There is some ambiguity that is there in the system. It didn't understand what I'm asking, so it is essentially giving me a clarification opportunity. Now here I have the option of giving it like a specific direction of how to clarify on it, or I can essentially give it like you go with the assumption that
[17:23] you want, and you give me the best system response that is there in the system. So, again, like I can just skip and proceed in this particular section. But, once it has essentially taken in all the context, and it has all the tool descriptions, it has broken down the
[17:39] plan, it will start the execution, after which you will start to see a plan that's coming up something like this. It takes about 10 to 15 minutes in terms of coming up with the full solution. So, again, which is where I'm just skipping to things that are already generated. So, here we modeled the output in a
[17:56] manner which is fairly similar to a how a media actually want to like Okay. So, here you start with an executive summary. So, this is like the TLDR to your problem. It's the most direct response. It is a summarization
[18:12] of all the analysis that it has done across like all the steps that it generated. It brings up like a tactical persona. These are the people that you can essentially target. Now, I have again like my audience profile is not always going to be like homogenate. So, we have like different
[18:28] variants of that persona as well of like their based on their browsing habits, based on their demographic attributes. So, you can essentially choose one or many of these profiles for targeting. Again like, as a media planner, one thing that we've seen is they never want to go through the full brief. So, what
[18:44] we do is we essentially summarize the key taking talking points out of the entire analysis, and give that a stick take key takeaways. And further like connect this to activation elements as well. Insights are great, but how do you how do I connect this to activation? That's
[19:00] always been a gap at least in our ecosystem, and I'm sure that's the gap in like most inside ecosystems. So, we've tried to moderate that by like giving direct recommendation of how this intelligence feeds into a campaign setup. Uh for data nerds like me, I always want
[19:15] to go a level deeper. I do want to understand like how did we come up with the things that we came up with. So, you have like deeper breakdowns of insights in here as well. Uh most of these charts are fairly interactive as well. So, again, it's looked at the
[19:30] browsing behavior of people. It's looked at say the demographic of audiences. It's looked at uh at the temporal habits, uh when they're active. Uh again, I'll switch to another one.
[20:00] Okay. So, yeah. Uh again, like I can understand like deeper insights by reviewing all these analyses as well. Sorry, I think I just lost my thread. Uh no worries. Yeah. So, again, you get the idea, right? Like we we are able to decompose that problem statement down into like
[20:16] smaller insights. And again, like get like a foundation of what do I need to do next? While the agent is going to do some like uh retrospective action where it's going to give you some indication of what are the other problems you need to be looking at. You can always do that yourself by looking at the insight. If
[20:32] you find something interesting or something off, something away from the hypothesis that you're essentially trying to establish, you can always go back and like the word uh conversational again start as well. There is like an in-context memory which will help you in terms of uh
[20:47] like persisting that context of a conversation. So, if you wanted to say like in the temporal chart, could you also show me like what browsing habits or what kind of like keywords are happening across like say between uh late night and early morning? You can even like
[21:03] deep down deep deep dive down to that that point as well. So, again, this gives you a starting baseline. You can always build on top of that and that's how like media planning frameworks work. Uh cool. So, this is like how we've set up the planning agent. Uh again, things are still in motion here. So, there will be
[21:19] things that we're going to be like uh updating on top of it. Uh the next step is in terms of bringing tactical media pre-framework. So, you while you start with this as a basic module and you get the strategies out, you're able to connect it to segment taxonomies, audience sizing, and build like
[21:34] full-fledged media plans and brief responses on top of it. Cool. So, with that, I'm just going to bring in Anish who's going to be giving a little bit more technical overview into what has gone behind like building this out.
[21:50] Tha- thanks. Thanks, Abhishek. Uh How was the demo? Good? Something that was different that you saw here was the data depth. If you go and ask it in normal areas, you will not see the the data depth that
[22:06] it comes up with. It is also not fast like what general LLMs and general charge there. It takes time. And it takes time because it is going inside. It is going inside all the data set that that 900 trillion signals. It
[22:22] is actually doing everything at that time when you fire that enter. Right? It goes through all of them and find out the trends that is relevant to that. So, it's not precomputed. It's on the go and that's why it takes time. So, it's important for us to sort of go back
[22:39] and see what is underlying that is powering this. Right? Uh I'll very quickly jump on to the high-level design. At the very center of it is a react structure. Right? It observes the problem, it reason the
[22:55] problem, and it sees what tool it can use to answer that question. That react loop in the center uh which you see between the planning agent service and the planning agent MCV server is doing all the hard work. Right? All the problems that it sees
[23:10] breaking it down, going to the tools that it has, and sort of finding out, "Okay, these are the few things that I can connect to answer that question." In last 6 months, which is when we started building this, we have integrated four different types of tool. Couple of them
[23:26] are powered by Data Bricks Genie. Right? And one of them is MIQ's historical data performance, right? Our campaigns, everything that we have done, the KPIs, which works, what doesn't work, everything is there in one queue. That is powering more the campaign side of
[23:42] thing. The second queue is all the watching browsing behavior of consumers, right? What are they watching? What are they buying? What are they browsing? And connecting these two links to come up with that plan. We also thought that why restrict to
[23:57] just internal system? Why don't interact with the external MCPs? The MCB ecosystem is evolving, right? Can we connect to them? Can we face the information and add that to our system? So, the third one, what you see on, which is GWI, which is a marketing intelligence panel
[24:13] driven intelligence that is available. So, we connected to their MCP. Right? If there are certain questions that are not there in our data, can we face those information from outside? We're still not going to open internet. We are not going to anything that is not verified. We are going to every signal that is
[24:29] verified, right? So, that's the GWI. And the fourth one is again moving towards activation, going to a audience persona. So, you saw in that Abhishek's demo, the top part, the persona, it looks into all the insights that it has
[24:46] come up with, and it creates that persona. So, these are the four tool it is now have access to. We want to make more and more added to that so that it can answer the vastness of all media planning, and it connects all of them to come up with this. Now, that's not the only thing.
[25:03] Couple of other features that you saw there was on the left-hand side the chat is persist the memory. If you have done something today, it will store that context. So, tomorrow when you are starting to start build something, you can build on top of that.
[25:18] You don't have to start from, right? So, all the queries, everything that has been done before, it stores in the memory. The bottom one, the curation server, is you have done something. Now, in media planning, it all happens that okay, these are the domains, these are
[25:33] the source, these are the things that we want to go after, the data is saying, but I don't want to go after. I want to exclude something. So, when you are sending the media, you also as a human, as a media planner, wants a control over it, what goes out, what restrictions. So, the curation
[25:49] server is where your local memories, local outputs are stored, which you can moderate. Right? This is the broad thing. Now, outside structure, I'll also talk about six key consideration that went behind it, all right? And this are the things we
[26:06] actually take quite a pride in actually, because we have bring them in our system. So, I'll go over all of them in details, but in this one, I'll just talk about little bit of them, like transparent thinking and intent. So, all the agent that is coming up
[26:23] with, it opens up to the human. Right? As a human, you need to audit it. You want to add, you want to modify certain thing when the thinking is happening before it goes for execution. So, in Abhishek's example, Abhishek approved and execute, right? He could
[26:38] have added more context. Because of the demo and the time constraint, we have to approve, but that's where human intelligence come. Agent has come up with one idea. Do you want to accept that idea? You want to modify that? And that's that part. Refinement and reflection, the agent in multiple stages
[26:54] actually defined. Okay, this is what I come up with. Does it align with what I was asked for? And if it is not, it will go for reflection refinement. The data that we had we had this data over the past. We have built on this data thing over the past 15 years.
[27:09] Whether all ready to consume by AI? The answer is no. We have to change the data. We have to add lot more context to the data so that agents can talk to it and understand. So, what humans while we are writing SQL with the entire context in our mind,
[27:26] what we can do, agents have to do it programmatically. So, we need to pass on those context. So, that's AI ready data. Output critic is very important, right? Is anything that comes up with, we can't just send it to insights and those things. We have to critique them. So,
[27:42] we'll talk about little bit on the output critic. All the everything that the agent is coming up from the data. How do you change the tone? It has to It need to be geared towards the end users. The planner, can it talk like a planner?
[27:58] Right? All the visualization, all the insights, can it make sense for the planner, the consumer? And the last but not the least, which you saw when Abhishek said that from day one we put observability at the back end. That anything that happens in each node
[28:14] of that agent, everything traceable. Right? What is the input? What is the output? What does the tool that has been called? What output came? Everything is traceable today. And that helps us to debug as well as that's the foundation to build self-learn model for tomorrow.
[28:30] Uh with that I'll jump quickly jump over to the first one. All right. What happens when the prompt goes from the user? So, first thing it try to break down the problem. All
[28:46] right? Because different data, different information is there in different system. If we just give that broad statement, it can't address. So, first it try to break it down. So, while breaking it down, it first looked into the context clues. That what are the things that are there in the
[29:01] prompt? What are the previous conversation this user has done? What is the user profile of it? Based on that, it try to break down the big statement uh into the decomposition, which is subatomic task. Okay. To answer this question, I need to break it down in
[29:17] this set 15 things. Now, it has broken it down. It again reflect. If I add these 15 things that I have broken, will I be able to answer the top question? So, it's not just one way down. It also try to add up. So, once it
[29:33] has done that reflection, it has find it try to find out categorize it in four things, which we call the intent. Right? The four things are direct, which is just a direct question, one-line answer, one data point answer. Right? It doesn't have to go through all the data sources.
[29:50] Focused, which is it need to do little bit more. Maybe it has to look into two or three data sets. And look into them, see multiple hypotheses in this, so which is moderate. Then is uh so sorry, the third one is moderate. The second one focus. Third one is
[30:06] moderate, where it need to go to more. And comprehensive, where it looks into all the data sources to create the narrative. So, this thing it has figuring it out before it run the any query to the databases. Right? So, once it has found out, uh then it will
[30:23] go to reflection and refinement. The first thing it has figured out, okay, these are the things that I have to do. This is my intent of the query. And this is my plan. The first reflection and refine comes as human in the loop. It pops up.
[30:38] This is what I my plan is. This is what I'm going to look. These are the tools that I'm going to call to answer this hypothesis. Are you in alignment with that or not? Do you want to look into some other places? Do you want to modify something? So, that's human in the loop. Based on the feedback that human has
[30:54] provided, it again refine the hypothesis. It refine the sources. It refine how the data need to come. So, once it comes, it lock in, you say approve and proceed, it goes for execution. It starts sequentially all of them and reflecting on it. So, once all
[31:09] that data comes, now the second layer of reflect and refine come, which is more about the context. The output is there. All the data is there. Does it answer the question that has been asked? Is the data statistically significant?
[31:24] Is the distribution that I am representing are there bias in the data? So, now the system is reflecting on the output. The first reflection was in conjunction with human. The second reflection is looking after the output that you have generated. So, once it agrees, it can
[31:41] summarize, okay, this is the output. Now, as a human, I can say, no, I'm still not liking it. Go deep dive. So, it can again go back, reiterate that process. Once we are happy with it, it goes to the output. Now, we spoke about how it moves. Now,
[31:56] the underlying thing is data. We had been having data in data bricks, unity catalog, everything. The first time when we connected that in Genie, we saw a lot of gaps. The data is not able to the agent is not able to understand and pick the data
[32:12] from the right places. It is sometime it was not even going to the right sources to face the information. So, what to do? So, here data bricks Genie's additional things that are there that helps us significantly. So, I'll talk about two things. One is tool and data description
[32:29] of the Genie space, which is the outward thing, which helps the planning agent, the agent to understand what our Genies MCP can answer. So, it's more like a outward thing. It tells, "Okay, I can do this this thing. These are the data I have.
[32:45] These are the things I can't do." So, those examples, those instructions we have to fit to that model. We have to add what each data sets mean, all right? What is user ID? What is DOD? What is DO So, everything we have to explain so that it understand
[33:01] what it is, right? And uh yeah, and all the columns. So, it Though we actually migrated lot of information from Unity Catalog, but we also added lot more extra information with what a particular column means. So, now from outside
[33:17] the agent is able to understand, "Okay, these are the things that are present in the data in the MCP, and I can do this this this thing." So, that's the outward part, which is tool and data description. Second part is all about the inside part of it. Now, the prompt has hit
[33:32] Databricks Genie. Now, how do we convert it to a sequel? Right? How do we sort of break that problem statement to that? So, that's all the part that where we say the instructions, the long list of instruction that we had. If these type
[33:48] of things come, this is how you break down your query. This is how you write. These are the few things you need to consider. We have to consider index. We have to consider percentage. We have to use raw numbers. So, those are the all induction that goes. We also give what are the good query look like. Right? If these type of things come,
[34:04] this is your sequel query. We write that, right? While writing this instruction, we also face a challenge from Databricks side, which is it is limited to 20,000 words. The instruction goes beyond 20,000, it starts hallucinating. And then comes the other things, which
[34:20] is sequel templates. The typical joins that we are doing we have we have moved them to the joins. The typical definitions that we are using, we are sending them to sequel template. And the queries, the good queries that those we move to this. So, we'll try to limit
[34:36] what is we providing to the instructions. So, that helps us overall to optimize the Genie MCP. So, the same data just with a lot of more context starts giving a brilliant results. So, we didn't change the data. We just had to feed in the
[34:51] context to it. Now, for data bricks and this summit which we felt is it's all about lot of context. Hopefully, the new advancement that are happening on ontology and those things, they'll power this thing up. Moving to the next one which is critique, right? Now, the output is out
[35:08] of data all the data sources. We look look into all those things like represent this quality. If the query is something do we need to send it a raw number? Do we need to send it as a percentage? Do I need to do it in index? Is there any bias in the data? Is the numbers insignificant? Is the number
[35:25] statistically significant? If the hypothesis are proven or not. So, all those checks happen when those things come out. We also saw certain time the sequel that has been generated, we ask another LLM to evaluate the sequel. Right? Whether the sequel has been generated properly. So, that's the
[35:40] sequel out of context. Sometimes the output that comes the data set is completely null. It's true that there is no data. But in the output we can't show that it's null. So, how do you filter things? If there are some things that are coming up out of context, how do you filter
[35:56] them before coming to the surface layer? So, that's all about out out of context. Statistical significance check I already spoke. Reasoning layer. We also see in the sequel if the proper logics are there. If if there are multiple main calculation that
[36:12] are happening in that query, are they correct or not? So, this all checks happen. It takes two action. Either is reduce the information that is flowing to the output layer or you send it back with a retry that again run this query with this this added information. So, that's
[36:28] the critical layer. Uh going to this uh next is once the data is out, now how do you make it to the customer relevant summarization, visualization, everything that comes the tone that the agent is responding to is exactly consumable by the end user. So,
[36:45] charts, the right charts, the data need to be represented in the right chart. The each chart level summary, we saw the executive summary, but also pillar-wise summaries. So, we do give that for the depth. Uh and every uh yeah, and end of it, it also gives media recommendation, that if
[37:01] you want to spend this way, how do you want to spend? So, with that, the last but not the least is the observability. This observability had been a foundation block for us. That every block the agent is going
[37:17] through, we wanted to know what is going on. We took help of obviously off the self land Smith to sort of set up everything. It looks into consistently every execution step. It It is looking after dev environment, integration
[37:33] environment, broad environment, so that anything that we are pushing in integration, how different it is from the dev or to the broad, right? So, we are able to identify the delta change in the output because of the delta change in the input, right? So, that's the part.
[37:50] Benchmarking, we have given enough golden prompts. Right? This is how good our output should look like. This is good, this is bad, this is the desired output. Every push that is happening as part of CICD, it looks after
[38:06] the things like executive summary quality, inside depth, uh recommendation quality, the pillar coverage, and the structural coherence. It looks after all of them. So, any push that we are doing, it will rate against this five pillar,
[38:22] and if it passes, then only it will get merged to the master. Otherwise, it will reject. That you retry, and so benchmarking and those things we have made it part of our CICD, so that nothing goes we are so that we are going to understand if the agent is behaving in certain way, why it
[38:38] is doing so. With that, I'll just summarize the overall technical thing beyond the normal agent creation uh that we learned in this journey is design for transparency, that everything that is happening can be
[38:54] openly share that, have human control over it. Second is, before we start building the agentic data solution, is our data ready for AI? That's the second. Third thing, eval and critique as part of the loop is not
[39:10] something that we want to retrospect. On the go when we are sending something, is evaluation, critiquing happening or not? And last but not the least, again, observability. From day one, we need to know each node, what are they doing, why are they doing certain things, which
[39:26] when you add up is what you see at the output of planning agent. So, these are not bare necessities what we felt, these are sort of fundamentals of how we build agentic solution. With that, uh we close our presentation. We'll be outside uh because of time
[39:42] crunch, we can't take any questions, but we'll be outside happy to chat on anything. Thank you.
Learn more about the Databricks Data and AI platform.
The information provided herein is for general informational purposes only and may not reflect the most current product capabilities or configurations.