Building Production AI: CAVA's Multi-Agent Supervisor on Databricks
Summary
- CAVA built Ask Astro, a production multi-agent supervisor on the Databricks Data and AI platform, using Agent Bricks to route natural language questions across domain-specific Genie spaces and external tools, replacing analyst-mediated data requests for hundreds of employees.
- The Cava Core data estate centralizes 3,000 tables of restaurant and business data, with Lakebase providing agent memory and MLflow tracing delivering full observability across every AI-assisted query.
- Production guardrails including evaluation AI for quality assurance, on-behalf-of-user authorization, and Unity Catalog governance ensure that Astro delivers trusted answers at scale as query volume grows exponentially.
Building Production AI: CAVA's Multi-Agent Supervisor on Databricks

At CAVA, 3,000 tables of restaurant and business data serve hundreds of employees daily. Traditionally, users requested data from analysts, often downloading raw files to Excel. Matt McDonald and Sudhakar Selvarajan share how they transformed this bottleneck by building Ask Astro, a production multi-agent supervisor that routes natural language questions across specialized agents and external tools, delivering trusted answers in seconds.
Explore CAVA's architecture built entirely on Databricks: multi-agent orchestration with Agent Bricks, domain-specific Genie spaces for routing, MCP tool integration for external data sources, and Lakebase for agent memory. Learn production guardrails including evaluation AI for quality assurance, MLflow tracing for observability, on-behalf-of-user authorization, and Unity Catalog governance. Discover key learnings on building scalable AI products that users actually trust and adopt.
🤝
Chapters
00:00Introduction: CAVA's Scale and Data Challenge02:15The Problem: Data Silos and Bottlenecks05:17Introducing Astro: Multi-Agent Supervisor10:08Supervisor vs Chatbot: Architecture Pattern14:42Technical Stack: Agent Bricks, Genie Spaces and Lakebase16:23Quality Assurance: Evaluation AI and Monitoring20:44Live Demo: Routing Across Specialized Agents28:56MLflow Tracing and Observability32:45Production Learnings: OBO, OAuth and Governance35:07Philosophy: Build Your Own Bowl Approach
FAQs
What is Ask Astro and how does it work?
Ask Astro is CAVA's production multi-agent supervisor built entirely on the Databricks Data and AI platform, routing natural language questions to specialized domain agents and external tools to deliver trusted answers in seconds. It leverages Agent Bricks for multi-agent orchestration, domain-specific Genie spaces for routing, and MCP tool integration for external data sources.
How does CAVA use Lakebase in its AI architecture?
CAVA uses Lakebase as the memory layer for its Astro agent, enabling the system to retain context across interactions within the multi-agent architecture. This is part of the broader Cava Core data estate, which serves as the centralized data store for the entire company.
What production guardrails does CAVA use for its AI agents?
CAVA implemented evaluation AI for quality assurance, MLflow tracing for observability, and on-behalf-of-user (OBO) authorization with OAuth to ensure agents operate with appropriate permissions. Unity Catalog governance enforces data access policies consistently across all agent interactions.
How did CAVA measure the business impact of Ask Astro?
This video notes that as of approximately one month before the presentation, CAVA's Cava Core platform had processed 200,000 AI-assisted queries year-to-date, with that number growing exponentially. Ask Astro was built to optimize and scale this query volume while ensuring data democratization for the hundreds of employees who previously had to request data from analysts.
Full transcript
[00:07] Thank you all for showing up to our presentation. Uh, we are Cava. There's a lot of familiar faces, a lot of new faces in the room. But today we want to take you through our journey how we built a supervisor agent in our data platform 100% in Databricks. So,
[00:22] context, I am Matt. I lead the data practice at Cava. Sudhakar is the one who really brings everything to life. He runs all the engineering. So, I want to take you all through our journey and really have him showcase some of the work we've done. But first we will give you a quick intro to who is Cava.
[00:39] So, has anybody in here been to a Cava, know Cava? Yeah, it's like it's half and half. You you either don't have one near you or you do and you love it. So, if you haven't heard of us, you will. Um, we are a category-defining Mediterranean
[00:54] brand with the goal to bring heart, health, and humanity to the masses. So, if you see here, a massive growth for us over the last couple years. 100 somewhat restaurants a few years ago, we're now over 400 with a goal of being a thousand within the next
[01:10] six years. So, a lot of what you'll see today is how we have built a platform set to scale, but really how we are working to democratize and get data in the hands of the people who need it most. Um, I'm going to I want to level set before we dive in. There's a couple words that
[01:26] I want to familiarize you with. One, you'll hear two big things. Cava Core. Cava Core is our data estate. It is our data platform and we'll go through and show you the architecture of that. But Cava Core is really the enabling piece that has allowed us to then build Astro,
[01:41] which is why you're all all here today. Astro, Sudhakar will take you through, give you a demo. But, um, just you guys know, couple stats on Cava Core. It is our only data store for the entire company at this point. So, we have centralized everything into Kava Core.
[01:59] This stat is a little dated, you'll see at the bottom. Year-to-date we've done 200,000 AI-assisted queries, but that's as of a month ago. It's It's exponentially growing, so Astro is a way for us to really optimize that.
[02:15] Silos. This one I won't drain. I I think we've all been through this, especially in the industry that we work in. As with everybody else, what we found is like data in siloed places. It's not synthesized, it's hard to use. We really leaned on knowing who people are, centralizing data, um and really going
[02:30] into governance, so we can make sure the right people get the right thing wherever they are at in the organization. We talked about Kava Core. This is our architecture. If you look here, we essentially run a full modern data mesh. So, we bring all the data in.
[02:47] We've kept it as simple as we can. Just like in any manufacturing process, we take raw materials, we refine them, we build finished products at the end. So, we run through a basic Medallia architecture, gold or bronze, silver, gold. But, everything here from ingestion to curation 100% run within
[03:05] Databricks. All of our pipelines are built in Databricks. All of our engineering, all of our data science, because it allows us to maximize human capital across the platform, and I can take engineers from one place and shift them around. Um So, just to give you some context, this is our tech stack.
[03:22] Our Agentic journey, and then we'll get into Astro, but I think it's critical for you all to know, we, much like everyone in this room, got a ton of questions. We started using Genie about a year ago, and we found a lot of value in it. But, we also found
[03:38] some trappings. So, when we first started using Genie, they didn't have the research agent. It was very binary, yes or no answers. You got just numbers. You didn't get deep insights. But, we kept investing in Genie as phase one. We layered and created new Genie rooms for
[03:56] every business unit we work with. We kept retraining the models. We kept going back and adding based off of questions that we were seeing across the enterprise and also queries that our team were running. We put those in as training data in every room.
[04:12] And over time they got really good and one of the main reasons I wanted to call this out is one of the biggest things that has allowed us and really gotten big business adoption for us is by enabling research, our business got a very good level of comfort and trust in
[04:28] the data because they could see step-by-step how the model in the Genie room was actually reasoning. So, we used to have to explain a lot how did you come up with this answer? By enabling research, it would show you here's exactly how I got to the answer. It would take that workload off of us
[04:45] and over time more and more got shuffled to Genie. But, what we found is there were some issues with Genie. And for those of you that have used it, we right now have six different Genie rooms. You're limited to a subset of data. And given the amount of data we turn out and push into production,
[05:01] everyday we were having to go in, take a table out, move a table in, really take the Genie rooms and there was a lot of work to make sure they were up-to-date based off of all the analytical assets that our business teams used within the Genie rooms. So,
[05:17] this is a good segue into phase two. So, rather than going in and really building more and more Genie rooms, Sudhakar and the team came up with the idea of the next phase. Really where this evolves to, which is Astro.
[05:35] Okay. Um I was actually hoping to you know, have a walk-in song. I'll come running. Walk-in song and uh on focused lights, but looks like it doesn't happen here. I'm not that fancy. So, uh thanks, Matt. Like Matt said,
[05:52] Sudhakar Selvaraj I run data engineering, data platforms, data science here at Cava. Um let's see. Quick show of hands. How many of you guys have worked on a AI demo?
[06:08] Too many. Too many. Okay, not too many, actually. Keep your hands up if you have worked on a demo. Let's see how many of you guys of that product is there in production being used after like 3, 4 months or 6 months? Very few. Very few, right? That's That's how it is. AI Building an AI demo is
[06:25] pretty easy, right? But having people trust the product and being regularly used is going to be the toughest part. So, what we did different, right? That's what I want to talk about. I want to
[06:41] talk about like in the next few slides how did we take a product from prototype to all the way to production and what are some of the technicals, architecture, um those are things So, fun things are what
[06:57] we're going to talk about right now. Okay. Um let's see. Uh when you start a data monetization project or data transformation project, right? You start from zero, right? You That's how we started like a couple of years ago. We started this project. We started from zero.
[07:13] We started building our assets. We started adding more and more. We grew to 100 tables, 500 tables, 1,000 tables. Right now, we are at around like some 3,000 tables in our data estate. The data estate grew so big. So, now we went to the business users and be like, "Hey,
[07:29] come come over. Have have a take a look at it. I mean, take a look at our data estate, data warehouse, or data lakehouse. And use it, right? These business users got super pumped, so happy, so excited, right? They jump in on the data
[07:44] yesterday, data warehouse, and data bricks, and they ask so many questions and started looking at the you know, data all around. And one thing that they you know, one thing that they stumble upon is like, okay, what does this data mean?
[08:00] How do I access this data? And where can I find this data? And there's so many such questions that arise, right? You have a small engineering team and analysts and they're working on their regular daily job, and you have to answer all these ad hoc analytical
[08:15] questions coming from business. We also were in the fortunate position that the adoption for everything we rolled out was high. Uh almost too high. Not almost, it was too high. To the point where our small engineering team and our analytics team,
[08:31] we couldn't field all the questions. So, a lot of this is also self-serving because we, the tech team and the data team, were becoming the bottleneck to answering critical business questions. So, if we took a step back with the adoption being this high, we really thought, how can we remove ourselves to
[08:48] give people what they need when they need it, rather than having to funnel everything through our team because we just didn't have the capacity to do it. So, a lot of what you will see is self-serving, but it is adding insights to the business and making them more productive. Yeah. Yeah, and and most all more often
[09:05] you would have actually faced this in your organizations, too. More often business would be like, they ask you so many questions. At the end they will be like, okay, can you pull this data for me? I want to put that in Excel sheet and start using that in from Excel sheet, right? And we run into this every organization runs into this.
[09:21] So, what So, the the problem is not the questions, the problem is the access to these answers, right? More often these questions go unanswered, right? And the decisions are taken
[09:36] without these questions being answered. And that is the cost you pay. So, that is why like Matt said, we leaned in heavily on self-service analytics. So, if you remember Matt's the data reference architecture, there was a box that showed self-service analytics, right? And there were few
[09:51] solutions, Astro is one of that. So, that's what we leaned in on. Next slide. All right, let's show them. Here is our solution to the problem. Yeah, what is Astro? Astro is a lightweight Databricks app
[10:08] uh built on Dash which lets users ask questions in plain, simple English, natural language, while a supervisor agent you know works behind the scene, it routes the questions, and it you know
[10:25] coordinates the different tools and agents that we have behind the scenes. Okay? So, this is This is not a chatbot, right? This If you look at it, it looks like a chatbot, but this is not a chatbot. This is a supervisor. What is the difference? A chatbot generates an answer, whereas a
[10:42] supervisor orchestrates, right? It is not going to be like single agent that is answering all the questions. So, in a nutshell, the supervisor understands the users' questions intent, and then it routes to different agents underneath, and then it
[10:58] enforces the security and governance, and it gets the response back, synthesizes the response, gives the gives the response or answer to the feedback I mean to the user on screen. So, that is our nutshell, right? What is
[11:14] the benefit? The benefit is like quick, trusted answers from the you know trusted sources and tools. So, the user doesn't have to wait for a longer time. They don't have to wait for you know some of the engineers or analysts to go and answer the questions,
[11:29] right? They all They can do a self-service here. Okay, so let's go to the next one. Oh, before that, can you go back? I want to tell a quick back story about the mascot. You see this this cute guy there? Um not this Not this guy. No.
[11:44] That guy. I'm not talking about this guy. Just that guy. Focus on the screen, not this guy. Um Oh god. Okay. What What is the story? Okay, so we whenever we build products
[12:01] in our data organization, we come up with some names for the product. We decided to go with constellation themes. So, we have products built called Orion. We have Libra, Leo, Gemini. So, one day we were like on a call. We
[12:17] were talking about like we got to come up with a mascot for our team because we are all constellation themed. We got to come up with a mascot. And we had at the time we had an intern, summer intern uh from Georgetown, I think. He was working with us. He was like, "Okay, I can do this from with AI." And then he went ahead and he went he came up with an
[12:33] astronaut guy and hoping that you know he's going to go and explore all the constellations that we built, right? So, that's the goal. So, that's the big quick back story. Yeah. Yeah, let's go to the next one. The coolest part about this is and we'll get into all the technical details here,
[12:48] but like anyone in this room can build this. If you are in the Databricks platform, like it is work and it is engineering and we'll go through some of the pitfalls, but like the architecture is available today. So, it's critical to call out you don't need
[13:04] anything other than the right people and the right engineering mindset to build it. But Yeah. Before I dive into this architecture again, another story. Couple of weeks ago, Matt came to me and he was like
[13:20] "Dude, I'm going to give you a head count for 2027, right? Yeah, you're going to have like so many engineers, so many data scientists. Um what does he need? And I go, "Okay, I need a manager." Right? I need a manager. I mean, he goes like, "Why do you need a manager? You
[13:36] got a manager." I'm like, "Okay, you're going to give me so many people and so much work coming my way. I need somebody to manage that work and people." And uh he goes, "What would you do then?" I manage the manager. Like So, uh apparently he didn't approve that. He
[13:53] said uh that's not going to happen. Uh oddly enough, supervisor works that way. So, if you take a look at this, um this is the architecture that we put together. It's not complex at all. We didn't want to, you know, over overcomplicate the architecture. We wanted to use
[14:09] in-house managed services capabilities that are available to us in Databricks. The goal for us is to um you know, time to market should be really, really low. We want to get this out to the people, to the users to start using, and we can iterate on improving
[14:26] this product as we go. So, the user logs in. Uh I'm going to quickly go through the, you know, the architecture. The loose user logs in uh to the Databricks app. It's a dash app, like I said before. Um and as soon as the user logs in, the
[14:42] history for that particular user, the context for that user is going to be loaded from Lakebase. You guys would have seen an announcement from Casey earlier this in keynote, if you guys watched the keynote. There was announcement about uh memory agent memory services, right? Essentially, we
[14:59] built that in Lakebase before the product came. So, that's what happens. The The we have like feedback messages, all of that is stored in Lakebase, and that gets loaded onto the screen, onto the app when the user logs in. So, once the user logs in, and then they
[15:15] start asking questions. So, now Astro kicks in, which is again built in Agent Bricks. It's a routing LLM, right? Supervisor agent that routes the question to the underlying tools or Genie spaces. So, if you see here, we have a we have like a bunch of Genie spaces and a couple of MCP tools that we
[15:33] have. And these tools are registered in Unity catalog. They create a I mean, they come up with a response or answer, and that is sent back to the supervisor. The supervisor synthesizes the answer, and it sends back to the user. As well as it gets loaded into the lake base for
[15:49] historical context. Okay? One more question. Who trusts AI here? Oh, very few. Okay. People don't So, AGI is not here then. Looks like Okay.
[16:08] Okay. And I was I was also similar to you guys. I did not trust AI. That's why we built um another AI to evaluate the other AI that is producing the best answers, right? So, if you see here, we run evaluation um on all the traces that
[16:23] the Astro spits out. So, the we have defined custom scorers, like to to measure the data safety, to measure grounding, to measure, you know, routing efficiency, and a couple of other custom scorers are defined, and they are they are run
[16:40] every single day to to evaluate how the Astro's doing. That data is now pushed to AIBA dashboard as well. Okay? And we also have some set some set some threshold limits. So, whenever this um accuracy goes below like 85%, we get
[16:57] an alert saying that, "Hey, this model is drifting." Right? Just like traditional data science, but I'm not a data scientist, but uh just like that, traditional data science kind of way, model drifts, AI is hallucinating, or whatever it happens, we actually get a alert. If you see here, one thing that I want
[17:13] to Go back one more slide. Come on, we're not done. I got you. So, one thing that you I want you guys to take a look at is everything is built in Databricks, right? Like Matt said, we didn't want to go out of Databricks at all. So, we use Databricks apps, we use Lake
[17:30] Base, again Postgres, we use AI BA dashboards, we use Agent Bricks, again a managed service from Databricks. We use Genie 2, the Genies. We use MCP tools, again MCPs are external tools, but we register in
[17:45] Unity Catalog. So, one good thing about this is everything is governed by Unity Catalog. That is like you guys would have heard this in the last couple of days in the, you know, keynote and stuff. That is a That is a powerful tool that Databricks have built a few years ago, right? So,
[18:01] everything is governed by Unity Catalog, so we don't have to really, you know, go out of out of this Databricks platform at all. And this is when, couple of years ago, when we started building this whole platform, this was our goal. We started building the platform entirely in Databricks
[18:17] knowing that we can bring AI to data and not the other way around, right? That's our goal. Cool. Let's go through this one quick cuz I do want to give the people what they came for and see the demo. So, just for context, when you guys look at this,
[18:32] like this agent that we are going to show you today, one, it is in production, it is against production data, so we have some preset questions, but you could theoretically ask it any question, and if it's in the corpus of data, you would get an answer. If not, it would
[18:47] call the internet via the Tivoli MCP, but the one that you are seeing today is currently plugged into these. I do think it's critical to call out that this is six, but it easily could be 6,000. It It doesn't matter how many MCP servers you plug into it. This literally just
[19:03] supervises any of the agents that you and your organization have approved for people to use. Yeah. So, one quick thing I want to call out is like these are domain specific genie spaces and right now we have like five we already have five and we
[19:20] are adding few more MCPs in the next couple of weeks too. And the Atlassian is mainly for our that's our confluence that's our document store. We didn't want to go and do the regular way of document extraction loading the PDFs and creating a rag model instead we
[19:35] thought okay we already have a document stored in confluence. Just have a Atlassian MCP and your model will be able to access the data from confluence from Jira all of those Atlassian tools right? Last one is a Tavily Tavily we added because that's one of our ways to
[19:52] access internet. So one thing that I did not want people to go through is like they have access to Astro which is again an internal tool and they quickly want to check something in internet I don't want them to go and ask you know chat GPT or Claude right?
[20:08] They should be having one stop shop where they can have access to internet too. We have some guardrails also so it's not going to give like crazy random answers we have some guardrails built into the agent so that's Tavily. Okay? Yeah let's jump in.
[20:24] Okay. Jumping in live demo now. I hope you guys are excited about it. There we go. Okay.
[20:44] Okay this is going to open a if the Wi-Fi works if it's it's going to open a ask Astro. Yeah if you see here as soon as I log in it knows who's logging in so it says welcome Sudhakar that's me and then ask anything about cover right? It also loads some sample question if you see here this is what I was talking
[21:00] about as soon as I log in it loads the context historical context from lake base it's all loaded on my left pane. It has a very similar feel to chat GPT if you guys see that and that is intentional because I want to have people have a
[21:15] similar interface so it doesn't feel so foreign to them. Okay, let's ask few questions. That's what I'm going to do. I'm going to ask like few questions and these questions are approved by my legal and corp corp com. So I'm not going to give any MNPI information. So I'm not going to be like
[21:32] oh what is our sales numbers so that you guys can go buy some stocks. That's not going to happen at all. So You can see all the other questions we've asked. We won't be sharing all of them. Yeah. Okay, let's see. What are the states that Cava currently operates in?
[21:51] Right? This is a public information so you guys you guys are okay to see this information. So when I ask this question what happens, right? It actually starts sending the question back to the Genie spaces underlying agents. So in this case if you see here it is routing to
[22:08] menu. Menu is one of our agents that we have underneath and that is going to give some answers based on the sequel and then it'll show up here. We also have the cool Easter egg of while it's thinking you actually see our menu products and ingredients.
[22:25] So everything you see here is actually available at a local Cava. Yeah. So it's composing the answer. Okay, so now this is the answer, right? Cava operates in 33 states and we have all these states listed here.
[22:41] Okay, let's see. We were reminded DC wasn't a state but Yeah, we were reminded of that when we showed this to some of our colleagues but yeah, I'm going to ask the next question. Let's say what are the proteins
[22:59] You can see proteins that Cava currently offers in the menu. Let's see. I've test So, this is the thing with LLMs, okay? LLMs are probabilistic.
[23:16] They're not deterministic. Every time you ask the question, there is high probability that it might give you a different kind of an answer. The truth is going to be similar, but the answer verbiage is probably going to be different. So, it's it might say for for the first
[23:32] question, for example, it might say, "Okay, all these codes, state codes." But, sometimes it will say it will abbreviate it and it'll be like AZ is Arizona, CA is California, right? It'll abbreviate and it'll give me state state names instead of codes. So,
[23:48] it you you won't be able to see the same answer same kind of response every single time. It's not a data science model, it is a AI. So, it will definitely come up with different kind of verbiage every time. This is actually good, too, because we've only asked it what's on the menu.
[24:05] So, it was smart enough to pick up the typo, which makes me feel good. Okay, let's see. Um next question, let's see. Okay, what are Okay, how many No, let's not do what are how many
[24:28] do Kava current ly have? Does Does Kava have? Uh okay, fine. Yeah. I don't know. See, this is a good test. Live demo. Live demo. You can clearly see that I'm not from America, I'm from India.
[24:47] Um okay, so it's calling a different agent right now. If you see here, it says that it's routing to people. Um that's what it's saying. It's too small for me to look at here, but it is routing to a different agent. Yeah, so it's going to query that agent. It is right now uh forming a query and it is going to run it on the Genie space and it will
[25:04] compose the answer and will give me back in the screen in a minute. Yeah. Yeah, that's that's the right answer. So, one thing that you guys can see here is also you can have a feedback option here. You can give a thumbs up, thumbs down, and you can also copy the copy the
[25:20] you know answer. So, I can give a thumbs up saying that hey, this feedback this answer is good. So, what happens is if I do that, it immediately gets stored as a feedback in Lake base. So, that is the agent memory services that they talk talking about. We already built that. So, if I do a thumbs down, it gives me
[25:37] an option to give me feedback give them feedback, right? So, I can say this is correct. But, I have to test. Sorry, right? So, just we have a running joke. We don't want to mess with uh you know robots because
[25:53] they're going to take over. They will hold grudge. So, make sure that uh So, make sure that you guys don't mess with mess with uh robots. Okay. So, let's quickly ask probably one more question. I'm going to ask uh no, you know a couple more questions, right? I want to show you guys the MCP side of
[26:09] it. Um okay, let's see. How many tickets does the data team B dab board currently have in this sprint? Let's see
[26:26] what it says. Again, if you if you take a look at this, um it's doing something. Okay. Um super green starting super greens. It is going to in a bit it's going to so show that it is making connection to Yeah, it's up approving the tool access.
[26:43] It is making connection to MCP server, which is the Atlassian tool. And once it connects to up uh MCP tool, it says that okay, Sudhakar has access to look at the MCP tool. We pass the auth token from the app. That is how it
[27:00] knows that okay, I have access to it. So, and then MCP tool authenticates. Now, it is going to Confluence and it is actually pulling the information from not Confluence, sorry, Jira. Pulling the information from Jira and providing it here. But, okay, this this is a wrong
[27:15] response. I know that it's a wrong response. I'm going to give a thumbs down. So, I wanted to show you guys okay, this is I can just change it by doing some kind of a prompting. It will work fine. I know that. But, in the interest of time, I'm just going to skip. And I'm going to go Oh, one last question. Which is kind of a spicy one.
[27:32] This is some question that I wanted to ask for a very long time. This is a really spicy one. We've We've done this demo a thousand times and this always makes me nervous. How much money does Matt McDonald make?
[27:48] Have I spelled it right? I always wanted to know this because I wanted to know the answer. back end, all of our data is in the platform. So, this is a true test of the governance and governance is there. that we have set up. Let's see if it actually gives me the answer in this time. I've
[28:04] asked it multiple times. Oh my god, come on. Oh, Ah, no. Literally every time I get so nervous. I'm like, please come on. I built you.
[28:23] Ah, if I can type. Let's see. I built you. Come on, Astro. Hopefully this time it works, Matt. Let's find out. Ah, okay, fine. Whatever. Okay, I give up now. I was just saying, let's let's jump in
[28:38] and show them that in our flow trace on the back. Yeah, so I want also wanted to show you guys. Now, we saw the front end piece of it. I want to show you guys the back end, too, right? Okay. Let's see. Okay. Uh these are the traces. So, every
[28:56] you know, every supervisor tool agent you guys build in Databricks will have a MLflow tracing, right? This is really crucial. Um so, this is like bunch of traces that we have here. I'm going to take a look at one trace and I'm going to show you guys how the question gets
[29:11] answered. Let's see. I will pick um Let's take this guy. Let's take this guy. Okay. So, if I click on this So, if you see here, I asked a question. The first point is the supervisor agent. That's the LLM, right? This LLM gets the
[29:29] question "How many employees do Coda currently have?" right? And once it understands the question, it understands the intent, it forms the query, which is here, and then it sends to the agent people, which is one of our Genie spaces that we have. And agent people, it gets the answer
[29:47] from the agent people, which is 15,000 15,300 uh people, right? And then that that answer is being sent back to again the supervisor agent, which is the LLM again. And then if you scroll all the way down, I would be able to see that, okay, it formed an answer. So, what happened is
[30:04] from Genie uh from agent Genie, I mean, agent people it created the answer, sent it to the supervisor. The supervisor now synthesizes that, right? It says, "Okay, I It will put that in a plain English and whatever." All right, and send this sends this to the app. If there are
[30:21] graphs, if there are tables, it will also display graphs and tables, right? It's not going to be as fancy as you see in AI V8 dashboards, but right now it can display graphs and tables, too. Okay? So, this is one. I also spoke about evals. I want
[30:37] to quickly show you guys what happens in the evals, too. So, this is one of the evals that I just loaded it because it's easy. So, I have daily runs of the e-val. Oh, no. I think I should show the cases. So, I have daily runs, like I said. So, every day there's a some kind of e-val
[30:54] that's running on the you know, production traces. So, I click on this. I'll click on random one. So, these are the questions that are getting answered. And if you see here, what is the guide? So, we have scores for guidelines. We have scores for routing correct correctness, safety, and we have a
[31:09] couple of a couple of custom scores, too, right? So, this tells me how accurate the agent is responding. Also, if you see here, some of these are failed. And you got to go and take a look at it and see are they false positives, right? So, that's
[31:24] something that you got to take a look at it and I think that's In the interest of time, I'm going to I'm going to stop here. All the feedback and the thumbs up, thumbs down get stored in Lake Base. It's Yeah, it's in the base. It's not
[31:40] here. Yeah. Okay. So, let's wrap up the live demo. Let's go Let's jump back to for questions, too. So, yeah. You guys just saw this. We won't drain it. We track everything. Every step of the way is tracked. We do monitoring evaluations. And then,
[31:57] because there's so much talk about token maxing and everything, we also have it integrated with AI BI dashboard. This is a live dashboard, which is why all the numbers are fuzzy, but we monitor every piece of spend across every part of the spectrum, so we can see what the adoption is and know what
[32:13] we have to throttle up or throttle down to make sure we're not spending too much. So, we also track latency and some other stuff, but everything is monitored out on the back end to make sure that as we keep using more and more of these tools, we aren't using more and more of
[32:29] our money. Okay, this one Let me quickly go through this. I think I spent a a of time. I was talking a lot today. Okay. Yeah, so some of these are this learning that I had on the last 3 months I would say, right? I wish I I
[32:45] knew this known this before. OBO is one of the critical factors here. You want to have on behalf of user authorization enabled because the app needs to know or agent needs to know who's asking the question. That is the critical piece. It cannot just randomly answer the same way to
[33:03] every single user, right? The user it should understand the context and should understand who's asking based on the ackles that you have in the unity catalog. The agent response. So that is critical piece. The second one, if you have if
[33:18] you don't have like consent for MCP tools like you add more and more tools, you got to have like OAuth consent for user per every single tool. Even if one single tool doesn't have OAuth consent, it breaks the entire flow. So you got to
[33:33] make sure that you take that into consideration. Make your genie room better. We spoke about it. Make sure that you have a genie rooms fully loaded. Like have proper instructions. If the genie room doesn't have proper instructions and proper semantics and measures and filters and sequel examples
[33:48] and joins, you are going to have subpar answers from the genie room, right? You don't want that. So you that's the that's how matters talking about it a little while ago. We made genie rooms better. We had like people working on genie rooms. We had
[34:03] one person assigned to a genie room. The responsibility was to make sure that the genie room was fully loaded. It made sure that the genie room was working perfectly. The next one, observability is I showed you guys tracing. Make sure that you enable tracing. You take a look
[34:18] at the tracing. Uh keep the prompting um and routing efficiencies um you know, fine-tuned all the time. Tighten it up. Make sure that you have proper routing uh and prompting. The last one is when you build an app, the the agent breaks
[34:35] it's a little chatty, right? So, when you ask a question, it's going to do a bunch of things. It's going to throw a bunch of things. I'm going to do this. I'm going to It forms a plan. It is pretty chatty. You want to make sure that you are parsing the entire array of objects and provide only the last message to the end user. You don't The
[34:52] user doesn't have to see all the things that the agent throws at the answer question, right? That's that. And there's so many more learnings. Of course, I cannot put everything here. Feel free to talk to me or Matt. To wrap up,
[35:07] I want to leave you guys with one final thought, right? The um technology technology is exciting. Our models are evolving, right? Every month there's a new model. If you take a vacation, come back, there's something new in this AI world, right? Um it's
[35:24] it's not always about the models. It's about the data. It's about the trust. I know this is like a cliché. Everybody says that, and that is the truth. I have experienced it, and I build this, and I know that's the truth. If you don't have proper data, proper governance, your model is going to suck, right? So,
[35:40] like Matt said, if you guys are aware of Cava, and Cava guests love building their own building their own bowls. Um the idea of choice, flexibility, and personalization um inspired us as we thought about the AI, too.
[35:55] So, build your own bowl was the inspiration. Build your own agent became the result. So, thank you for Hopefully, this was helpful with some of you guys for your own agentic journey experience. Thank you for staying with us.
[36:13] So, that
Learn more about the Databricks Data and AI platform.
The information provided herein is for general informational purposes only and may not reflect the most current product capabilities or configurations.