Skip to main content

AI observability and evals: from POC to production

Summary

  • This video covers the shift from AI proof-of-concept to production, where enterprises now run thousands of agents simultaneously and require systematic observability, governance, and evaluation frameworks rather than informal spot checks.
  • Calibrated LM judges replace manual evaluation and the 3x3 framework measures must-have behaviors, must-not-have behaviors, and actual user behavior patterns, with teams that invest in proper governance reported to be 12 times more likely to reach production successfully.
  • Agent sprawl — the hidden cost of unmanaged agents growing across the organization — is identified as a key risk, addressed through data-centric tracing, ROI measurement, and centralized governance including MLflow AI gateway.

AI observability and evals: from POC to production

Watch: AI observability and evals: from POC to production
Enterprises are rapidly evolving into multi-agent systems running thousands of AI agents simultaneously. To move from proof-of-concept to production, you need robust observability, governance, and evaluation frameworks. This keynote examines the pillars of AI governance, defined boundaries, data-centric tracing, clear ROI measurement, and interoperability, plus the critical gap between current practice and systematic quality measurement.
Learn how to build production-grade evaluations using calibrated LM judges instead of manual spot checks, implement the 3x3 framework for measuring must-haves, must-not-haves, and actual user behavior, and navigate the path between proof-of-concept and production with proper observability tooling. Discover why 12 times more AI projects reach production when teams invest in governance and systematic evaluation.
🤝

Chapters

FAQs

What are the pillars of AI governance for production agents?

The pillars covered in this video are: defined boundaries for what agents can and cannot do, data-centric governance using audit logs and traces, clear ROI measurement to justify AI investments, and interoperability so governance controls work across different tools and platforms. MLflow AI gateway is highlighted as a centralized governance layer supporting these pillars.

What are calibrated LM judges and why do they replace manual spot checks?

Calibrated LM judges are LLM-based evaluators aligned to human judgment through calibration, making them consistent and scalable replacements for manual review of agent outputs. As enterprises run thousands of agents simultaneously, manual spot-checking becomes impractical, making systematic LM judges the only practical path to quality measurement at scale.

What is the 3x3 evaluation framework?

The 3x3 framework evaluates AI agents across three dimensions: must-have behaviors the agent should always perform, must-not-have behaviors the agent should never perform, and actual user behavior patterns observed in production. This framework helps teams move from subjective assessments to systematic metrics that can be tracked over time.

What is agent sprawl and why is it a risk?

Agent sprawl refers to the uncontrolled growth of AI agents across an organization, with teams deploying agents independently without centralized visibility or governance. This creates hidden costs in compute, tokens, and security exposure, and makes it difficult to audit agent behavior or attribute failures to specific systems.

Full transcript

[00:03] Hey everybody, how's it going? Um, we're going to go through some slides. I probably have a little too many, but um, I did want to start off with some of the key concepts and learnings. Uh, so one of the cool things about my job is that I get to work very closely with customers on building production AI agents, okay? Uh, it's
[00:19] first start off with data engineering systems, but then we obviously shifted over to AI agents. And so the biggest concern is why I called called the keynote the way I did is that where how do you architect intelligence without the chaos? Unless you all love chaos. Yes? No? Okay. Well, I do, but that's a whole
[00:35] other problem. Okay, now uh, we're going to break this down into sort of four points. State of AI agents, uh, then we're going to talk about the pillars of AI governance, uh, a little bit about agent sprawl, but then of course the key concepts here is about the importance of observability. So if you read nothing else, it's these
[00:51] four concepts. I'm going to actually going to go through these concepts right now, uh, which is basically this. All right. Enterprises are quickly becoming multi-agent systems. They're running multiple agents multiple times. This is actually from our state of AI agents report. These are some of the key
[01:06] learnings that in the end we're talking about not one or three or five, but we're often talking about thousands of agents that are now running in the systems, okay? So this gets really complicated really fast. Why do we care about observability? Because we actually
[01:22] have to make sense of all of that, okay? Um, AI agents are driving database functionality, database features, okay? This is an important aspect. You you want to cache whatever you're running, you often need database systems, okay? And so you'll
[01:37] notice that this massive jump from October of 2023 to 2024 to 2025 in which basically, uh, we've got people who are creating these databases, but even more so agents that are creating databases autonomously, okay? So that's an important factoid. Um
[01:53] AI is also now critical part across workflows for many different enterprises. It doesn't matter which one it is. Whether you talk about media, energy, financial services, all these different things, each one of them have slightly different. I'm going through these slides really quickly just because I want to talk more about observability.
[02:10] That's the reason why we're here. The final tidbit, AI evaluations and governance are the building blocks of production. The main reason why we're here today, okay? And so there's basically a 7x more systems, more agents that go in production when they actually start building systems
[02:26] like Evals. Without them, basically have a problem getting these systems into production. Okay? So this is the reason why we actually really care about this agent tonight to be talking about evaluations, okay? I'm going to flip on ahead. Now, pillars of AI governance. Now,
[02:44] if I actually go through just the section, this would be its own keynote, honestly. But so I'm going to go through this relatively quickly, but if you learn sort of nothing else about this, you actually have to have defined boundaries, data-centric AI governance, ROI intelligence, and open and interoperable. I'm going to go through
[03:00] the four concepts relatively quickly, okay? All right, defined boundaries, okay? Agents actually have to be told what they're allowed doing, what they're not allowed doing. If you don't define the boundaries, for example, I have a link here which talks about how Meta accidentally let a rogue AI agent reveal
[03:15] confidential information. Not picking on Meta. I've got lots of friends there, so this is not me trying to pick on those folks, okay? This is just something that had popped up when I was writing creating the slides, okay? But the reason I'm calling this out is because we we're all one step away from making a mistake and letting our agent
[03:30] do things that they shouldn't do. Have people playing with OpenCL? Yes? Yeah, okay, so are you actually running OpenCL on your own separate Mac mini that is not connected to your network at all? If you're not, guess what? You just revealed a whole bunch of your info, okay? So that's the call outs we have to
[03:47] make. We actually have to have those defined boundaries. All right. In that approach, you have to have to have data-centric AI governance. So, in other words, very much in the conversations that we're having about observability, you actually do need things like audit logs and traces. The
[04:02] ability to analyze those logs and traces, and ingest fresh, golden, trustworthy data. If you're not doing these things right from the beginning, what you end up doing is if you think about this process as an afterthought, you're not going to get all the right
[04:17] data, you're not going to get it on time, and you're going to come to conclusions that actually are not quite right, okay? So, this is an important aspect. A lot of people forget about that, and so that's why I want to call that out. And then, a really important statement here is that, "Okay,
[04:34] the question people often ask is, 'How much are we spending on AI?'" And of course, I'm guilty of spending way too much on tokens. Yes, you are? Okay, cool. So, you I don't know if you've had your manager getting on your case, but I definitely have. Now, the answer I often give is like, "But
[04:50] what are we getting from it?" So, for example, I spend $100 on an agent, and I get 20,000 marketing leads if I happen to be a marketer. That sounds awesome. That is excellent ROI. But, the 20,000 marketing leads are from last
[05:06] year's data. How useful is that? Okay? So, that's the whole point of observability that we want to know what we're doing, what we're getting, what we're getting out of it. If you don't start with that process right from the get-go, you sort of end up failing. What most people are not realizing is that
[05:21] you do need some form of centralized governance. I'm I happen to be talking about MLflow in this scenario just because it's an open-source software, but anybody that's coming to this event, any of the vendors that are here at this event, any other open-source system, all of this concept applies. Okay? That's
[05:38] what I'm trying to get across. Like, you need some form of central governance to understand what's your cost and capacity views. What is your compliance violations? What are the integration bottlenecks in your systems? Okay? And so, for example, if you have an AI gateway, that gateway allows you to go
[05:53] ahead and manage permissions, rate limits, input guardrails, output guardrails. Okay, you need systems like this. So, my quick, because I'm from Databricks, MLflow pitch is that MLflow AI gateway happens to be able to do a lot of that stuff really well. Okay? The only reason I'm calling that out, and I
[06:08] don't mind calling out here, is just because this is an open-source project. So, it's anybody can integrate with this. We make it available for free. We take a lot of advice from the community to try to make it better. But, the reason I call this out is because from that state of AI agents report, what you'll notice that
[06:24] what we notice actually when it comes to customers of Databricks, they are often doing 7x more investment in AI governance and security since 2025. Okay, they've massively jumped up this idea of governance and observability. They had to because they're the ones that do that are 12 times more likely to
[06:42] have AI projects in production. Okay, so everything we're talking about here today, every single one of the vendors that are here today, the reason why evaluation is so and observability is so important is because that of that stat right there. Okay? All right. So, one little quick thing,
[06:59] in the process of doing 10,000 or 20,000 agents or however number of agents, you get the scenario of agent sprawl. Okay? Now, I'm not going to dig too deep into it, but the whole concept is that you've got low quality of reasoning, no visibility, audibility, too many AI vendors, no way to measure quality.
[07:15] Okay? So, everything we're trying to talk about here today is an attempt to get around that, or at least make sense of what you have in your systems. Okay? And so, let's talk about the importance of observability. Boom. Almost. Let me switch gears just a tiny
[07:33] bit. Okay? I'm going to do a segue here. What is software? Don't worry, I presume most of us know that answer. Okay? But, I did want to call out like it's a means of translating human intention into operational behavior. Okay? That's that's the definition right there. Okay?
[07:49] So in other words, human intention, instructions to a computer, computer action, verify it works. Okay? That's what we all do as software developers. Key steps in building software is that you have to create a specification of some type. Okay? Hopefully the specification's readable and hopefully the specification isn't 300 pages long
[08:06] cuz in fact but you do need a specification so that way you can actually figure out what you're building. It's really that simple. Okay? And then you take the behavior of that implementation to match that specification. All right? That's easy in code.
[08:22] Code has semantics. Code has a process. When you have two or three different people looking at the code, they can usually interpret the exact same way what the code is supposed to do. Cool. This is hard in AI. Same prompt can be read by three
[08:39] different developers with three very different answers. You send the same prompt to the same uh LM, they also might give you different answers. Okay? So this is why the only thing we can really do is look at the behavior of the system and compare that.
[08:55] So because of that, again, why is observability so important? Why is evaluation so important? Okay? So the challenge is how do we build AI systems that are equivalent to million-line programs? Okay? How that multiple people can work on simultaneously. All right? And I think
[09:11] if nobody if nothing else from the theme of what I'm talking about so far, it has to be about observability, has to be about evaluations. Okay? So we only know one way to do this. Okay? So there we go. It's basically building that specification. So what do I mean by specification?
[09:28] Evals. Okay? The purpose of today's event. The specification for us to get the LLMs to actually understand what we're trying to build is building solid evaluations, okay? So, the final tidbit of this particular session, just to make sure I have plenty of time. Good.
[09:45] Here's our challenges, our framework, what the potential solution is, and what are the takeaways. Okay? This is the meat of today's topic. All right. So, the challenge is measuring quality is important, but it is difficult. Everybody agree with that? Okay, I hope we're hoping that the six
[10:01] companies that are here, plus others, are starting to make that easier for everybody, but the reality is it's still difficult. So, let's talk about that. There's tons of use cases. Everybody knows that we have these use cases. Coding agents, we have delivering cool restaurants, personalization, things of that nature.
[10:17] Okay, that's great. But, each use case has different quality requirements, right? Just because it's super simple or super straightforward in the coding agent scenario, because it's pretty much straightforward. If you're talking about personalization, that often can be very subjective, right?
[10:35] Okay. So, the GenAI measurement gap is basically what do we need? Reliable quality signals, scalable evaluations, and consistent standards. But, what do we get? We what we actually have is manual spot
[10:53] checks. Anybody guilty of those doing those? Okay? Yeah, I'm if you're running these systems, that's what you're doing. You're trying to run these manually. You need subject matter experts, right? So, for example, if I'm training an agent on coffee, then sure, you can probably I can probably be that right person. I'll
[11:08] literally drown you all about coffee facts, okay? But, that's not what you're here for. And the fact is you don't want to be guilty of basically being dependent on just a few people, a few subject matter experts, to define everything. And the worst one is vibes,
[11:24] right? We vibe our our evals. It's already bad enough that we're vibing our code, now you want to vibe our evals to and hope the heck that coming out is actually makes sense. And so, an approach that we often find as a solution, okay, is calibrated LM
[11:40] judges. What do I mean by caliber? I mean LM judges basically, you know, LM as a judge, it's you're using LMs to go ahead and judge the output that's coming. But, you calibrate them. You have to fine-tune them. So, that way they're designed for the subject matter that you're trying to hold at. So, if I'm
[11:55] building a coffee agent, they happen to know all the information that you knew about know about espresso from Vashon via Seattle. That's my bias, okay? Because they have all the information to actually assess coffee correctly. All right? So, why calibrated LM judges? There are many
[12:10] ways for an LM to basically be fuzzy. Sound like they're saying the right thing. Sometimes they're saying, sometimes they're not, okay? Uh they often have fuzzy behavior requiring fuzzy verification, okay? This is the basically a fun way of saying it's not quite
[12:27] right, okay? Uh LM judges can emulate human judgment though, if you calibrate them correctly. So, if you're leave nothing else from this particular session, with all the complexities on how to build evals, this is one of the ways that we found to
[12:45] basically using LMs to actually help us do Sorry, excuse me. Create proper evaluations, okay? Now, what's the framework for this though, okay? The what I'm about to show you is a 3 by 3 approach for understanding your
[13:02] evaluation needs, right? Because right now, I'm still talking in terms of abstraction. You need an evaluation. How do I actually build up my framework for what needs to be evaluated, okay? So, let's simplify. You have the must-haves, and you have the must-not-haves,
[13:19] and what users actually do. How many people have had a fair scenario where you spec out whatever feature you're trying to create, you think you've got it all figured out and then you realize the users managed to figure out how to completely break everything you just did. Is that true? Yes? Come on, everybody's
[13:34] got to raise their hand. There isn't a single person here that hasn't suffered from that if you're in the coding, okay? So. All right. So, what are the must-haves in this scenario? They're inputs that the application is uniquely trying to answer. In other words, I have an agent
[13:49] that's on processing video information or processing PDFs. It must be accurate 100% of the time. Okay? That's my definition of the must-have. I can't have it hallucinate
[14:04] receipt information. I can't have it hallucinate invoice values. I need it to be accurate 100% of the time. Otherwise, I can get away with expense reports that are crazier than normal. I wish I could, but I can't, okay? But must not have
[14:20] inputs that could possibly, but should not, lead to the application responding erroneously. Like the fact that the application is processing my receipt and it accidentally says you just spent $100,000 at Lamar. I do spend money there, but not that
[14:36] much, okay? We cannot have an accident here. It needs to be the correct value, the correct amount of money, okay? And then, like I said, what the actual usage is, what users actually do. Okay. So, using that frame, this is your 3x3.
[14:51] Okay? The must-haves, must not haves, actual usage, okay? So, if you're doing a POC, your must-have is the proof of concept, often focusing on the good enough. This is your vibe checking scenario, okay? Which is fine. I didn't say you can't do vibe checking. I just said you can't do
[15:08] vibe checking all the way to production, okay? Understand that. All right? The must not haves, basic safety checks. You do want things like that, okay? Actual actual usage, minimally, if at all, addressed. Okay, cool. Simple straightforward.
[15:23] If everything goes correctly, everything goes right, then you can get the sucker into production. Okay, we just the farthest right here. Okay? And so then must have online monitoring of performance of these targets. Must not have online monitoring and mitigation of for safety critical
[15:39] inputs, and then actual usage intelligent updating of offline evaluation sets. More times than not when you're trying to build your systems, the left one is straightforward and easy and everybody's experimenting. If you have the discipline and the processes put in place, you can get to the right one.
[15:57] Okay, the production. I wasn't able to figure out how to make a triangle out of this, but this is the Bermuda triangle of GenAI. Okay? How do I get from POC to prod? So, the pilot. Some people will call it slightly different, so that's fair, but the point is how do I get that middle section? In
[16:14] other words, the must-haves. More comprehensive coverage potentially with the help of additional stakeholders, because we all love talking to our stakeholders, right? No, you're No, you're you're all supposed to disagree with that, okay? We We We all love talking to our
[16:30] management, don't we? Yes? Okay, you're lying to me if you did say that yes, okay. All right. Must not have more comprehensive coverage potentially with the help of additional stakeholders. Exact same problem with your must-haves and your must not
[16:46] haves. And then what the actual usage what users are actually doing. Can they use the POC to collect real inputs from a small set of invited users? Okay? So, that's the Bermuda triangle, and again, my apologies for not figuring
[17:01] out how to create a triangle out of this, okay? But Bermuda triangle of GenAI. Okay? So, evals is how we solve that problem. There you go. Solution recipe for building goal standard evaluations. Basically, if you look at the concept of what we
[17:18] have to do, you start have a POC running, you collect inputs from the target audience, like the people that are testing the things for you, you define a specific annotation or guidelines so the SMEs can go ahead and actually review it, and then you go ahead and turn the SME ratings into the golden or oracle
[17:33] rating. So that way you can constantly have the judge, the LM, go ahead and constantly reference that information. So that way it can accurately calibrate, actually evaluate what's going on. But don't forget, Evals are an iterative process. Everybody that's going to be
[17:49] talking here today is going to call about that. So we're going to constantly go back and forth and back and forth and back and forth. But ultimately, that's what comes what it comes down to. Unless you're doing this process as a regular software engineering life cycle, and you're constantly evaluating, you're constantly doing iterations, you're not
[18:06] going to succeed on how to get out of that Bermuda triangle gen AI to go from POC to production. Quality measurement is necessary for building and improving gen AI systems. This is the reason why we created this event. We wanted all of our community members to come together to speak, all the people who are coming here. Why? We
[18:22] all agree on this absolute context. You have to evaluate across the three dimensions. The must-have, must-not-have, and actual usage. You're trying to avoid that Bermuda triangle, okay? By using a very principled measured approach.
[18:39] Bottom line. Start measuring systematically, okay? Start improving systematically.

Learn more about the Databricks Data and AI platform.

The information provided herein is for general informational purposes only and may not reflect the most current product capabilities or configurations.