Skip to main content

Databricks Platform: From Lakehouse to AI Agents

Summary

  • Databricks has grown to nearly 10,000 employees across 30 global offices serving more than 20,000 customers, with continued investment in open-source projects including Apache Spark, Delta Lake, MLflow, and Unity Catalog.
  • The Databricks Data and AI platform unifies governance and security across data, models, and AI agents through Unity Catalog, with composable Agent Bricks enabling multi-agent systems and Genie providing natural language analytics for every employee.
  • CMS demonstrates how government agencies can use the Databricks Data and AI platform to detect fraudulent healthcare claims, modernize provider directories, and close the last mile from data to mission-critical action through Databricks Apps.

Databricks Platform: From Lakehouse to AI Agents

Watch: Databricks Platform: From Lakehouse to AI Agents
Databricks' data intelligence platform unifies governance, security, and access across your enterprise. This keynote demonstrates how to break down data silos using Unity Catalog and Delta Lake, govern data consistently across tables and models, and make sophisticated AI accessible to every employee through Agent Bricks and Genie. See live demos of multi-agent systems, AI/BI dashboards, and production fraud detection platforms built on Databricks.
CMS shares how government agencies operationalize AI at scale: detecting fraudulent claims, modernizing provider directories, and accelerating digital transformation. Learn how companies and agencies use composable agents to access all data with unified governance, Genie for natural language analytics, and Databricks Apps to close the last mile to mission-critical value.
🤝

Chapters

FAQs

What is Agent Bricks and how does it enable multi-agent AI systems on Databricks?

Agent Bricks is a composable agent framework on the Databricks Data and AI platform that allows organizations to build multi-agent systems where individual agents access enterprise data under unified governance. Teams can combine agents into orchestrated workflows while maintaining the access controls and lineage tracking provided by Unity Catalog.

How does CMS use Databricks to detect fraudulent healthcare claims?

CMS has built a fraud detection platform on Databricks that analyzes claims data to identify fraudulent patterns, forming a core part of its 2026 strategy for fraud prevention and cyber defense. The Databricks Data and AI platform provides the unified data foundation and AI capabilities needed to operate this system at government scale.

What is Genie and how does it enable natural language analytics?

Genie is a Databricks capability that allows users to query enterprise data using natural language, translating questions into results without requiring the user to write code. It is designed to democratize data access so that executives, analysts, and non-technical employees can all explore data through conversation.

How does Databricks Unity Catalog unify governance across data, models, and applications?

Unity Catalog provides a single governance layer that applies consistent access controls, lineage tracking, and compliance visibility across data tables, ML models, and AI applications on the Databricks Data and AI platform. This unified approach means organizations do not need separate governance systems for each type of asset, reducing fragmentation and audit complexity.

Full transcript

[00:10] I'm really thrilled to welcome you to Data Bricks World Tour. Now, we're living in interesting times, new and unprecedented era of innovation. It's powered by data and AI that's open and accessible to all.
[00:26] And really because of that, this is the generation that will transform every industry and give rise to entirely new industries. And I believe this is the generation
[00:41] that will solve previously thought to be impossible problems. So whether you're an executive or an engineer or an architect or an analyst, we all have a big role to play and
[00:57] that's why we're all here today. And so again, you know, on behalf of data bricks, thanks for uh attending data and AI world tour. You know, we've done these at 19 cities spanning the globe. And honestly, we've saved the best for last right here in Washington DC. Please
[01:13] give yourselves a big round of applause. So now whether you're a technical sort of expert or executive new to data bricks or a longtime user I think we've got something here for you today. I'm
[01:30] going to hold here for just a moment so you can scan the QR code to view today's agenda. There's also uh a link sent to in your email as well. So, as you look through that agenda, I encourage you to find speakers you're interested in
[01:46] hearing from, topics you're interested in learning more about, and dive in with us for a full day of data intelligence. So, a little bit about data bricks momentum. It has absolutely been incredible over the past year. We have
[02:02] grown every facet of the business. We have approaching 10,000 employees spread across 30 major global offices worldwide. We have continued to invest heavily in open-source with things like Apache
[02:18] Spark, Delta Lake, MLflow, and Unity Catalog. And we've brought new strategic partnerships that bring models to your data with Open AI, Google Cloud, Palunteer, Enthropic, and SAP.
[02:37] Along with that, there's been a huge expansion of our ecosystem. We continue to grow both the number and depth of our SI practices. Our built-on business has grown to hundreds. We have significantly expanded our technology partnerships and data providers are
[02:54] embracing our marketplace. And all of that's in place to serve 20,000 plus customers around the world. in every industry leveraging data bricks every day, democratizing data and AI at
[03:09] their organizations, building AI products and agents directly on top of their data. We're grateful to have many of our government speakers with us today sharing their stories of how they're driving innovation with data bricks. This morning, we'll hear from Patrick
[03:26] Newold from CMS here on the keynote stage. And I want to take a minute to thank all of our speakers for sharing their inspirational stories and allowing us to learn from you. So, let's give our speakers a big round of applause.
[03:44] Thank you. I'd also offer a big thanks to our our sponsors. This event really wouldn't be possible without them and they play critical roles in delivering technology to our customers. So, we really appreciate their support. Be sure to visit them in the expo hall next door
[04:00] for help or discussion around accelerating your data and AI projects. And with that said, let me pass it over to my good friend and chairman of Data Bricks Federal, Rory Patterson. RORY, I FEEL GOOD.
[04:22] All right, welcome. I feel good. I don't know who picked the music, but this is actually a song that I love. So, it's very I don't know why they would know that, but um well, welcome. Thanks for thanks for taking some time and spending it with us this morning. I'm sorry it was so late. You may not know this, but we actually tripled the number of people we thought we could fit in
[04:37] this event. So, like there's so much interest and you know, I can't say say thank you enough for all of you that showed up and all the people that are in the overflow room. And then I think there's also people still waiting outside trying to get in or that we had to turn away. So, I really appreciate everyone that showed up today. Uh, you guys may not know this, but uh, I was
[04:53] actually sitting where you are nine years ago as a government employee, uh, Blue Badger up at Fort me and, uh, wanting to bring commercial technology into the government. And I was struggling to find organizations that I could partner with that would actually allow me to bring their tech in and that we would actually be feel good about
[05:10] using it. Um, and so it's great to be on the other side of that and and being able to bring this to you. I did want to talk a little bit about before we get into this like personal use of AI. How many people in here like on a daily basis are already using LMS?
[05:25] Perfect. Just about everybody. I mean that's like point number one is like everyone has to start using the technology in your personal life. How many people are using it to like write emails? How many people are using it for uh like instead of like Google deep diving, you're now like LLM deep diving. I'm
[05:42] using it for like medical like questions I have every time I see something on my face and you know it's like I want to find a new doctor. Tell me which doctor is the best one. I started using it recently for interviews. Anyone using it for interviews where uh you're like hey I'm about to interview this person.
[05:57] Here's the job description. Give me five questions I should ask them. Right. Um hopefully there's only like three members of my team here. last quarter I wrote uh all of their performance evaluations uh in the product and and actually like it was the best work I've
[06:13] ever done uh to be totally honest. U but what was really interesting is it was like a clarity of what I was already thinking with all the documentation I already had and it simplified a lot for it and I think it made me more effective. Um so I used LLM's today uh
[06:29] yesterday actually when I was making this to and I asked uh Gemini two questions and I said uh describe to me the challenges in a graphic that uh the government has with bringing commercial technology uh in into their environment. And the second question was hey tell me
[06:46] the difference between uh government use of technology and private sector use of technology. And in two seconds, Gemini produced these two graphics. And they're like really clear. And I actually think they're pretty consistent with what like most people think are the challenges,
[07:01] which is like if you go on the right first, there's like a 20-year divide. I don't know. I haven't been in the government in nine years, but when I was there, we were using like nine years ago, we were using like Office XP, so it definitely felt like 20 years behind. Um, and then what are the challenges? I
[07:18] think our job is to break those things down. Like Gemini is just a reflection of what we all think. It's just using what available information is on the internet that it's been fed and trained off of to tell us this is what I think of your problem. And I think it's pretty accurate. I'm hoping that we can break
[07:35] both of these paradigms. Uh now I'm hoping that like the next time I ask this question of Gemini a year from now, it is going to give me a different answer. So here we are. Data intelligence in the enterprise is the the pitch I was supposed to give you guys. Uh, I don't know. Has anyone been
[07:52] to one of these before? Like to the this one in DC? Only one person. Yeah. So, I'll tell you a quick story. Two years ago, I gave this pitch uh here uh in DC and uh they had the same deck. They're like, "Hey, we had the big tour. At that
[08:09] time, it was like 14 sites or 13 sites." And they gave me the deck and they're like, "Just give this pitch." And I looked at and I said, "I can't talk about any of these things. I can't talk about a single thing in this deck because none of those things were available available to the government. Not a single thing was available. So I
[08:24] was able to talk about what we refer to as jobs in this company or like Spark ETL pipelines and I was able to talk about machine learning and then everything else would have pissed everyone off because they couldn't actually use any of it. So this year I'm happy to say that all I had to do was cross out enterprise uh
[08:41] and put in government and everything in here should be available to you guys. Now, I want to break the paradigm that the government doesn't get access to commercial technology in a timely manner. You guys have access to it today. And it's my goal to make sure you continue to have access to it. That's at
[08:56] Fed Ramp, GovCloud, IL5. Like, it's my mission in life to make sure you guys get it. So, um, all right. So, even with all that, our mission at Data Bricks is to democratize data and AI. What do we mean by that? We
[09:13] mean everyone should be able to have access to your data. It's yours. It belongs to you as an organization, as a government. It belongs to you. It doesn't belong to tech companies. It doesn't belong to your partners. It's more valuable than ever, but it's only valuable if you have access to it and if
[09:28] the people in your organization know how to take advantage of it. We want to democratize AI, which means it shouldn't just be a pilot program that sits in one part of your organization or takes forever to get out. We think that everyone should be able to play with AI. the way you're playing with it in your home life, the way you're rewriting
[09:45] emails, you should also be able to build agentic systems. Everyone should be to be able to build aic systems. Not a small few, not the super technical, everybody. And that's the only way we're really going to revolutionize what's happening right now. But as you guys know, there's a lot of like roadblocks
[10:02] to making that happen. Oh, skipped my own slides here. So, what's happening is we're seeing this happen in every industry already. Mike talked about 20,000 customers across every industry across the globe is already using data bricks to unblock the value of their a
[10:18] their data with AI and public sector is no uh is is not lost from that. Department of defense, state department, treasury, we'll be helping you to process your tax returns this year. You're welcome. Uh I hope everyone does well. um veteran affairs, postal
[10:35] service, health and human services. We're we're 80% we're in 80% of federal government uh departments and 400 uh government customers in the United States. So, uh hopefully everyone's getting to take advantage of it, but
[10:50] it's still challenging and the challenges still exist, right? Which is like your data is still locked up in these estates. It's the way you bought the data. It's the way you govern the data. It's the way that the proc procurement process works. It's the way that the funding process works. It's the
[11:06] way that we wrote RFPs 10, 15, five years ago that forced you to buy a silo and build a silo. It's the fact that companies locked you into proprietary data formats. It's the fact that sometimes you don't own your own data. I
[11:22] think someone was telling me that the government doesn't own the data coming off of the F-35. That's incredible. I think our tax dollars went to that. Um, it's incredible to think that like the government is paying for something and doesn't or can't access their own data.
[11:38] Um, so we want to remove those barriers, but you're also being locked in because if you've built a silo, you have siloed security policies, you have siloed governance, or you've built something and you've bolted on governance. Um, and finally, the new lockin, which is siloed
[11:56] AI or siloed automation. AI that's only applicable to whatever application it has access to. So, how many people are only going to go to work today and log into a single application? So, your AI is only as powerful as
[12:12] whatever application you're using at that individual moment. You need to have a platform that's able to access all your data in all of its in in every uh application, every system, every data set in order to truly operate like a human operates.
[12:29] So we took a little bit of a different approach. You may or may have not have heard of a data lakehouse. And you know in 2019 the the founders of the company uh put together this architecture of a
[12:44] data lake house and the architecture was simple open data storage formats. The data belongs to you. We started off with an open- source project called Delta Lake. We've merged it with another one called Apache Iceberg and now we have something called Delta Uniform. Making open data
[13:02] open means you can read and you can write. It's not enough if another vendor tells you that they can read open formats. Of course they can. We've made it easy for them to do that. But do they write back to proprietary formats? Does that data belong to you? Do you own it
[13:17] in your own S3 or ADLS bucket? You should be asking yourself these questions. If you don't, it's not a lakehouse. Second, unified governance, which means we've removed all the individual application lockins that you have and silos that you have into
[13:33] individual governance that forces you to do more administrative tasks just for others to share and unlock access to your data. But it's not just enough. Unified governance isn't enough just for access control and auditing, which is generally what's available to you. You
[13:50] should expect unified governance to be able to access more than just your tables. Your files, your PDFs, your models should all have the same governance framework. So that if you have access to that data and I want your model to have access to just the data
[14:05] that you have access to, you should be able to govern that. But there's other capabilities that you need and they and they're not available to bolt-ons. And those capabilities include auditing, cost control, lineage, business semantics. You should probably replace business semantics with ontology. We use
[14:22] them kind of like synonymously uh at data bricks. Um and you should be able to share your data outside and we've built connectors to allow you to easily and freely share that data. It doesn't require someone else to have data bricks in order for us to share data with them
[14:39] across platforms. All right. So you should expect all these things from your platform. It should be easy to use. And to help us understand how this is going to work, I'm going to ask my friend Jonathan to join us on stage and he's going to give a quick demo of how unified governance
[14:54] helps you unblock your enterprise. Jonathan, thanks Rory. Now, let me show you all what unified governance looks like in practice through the lens of fraud detection at a government agency.
[15:10] Meet the Services Bureau. It's a fictional agency that processes cases for businesses and individuals. They pro they send data in through our services bureau lands directly into our lakehouse where
[15:27] machine learning algorithms and business rules flag potential fraud. Then our fraud analysts will go through and mark cases to investigate or to process. From there, they will go through and
[15:45] be sent to financial crimes or any other industries that need to happen. So, let's take a look at what that looks like within Unity Catalog. So, within Unity Catalog, we're going to drop into our operations dashboard.
[16:01] Sorry, our operations schema. Within our operations schema, we are going to look at our gold fraud investigations table. Within our gold fraud investigations table, we are going to see that we have several columns here
[16:19] that are tagged with govern tags that label this as PII. And what this means is that these tags can be leveraged to create attribute-based access control policies. These policies allow us to
[16:35] prevent our junior analysts or those with less trust from seeing sensitive data. So if I put on my junior analyst hat and we jump into our plat into our table here to see the data, you will see that we have these columns masked from
[16:51] our junior analysts. These junior analysts are unable to see these sensitive data. However, they're still able to perform their work. And when they need to elevate this to a senior analyst to make decisions, our senior analysts, they're governed by different
[17:06] policies. So, they're able to see these data. Now, while we're here, we should take a look at the lineage for this table. So, I jumped over to the graph over here. And within this graph, you can see the full lineage of this table from left to right. Everything is auditable and
[17:23] traceable. Our compliance teams love this feature and our audit teams do as well. We can even see the column lineage here all the way from the left to the right and see our outputs. And we can even see on the far right side we have
[17:38] outputs here a new materialized view table. This table has been selected by me as a senior analyst now to be sent for investigation. I have selected a few columns as well as a subset of data such that I can go in and just send this to
[17:56] our investigation team for them to take action on. So let's take a look at this materialized view. I'm going to head back to Unity catalog and right here we see our high-risk immediate cases. This is what we need to send to our folks at
[18:12] the financial crimes division. Now we see in this table we can quickly share this data using delta sharing another open protocol to securely share data with the financial crimes division. Now
[18:27] this financial crimes division we can securely share this with them with no copy of data. We're not exporting data and sending it to them through email or anything like that. All I have to do is add to their share and you'll see I've created this share and they have access
[18:43] to these data in a secure open way. Now when we go and send this to them they can use the credentials that I've provided to them in the share to access via data bricks or any other tools that they're familiar with such as PowerBI or Python or even Excel. So, this is truly
[19:02] the power of having your data all in the lakehouse governed with Unity Catalog. No more emailing Excel files, no more audit scrambles, just governed accountable data access with Unity Catalog and Delta Sharing. Thanks. Now, back to you, Rory.
[19:18] Thanks, Jonathan. It's it's incredible. he can share that file and then he can decide when he wants to stop sharing it and he can just revoke access and if the the file changes he can just reshare it so that they have the most updated copy and so
[19:34] you stop losing copying copying and copying your data across your enterprise. So what's the next layer? Composable agents. agents that can see across your whole enterprise and that have the same access that you have because otherwise
[19:50] you're just building these agents in silos. Maybe in the IT department, maybe in the CIO's office, maybe in a co's office, maybe in an AI officer's office, but really it's everyone that's going to be able to take advantage of this revolution. And the faster you give these tools in the hands of individuals,
[20:06] the faster I think that this uh innovation is going to hit the market. But uh and then on top of that, that's what makes the Ada intelligence platform. But I do want to spend a little bit more time talking about what we call agent bricks as a company. I think everyone is
[20:23] familiar with this uh MIT study that came out that said like 95% of AI projects fail. I I don't disagree with that study, but I do think it was highly skewed during a time where everyone was uh super bullish about AI, but not
[20:39] necessarily super focused. Um, you know, you have these tools, they can pass the MCAT, they can do math Olympian level math, but I really have this problem of like I want to parse this PDF. Are they really good at that? I really have this
[20:54] problem that I want to understand uh the likelihood of an outage inside of my organization. Are they really good at that? I really want to condense all my policies and be able to answer questions about uh you know uh violations of those policies. Are they really good at that?
[21:10] So I think we would kind of went off the deep end and assume that just because an LLM can answer math olympian level questions that it could also help us in every task inside of our organization. But they weren't built to do that. Um, and so they're hard. So I think people had a lot of questions about like, hey,
[21:26] why is it that my pro projects are failing and what can I do to help? These are some of the biggest ones that came up for us. It was the quality of the output. How do you help me manage the quality of the output? The second was, hey, what technique should I use and which tools do I need to use? Do I need
[21:41] to use ML flow? Should I use a vector store? Should I use uh, you know, should I have governance inside? What data should they have access to? And those technologies, they're coming out so fast. Yeah, I'm able I may be able to pilot something today, but then when I
[21:57] go through the security review process to actually get that into production, it doesn't meet any of my standards. So, it was a great pilot, but it was a waste of money. And finally, how do I think about cost versus quality? How do I make those trade-offs? Do I want to use the most recent model at what cost? Um, or yeah,
[22:16] it was a great pilot that we made. However, that pilot, it was totally worth it at a dollar an outcome, but I'm going to do this thing a million or 10 million times. Is it still worth the cost? Um, so how do I think about those things? So, what did we do as a company? We we abstracted it away. By the way,
[22:31] our first rev of this was to give you guys the Lego pieces and let you build. And we gave you guys, you know, MLflow and we gave you a vector search and we gave you model serving and we gave you um a bunch of tools that you could use and you could push them together and it
[22:46] was not making things into production. So, we took a step back and we said what are the what are the most common use cases that our customers are trying to solve and let's just make agents for that. And so, we have an agent factory and you you guys can contribute to it. If you have a great workflow that you think other people are going to use or a
[23:03] great tool that other people are going to use, we will help you build the workflow and we'll put it in part of the factory and you guys have access to all these today. This means that every one of your employees has access to this today. So it's not like it's caught up in some department and they only have access because of Unity catalog and
[23:19] Delta to the data that they have access to. So their agent will work exactly as it should within the data that they that they can leverage. So this is when we talk about democratizing AI, this is kind of like the direction we're going in and this is available to in the console today. Everyone that has agent bricks, but not every agent have we
[23:38] thought about not every workflow has been defined by data bricks. There's probably a million out there and we've found 16 of them or 24 of them. So what do we do in that case? We've given you the infrastructure and abstracted away all the complexity. Don't worry about what tools you need to use to create
[23:54] these yourselves. And because of our framework that allows you to create LLM judges, optimize your infrastructure, and make trade-offs versus quality, we have a system that will continually allow that continually allows you to build those widgets that you saw on the
[24:09] last slide for yourself. And these are simple and easy. Um, the infrastructure exists. We've abstracted away all the technology. And uh, this is a great example of uh, Astroenica who's already using it. They took the agent bricks infrastructure. We didn't have a widget for them. They
[24:26] wanted to parse through a bunch of documents for um uh drug discovery trials and they were trying to figure out okay which LLM should I use? Well, data bricks allowed them to both optimize the the three different LLMs that they were using at the time. I think they were looking at chat GPT 5
[24:42] mini and claude. Uh and maybe they were looking at llama at the same time too. and they decided, we don't tell you which one they decided to go with, but we were able to figure out for their specific use case which model was actually the highest quality and the best performance. Uh, so they actually
[24:57] used a lowerc cost model and got the same quality as the other ones because we allow you to do those judging in your infrastructure against them simultaneously. Um, so it's a great use case for Astrogenica. I think that almost anyone could be doing building things like this. Uh, but to show you how easy it is, maybe Jonathan can come
[25:12] back on stage and show us how we can build an LL or a Gentic system real quick. All right. So, we just saw Rory show us that it's possible to build quality agents within data bricks. So, let's see how our fraud team is able to do that. Here we can see that our fraud
[25:30] team is leveraging a multi- aent supervisor. And what that means is it's able to leverage other agents such as a genie space for SQL queries using natural language, a knowledge assistant for querying policies and procedures as
[25:46] well as an external MCP server. I'm going to grab this query here while we jump into the UI and see what this looks like. We're going to use this later. All right. So here we see our agents panel within datab bricks. You can see at the top we have all of the agents that Rory
[26:03] has mentioned as well as the ones that our services bureau are leveraging in our fraud team. And you can see down here below these are the agents that were created our multi- aent supervisor as well as our knowledge assistant. And then Genie lives here off to the side as
[26:18] well. So let's jump into our supervisor agent and see how that is configured. You can see all through clicking through the UI, all that's needed for our supervisor is to provide some context as to what it's supposed to do. What type of questions can it answer as well as what agents does it have to leverage? We
[26:36] can see it's going to leverage our genie space. It's going to leverage our knowledge assistant as well as an external MCP server to find emerging fraud trends that are happening out in the world. So, let's bring back that question that I brought in and this is from the perspective of an executive.
[26:52] I'm going to wear that hat right now. So, our executive I I want to know from last week until now, what has happened in my fraud program? And what are we going to do about this in the coming weeks? What should my priorities be? So,
[27:07] we're going to see that now our supervisor agent is going to query our genie space. It's going to ask a natural language question and get back this result of tables or this result in a table, excuse me. And then we see our knowledge assistant is going to bring
[27:22] back policies from our documents that are within our lakehouse. And we get all of our guidelines and procedures here. A lot of good information. And you can see as I keep scrolling, there's even some footnotes. We can click on these and see PDFs if we need to reference those. And
[27:38] then here is our external MCP server. We're reaching out to the web to do a search of emerging fraud trends. We want to know what's happening out in the world. That way I can plan for my team to work on new models to detect these new fraud schemes.
[27:54] After all of that, after quering all these agents, our supervisor is going to synthesize that and bring back a briefing for me to make decisions. And you can see from the past week, we are in a critical situation. We have some
[28:10] overdue cases and we need to do something about it. So, not only did we get this information from this past week, but our supervisor is going to tell us about what we should prioritize and what's coming in the future. Something that would have taken days before from my team. Now, if I want
[28:27] to improve the quality of this agent, I can quickly jump into our labeling session. And what this means is that I can look at past questions or questions that I want to improve and provide guidelines to this agent to improve the
[28:43] quality of the answers in the future. So perhaps in this case I think this is a bit too long. So I want to reduce this to about three sentences. And now that I've added these guidelines and I can share these guidelines with my experts, these quality this quality enhancement to our agent
[28:58] will come through into the future in future queries. Now, this is AI in production. This isn't merely just using an LLM and throwing our data in there. This is bringing insights directly into your lakehouse to leverage with your data. All right, let's get back to the
[29:15] keynote. Thanks, Jonathan. Isn't that cool? This is this is everyone. This is the unblocking AI for everyone in the enterprise. Whether you're using a template or you're building it from scratch, everyone should have access to it. And it's easy to use because you
[29:31] already have the lakehouse. We've already done all the governance. We've already done all the security. We've already done all the data management. And now you can just start rapidly evolving and piloting, but everyone can do it. All right. Now, I'd love to say that AI is everything and that everything you're going to do tomorrow is an agentic workflow. Uh but you're
[29:48] still going to go to work and you're still going to have uh dashboards that people ask for and metrics views. And so I want to talk a little bit about our AIBI product. I think one of the reasons we're seeing such amazing growth in the last year, we've seen 500% growth in
[30:03] this product. By the way, it's free. It's available to everyone. So, you can just use it on top of data bricks. Um, why do we see so much use happening in this product? Because you don't need to code in Python or SQL anymore in order for you to use our product. So, this is just natural language. Ask questions in
[30:19] English. It'll make you a graph. if it's your data you have access to you it'll build you whatever graph whatever trend line whatever scatter plot you want uh similar to Tableau or PowerBI but all that definition all those business semantics remember they're in Unity
[30:34] catalog right so you're not locking in the definitions that run your business into your BI tool you're just using your BI tool to query your data um so if you update one BI layer and you have like how many people have the I don't know how many uh uh PowerBI and Tableau users
[30:52] we have here. But how many times you have conflicting metrics across your business across different dashboards? It's because you've locked those definitions into those dashboards. Build that into your ontology into your business semantics into uh Unity catalog. And then everyone leverages the
[31:07] same definitions. This is how I calculate revenue. This is how I define business units. This is how I think about geographies. This is how I think about headcount. Is it people that are hired? Is it job offers accepted? Is it butts and seats? Is it people with terms? Like those definitions range
[31:23] across all of our enterprises, but they should be in one location and then your BI tool should be able to access that so everyone's getting the same definition. So you can do that with uh with our AIBI tool. The AI part of it is hey, we're using AI to make it simple. We're text
[31:38] to SQL, writing statements, making your graphs, but more importantly, we're allowing you to do the thing that everyone asks for. I don't know if you got this. I definitely got this a lot in the government which is like this is great but I need a double click on this and then you're like great but usually
[31:54] that double click requires someone to go back and doubleclick and know the questions you wanted to ask and then go get that information reorganize it print it out put it on another PowerPoint slide and then bring it back to you and so the double click ended up taking
[32:10] another two or three hours or how many more days or how many more meetings by the time it got back on your schedule. We want to allow you to have that double click immediately. So, we can't anticipate every question that you're going to be asked. So, we just want to give you the tools in English to be able to ask that next question right there in
[32:26] the meeting. Bring the bring the technology with you. As soon as you have another question, a follow-up, you should be able to ask that question immediately in real time. And we just ask it and we change it right there. We want to give that power to everybody in the enterprise because sometimes the smartest person doesn't raise their hand in a meeting but has an a great
[32:43] question. And sometimes you have a stupid question that you don't want to ask in the meeting. And so we're going to give you the tool for that, too. Jonathan, why don't you come out and demonstrate how they can ask smart and stupid questions using Genie. All right, let's take a look at what our fraud team is going to do with AI, BI,
[33:00] and Genie. All right, so our team, both our executives and our analysts, they're going to have these questions. They're going to need insights into the program at a broad level. So they're going to leverage Genie for those natural language queries as well as dashboards
[33:16] for that top level overview. And of course our executives, that's the first place they're going to go. So let's actually go look into a dashboard built for an executive. All right. So here we are. This is our services bureau executive dashboard. And you can see we have a variety of
[33:33] visualizations. We have some KPIs at the top that are important for us to get a broad overview of how our program is doing, how many cases are coming in and what's our workload. So the first thing we see because this is important to us is the workload distribution. We can see all of our examiners here and even zoom
[33:49] in. So I there's something fishy going on here. So let's zoom into this part and we can see that Jennifer has a huge case load. So, I'm actually going to double click here and I'm going to select all these pieces of this chart for Jennifer. And as I do that, there's
[34:05] some cross filtering going on. So, all the other graphics are going to change and we see our KPIs updating. And now that we did that, we see there's a problem. Jennifer is overloaded. Has 445 overdue cases. That's a problem. We need
[34:20] to get Jennifer some help immediately. Looking at our chart here, we can see that there are some folks who have a lighter load like Jonathan here as well as Jaden, but I don't know if they can actually help Jennifer. I don't know the level of these people. I don't know if they're
[34:37] able to handle these complex cases. So, I'm going to ask Jeie a question here. I want to see the workload of all of my analysts as well as their level. I want to see can they actually handle this stuff and who can jump in. So, I'm gonna ask this question to Jeie.
[34:59] All right. And here we go. And Jeanie is gonna go and create a query using my English question and write some SQL. And we could even look at that. And we can see here, this is what I would do as well. I would go and query this gold table to get back these results. And Genie has brought back a table for us to
[35:16] look at the raw data as well as this chart. And this is where we get the insights where we can take actions. We see that our senior analysts are overloaded and we have this long tale of trainees waiting in the back end. So that tells me as an executive I need to accelerate our training program to get
[35:33] these folks in to help Jennifer out and help our other senior analysts. So this is the power of AIB genie. It allows us to take our insights to action both for our analysts as well as our executives just by asking questions in plain
[35:49] English. All right, over to you Rory. Thanks, Jonathan. This is really it. The the ability to get to the point where you can actually make a high quality decision generally doesn't come in the first review of
[36:04] data. Generally those dashboards were built for some other purpose not to answer the question of the day. They were me meant to give a high level view of the business. But how often do we not have the double click or the triple click or the you know absurd question
[36:19] that we have to ask that actually is the thing that drives a high quality decision in our business. Giving everyone access to this I think is incredible. So uh I know we've talked about a lot. I still have one more product that I want to uh go over real quick. It's data bicks apps. Uh I think
[36:35] that the future of access to data inside your enterprise will come through apps. I love data bricks. It's the most robust data platform out there. Um but it's everything. It's ETL and data pipelines and governance and a data warehouse and
[36:52] an OLAP engine and a streaming capability and you know you name it. All those capabilities are there. And when you log into data bricks you have access to it. But I see a future where most people actually are getting the power of data bricks but with the benefit of uh
[37:08] simplification of apps. So we're going to democratize agents uh through applications. But the challenge with doing that on another platform is it's hard to bring these applications into your environment. You have to go through just to bring in the the new the latest
[37:23] app that you want to finish that last mile between all the data that you have and the use case that you're trying to trying to solve. Bringing that application into your environment is like starting over. You have to go through all the security, the atto process, the governance, the data control. But data bicks already has all
[37:40] that in our platform. And because of the way that we've organized it, it makes it super easy to take advantage of the data intelligence infrastructure to build the next app. You already have the governance, you already have the security, you already have the Fed ramp, GovCloud, IL5 compliance, and now you
[37:55] can just use the open ecosystem of apps to finish the last mile. Um, this is an incredible innovation. I think it's actually our fastest growing product as a company, but I think it's the most applicable in the public sector. Why? Because that last mile is the thing that's defined your requirements for
[38:12] years. It's the last mile that makes you unique compared to the commercial world. And if we can help you finish that last mile but in a an accelerated way that because you already have access to all of your data, I think the innovation and the impact is just going to rapidly
[38:27] evolve here. So I know we have been doing a lot of demos. Uh we have a bunch of use cases out there for apps just like we have an ecosystem of aentic workflows. The apps use cases are out there. I think there's actually thousands of them now um that have being built. I'd love to see thousands of them
[38:43] coming out of the public sector. But I want to have uh Jonathan show us how to abstract away all the amazing things of data bricks and just show us data bricks through the eyes of an app. All right. So let's bring everything together with data bicks apps.
[38:58] Our services bureau has built the fraud platform that you see on the screen here. This is a data bicks app hosted within data bricks built using open source tools from GSA like the USWDS. And this is for our analysts. So let's
[39:14] dive right in. Here I am logged in as a senior analyst and I can see all of my cases and action on them one by one as I need. And when I dive into a case, you can see I have all of the information in one location for
[39:30] me to make a decision on whether or not this is fraud and I need to send it off or I can approve this case. You can see I even have all of my machine learning flags as well as supporting documents that are hosted in our lakehouse. We can even see I'm use using delta sharing to
[39:46] get thirdparty verification data. And of course, we can't forget we have our agent that we built from before. Our multi- aent supervisor is running in the background evaluating cases as they come in and making suggestions to me to say, "Hey, I think this is fraud or I think
[40:02] this is okay and this is what you should do next." Now, for demonstration purposes, I am gonna deny this and send it off to investigation. So, I'm going to click this and then I need to give a reason. So, I think this is perhaps fishy or there's some type of fraud
[40:18] going on. So, I'm going to type this out and then I click this button here. And now I go to this sharing widget that I've built on top of Delta sharing. So here what I can do is I can select from the departments like the financial crimes division or investigations and
[40:33] send the data to them and select a subset just like I did before with a materialized view that I created. And then all I have to do is click one button and because I'm a senior analyst I can do this. I have the permissions and I've shared this data securely with
[40:49] this department and I have all my shares here. Now what about our executive? Our executive can come in here and see the same dashboard that we saw within the data bricks UI without having to log into data bricks at all. They can just come here and interact with the dashboard just as they did before. We
[41:05] can zoom in and do the same investigation. And we even have Genie here at our fingertips to answer questions when we have more. This is the power of bringing all of your data within the lakehouse governed with Unity catalog and getting insights with
[41:21] natural language and genie and building this all into one data bricks app. All right, thank you all. So hopefully you see where we're going as a company. We're trying to make sure
[41:37] that you get to democratize your data and AI. uh and we're making it as simple as possible for you to leverage the all of take advantage of all the data inside your enterprise but leverage it in the ways that you want to le leverage it whether it's through AI workflows AIBI
[41:54] or apps we think that bringing this technology to you guys is like our mission in life and uh we're super thrilled to be able to uh talk about it today but the thing that brings most things to life is having customers talk about it so I'd love to have Patrick come to the stage from CMS he's is the
[42:09] CIO. He's been building things on top of data bricks and I'd love for him to share it with you. Fall out. Honestly, I want to see you.
[42:33] Good morning, world tour. I look I am so thrilled um to be in what I believe is the most exciting place to be right now in the public sector today and uh the data and AI world tour. So just give your hands round of applause.
[42:49] Um and I also have a the honor uh to represent what I think is the hottest federal agency in all federal government, CMS.
[43:07] Look, today uh I want to share about how CMS is building an AI native future. We're going to democrize, democratize AI and that with the goal of delivering better care,
[43:23] strengthen the trust and safeguarding our resources, our taxpayer dollars that we all um chip in on. Um but CMS is not just modernizing it. Uh we are impacting lives each and every
[43:41] day. And I'm excited to be here today to talk about where CMS is heading with our IT initiatives and how you all are essential partners in our journey. Uh we have a very special relationship with
[43:56] data bricks that helps us harness the power of our data in ways that matter including helping our beneficiaries and crushing fraud. So let me ask you to picture this.
[44:14] You're in a small rural clinic. A waiting room is full. A doctor is doing the very best with limited resources. A patient badly needs a referral to a specialist
[44:34] that doctor reaches for the provider directory. It's a thick binder that hasn't been updated in months, years, or even decades. They flip through the pages, find a number, they dial it. But here's the
[44:51] problem. The number doesn't belong to a real provider. It's fraudulent. Someone inserted a fake number, fake listing along the way, hoping to bill
[45:06] Medicare for services never provided. The doctor doesn't know that yet. the the patient is waiting and care is delayed all while frosters
[45:22] are waiting for a payday. Now imagine this. Imagine that same clinic but the tools we're building in CMS. Instead of old dusty binder,
[45:38] uh the doctor pulls up a unified digital directory. Every provider is instantly accurate, instantly verified and secure. Fraudulent entries are flagged
[45:54] before they even get into our system. The right specialist identified in seconds. The patient gets the appointment for that day. care is faster, trust between the patient and provider is stronger and CMS grows stronger all together.
[46:12] That difference between that binder, that's the old binder and that shiny digital directory. It may feel small, but in reality,
[46:28] it's the difference between delay and timely care. the difference between fraud slipping through the cracks and fraud being stopped at the door. Between a system that frustrates and a system that protects.
[46:48] This is what we mean about digital transformation at CMS. It's not just about it projects. is real people, real clinics in real moments, moments that matter with people who feel the impact of whether we modernize or not.
[47:06] It's also about protecting our taxpayer dollars from those and we have a lot of them who want to exploit it. This is the kind of transformation CMS is driving for each and every day. So, we are not one of those 20-year
[47:21] government agencies. Now, this was only one clinic in my example, one patient and one doctor, but we operate at a staggering scale. Check this out. CMS, we manage over $1.7
[47:39] trillion in health spending each and every year. larger than Apple and AWS if we have anybody from Apple and AWS in here. By revenue, we serve over 160 million Americans. That is one out of every
[47:54] American in our country. We rely on a workforce of just 5,000 and partner with tens and thousands of contractors. And each and every week
[48:09] we defend against billions and billions cyber attacks against our systems from adversaries, bad doctor actors who would love to get our hands and their hands on our sensitive data. CMS, we not just a pair of claims.
[48:28] We are a national engine of healthcare. We set insurance rules. We establish quality standards. We make payments to providers. Safeguard sensitive health information and protect taxpayers dollars.
[48:43] We balance the need for speed with responsibility to protect because every claim, every login, and every data transfer carries enormous weight.
[49:00] This makes us one of the most consequential agencies in all of federal government. And with this scale, standing still is just not an option for us. CMS faces a critical challenge. We
[49:15] must crush fraud. We must secure our data. And we must deliver faster all by relying and providing more reliable care. That is a daunting task. So to meet these demands, I want to talk to you about some ways that CMS is looking to transform.
[49:31] And that's exactly what we're aiming to do in 26 uh calendar year 26 all while being anchored in a AI native operating model. Our 2026 strategy focuses on nine key IT priorities and you see them
[49:46] listed here but these are not just technology projects that I mentioned earlier. They are foundation or transforming healthcare nationwide. Each priority accelerates operational excellence, protects public trust, and
[50:03] modernizes how we do our programs. We're taking and talking about advance advanced cyber defenses, smarter claims processing, unified provider directory, and yes, a
[50:20] workforce modernization to help do it all. Let me dive deeper in just a few of what I find as our gamechanging priorities. Here you will see four that's going to reinforce another to deliver faster, safer, smarter health
[50:35] care for over those 160 Americans I talked about earlier. Each reflects not just what we're building, but why it matters to the people we serve. First, we are embedding AI across CMS to prevent fraud and we want to detect
[50:52] anomalies and we want to support smarter cla decision-m CMS as I mentioned is a prime target. Billions of cyber attacks hit us each week. Cyber defense is not a project for us. It's a living discipline.
[51:07] Protect protecting CMS's sensitive data is my top priority. And this is why recently I issued a memo that requires all the contractors do business with us to complete vetting prior to accessing our systems. No work that requires
[51:24] access to our systems can begin until everyone is cleared. But that's exactly why we built um security into everything we do. Our approach includes zero trust architecture, advanced monitoring, automated responses that adapt to new
[51:41] threats. We expect those to kind of be built in the tools and platforms that we invest in. Now, let me touch on fraud prevention because that is very equally important in protecting the taxpayer dollars that I've been mentioning and preser preserving the trust in our system.
[51:57] Fraudulent claims waste resources. We all know that it delays care and erodess the confidence in Medicare and Medicaid. So how do we tackle this challenge? Our approach combines three key AR AI cyber
[52:13] security tools. One, AIdriven anomaly detection to flag suspicious activities and for our claims instantly. data matching algorithms that prevent duplicate building to uncover fraud,
[52:29] such as the $106 million in improper payments that CMS and VA found together earlier this year. And three, automatic triage to route cases to investigators quickly. We don't have time to wait once we those things are flagged.
[52:46] Every fraudulent dollar stopped becomes a dollar invested in care and trust earned in our citizens. That is why we have partnered with data bricks to identify improper payments
[53:02] through predictive models by comparing peer groups. This cyber defense is foundational and essential because it protects the integrity of our claims and our programs. Which brings me to a second priority.
[53:19] We are modernizing our claims processing platforms so providers and beneficiaries can get faster and more accurate service. CMS processes over a trillion dollars in claims each year as I mentioned earlier.
[53:35] Yet legacy platforms are limiting our ability to adapt and slow up service. Our modernization approach is module, resilient, and cloud enabled. So with an AI mindset throughout,
[53:50] we're applying AI in four key ways. One, smart routing that gets claims to the right system faster. predictive analytics that catches errors before payments even go out.
[54:08] Chat bots and virtual assistants that handle routine task. Freeing our staff for the more critical mission work. And this results in faster, more accurate payments to providers, which is very important to our providers.
[54:24] Better expenses to our beneficiaries. We want to bring the the best experiences possible and the stronger stewardship of of the taxpayer funds that we oversee. But to accomplish all this, we must migrate off mainframe and
[54:39] modernize. And to achieve this transformation, we need the right infrastructure and the right talent foundation in place. Let me talk a little bit about that. So, we're building cloud platforms, shared services
[54:55] across the silos of our organizations, and a skilled in-house workforce in AI, cloud, cyber security to sustain these capabilities. None of these transformations succeed without the infrastructure and the
[55:11] people that power them. CMS is embracing a multicloud environment for resilience, scalability, and flexibility. We are building enterprise shared services and platforms. They're no longer going to have several different
[55:27] platforms that's doing the same thing in our ecosystem. We need to reduce the duplication, accelerate delivery, and improve security by doing that. Most important is our workforce.
[55:43] No matter what technology we bring in house, it doesn't alone transform government. Our people do. So we are investing in training and upskilling in AI engineering, cyber defense, cloud security and full stack
[56:00] development and data science. This ensures the teams can design, manage and sustain the platforms that we invest in. The outcomes that we hope to see is faster time to market to our tools. We
[56:17] can't wait more consistent services across CMS and a workforce that has prepared for these technology shifts that you've been seeing this morning. This infrastructure and the workforce foundation
[56:34] enables my fourth and last priority that I want to share with you today. I have none and that's making sure that data flows seamlessly where it's needed the most. This brings us full circle to that royal clinic
[56:50] that I described in the beginning of this discussion enabling secure data interoperability and creating a unified national provider directory so that providers have never have to rely on an outdated binder again.
[57:07] I believe that data interoperability is not a technical luxury for us. It's a healthc care necessity. Our patients, our providers, our payers all depend on accurate, secure, timely
[57:24] data. One of our flagship initiatives that I talked about earlier is the national provider directory. What we aim to do with this initiative is create a unified digital directory that eliminates errors
[57:40] and referrals, reduce fraudulent billing by validating provider identity, and improving patient experience so sharing care is not delayed due to inaccurate data. An AI native approach, one that we're
[57:58] taking, makes interoperability stronger in three key ways. One, it validates provider data in real time so the information stays accurate. Two, it maps different data sets into
[58:15] reliable source so everything connects. And three, it spots gaps and inconsistencies early before they are a big issue and affect our care. The results,
[58:32] fewer errors, faster referrals, and a stronger trust between patients and providers and the payer. Together, these priorities, cyber defense, modernizing claim, scalable cloud and interoperability
[58:50] form a foundation for an AI native CMS deliver faster, safer, and smarter healthcare. So, at CMS, uh this impact for transformation is real. Um the things I'm talking to you
[59:06] today we are doing now. First um the adoption of AI tools is strong at CMS. As of this morning when I checked my dashboard 83% of our workforce is using AI every
[59:23] day. That means our employees are saving on an average of six hours a week. That's time that I can spend on high work. Second, uh, fraud analytics is already paying dividends. Now, we flagged, as I
[59:40] mentioned earlier, already 106 million in improper payments between, uh, VA and CMS. That's money recouped. That's money red redirected back to into patient care.
[59:55] And third, our shift to the cloud is enabling resilience and innovation across EMS. It's not just about the technology, but it's about creating the flexibility foundation that lets us deliver services faster, scale when needed
[01:00:13] to keep our system secure. For a specific example, we use data bricks assistant to migrate legacy SAS code and convert
[01:00:28] that over. This allowed us to modernize our systems and deliver better healthcare. This example shows that we cannot do it alone. We need strong industry partners like you.
[01:00:44] So, as members of the data and AI community, I want to leave you with a few questions. One, are we building AI that solves yesterday's pain points or AI that reshapes how government delivers value for the next decade? I hope it's the
[01:01:00] latter. Let me know. Two, what would our agencies look like if every mission decision, every policy, and every operational workflow were powered by real time trustworthy data?
[01:01:17] Why isn't that our reality today? The third question I have for you, if the public trust us with their most sensitive data, what must we do to earn that trust in
[01:01:32] how we use AI? As we look ahead to these priorities will take all of us, government, industry, working in partnership. We need mission specific innovations that drive outcomes, not flashy promises of
[01:01:48] new products and tools. Together, we have the power to transform healthcare. No doubt about that. By uniting our talented people with cutting edge tools, we'll deliver exceptional care to that
[01:02:04] every citizen, including you and me and our loved ones uh expect and truly deserves. This is our mission. and together we will make it a priority. Thank you everyone.

Learn more about the Databricks Data and AI platform.

The information provided herein is for general informational purposes only and may not reflect the most current product capabilities or configurations.