Public Sector Data Modernization: From Legacy Systems to AI
Summary
- This video covers how public sector organizations including Amtrak, California health agencies, the FDA, and the U.S. Department of Defense are modernizing from fragmented legacy systems to unified data platforms on Databricks.
- Amtrak is building a unified data platform to support its first major fleet and infrastructure transformation in 50 years, integrating multi-source data across a 21,000-track-mile national network.
- Genie, Unity Catalog, and enterprise AI governance enable government-scale operations across use cases ranging from infectious disease tracking to deploying AI agents across 16,000 FDA staff.
Public Sector Data Modernization: From Legacy Systems to AI

Public sector organizations manage mission-critical operations across government agencies, from healthcare to defense to transportation. Yet most rely on fragmented legacy systems, manual processes, and disconnected data silos. Learn how leading agencies are modernizing with enterprise data platforms to enable real-time decision-making, AI-driven insights, and secure data governance.
Hear real-world stories from Amtrak building a unified data platform across 21,000 track miles, California's health agencies transforming infectious disease tracking and behavioral health programs, the FDA deploying AI agents to 16,000 staff, and the Department of Defense managing 18,000 Databricks users. Discover how Genie, Unity Catalog, and enterprise AI governance enable government-scale operations while maintaining security and compliance.
🤝
Chapters
00:00Public Sector Data Modernization Forum01:07Amtrak: Fleet and Infrastructure Transformation06:36Lakehouse Architecture: Multi-Source Data Integration10:44Experience Layer: Making Data Accessible with Apps12:20California Health: Transforming Disease Tracking Operations19:09Government AI: Genie Security and Governance Approval28:52Lessons Learned: Starting Your Data Modernization32:25Department of Defense: 18,000-User Enterprise Platform41:45Anduril: Real-Time Manufacturing Dashboards46:08Auto CDC Flows: Unlocking Real-Time Data at Scale50:26FDA: Deploying AI Agents Across Government53:58Government AI Trust: Building 85% Adoption01:00:19Enterprise AI Governance: MCP Servers and Security
FAQs
What data modernization challenges do public sector organizations face?
Public sector organizations rely on fragmented legacy systems, manual processes, and disconnected data silos that prevent real-time decision-making. This video shows how agencies across healthcare, defense, and transportation are consolidating onto enterprise data platforms to enable AI-driven insights and secure data governance.
How is Amtrak using Databricks as part of its digital transformation?
Amtrak is building a unified data platform to support one of the largest digital transformations in the organization's history, coinciding with its first major physical fleet transformation in 50 years. The platform integrates multi-source data across Amtrak's 21,000-track-mile network, which runs 97% on freight railroad tracks owned by host railroads.
How is the FDA transforming document review with AI on Databricks?
The FDA deployed Elsa, a generative AI platform, to approximately 16,000 staff across its centers, consolidating fragmented AI systems onto Databricks with MCP server connectors and Unity Catalog for secure, governed agent access to regulatory documents. Document searches that previously took weeks can now be completed in approximately 3 minutes.
How does the Department of Defense use Databricks at scale?
CDAO, the Chief Digital and AI Office, serves 18,000 Databricks users across 98 petabytes of data in unclassified, classified, and top-secret environments. The platform powers mission-critical use cases including supply chain management, audit, and operational planning for the entire Department of Defense.
Full transcript
[00:50] Nat. Hey. Hey.
[01:07] Please help me welcome Prativa from Antra to the stage. I don't know how many of you kind of uh know that Amtrak is about five decades old back from the guilded age private
[01:25] railroads. I learned this history when I joined Amtrak about two years ago is when the NRPC uh Amtrak was formed as a national can you hear me now?
[01:41] Okay, I'm gonna I'm not yelling, but I'm trying to reach the back of the room. Um, Amtrak was formed in 1971 after a bunch of private railroads and then the airline industry took off. We
[01:59] built the interstates and, you know, auto lobby whatnot um to create national passenger, the only passenger railroad in the US. We have one more uh okay one more passenger railroad in Florida. I'll
[02:15] not name them. Um but this is the the largest uh network. We have about 21,000 track miles. 97% of those track miles are not owned by Amtrak. It's all
[02:30] running on freight railroad tracks, what we call as host railroads. Um this is a significant moment for Amtrak and this is kind of what got me into Amtrak. Um uh the last in the last couple of years um Amtrak is going through a massive
[02:47] physical transformation. Everything that you saw in that video it's not just a concept. It's happening. It's real. Um there is uh not only investment in infrastructure, there is new fleet. Ailla. How many of you have ridden an
[03:02] Ailla? Um, we have new Asella fleet now running between Boston and DC. Uh, 28 train sets and for those of you here on the west coast, especially, you know, if you're in Portland and Seattle area, coming
[03:18] this fall is our aero train sets uh built by Seammens. We are acquiring 83 of those train sets. Um and then very recently there is an um a bit to kind of go uh get new long-distance fleet as
[03:35] well, right? Kind of long-distance travel. Think about New York to Chicago, Chicago to LA, those types of lines. So there is significant physical transformation that's happening at Amtrak in, you know, this is the first
[03:50] of its kind in the last 50 years. Um and not only are we doing physical transformation, we are in the midst of one of the biggest digital transformations in the organization. um which is how
[04:05] 15 years uh doing data and intelligence uh initiatives at financial services companies and uh when Amtrak kind of uh got my attention a couple of years ago uh kind of what I believe the the possibilities
[04:22] here um to be able to actually be part of this massive transformation. This is an important moment uh for passenger railroad here in the US. So the hardware right thisware the hardware
[04:38] they have a lot of uh fleet rolling stock about units 500 of them are um what we call as locomotives the you know the power cars and about 2,000 of them are um the
[04:56] um each one of them including our new train sets are data generating assets sets 100 plus sensors in each one of these trains and the modern train and the aerot train sets um have sensor mesh
[05:12] that send real telemetry to our back office. Uh our mechanical teams are not reacting to uh issues but they get predictive signals right monitoring. Um these train sets are equipped with you
[05:28] know things like uh refrigerator refrigerator temperature sensors right food safety checks so on so we real time capture and to process this data we needed a platform such as data bricks
[05:44] right and to be able to combine that with the rest of our to be able to kind of make it available for our uh mechanical teams safety teams um infrastructure organizations
[06:03] collecting data are trains. It's not just the trains. We are also the efforts going on an infrastructure perspective. Um than $50 billion in investment uh in capital projects. These this pipeline
[06:20] the $50 billion pipeline needs it's data intensive. We need the data to plan, build, maintain, and prioritize these efforts properly. These are longunning projects. These are not like three month, six month efforts in order to kind of keep pace with all of these
[06:36] investments. It's data intensive. We need to be able to kind of capture that and make it available to our capital planning organization. So, and we are doing this with Lakehouse. So, I'm going to quickly touch upon like what the architecture looks like. I'm not going
[06:52] to drain the slide a whole lot. It's it's a very simplified architecture uh for all the intents and purposes here. Um our LOS platform selection was a strategic one. Right? This is what is
[07:08] going to help us scale and grow as we going through this massive physical transformation. Right? So, we need the ability to bring data from all of our key sources, whether that's the fleet, the infrastructure assets, um whether that's kind of our corporate
[07:24] systems, reservation systems, um marketing engagement data, bring it all together. Lake Connect is part of our kind of the integration pipeline, simplifying the integration. We are taking advantage of all of the leak blue connect capabilities and we will I can't
[07:42] spot Rory here but his commitment to support this space we're going to keep taking advantage and pushing the product teams here at data bricks to kind of enable even more capabilities for us right what we have taken advantage of is still probably uh you know we scratching
[07:58] the surface here u but still there are some nuanced and unique challenges uh especially in our real time space where we'll continue to partner with data bricks to kind of take advantage of it. So we are leveraging the medallion architecture kind of in kind of uh
[08:15] enriching the data curating the data make it making it consumable and we are using unity to kind of govern the data um and making advanced analytics possible.
[08:34] So the a glimpse of the use cases or how we think about that as data and analytical products that we are building through the lakehouse platform. The first uh that we have launched is around the fleet intelligence which is around our uh SLA and the aero data sets. um also bringing in some of our legacy fleet health information and making it
[08:51] available uh for the various uh organizations. But we are continuing to build a significant pipeline of those data and analytical products within the platform. Whether that's for operational readiness to kind of plan and you know schedule the work around our fleet crew
[09:06] cons maintenance windows or whether we are bringing in and modernizing our reservation uh platform and bringing the booking and the reservation data and providing data products and analytics analytical products supporting our commercial organization or to help with
[09:21] our capital prioritization um processes. Um and I know I'm running a little little short on time here. So we have taken advantage of the platform to make it connect the data sources that we need to bring together. Now we have the
[09:38] governance framework around it. We are building the observability into the platform but we want to continue taking advantage of the predictive and the compounding capability the platform has to offer. Right? So including taking advantage of uh the Genie capabilities and all the announcements that came
[09:54] about in the last couple of days right but I want to uh this is kind of the holy grail of any data and analytics journey um it's hard to touch and feel data right in an organization the only
[10:11] way people have been used to consuming data is dashboards. How many of you still have dashboards? Reports. Dashboards. How many of you have gotten rid of reports and dashboards?
[10:28] Not a single one of them. I don't think so. We're going to get rid of them. Um, Excel macros. How many of them love it? Yeah, we all do, right? Um, so we are trying to create an experience layer that makes data and intelligence more
[10:44] consumable and something that you can touch and feel and make it less abstract for our consumer. This is part of our mission to make it self-service to democratize data and intelligence across Amtrak. And here is a little bit of a concept that we are kind of trying to
[10:59] embark on. We see there is no audio to this one. So this is uh a data bricks apps experience a concept which will pull all of the
[11:15] data and the analytical products into this experience layer helps you navigate and look at all the the metadata that's coming out of Unity as well as be able to see the lineage uh in here uh for the
[11:31] the data products that we have created and it is one-stop shop right whether you're a developer uh uh whether you're an analyst uh whether you're an executive you have kind of one place one experience layer to come in to be able
[11:46] to kind of shop for your data analyze discover and to be able to work with this. So there is um a piece of uh the uh the Genie engine that's embedded embedded into this experience uh layer as well and we're going to continue to
[12:03] kind of build this uh vision out for Amtrak and to kind of democratize data across Amtrak. With that, I will at this point
[12:20] introduce Ryan Mosley, Robert Snder, and Molly just bear for the panel discussion. Thank you. Hi everybody. I'm Molly Jpair. I lead our global go to market here at Data Bricks. We've got a great panel for you here today. We have two of California's most senior leaders who are pushing forward Oh, I'm here now. pushing
[12:36] forward some of the state's most important healthcare initiatives. Um, so before we dive into the technology and the how, could you guys ground us in your organizations and your roles and the Californians that depend on your data programs.
[12:56] Where we at? Afternoon. Afternoon everyone. My name is Ryan Mosley. I am the division chief for modernization at the Department of Healthcare Services. Uh I am also the project director for behavioral health transformation uh connected to the governor's proposition one that is facilitating uh
[13:12] substance use disorders and uh behavioral health um transformation. So Rob Snyder work at the California Department of Public Health and the Center for Infectious Diseases where our team manages a database that is responsible or houses about 80% of the
[13:29] departmental data assets. um to ground that in like the actual data that we have. When someone tests positive for an infectious disease or visits their provider and they have an infectious disease, there's information that is sent to us. So all of those data are housed in our data warehouse so that we can draw insight and develop public
[13:46] health interventions from them. Terrific. So that's very important. Um Robbie, I want to start with you. Um before we were chatting uh we talked fax machines and an 11 days kind of between uh an out an outbreak signal and that
[14:02] data getting to you. Can you talk to us about what it was like before data bricks? Um and please don't leave any details out. I feel like the fax machine stories everyone here can probably relate to. Uh yes. So I mean in the last 10 years we've come
[14:18] thousand years lifespan changes since co and um I was thinking about how to respond to this question and some of the people that I know that work here probably heard this story before but I feel like it's very tangible um and people can relate to it. So like I said
[14:35] we get reports of every person who is tested or has an infectious disease in the state of California. Some of you might remember what happened in 2020 with COVID. Some of you might not want to remember like me. Um, and what happened was that we had these different
[14:52] pipelines through which we received data. We had some electronic data receipt but we still received a large number of faxes and um depending on how much time we give I can give another example of how fragmented our data assets were before. I love a good embarrassing please. Yes.
[15:09] So, the first one for faxes, um, I would have to go in on the weekend to refill the paper ream because we were receiving so many fax reports. And then we would take those faxes and then we would put them in an Excel sheet and then we would
[15:24] ingest them into our databases. Um, so I did not really like going in on the weekend, but part of the reason that I had to go in on the weekend was because we had another data process that was set up. Um, without going too far into the weeds about some of our specific
[15:40] infectious diseases, we had this data process that literally took six days to run on a desktop and it would time out and have to go in on the weekend to restart the computer sometimes. Might as well refill the ream of paper on the fax machine while you're there to reset it. And um, since that time, we've taken that six day process and it now runs in
[15:57] about 30 minutes. So um, a huge change. So and it's about it was related to hepatitis C, which I think is super interesting. one of my pet interest uh conditions and we can do a lot more with that than we could before.
[16:12] All right, Ryan, over to you. Um, so the behavioral health transformation, a huge governor initiative. I imagine the whole state is watching. Lot of pressure. What were the data limitations that made it hardest to track whether the state's investments
[16:27] were making an impact? Pretty much in between the check and the results. Anything in the middle there. So um when I was asked to uh lead out behavioral transformation from a technical side um there was the idea of
[16:46] DHCS uh not as fun as infectious diseases but really handles a lot of money that goes out to counties and so forth. And so as we are, you know, sending out that money, the counties are putting together a plan and they're going to tell us
[17:03] about how they're going to spend that money. And I was like, great, let me see what that's all about. And someone directed me to a website that a PDF was published. And it was about a 50page document that had a three-year plan in
[17:18] every different format, every different style. And that was their plan. And I go, "Okay, well, how do we know that that's act?" Yeah, we don't really do that. So, like I said, everything in between from we started with a PDF and
[17:34] that was how the process uh went before. And one of the biggest things about uh the governor's kind of initiative here is that accountability. And so, obviously, we had to stand up the beginning factors of that, which is the
[17:49] intake of that data. And then we had to process that data and understand their plan, be able to segment it against different uh you know um policies, procedures uh and practices. And so as
[18:04] we move through that process, we are trying to come out the other side to be showing the accountability. And we just launched the county profile which is starting to actually visualize the use of those funds. Now we haven't got to that piece of it yet. we've got demographics live and things like that.
[18:21] So, we've started to break down and and that intake and then have the output and actually have the infrastructure in between. But yeah, as the uh as we started, we went from a basic PDF, not a fax machine. Uh so, he's got me there. But, uh we have our share of PDFs.
[18:38] Uh but yeah, that again, it was just so limited and it was really just kind of by feel how the department was doing that. And now that the accountability is needed, that infrastructure is where we're laying the tracks to drive towards that. Makes sense. All right. So, I feel like we've got a good understanding of kind
[18:54] of the before. Let's move into how everything looks now with data bricks and kind of moving forward. Um, Robbie, this is I'm like incredibly impressed by this. You have somehow managed to get Genie approved by your organization in
[19:09] months. Um, as someone who was in the federal government, I'm like in awe by that. Um I assumed that it was like cupcakes and like cakes and he's like no it was just it was a strategy Molly like everyone obviously you have to get governance and security involved in the beginning and I was like I want you to talk about this. So can you talk to the
[19:26] audience who I know has very similar challenges around atto and security and kind of talk them through how you were able to get a new feature approved so quickly and kind of what what Genie has enabled your teams to do. Yeah. Yeah. So I will say that it was
[19:41] not just me. Obviously there's a big team involved in part of this wonderful team. Some of my team is here but our our ITSD um group has been in charge of a lot of that approval. There is a rigorous approval process a number of forms. I think it's like a 5305F
[19:58] that's embedded in my brain. Um I didn't make that form by the way. Um and yeah, I think just conceptually the really important thing for us given that we have so much data and so much PI and
[20:16] so much sensitive data is ensuring that whatever happens just remains on site and building the safeguards to ensure that right like the privacy things that we have in place it's all native is only done there nothing feeds back into the LLMs um from a policy and bureaucracy
[20:32] perspective I think it's like be persistent in talking to people a lot and working together and really I think articulating the value of having these modernized data systems and um maybe just to really call back to what I was saying before co has really
[20:48] demonstrated the need for those kinds of things. So I think we have a lot of support um throughout the organization to help move towards and move off of these legacy systems that hindered us before in excess and give us the opportunities to use them. So I think that's the first part of your question in terms of like this is how we did it.
[21:05] It's probably not sufficiently detailed, but every organization probably has their own challenges and bureaucracies that they're going to need to go through. We were just talking before y'all don't have Genie yet set up and DHCS and your sister department. So I think it just varies a lot and we feel
[21:20] pretty lucky honestly to have people who are helping to take care of it for us. Um, the second part in terms of how we use it, I am optimistic, but I'm actually not sure exactly how we can use it yet to derive some of these advanced analytics.
[21:36] Some of the things that we've already used Genai for have been just simple product simple productivity tools. Um, a lot of the things that my team uses it for involve, hey, I wrote this thing in SAS 15 years ago. How do I get it into
[21:52] SQL so that I can use it in data bricks? I do this thing in R. I want to learn how to use this package. You write this like skeleton framework using the scripts that you have. There is a lot of interest and I am tentatively optimistic that we can use it to draw insights from our data. I was
[22:08] just talking to to some of my colleagues in the um in the audience before about some of the risks that are involved in drawing incorrect insights from our data at times. So I think there are opportunities to do a lot with it. Um, something that we have talked about
[22:25] that I am very interested in learning more about and I don't know if if Genie or any AI is going to be able to do it is how to parse unstructured health data. So like health notes that we get where you get these complicated doctor's visits. Um, my wife works in a tech space and we're
[22:41] always talking about it too and they're trying to work on it and we're trying to work on it. I don't know if I'm I don't think I'm going to be the one to figure it out in public health, frankly. But I think there's a way that we can use to kind of winnow down some of the work that we do and like optimize the time
[22:56] that we spend. So maybe we'll spend less time manually reviewing everything only reviewing some. The example I see we still got 12 minutes, so maybe I'll go a little bit down the rabbit hole for another one of my favorite pet projects around um syphilis.
[23:12] I feel like dinner. It's a spyroet. I mean, this is a dinner time with Robbie. It's a spyroet. It's super cool from that perspective. It's an interesting uh disease, but from a diagnosis and intervention perspective, it's very
[23:28] complex. For those of you that aren't familiar with it, there isn't just like a yes or no test that you can do to determine whether or not someone is infected. At the other end of things, it's very challenging to determine whether or not someone has been adequately treated. If you want to talk about it later, happy to talk about the
[23:44] details, but I'll spare you the gory details now and just mention that the treatment for it is something that is pretty non-specific, right? It's penicellin, but it has to be dosed adequate amounts over an adequate time based on how long someone has had the disease. And it's just very difficult to get that for a human reading those notes
[24:00] themselves. We often have to call the doctor's office to check. And so I would love if something could be done like that with a Gen AI tool and maybe it's going to be around certain populations and whatever, but um one of the most important populations that we think about with syphilis is children who get
[24:17] syphilis, right? So congenital syphilis um and the risk of saying that someone was adequately treated and being wrong about that is too great. So I think I want to be able to use it for things like that, but we're a ways away. But
[24:32] for now, like the productivity aspects are making it so that we can spend more time on doing those kinds of interventions to actually prevent these really severe outcomes like kids born with syphilis. That's terrific. Um Ryan, turning to you, can you talk about some of the meaningful changes that have happened in
[24:49] your ability to track and improve behavioral health? Um yeah, so we are right at the kind of precipice of behavioral health as I talked about. we were really facilitating the intake of data from
[25:04] counties and and again digitizing a form. It's not rocket science, but it's transformational in the government space and and now we're looking at better ways to intake data from the counties. And so as we and behavioral health is changing the way
[25:20] that money is being sent out and so we are really um kind of at the beginning of that measure and but again this is the governor's uh agenda to show those outcomes and show that that money has put more beds I mean the improvisation already has put uh I think it's 60
[25:39] billion six I'm going to misquote it I should know it but that's programs responsibility billions of dollars into counties that have put, you know, hundreds of thousands of beds and facilities actually on the ground, brick andmortar facilities. Uh, and so we're really doing that. I think the one of
[25:56] the impacts though that I will say from an infrastructure out is just our ability to exchange data with partners, right? Um, you know, homeless data is not captured by uh Department of Healthcare Services. And so we had
[26:11] something like 15 different organizations that we had to uh exchange data with and do that on a regular basis. And to be able to do that now uh and again talking about how things used to happen uh probably no one in this room remembers SFTPs
[26:29] uh but we still live in that world. And so uh to be able to actually exchange data in near real time with other government organizations uh specifically ones that are on the same platform have been a just a monumental change on our
[26:44] infrastructure outcomes. So uh and I'm I'm hoping uh I know I'll be able to show those from those impacts from the technology lens will directly correlate to the you know demonstration of the outcomes for behavioral health that we
[27:00] wouldn't be able to do because previously we would have been moving data around by hand or you know waiting for a batch job to run. Um, so this was not scripted before, but given Robbiey's syphilis example kind of of the things that he's most excited about in terms of
[27:15] outcomes, is there something that you're looking forward to similarly where you're like this actually whether it's AI or just improved data practices that it's actually going to drive huge impact whether it's for Californians receiving behavioral health or internally?
[27:31] Yeah. So there's a something that we're rolling out uh actually the end of the month is the ability for counties to submit individual service level data. So encountered data, put hands on people, help people out in the field and so forth. And and again that circulates
[27:48] back to how money is paid and and so forth. But what we're looking to do is is as that data comes in, we're giving the counties that real information and a snapshot of what DHTS knows about you today. And then we're actually going to publish that information. And so for the
[28:04] first time that I know of at Department of Healthcare Services, we're able to actually give them a profile of what they're doing and an understanding of what we perceive as going on in their community. So if they're uploading claims, we can tell them how many
[28:20] claims, you know, historically we have for you for the last six months and and build that story as as opposed to, hey, give us a monthly data drop, right? And so just to build those things and that transparency and show that history and the lineage, all that comes down to how
[28:37] we handle the data in a more modern way. And that's what I'm really looking forward to is be able to show them that back to them. Yeah, absolutely. Um so I'm always cognizant when we have big audiences like this um that everyone is kind of at a different point in their journey. So you guys kind of have already deployed
[28:52] data bricks with the knowledge that folks haven't started in some of these cases. What would you want to tell other public sector leaders? There's something that you wish you had known before you started out um or how to move faster.
[29:09] any advice that you have for audience members that are maybe just starting out on their data bricks journey. So I think I mentioned this I'm I'm only six years at the state I came from private sector and with Prop One came an
[29:26] exemption uh to move very quickly. Uh, and so I've been very fortunate to be able to partner with Data Bricks and bring them in in in a very quick and short timeline to stand up a lot of what we're doing. Um, I probably moved too
[29:43] fast, right? Uh, in a government environment and so I I probably could have moved slow to move fast, but I do not want that to be the tagline or when you leave here, don't associate that, you know, as my ultimate advice. I think really just if I had something that was
[29:59] valuable across the board, maybe not go slow or go fast. Um because it starts small, right? In government there's tons of data silos. There's tons of opportunity to make things better. And what I have learned in my short time in government is usually there's a huge plan and we try to solve all the things
[30:16] at once. For me, it's about go solve one little problem and gain a partnership with your program people. Start small and then it'll evolve and pick up momentum and then incrementally you will show that value to your program partners. And so that's really again if
[30:32] you're going to take something I say it's not the the slow part. It's the incremental and show value to your partners is really where you're going to benefit. Absolutely Robbie. Yeah. A lot of things you said that resonated with me at at towards the end. You were talking
[30:48] about how important it is to like work with your partners and really understand the problems that you have. Can also relate to the like, oh, you need to like wait to do this thing because the time's not now. And I think like my advice and my experience, I've only been there a couple more years at the state than you have. There's just like never the right
[31:06] time. It's never like the correct time to do it. just like start to take some of these pieces and move forward on it. Listen to the people who are actually doing things. There are a lot of things that we've been able to achieve and build trust from just like automating processes that took tons of time, right?
[31:22] Like we had multiple epidemiologists who had to babysit a script for six days. Like we talked to them and we just like rewrote it and you know like there's a lot of nuance in there, but like there's never a wrong time or right time to start on this. And I think the other piece that I've come to terms with now
[31:38] uh is that like it's it's never you're never done modernizing your data infrastructure. Like it's always going to evolve. You're always going to have some plan and you're never actually going to achieve what is entirely laid out. And so there's like this balance of communicating what you're actually able to do. But as long as you're able to
[31:53] articulate the value that you gain from doing these things and the way that what you're doing impacts the people that you're trying to serve, I think that that that's really helped us be able to continue to push forward some of these efforts. Terrific. Thank you guys both for for joining us today. Um I feel like just as
[32:09] an American, it makes me feel very glad that people like you are are looking out for all of the citizens health so and pushing forward on like when things are easy. like I feel like you're both did not go through all of the hard things, but I know that there was lots of them and lots of obstacles. So, thank you. If everyone will join me in thanking our
[32:25] panelists. Hi, good afternoon everyone. My name is Aaron Mills and I'm here representing the Chief Digital and AI Office or CDAO. In a rapidly modernizing world, our
[32:40] mandate is to continue to stay at the forefront to deploy secure and scalable infrastructure and to leverage data to accelerate decision advantage. Delivering solutions at department scale
[32:55] introduces a bit of complexity. Understandably, our users are distributed globally located across the combatant commands. Our mission set is vast. We manage a supply chain with three times as many
[33:12] suppliers as Walmart and own more ground vehicles than FedEx. We employ more people than the entire population of Philadelphia. All love to Philly. I hear they're getting a new rail project. Instead of building disperate solutions
[33:29] to manage these disperate domains across supply chain, logistics and people, we've built an enterprise data platform and a top that enterprise data products. We now rely on common, trustworthy, accessible data.
[33:45] We serve 85,000 active users across tools and 18,000 data bricks users accessing data from 700 authoritative data sources at 98 pabytes of total storage.
[34:01] We've deployed data pricks across four impact levels and operate three production environments at UNCClass, classified and TS domains. We've unlocked the power of our data by serving as a core data layer to the
[34:17] department. Everything we build must be interoperable. Pri prioritizing open formats and integrations with other systems. We are setting design patterns that allow our distributed user base to ingest, govern, and share their data
[34:34] rapidly and securely. Evolution is a common thread across stories in this room. As technologies become available, we adopt. In 2020, we at CDAO began our journey with data bicks. And in 2025, we migrated from PVC
[34:52] to data bick SAS offering E2. Migrating 10,00 clusters and 8,500 jobs is not for the faint of heart, especially during a government shutdown. Adding on a layer of complexity, it was also time for a platform rearchitecture.
[35:10] We had grown to a size where we could no longer operate from a single AWS account and a single data bricks workspace. In eight months, we migrated to a multi-tenant architecture, deploying separate infrastructure for our largest communities, Department of Navy and Air
[35:27] Force A4. Today, our customers are using natural language to talk through to their data via Genie. Since we enabled Genie in February, we already have 1,000 Genie spaces and 70,000 Genie queries,
[35:43] reducing the time required to derive actionable insights from data. Today in summer 2026, I hope to look back on this time and remember three cultural touchstones of equal importance. The
[35:59] Knicks winning the NBA finals. Go Knicks. Um the World Cup, of course, and our Unity catalog migration. Unity Catalog is helping us provision granular data access and distribute data
[36:16] stewardship. As you can imagine, we have many sensitive subtypes of data. We've applied the CUI framework for controlled and classified information as govern tags on assets. These tags are mandatory at asset creation. Data stewards approve
[36:33] CUI classifications and gate the release of control tables. To take on a little bit more during this migration, we are also enabling serverless compute. Serverless is required for our tenants to modify and modernize legacy access
[36:49] strategies. It also eases our path to future feature adoption as a lot of the capabilities we've been hearing about here require it as a prerequisite. Pre-Unity catalog, one of our use cases
[37:05] created 6,000 distinct views to customize which ERP data their different consumers were able to access. Now with column masking and row filtering, we maintain stringent data access policies, validate compliance
[37:22] through a policy engine, and minimize and reduce complexity. Through Unity catalog, we provide a unified entry point to consumers and reduce the amount of data copy when sharing with partner systems. We grant
[37:38] access directly to the managed tables and to their associated metadata. This means the CUI tags, limited dissemination controls and distribution statements are read alongside the data and access is governed appropriately in downstream systems. We've completed a
[37:56] successful pilot with Foundry reading Unity catalog data at lower latency, reading the tags and writing back to our core data layer. This governance layer provides maximum flexibility to distributed organizations. We at CDAO are able to
[38:13] focus on democratization and interoperability. We don't need to be opinionated on where our users are accessing data as we are able to securely share it. Four foundational components have enabled us to operate at the scale we do
[38:30] today. Mission spaces allow our largest consumers to deploy and operate their own instances of tools and services in a dedicated AWS account. Our core data layer is accessible across accounts.
[38:46] Infrastructure is code. As we modernized our platform architecture and our data storage strategy, we needed a declarative version controlled approach to infrastructure deployment. As an example, our S3 GitHubs pipeline
[39:02] enforces platformwide standards provides a complete audit trail when provisioning buckets. We built a suite of data movement services as common services all of our users are able to access. These are AWS
[39:18] native event driven. But in 5, our unclass environment, we needed to modernize those to read from Unity catalog. We maintained backwards compatibility in our higher level environments on the high side where E2
[39:33] is not yet available. We are actively working with data bricks on that partnership to make E2 available in zipper and JWIX. Lastly, observability. We allow our c customers and consumers access to see usage and spend and do
[39:50] this through cost center tags and enterprise dashboards with drill down ability into high-cost workloads. I want to leave you with mission impact. We have built this data foundation. Now what problems are we using it to solve?
[40:08] Commonly understood data improves operational planning. We see this in common operating pictures developed across classes of supply. These applications illuminate supply chains and provide scenario planning capabilities.
[40:24] While products built around parsed naval message traffic provide Navy leadership with real-time visibility into ship status, which then improves short and long-term planning. Trustworthy data promotes a clean audit.
[40:40] Comproller is establishing authoritative lineage, resolving gaps in transaction level visibility and improving reconciliations. These applications save time. Headquarters Marine Corps streamlined a
[40:56] manual quarterly reporting requirement. In doing so, the Marine Corps recovered approximately 200 manpower hours from uniform service members each quarter. And most importantly, these products have real world impact. CDAO supports
[41:14] non-combatant evacuation operations in times of crisis alongside partner agencies. These are just a few of the examples of the breadth of problems we are solving today. Now that we have established and laid the foundation with a core data
[41:30] layer, I can't wait to see the problems we solve tomorrow. And now, please welcome to the stage, Callum McCann, head of data, Arsenal OS at Anderl. I want to begin by asking you to imagine
[41:45] that you work at a defense company that is rapidly scaling production on a range of technically complex systems. Alongside the actual engineering that needs to happen, you need to optimize the systems that comprise your industrial workflow. And the timelines to deliver outcomes are only getting
[42:01] shorter by the day. To accomplish this means connecting outputs from dozens of systems in countless shapes and taking that data and making it useful all while dealing with security controls and federal constraints and then once you have that data uh taking it to the floor
[42:17] and making it useful to all the folks that work on the floor. Welcome to a regular day at Andre. My name is Callum McCann and I'm the head of Arsenal OS data at Andre. There we go. Uh, Anderl is a defense technology company whose mission is to transform national security uh, with
[42:35] advanced technology from autonomous submarines to fighter jets to mixed reality headsets. And Arsenal OS is the way that we build that hardware at software speed. One system connecting engineering, production, factory machines, and field feedback so that
[42:51] nothing waits for humans to move information. In January of this year, my team was tasked with delivering submittency reporting to a new manufacturing line, producing hardware important to national security. We' never done it before, and the need was urgent. Every week without
[43:08] visibility made it harder to meet our deadlines. This is the story of how we accomplished this in 8 weeks with data bricks. So the fun thing about building data infrastructure at Andreal is probably familiar to most folks in the room and that most of the standard choices don't
[43:25] work for us. Everything we use has to be deployed under GovCloud under strict federal requirements which means that from day one most commercial SAS tools are out. Our team had already standardized on SQLbased frameworks in the past, but they could never meet the requirements that we needed for this
[43:40] project because they were a batch. They'd never meet that one minute latency. We considered spinning up a custom streaming stack, but that would have meant building a parallel universe. And given the deadlines, we just didn't have the time. So, what we needed was simple for our stakeholders to describe
[43:57] live dashboards or displays on the factory floor showing the state of the manufacturing line in near real time. Inventory levels, equipment effectiveness, issue on the line, all updating within the minute. Simple to describe, not so simple to build in the environments that we operate in. So this
[44:14] is when we started talking with data bicks. We kicked off the conversation with them in January and wasted no time getting started. With support from their team, we had a developer instance deployed on our infrastructure in just under 3 weeks. That alone was an accomplishment
[44:30] by both teams. Most vendors take months to cure to clear our security reviews, let alone weeks. From there, the architecture came together pretty quickly. One of our data engineers set up a change data capture or CDC connection to our manufacturing
[44:46] execution system which enabled us to capture every change as it was happening. We streamed that raw change data into data bricks used spark structured streaming broke it out into individual data sets and within 3 weeks we had data flowing at the latency that we needed.
[45:03] But the thing about CDC data is that it tells you what changed not what the current state is. Uh we had a deep log of inserts, update, and deletes, but that actually didn't matter to the folks on the floor. What they needed were actual dimensional
[45:18] tables, work orders, inventory counts, work center statuses, all enriched with other data sources. And to build what they needed, we had to join across dozens of relationships. It was these joins that gave us the most problem. Every stream we tried to set up that
[45:34] joined together all of the necessary dimensional data ended up breaking our SLA almost immediately. to say nothing of the technical complexity of managing all of those checkpoints. The business needed results and we didn't have time to experiment. At this point, we were about six weeks in and it looked like we
[45:50] didn't have a clear path forward to delivering outcomes. That's when we discovered auto CDC flows. It's a capability within data bricks that takes raw CDC data, applies primary keys, and then automatically reconstructs the current state of the table continuously in real time. That
[46:08] was the unlock. Suddenly, we had the current state of every table without needing to manage it ourselves, without needing to deal with those checkpoints I mentioned. Once we had these, we were able to layer continuous materialized views on top of them to handle all of the joining across the different data
[46:23] sets that I mentioned running every 10 seconds. Uh, two more weeks of development, testing, and validation followed. And then we had it event to dashboard under a minute. Eight weeks after kickoff, the dashboard was live on the production floor.
[46:40] Operators are now able to glance up and see all of the context that they need. Uh quality issues, overall performance, status of the line, and more. And this knowledge is critical in a very fastmoving environment. But as the foundation shifts underneath us in this
[46:56] fastmoving environment, we recognize that not every solution we are going to build is going to be the most technically elegant. Uh there may be streaming purists out among this crowd who are looking at what I'm describing and recognizing that there are dozens of things we could have optimized or
[47:12] savings we left on the cutting room floor. And you're right. But more important than getting to the perfect technical solution is getting to the solution that works for the business fast. And that's what we've been doing. And it works. We've been running this for three months at this point. And most
[47:28] importantly, it met the timelines that the business needed. Taking an approach approach rooted in rooted in practicality also gave us muchneeded flexibility. We didn't have to spend weeks perfecting one pipeline before we even knew what the most important metric on the floor was going
[47:44] to be. We could bring in late arriving data, handle any degree of normalization, and add new sources without needing to rearchitect or spend weeks on each new data set. When you work in an environment as fastmoving as Andrew, where the product and the process are evolving constantly, you
[48:02] learn very quickly that assumptions are very expensive. Flexibility is what lets you keep pace. An old manager of mine was fond of the adage, make it work, make it right, make it fast. Getting real data in front of real operators on the floor was more
[48:19] valuable than building the perfect technical solution in isolation. You can't improve what people can't see. Now, everything I've described is only possible because of the androll team behind it. Uh Morris Lee, Thomas Hill, Juliana Alabelli, and Brian Crant. They
[48:35] built this under real pressure and they delivered. But it was also possible because of the data bricks platform. It had depth when we needed it. Every time it hit or it felt like we hit a dead end, there was another capability behind a door that ended up solving the problem that we were running into. And what did
[48:52] we learn from all of this effort? There's a world of difference between go run three queries and piece the data together and glance up at the wall to know what's happening. 2 hours and 1 minute are not the same thing. And the approach we landed on is a scalable one.
[49:07] Other production lines across the company have seen the success that we've had and are breaking down our doors to get this rolled out. Literally as I stand here, deployments are happening at other product lines and we're just getting started with what we're planning to do with the data bicks platform. Uh
[49:23] we're going to keep moving forward with rapid iteration. The team has spent the last two months uh migrating all of our legacy pipelines over to data bricks so that we can begin to utilize capabilities like Genie and apps all resting on top of fine grain ac uh fine
[49:38] grain access controls implemented with Unity catalog. And as soon as the data bricks team makes lakebase available to govcloud customers, we plan on throwing dozens of workloads against it. If there's one thing I want you to take away from this talk, it's this. A deep
[49:54] platform plus a motivated team can deliver on timelines that others would call unrealistic. Stability and flexibility let you serve the business, not just design a perfect architecture. And when you're building hardware that the war fighter depends on, serving the
[50:11] business isn't optional. It's the mission. Thank you. And with this, I'd like to introduce Venu Bopana and Molly from Data Bricks.
[50:26] So, almost every American touches the FDA before breakfast, and we rarely notice it. Whether it is the food we eat, the medicine we take or the health devices we interact with, the trust we have in the FDA is invisible. But behind that trust is a extraordinary amount of
[50:42] data. Um, and for a very long time that data was in disconnected systems that didn't talk to each other. And my guest here has spent the last decade changing that. He's led some of FDA's largest data programs and is now currently deploying Agentic AI at scale across all
[50:59] the centers of the FDA. Benu, I'm so excited for this conversation. Thank you for joining us. Thanks, Molly. To start off, could you talk a little bit about FDA's mission and your role implementing AI solutions across all of the FDA centers um within the office of
[51:14] digital transformation? Sure. Um FD's role is ensuring safe and effective um drugs then veterinary medicine biologics devices and ensuring that food supply chain safe and
[51:30] effective food supply chain as well as radiation emitting devices. It's pretty broad uh um uh remarkable and broad mandate uh that the Congress gave us and um this is something that touches every
[51:46] human I mean American um I mean every single day every 20 cents spent on and uh by US consumer it's regulated by FDA. Now in between all of this my role sits in office of digital transformation
[52:03] where I focus on AI initiatives and ensuring that to meet the mission we bring in the right tools technology and to keep up with the demand uh you have uh thousands of submissions that comes
[52:18] in every month and how do we keep up with that right now um with all of this one of our most recent uh transformation is we have deployed um an AI generative AI tool called Elsa uh to about 16,000
[52:36] FDA staff. Um anyone can just go and uh spin up Elsa they can choose different models multiple models of a choice um and then they can just uh conduct their business. Um one of the most exciting
[52:52] thing is uh we've passed uh it's been about one year two months I believe that we since we launched we've passed be beyond um the regular questions and answers now our staff I'm talking about medical doctors scientists they're
[53:10] creating agents and at scale and we have about hundreds of agents created per week and uh it's remarkable I It it is amazing to see how quickly the FDA staff is ready to adopt and this is something
[53:27] not just for um the regular uh you know data scientists uh it it starts from administrative staff to all the way to um you know the medical research scientists who are really using the platform uh extensively. Yeah,
[53:43] that's a background about FDA. Yeah, that's very impressive. Um I read somewhere that you went from less than 1% of everyone at the FDA using Elsa and Halo to 85%. Yeah. How did you create an environment where
[53:58] there was so much trust in the data platform? So um before we built ELSA uh there were multiple of course FDA has uh multiple centers uh uh let me just uh make sure that everybody has a little bit
[54:15] background of those centers uh for drugs cedar then biologics then uh devices we also have uh CVM veterinary medicine then CTP tobacco products and in between all of these OI which is inspections
[54:31] uh they ensure that all of these products that are being reviewed and approved and the facilities that manufacture they're thoroughly inspected and ensuring that the safe and effective uh products are uh uh released into the
[54:46] American public right and also OIS role also involves all the imports that come in uh the containers uh uh manufactured through various countries that could be drugs devices they It's their responsibility to en ensure that um you
[55:04] know all of these products are vetted, approved and safe. Right now coming back uh each of these center were kind of siloed. They had their own u uh AI capabilities. They all built u uh their own chat bots and this was causing
[55:20] significant uh uh uh cost as well as you can really not get the true picture of any of this data that you want to really use for AI right um so as part of that uh you know evaluation uh the the
[55:39] leadership from the IT leadership they did realize the fragmentation across the centers and That's where they initiated this uh uh you know how do we bring all of this and consolidate into a single platform and of course that's all within data bricks and we were able to do it
[55:54] within uh about a 3 to four months I believe with about uh 50 to 60 data sources all eight centers brought in and still there is more need to do of course some of the products that Ali was talking about those are the ones that we
[56:10] are looking for because we really need those um not just uh just the traditional data platform but more on the agent AI capabilities and all of those I think the and now coming back the 85% so if you look at it lot of our
[56:28] staff like be it of course FDA uh the organization uh works heavily on documents we have about a pabyte of a data and then every day we get uh hundreds of uh gigabytes of data so if
[56:44] you look at Um lot of these uh centers they have to evaluate this data submissions that come in and also there are a lot of SOPs because of the uh federal mandates you will have to follow those SOPs as well as guidelines uh
[57:01] regulator regulations and all of those. So they used to have go through every time there is a question somebody has to go through uh 500 pages document try to search and figure it out. Now the staff took all of these created workspaces within Elsa and then uh you know they
[57:19] they created agents anybody can just go and ask a question with a simple chatbot you get grounded and uh more like u I would say uh you know uh data that is relevant to FDA right so that's how uh
[57:35] you know the adoption we've seen at scale and uh this is really the story of where we got in you know, on that we were talking about kind of the importance again what Ali was talking about before was um the importance of enterprise context and how you've done that so well with data bricks at at FDA
[57:52] and that's helped dive that 85% which super impressive. You talked about kind of the the fragmentation across the FDA centers. I'm curious as you were consolidating all of the data platforms and solutions um kind of under Halo and data bricks, what were the pain points
[58:07] that you saw and how did you work around them? So um one of the I I think uh we were able to convince the rest of the centers um because we have a success story. For example, um Cedar, which is
[58:23] the drugs, um was probably I would say five years into the journey with the data bricks, right? Uh we were we sponsored uh data bricks to FSMA high five because there is a need because lot
[58:38] of our data is trade secrets and we have to maintain those uh um critical security measures that we need to maintain. So for the past 5 years uh we've built this massive data platform to support cedar and we showed the value
[58:55] of how we can really make data available for example data sharing what used to take uh probably like uh four or five days between centers we were able to show uh uh you know cut down that to couple of hours right um similarly with
[59:12] the data processing data streaming making data available real time so we've got a success story and on top of that we also had Elsa which was really tapping into the data and showing the value of having a foundational data platform and that success story actually
[59:30] became kind of a contagious I would say uh the rest of the centers saw the value and they were very quick to adopt. I mean there are some concerns outliers but I think uh overall uh and those were genuine concerns in terms of security and uh how data bricks can handle the
[59:47] unity catalog. I think Unity catalog gave us a very good story where we could prove it out how data can be contained and how it could be u you know the data assets uh not be uh shared without uh
[01:00:02] the proper approvals right guardrails and all of those. Yeah, that's a good segue. Um so as you created the the data governance foundation what role does security and governance and unity catalog play in really creating um the ecosystem that AI could come in and provide that value.
[01:00:19] So very good question. So um again uh so we we we have we have the uh AI uh uh like the agents. Um now we started off with the initial chatbot where you could
[01:00:35] ask general questions but uh but now we need to take it to the next level. how do you make this data available through AI right and that's where we were able to build MCP servers uh connectors and then we were able to layer this MCP on
[01:00:53] top of the Unity catalog and then we were able to get to the of course using arbback abacks and getting into the most granular level of a table level access or whatever the most granular level you could think of and uh that's what made
[01:01:08] is uh and then we when we exposed this MCP tools via Unity catalog which gives us a very clear structured organized data I think that really helped us. I mean uh uh so so now the the doctors or
[01:01:25] scientists they actually can uh you know go to the unity catalog and then identify what uh uh their their their particular uh uh data that they're looking for and they're able to tap into it and then uh you know all they need is
[01:01:41] a good prompt and then the attached with the MCP server uh and then or connector and then also the knowledge base and and uh they were able to convert their SOPs into agents and that's the success story.
[01:01:56] That's terrific. Um just to make I want to double click on that to make it real for everybody here. What are some example use cases of Elsa that FDA could not do before that they can do now. So um a couple of things with Elsa. Um
[01:02:14] getting data access to the data is a challenge right for example um we do have regulatory submissions that is the starting materials of how drugs are manufactured and that data is sitting in about 3 to 4 million pages of documents.
[01:02:30] Similarly, there are other uh data assets like for example the relationships between the product and the supplier or the manufacturer that's against buried in probably another four five million pages of documents we have. It's it's almost impossible to get
[01:02:46] access to this data. I mean the reviewers had to really go open those documents do a keyword search find that information. So through data bricks using um the the NLP from the ML flow uh we were able to extract uh uh the the
[01:03:04] key assets like be it starting materials or the broken relationships right we were able to extract them and expose um and that's where Elsa was able to tap into it and with a simple prompt they are able to all they have to do is ask
[01:03:19] by application number or submission number okay give me what are uh starting materials for this manufacturer and within 3 minutes versus what they were spending like uh almost like two weeks. Yeah. Yeah. That's incredible. So I know that you have a large number of AI use cases and
[01:03:36] I I I think it was the Wall Street Journal and the Washington Post described what FDA has done as like the enterprise AI blueprint for government. Um what use case are you most excited about kind of moving forward? I think uh the most exciting use cases are where uh
[01:03:53] wherever our review staff can spend less time trying to hunt for information or searching for information rather focus on their core uh you know job right be it reviewing applications their subject
[01:04:08] matter expertise in in evaluating drug products right I think if we can get to uh them not spending the time I think that is where um uh we really see some success. I think um um I we we we've
[01:04:25] seen that uh happening like where uh lot of this uh manual processes the review staff are really looking at taking advantage of this but we are still growing there is lot of ways to grow. I think we also the other areas are uh um
[01:04:44] how do you make this uh MCP tools not just for one center what like for example Cedar if you can scale it up to the rest of the centers right which we are I think we are in that path and also making the data contest context more for
[01:05:01] each center right readiness I think that is where our journey is if we were to make this at scale and make it available for rest of the centers. I think that is a very good uh success story for us. Absolutely. Um so I feel like I could speak to you for hours. I'm a huge fan
[01:05:18] girl of what you've done at FDA. Um but we are running out of time. Um I feel like just to to summarize what I've heard is that I feel like all of the headlines I see is like FDA does AI, but the reality is that you created a strategy, you laid the groundwork, you
[01:05:33] have a governed data foundation, and then you brought AI to that to provide the enterprise context. And then you created an ecosystem where you know more than 80% of the scientists there are now using it which is so impressive. Um so thank you so much for sharing your story with everybody. Um and please everyone
[01:05:49] please please join me in welcoming. Thank you.
Learn more about the Databricks Data and AI platform.
The information provided herein is for general informational purposes only and may not reflect the most current product capabilities or configurations.