FDA's Enterprise AI Blueprint: Deploying Elsa at Scale with Databricks
Summary
- The FDA deployed Elsa, a generative AI platform, to approximately 16,000 staff across 8 centers on the Databricks Data and AI platform, consolidating previously fragmented AI systems and achieving 85% adoption.
- Unity Catalog provides fine-grained data governance over a petabyte of regulatory data, while MCP server connectors enable secure agent access to regulatory documents, reducing document review from weeks of manual searching to approximately 3 minutes.
- FDA staff including medical doctors and scientists are now creating hundreds of agents per week, with usage scaling beyond data scientists to administrative staff and medical reviewers across all FDA centers.
FDA's Enterprise AI Blueprint: Deploying Elsa at Scale with Databricks

The FDA deployed Elsa, a generative AI platform, across 16,000 staff in 8 centers, achieving 85% adoption and reducing document review time from weeks to minutes. With a petabyte of regulatory data and hundreds of gigabytes arriving daily, FDA staff previously spent weeks searching for information through disconnected systems.
This conversation explores how FDA consolidated fragmented AI systems across all centers onto Databricks, leveraging Unity Catalog for fine-grained data governance and MCP server connectors to enable secure agent access to regulatory documents. Learn how they scaled adoption beyond data scientists to medical reviewers and administrative staff, transforming document discovery from weeks of manual searching to 3-minute AI-powered queries.
🤝
Chapters
00:00FDA's Mission and Data Challenges02:02Deploying Elsa Across 16,000 FDA Staff03:23Achieving 85% Adoption and Building Trust04:44Consolidating Fragmented AI Systems Across Centers08:32Data Sharing Improvements and Success Stories09:54Unity Catalog, MCP Servers, and Agent Architecture11:49Real Use Case: Document Intelligence for Regulatory Submissions13:27Future AI Use Cases and Scaling Across Centers
FAQs
What is Elsa and why did the FDA build it?
Elsa is a generative AI platform deployed across the FDA to help staff access and analyze regulatory information more efficiently. Before Elsa, FDA staff spent weeks manually searching through disconnected systems, but a query that previously took weeks can now be completed in approximately 3 minutes.
How did the FDA achieve 85% adoption of Elsa?
The FDA achieved 85% adoption approximately one year and two months after launch, with usage evolving well beyond basic question-and-answer interactions. Staff from administrative roles all the way to medical doctors and scientists are now creating agents at scale, with hundreds of agents being built each week.
How does Databricks support the FDA's data governance requirements?
The FDA uses Unity Catalog on the Databricks Data and AI platform for fine-grained data governance over a petabyte of regulatory data with hundreds of gigabytes arriving daily. MCP server connectors enable secure agent access to regulatory documents while maintaining compliance across all 8 FDA centers.
What types of use cases does Elsa support at the FDA?
Elsa supports use cases ranging from document discovery and regulatory submission review to agent-driven workflows for medical reviewers and scientists. Staff can choose from multiple AI models and create customized agents for their specific workflows, transforming how the FDA processes the thousands of submissions that arrive each month.
Full transcript
[00:08] So, almost every American touches the FDA before breakfast and we rarely notice it, whether it is the food we eat, the medicine we take, or the health devices we interact with. The trust we have in the FDA is invisible, but behind that trust is a extraordinary amount of data. Um, and for a very long time that
[00:24] data was in disconnected systems that didn't talk to each other. And my guest here has spent the last decade changing that. He's led some of FDA's largest data programs and is now currently deploying a genetic AI at scale across all the centers of the FDA. So, I know
[00:41] I'm so excited for this conversation. Thank you for joining us. Thanks, Molly. To start off, could you talk a little bit about FDA's mission and your role implementing AI solutions across all of the FDA centers within the office of the digital transformation?
[00:56] Sure. Um, FDA's role is ensuring safe and effective um, drugs, then veterinary medicine, biologics, devices, and ensuring that food supply chain, safe and effective food supply chain, as well as radiation
[01:13] emitting devices. It's pretty broad uh, remarkable and broad mandate uh, that the Congress gave us. And um, this is something that that touches every human, I mean, American um,
[01:29] I mean, every single day. Every 20 cents spent uh, by use consumer, it's regulated by FDA. Now, in between all of this, my role sits in office of digital transformation where I focus on AI initiatives and
[01:46] ensuring that to meet the mission, we bring in the right tools, technology, and to keep up with the demand. Uh, you have uh, thousands of submissions that comes in every month and how do we keep up with that, right?
[02:02] Now, um with all of this one of our most recent transformation is we have deployed an AI generative AI tool called Elsa to about a 16,000 FDA staff.
[02:18] Anyone can just go and spin up Elsa. They can choose different models, multiple models of a choice and then they can just conduct their business. Um one of the most exciting thing is
[02:33] we passed it's been about 1 year 2 months I believe that we since we launched we've passed beyond the regular questions and answers. Now, our staff, I'm talking about medical doctors, scientists, they're
[02:50] creating agents and at scale. And we have about hundreds of agents created per week and it's remarkable. I mean it it is amazing to see how quickly the FDA staff is ready to adopt. And this is something
[03:06] not just for the regular you know data scientist it starts from administrative staff to all the way to you know the medical research scientist who are really using the platform extensively. Yeah.
[03:23] That's the background of what FDA Yeah. That's very impressive. Um I read somewhere that you went from less than 1% of everyone at the FDA using Elsa and Halo to 85%. Yeah. How did you create an environment where
[03:38] there was so much trust in the data platform? So, before we built Elsa, there were multiple of course FDA has multiple centers. Let me just make sure that everybody has
[03:54] a a little bit background of those centers. For drugs, cedar, then biologics, then devices. We also have CVM, veterinary medicine, then CTP, tobacco products. And in between all of this, OII, which is inspections.
[04:11] They ensure that all of these products that are being reviewed and approved and the facilities that manufacture they're thoroughly inspected and ensuring that the safe and effective products are released into the American public,
[04:27] right? And also OIS role also involves all the imports that come in. The containers manufactured through various countries that could be drugs, devices. They it's their responsibility to ensure that
[04:44] you know, all of these products are vetted, approved, and safe, right? Now, coming back, each of the center were kind of siloed. They had their own AI capabilities. They all built their own chatbots. And this was causing
[05:00] significant cost as well as you can really not get the true picture of any of this data that you want to really use for AI, right? So, as part of that
[05:15] you know, evaluation, the leadership from the IT leadership, they did realize the fragmentation across the centers. And that's where they initiated this you know, how do we bring all of this and consolidate into a single platform.
[05:31] And of course, that's all within Databricks. And we were able to do it within about a 3 to 4 months, I believe, with about 50 to 60 data sources. All eight centers brought in and still there is more need to do. Of course, some of the products
[05:47] that Ali was talking about, those are the ones that we are looking for because we really need those um not just just the traditional data platform, but more on the agent AI capabilities and all of those. I think the And now coming
[06:03] back, the 85%. So, if you look at it, a lot of our staff, like be of course FDA the organization works heavily on documents. We have about a petabyte of a data and then every day we get
[06:20] hundreds of gigabytes of data. So, if you look at it, a lot of this centers, they have to evaluate this data submissions that come in. And also, there are a lot of SOPs because of the federal mandates, you
[06:37] will have to follow those SOPs as well as the guidelines, regulations and all of those. So, they used to have go through every time there is a question, somebody has to go through 500 pages document, try to search and figure it out. Now, the staff took all
[06:54] of this, created workspaces within Elsa, and then you know, they they created agents. Anybody can just go and ask a question with a simple chatbot. You get grounded and more like I would say you know, data that is
[07:11] relevant to FDA, right? So, that's how you know, the adoption we've seen at scale and this is really the story of where we got in, you know, all that. We were talking about kind of the importance again, what Ali was talking about before was the importance of
[07:28] enterprise context and how you've done that so well with Databricks at FDA and that's helped drive that 85% which super impressive. You talked about kind of the the fragmentation across the FDA centers. I'm curious, as you were consolidating all of the data platforms and solutions um kind of under Halo and
[07:45] Databricks, what were the pain points that you saw and how did you work around them? So, um one of the I I think uh we were able to convince the rest of the centers um because we have a success story. For
[08:00] example, um Cedar, which is the drugs, um was probably I would say 5 years into the journey with the Databricks, right? Uh we were re-sponsored uh Databricks to FISMA high IL5 because
[08:16] there is a need because a lot of our data is uh trade secrets and we have to maintain those uh um critical security measures that we need to maintain. So, for the past 5 years, uh we've built this massive data
[08:32] platform to support Cedar, and we showed the value of how we can really make data available. For example, data sharing. What used to take uh probably like uh 4-5 days between centers, we were able to show uh you know, cut down that to a
[08:49] couple of hours, right? Um similarly with the data processing, data streaming, making data available real time. So, we've got a success story, and on top of that, we also had Elsa, which was really tapping into the data and showing the value of having a
[09:05] foundational data platform. And that success story actually became a kind of a contagious, I would say. Uh the rest of the centers saw the value, and they were very quick to adapt. I mean, there are some concerns, outliers, but I think overall, uh and those were genuine
[09:22] concerns in terms of security and uh how Databricks can handle the Unity Catalog. I think Unity Catalog gave us a very good story where we could prove it out how data can be contained and the it could be, uh you know,
[09:38] the data assets uh not be uh shared without the proper approvals, right? Guardrails and all of those. Yeah. That's a good segue. Um so, as you created the the data governance foundation, what role did security and governance and Unity Catalog play in
[09:54] really creating um the ecosystem that AI could come in and provide that value? So, very good question. So, um again, so we we we have we have the uh AI uh like the agents. Um
[10:11] now, we started off with the initial chatbot where you could ask general questions, but but now we need to take it to the next level. How do you make this data available through AI, right? And that's where we were able to build MCP servers uh connectors, and then we
[10:30] were able to layer this MCP on top of the Unity Catalog. And then we were able to get to the uh of course, using our back A backs and the getting into the most granular level of a table level access or whatever the most granular
[10:45] level you could think of, and that's what made us uh and then we when we exposed this MCP tools via Unity Catalog, which gives us a very clear, structured, organized data. I think that really helped us. I mean, uh
[11:01] uh so so now the the doctors or scientists, they actually can uh you know, go to the Unity Catalog and then identify what uh their their their particular uh data that they're looking for, and
[11:16] they're able to tap into it, and then uh you know, they all they need is a good prompt, and then the attach with the MCP server, uh and then our connector, and then also the knowledge base, and and uh they were able to convert their SOPs
[11:33] into agents and uh that's the success story. That's terrific. Um just to make I want to double click on that. To make it real for everybody here, what are some example use cases of Elsa that FDA could not do before that they can do now? So um
[11:49] a couple of things uh with Elsa um getting data access to the data is a challenge, right? For example, um we do have uh regulatory submissions that is the starting materials of how drugs are manufactured. And that data is
[12:06] sitting in about a 3 to 4 million pages of documents. Similarly, there are other uh data assets like for example, the relationships between the product and the supplier or manufacturer. That's again buried in probably another 4 5
[12:21] million pages of documents. We have it's it's almost impossible to get access to this data. I mean, the reviewers had to really go open those documents, do a keyword search, find that information. So through Databricks using um
[12:37] the the NLP from the ML flow, uh we were able to extract uh the the key assets like the be it starting materials or the broken relationships, right? We were able to extract them and expose um and that's where Elsa was able to tap
[12:54] into it. And uh with a simple prompt, they're able to all they have to do is ask by application number or submission number, okay, give me what are the uh starting materials for this manufacturer? And within 3 minutes versus what they were spending like almost like 2 weeks,
[13:11] yeah. Yeah. That's incredible. So I know that you have a large number of AI use cases and I I I think it was The Wall Street Journal and The Washington Post described what FDA has done as like the enterprise AI blueprint for government. Um what use case are you most excited about
[13:27] kind of moving forward? I think the most exciting use cases are where where wherever our review staff can spend less time trying to hunt for information or searching for information, rather focus on their core
[13:44] you know job, right? Be it reviewing applications, their subject matter expertise in in evaluating drug products, right? I think if we can get to them not spending the time I think that is where
[13:59] we really see some success. I think I we we've seen that happening like a where uh lot of this uh manual processes the review staff are really looking at the taking advantage
[14:15] of this. But we are still growing, there is a lot of ways to grow. I think we also the other areas are um How do you make this MCP tools not just for one center? What
[14:30] like for example CEDA, if you can scale it up to the rest of the centers, right? Which we are I think we are in that path. And also making the data contest context more for each center, right? Readiness. I think that is where our journey is. If we were to make this at
[14:47] scale and make it available for the rest of the centers, I think that is a very good success story for us. Absolutely. Um so I feel like I can speak to you for hours. I'm a huge fan girl of what you've done at FDA. Um but we are running out of time. Um I feel like just
[15:02] to summarize what I've heard is that I feel like all of the headlines I see is like FDA does AI, but the reality is that you created a strategy, you laid the groundwork, you have a governed data foundation, and then you brought AI to that to provide the enterprise context.
[15:18] And then you created an ecosystem where, you know, more than 80% of the scientists there are now using it, which is so impressive. Um, so thank you so much for sharing your story with everybody. Um, and please everyone, please please join me in welcoming Thank you.
Learn more about the Databricks Data and AI platform.
The information provided herein is for general informational purposes only and may not reflect the most current product capabilities or configurations.