California Public Health Data Modernization with Databricks
Summary
- California's Department of Public Health replaced fax-based infectious disease reporting workflows on Databricks, reducing a hepatitis C data processing pipeline that previously took six days down to 30 minutes.
- The Department of Health Care Services uses Databricks to support Proposition 1 behavioral health transformation, enabling near-real-time data exchange and county-level reporting on substance use disorder and mental health services.
- Both departments received governance, privacy, and security approval for Databricks Genie, and analysts are using GenAI for code generation in SAS, SQL, and R to boost productivity while exploring AI for unstructured health data analysis.
California Public Health Data Modernization with Databricks

California public health data modernization with Databricks is replacing fragmented workflows with faster processing, stronger accountability, and near-real-time data exchange. Leaders from the California Department of Public Health and Department of Health Care Services discuss modernizing infectious disease and behavioral health programs.
Learn how a 6-day hepatitis C process was reduced to 30 minutes, how Proposition 1 data supports county reporting, and how Databricks Genie was approved with governance, privacy, and security safeguards. The panel also examines GenAI productivity with SAS, SQL, and R, potential analysis of unstructured health data, service-level reporting, data lineage, and practical strategies for modernizing public sector infrastructure incrementally.
Proposition 1 fact sheet: https://www.dhcs.ca.gov/behavioral-health-transformation/proposition-1-fact-sheet/
Behavioral Health Transformation: https://www.dhcs.ca.gov/behavioral-health-transformation/
DHCS Data Platform case study: https://www.improving.com/case-studies/dhcs-data-platform/
Chapters
00:00California Health Data Programs and Leaders01:14Legacy Infectious Disease Data and Fax Workflows03:38Behavioral Health Data and Proposition 1 Accountability06:18Databricks Genie Approval, Governance and Security08:47GenAI Productivity with SAS, SQL and R09:51Unstructured Health Data and Syphilis Use Cases12:14Modernizing Behavioral Health Data Intake13:19Near-Real-Time Data Exchange Across Agencies14:42County Service Data, Profiles and Lineage16:03Advice for Public Sector Data Modernization
FAQs
How did Databricks improve infectious disease reporting at California's Department of Public Health?
Before Databricks, California's Department of Public Health received disease reports through fragmented pipelines including faxes, with processing for some workflows taking up to six days. With Databricks, a hepatitis C reporting process that previously took six days was reduced to 30 minutes.
What is Proposition 1 and how does Databricks support it?
Proposition 1 is a California governor's initiative facilitating substance use disorder and behavioral health transformation across the state. The Department of Health Care Services uses Databricks to support county-level accountability reporting for the behavioral health services funded by this proposition.
How did California public health agencies get Databricks Genie approved for use?
Databricks Genie went through a governance, privacy, and security review process at California health agencies before it was approved for deployment. Establishing these safeguards was a prerequisite to using AI-assisted querying on sensitive public health data.
How are California health data teams using GenAI to improve analyst productivity?
Analysts at California health agencies are using GenAI for code generation in SAS, SQL, and R, reducing time spent on routine coding tasks. The agencies are also exploring potential analysis of unstructured health data and automated service-level reporting as future use cases.
Full transcript
[00:08] So before we dive into the technology and the how, could you guys ground us in your organizations and your roles and the Californians that depend on your data programs. >> Where we at? Afternoon. Afternoon everyone. My name is Ryan Mosley. I am
[00:24] the division chief for modernization at the Department of Healthcare Services. Uh I am also the project director for behavioral health transformation uh connected to the governor's proposition one that is facilitating uh substance use disorders and uh
[00:40] behavioral health um transformation. So >> Rob Snyder work at the California Department of Public Health and the Center for Infectious Diseases where our team manages a database that is responsible or houses about 80% of departmental data assets. um to ground
[00:59] that in like the actual data that we have. When someone tests positive for an infectious disease or visits their provider and they have an infectious disease, there's information that is sent to us. So all those data are housed in our data warehouse so that we can draw insight and develop public health interventions from them.
[01:14] >> Terrific. So it's very important. Um Robbie, I want to start with you. Um before we were chatting uh we talked fax machines and and 11 days kind of between uh an out an outbreak signal and that data getting to you. Can you talk to us
[01:30] about what it was like before data bricks? Um and please don't leave any details out. I feel like the fax machine stories everyone here can probably relate to. >> Uh yes. So I mean in the last 10 years we've come thousands of years light speed changes
[01:46] since co and um I was thinking about how to respond to this question and some of the people that I know that work here probably heard this story before but I feel like it's very tangible um and people can relate to it. So like I said we get reports of every person who is
[02:04] tested or has an infectious disease in the state of California. Some of you might remember what happened in 2020 with COVID. Some of you might not want to remember like me. Um, and what happened was that we had these different pipelines through which we received
[02:19] data. We had some electronic data receipt but we still received a large number of faxes and um depending on how much time we give I can give another example of how fragmented our data assets were before. >> I love a good embarrassment please. Yes.
[02:35] So, the first one for faxes, um, I would have to go in on the weekend to refill the paper ream because we were receiving so many fax reports. And then we would take those faxes and then we would put them in an Excel sheet and then we would
[02:50] ingest them into our databases. Um, so I did not really like going in on the weekend, but part of the reason that I had to go in on the weekend was because we had another data process that was set up. Um, without going too far into the weeds about some of our specific
[03:06] infectious diseases, we had this data process that literally took six days to run on a desktop and it would time out and have to go in on the weekend to restart the computer sometimes. Might as well refill the ream of paper on the fax machine while you're there to reset it. And um, since that time, we've taken that six day process and it now runs in
[03:22] about 30 minutes. So um, a huge change. So, and it's about it was related to hepatitis C, which I think is super interesting. one of my pet interest uh conditions and we can do a lot more with that than we could before.
[03:38] >> All right, Ryan, over to you. Um, so the behavioral health transformation, a huge governor initiative. I imagine the whole state is watching a lot of pressure. What were the data limitations that made it hardest to track whether the state's investments
[03:53] were making an impact? >> Pretty much in between the check and the results. Anything in the middle there. So um when I was asked to uh lead out behavioral transformation from the technical side um there was the idea of
[04:12] DHCS uh not as fun as infectious diseases but really handles a lot of money that goes out to counties and so forth. And so as we are, you know, sending out that money, the counties are putting together a plan and they're going to tell us
[04:29] about how they're going to spend that money. And I was like, great, let me see what that's all about. And someone directed me to a website that a PDF was published and it was about a 50page document that had a three-year plan in
[04:44] every different format, every different style. And that was their plan. And I go, "Okay, well, how do we know that that's Yeah, we don't really do that." So, like I said, everything in between from we started with a PDF and
[05:00] that was how the process uh went before. And one of the biggest things about uh the governor's kind of initiative here is that accountability. And so, obviously, we had to stand up the beginning factors of that, which is the
[05:15] intake of that data. And then we had to process that data and understand their plan, be able to segment it against different uh you know um policies, procedures uh and practices. And so as
[05:30] we move through that process, we are trying to come out the other side to be showing the accountability. And we just launched the county profile which is starting to actually visualize the use of those funds. Now we haven't got to that piece of it yet. we've got demographics live and things like that.
[05:47] So, we've started to break down and and that intake and then have the output and uh actually have the infrastructure in between. But yeah, as the uh as we started, we went from a basic PDF, not a fax machine. Uh so, he's got me there. But uh >> we have our share of PDFs.
[06:03] >> Yeah. >> Uh but yeah, that again, it was just so limited and it was really just kind of by feel how the department was doing that. Now that the accountability is needed, that infrastructure is where we're laying the tracks to drive towards that. >> Makes sense. All right, so I feel like
[06:18] we've got a good understanding of kind of the before. Let's move into how everything looks now with data bricks and kind of moving forward. Um Robbie, this is I'm like incredibly impressed by this. You have somehow managed to get Genie approved by your organization in
[06:35] months. Um as someone who was in the federal government, I'm like in awe by that. Um, I assumed that it was like cupcakes and like cakes and he was like, "No, it was just it was a strategy, Molly." Like everyone obviously you have to get governance and security involved in the beginning and I was like, I want you to talk about this. So, can you talk
[06:51] to the audience who I know has very similar challenges around atto and security and kind of talk them through how you were able to get a new feature approved so quickly and kind of what what Genie has enabled your teams to do. >> Yeah. Yeah. So, I will say that it was
[07:07] not just me. Obviously there's a big team involved in part wonderful team >> some of my team is here but our our ITD um group has been in charge of a lot of that approval there is >> a rigorous approval process a number of forms I think it's like a 5305F
[07:24] that's embedded in my brain >> um >> I didn't make that form by the way >> um and yeah I think just conceptually the really important thing for us given that we have so much data and so much PI and
[07:41] so so much sensitive data is ensuring that whatever happens just remains on site and building the safeguards to ensure that right like the privacy things that we have in place it's all native is only done there nothing feeds back into the LLM um from a policy and
[07:57] bureaucracy perspective I think it's just like be persistent and talking to people a lot and working together and really I think articulating the value of having these modernized data systems And um maybe just to really call back to what I was saying before co has really
[08:14] demonstrated the need for those kinds of things. So I think we have a lot of support um throughout the organization to help move towards and move off of these legacy systems that hindered us before and access and give us the opportunities to use them. So I think that's the first part of your question in terms of like this is how we did it.
[08:30] It's >> probably not sufficiently detailed, but every organization probably has their own challenges and bureaucracies that they're going to need to go through. We were just talking before, y'all don't have Genie yet set up and DHCS and your sister department. So, I think it just varies a lot and feel pretty lucky
[08:47] honestly to have people who are helping to take care of it for us. Um, the second part in terms of how we use it, I am optimistic, but I'm actually not sure exactly how we can use it yet to derive some of these advanced analytics.
[09:02] Some of the things that we've already used Genai for have been just simple product simple productivity tools. Um, a lot of the things that my team uses it for involve, hey, I wrote this thing in SAS 15 years ago. How do I get it into
[09:18] SQL so that I can use it in data bricks? I do this thing in R. I want to learn how to use this package. You write this like skeleton framework using the scripts that you have. There is a lot of interest and I am tentatively optimistic that we can use it to draw insights from our data. I was
[09:34] just talking to to some of my colleagues in the um in the audience before about some of the risks that are involved in drawing incorrect insights from our data at times. So I think there are opportunities to do a lot with it. Um, something that we have talked about
[09:51] that I am very interested in learning more about and I don't know if if Genie or any AI is going to be able to do it is how to parse unstructured health data. So like health notes that we get where you get these complicated doctor's visits. Um, my wife works in the tech space and
[10:07] we're always talking about it too and they're trying to work on it and we're trying to work on it. I don't know if I'm I don't think I'm going to be the one to figure it out in public health, frankly. But I think there's a way that we can use to kind of winnow down some of the work that we do and like optimize the time that we spend. So maybe we'll
[10:23] spend less time manually reviewing everything, only reviewing some. The example, I see we still got 12 minutes, so maybe I'll go a little bit down the rabbit hole for another one of my favorite pet projects around um syphilis.
[10:38] >> I feel like dinner. I mean, this is a dinner time in our house with Robbie. >> Uh, it's a spyroet. It's super cool from that perspective. It's an interesting uh disease, but from a diagnosis and intervention perspective, it's very
[10:54] complex. For those of you that aren't familiar with it, there isn't just like a yes or no test that you can do to determine whether or not someone is infected. At the other end of things, it's very challenging to determine whether or not someone has been adequately treated. If you want to talk about it later, happy to talk about the
[11:10] details, but I'll spare you the gory details now and just mention that the treatment for it is something that is pretty non-specific, right? It's penicellin, but it has to be dosed adequate amounts over an adequate time based on how long someone has had the disease. And it's just very difficult to get that for a human reading those notes
[11:26] themselves. We often have to call the doctor's office to check. So I would love if something could be done like that with a genai tool and maybe it's going to be around certain populations and whatever but um one of the most important populations that we think about with syphilis is children who get
[11:43] syphilis right so congenital syphilis um and the risk of saying that someone was adequately treated and being wrong about that is too great so I think I'm I want to be able to use it for things like that but we're a ways away but for now
[11:58] like the productivity aspects are making it so that we can spend more time on doing those kinds of interventions to actually prevent these really severe outcomes like kids born with syphilis. >> That's terrific. Um Ryan, turning to you, can you talk about some of the meaningful changes that have happened in
[12:14] your ability to track and improve behavioral health? >> Um yeah, so we are right at the kind of precipice of behavioral health as I talked about. we were really facilitating the intake of data from
[12:30] counties and and again digitizing a form. It's not rocket science, but it's transformational in the government space and and now we're looking at better ways to intake data from the counties. And so as we and behavioral health is changing
[12:46] the way that money is being sent out and so we are really um kind of at the beginning of that measure and but again this is the governor's uh agenda to show those outcomes and show that that money has put more beds. I mean the improvisation already has put uh I think
[13:03] it's 60 billion six I'm going to misquote it I should know it but that's programs responsibility billions of dollars into counties that have put you know hundreds of thousands of beds and facilities actually on the ground brick and mortar facilities uh and so we're
[13:19] really doing that I think the one of the impacts though that I will say from an infrastructure out is just our ability to exchange data with partners ERS, right? Um, you know, homeless data is not captured by uh Department of
[13:35] Healthcare Services. And so we had something like 15 different organizations that we had to uh exchange data with and do that on a regular basis. And to be able to do that now, uh, and again talking about how things used to happen, uh, probably no one in
[13:52] this room remembers SFTPs, uh, but we still live in that world. And so uh to be able to actually exchange data in near real time with other government organizations uh specifically ones that are on the same platform uh have been a just a monumental change on
[14:10] our infrastructure outcomes. So uh and I'm I'm hoping uh I know I'll be able to show those from those impacts from the technology lens will directly correlate to the you know demonstration of the outcomes for behavioral health that we
[14:25] wouldn't be able to do because previously we would have been moving data around by hand or you know waiting for a batch job to run. Um, so this was not scripted before, but given Robbie's syphilis example, kind of those the things that he's most excited about in terms of outcomes, is there something
[14:42] that you're looking forward to similarly where you're like, this is actually whether it's AI or just improved data practices that is actually going to drive huge impact whether it's for Californians receiving behavioral health or internally. >> Yeah. So there's a something that we're
[14:59] rolling out uh actually the end of the month is the ability for counties to submit individual service level data. So encountered data, put hands on people, help people out in the field and and so forth. And and again that circulates
[15:14] back to how money is paid and and so forth. But what we're looking to do is is as that data comes in, we're giving the counties that real information and a snapshot of what DHTS knows about you today. And then we're actually going to publish that information. And so for the
[15:30] first time that I know of at Department of Healthcare Services, we're able to actually give them a profile of what they're doing and an understanding of what we perceive as going on in their community. So if they're uploading claims, we can tell them how many
[15:46] claims, you know, historically we have for you for the last 6 months and and build that story as as opposed to, hey, give us a monthly data drop, right? And so just to build those things and that transparency and show that history and the lineage, all that comes down to how
[16:03] we handle the data in a more modern way. And that's what I'm really looking forward to is be able to show them that back to them. >> Yeah, absolutely. Um, so I'm always cognizant when we have big audiences like this, um, that everyone is kind of at a different point in their journey. So you guys kind of have already deployed data bricks. With the knowledge
[16:20] that folks haven't started in some of these cases, what would you want to tell other public sector leaders? Either something that you wish you had known before you started out um, or how to move faster. Any advice that you have
[16:36] for audience members that are maybe just starting out on their data bricks journey? So, I think I mentioned this. I'm I'm only six years at the state. I came from private sector and with Prop One came an
[16:52] exemption uh to move very quickly. Uh and so I've been very fortunate to be able to partner with data bricks and bring them in in in a very quick and short timeline to stand up a lot of what we're doing. Um, I probably moved too
[17:08] fast, right? Uh, in a government environment and so I I probably could have moved slow to move fast, but I do not want that to be the tagline or when you leave here, don't associate that, you know, as my ultimate advice. I think really just if I had something that was
[17:25] valuable across the board, maybe not go slow or go fast. Um, because it starts small, right? In government, there's tons of data silos. there's tons of opportunity to make things better. And what I have learned in my short time in government is usually there's a huge plan and we try to solve all the things
[17:42] at once. For me, it's about go solve one little problem and gain a partnership with your program people. Start small and then it'll evolve and pick up momentum and then incrementally you will show that value to your program partners. And so that's really again if
[17:58] you're going to take something I say it's not the the slow part. It's the incremental and show value to your partners is really where you're going to benefit. >> Absolutely, Robbie. >> Yeah. A lot of things you said that resonated with me at at towards the end you were talking
[18:14] about how important it is to like work with your partners and really understand the problems that you have. Can also relate to the like, oh, you need to like wait to do this thing because the time's not now. And I think like my advice and my experience, I've only been there a couple more years at the state than you have.
[18:30] There's just like never the right time. It's never like the correct time to do it. Just like start to take some of these pieces and move forward on it. Listen to the people who are actually doing things. There are a lot of things that we've been able to achieve and build trust from just like automating
[18:45] processes that took tons of time, right? Like we had multiple epidemiologists who had to babysit a script for six days. Like we talked to them and we just like rewrote it. you know, like there's a lot of nuance in there, but like there's never a wrong time or right time to start on this. And I think the other piece that I've come to terms with now,
[19:04] uh, is that like it's it's never you're never done modernizing your data infrastructure. Like it's always going to evolve. You're always going to have some plan and you're never actually going to achieve what is entirely laid out. And so there's like this balance of communicating what you're actually able to do. But as long as you're able to
[19:19] articulate the value that you gain from doing these things and the way that what you're doing impacts the people that you're trying to serve, I think that that that's really helped us be able to continue to push forward some of these efforts. >> Terrific. Thank you guys both for for joining us today. Um I feel like just as
[19:35] an American, it makes me feel very glad that people like you are are looking out for all of the citizens health so and pushing forward on like when things are easy. like I feel like you're both did not go through all of the hard things, but I know that there was lots of them and lots of obstacles. So, thank you. If everyone will join me in thanking our
[19:50] panelists.
Learn more about the Databricks Data and AI platform.
The information provided herein is for general informational purposes only and may not reflect the most current product capabilities or configurations.