Optimizing Government Data Strategy with Databricks and Unity Catalog
Summary
- Robin Sutara of Databricks defines five requirements for an effective government data strategy: timeliness, trustworthiness, repeatability at scale, affordability, and agility to answer new questions as mission needs evolve.
- USDA used the Databricks Data and AI platform and an AI hackathon to rapidly address an environmental permitting challenge, while DC's Office of the CTO modernized data analytics across multiple city agencies.
- Unity Catalog enables government agencies to unify data across silos without requiring massive migration cycles, allowing rapid prototyping that empowers analysts and mission specialists to answer questions independently.
Optimizing Government Data Strategy with Databricks and Unity Catalog

Government agencies face unique challenges in transforming fragmented data ecosystems into strategic assets. this video outlines how to align data governance with operational mission needs, break down silos, and democratize access across your organization. Learn how USDA and DC's Office of the CTO leveraged Databricks to unify data where it exists, accelerate decision-making, and empower employees from analysts to mission specialists.
The panel addresses the gap between data collection and actionable insights, demonstrating how a platform approach with Unity Catalog enables rapid prototyping without requiring massive migration cycles. Discover how infrastructure becomes invisible when it's standardized, how workforce empowerment drives scale, and why data access is the actual bottleneck limiting agency velocity.
🤝
Chapters
00:00Optimizing Your Agency's Data Strategy: Overview00:25Data Strategy Framework and Requirements02:02Gap Analysis: From Question to Answer04:41People Empowerment and Workforce Scaling07:05Consolidation Challenges in Government08:08Unity Catalog: Unifying Data Across Silos14:09USDA Case Study: Environmental Permitting Challenge18:17AI Lab and Hackathon: Rapid Transformation24:48DC OCTO Data Modernization Journey30:41Multi-Agency Deployment and Results35:45Future Vision: Security Analytics and AI Agents
FAQs
What are the key requirements for a government agency's data strategy?
According to this video, an effective government data strategy must be timely, trusted, repeatable and scalable, affordable, and agile. If data cannot be delivered quickly, if employees do not trust its accuracy, if generating each answer requires excessive manual effort, or if the strategy cannot adapt to new questions, the organization cannot extract sustainable value from its data.
How did USDA use Databricks to address its environmental permitting challenge?
USDA leveraged the Databricks Data and AI platform along with an AI lab and hackathon approach to rapidly prototype and deploy a solution for its environmental permitting challenge. This approach allowed the agency to demonstrate value quickly rather than waiting for a lengthy traditional implementation cycle.
How does Unity Catalog help government agencies unify data without massive migrations?
Unity Catalog allows organizations to govern and access data where it already lives rather than requiring all data to be physically moved into a single repository. This makes it possible to unify governance, access controls, and lineage across existing systems, enabling rapid prototyping without the time and cost of large-scale data migration.
What did DC's Office of the CTO achieve through data modernization with Databricks?
DC's Office of the CTO undertook a data modernization journey using Databricks that enabled multi-agency deployment of shared analytics capabilities. The initiative focused on democratizing data access across the city's agencies and building toward future capabilities including security analytics and AI agents.
Full transcript
[00:10] My name is Robin Sutara. I'm the field CDO at Databricks. What that means is I get to work with some amazing organizations and across the world as they think about how to actually execute against their data strategy. So I'm going to share a little bit about some best practices, factors to consider, and then we're going to
[00:25] hear from Freddy from the USDA as well as Matt from the office of the CTO with the city of DC. So super excited to let them share their stories. Let me start with a little bit of what is a data strategy. I think for lots of organizations they think about their
[00:41] data strategy in a couple of components. Most often I see them focus on things like the architecture, the use case, the data governance strategy, the master data management strategy. But really your data strategy is meant to help the organization achieve the mission by
[00:57] answering the right questions with the right answers. Right? And it's not just about speed, it's also about accuracy. How many of us have ever had somebody within the organization say I don't trust the data? I'm not sure the data is of quality, I'm not sure that I can actually use it. So it's not just about speed but also how are we thinking about
[01:13] accuracy because if people don't trust your data, they're not going to use it. So then there's no value in the data strategy to execute. It has to be repeatable or scalable, right? You It cannot require excessive effort on the part of your teams to be able to deliver every answer. If every answer is taking
[01:29] you weeks or hundreds of people hours to be able to deliver, you can't scale that across the entire organization. And it can't be too expensive. It can't actually break the bank. In today's world we're all being asked to justify our budgets and if every answer requires a significant investment, we're going to
[01:46] run out of funds before we can actually deliver. The fourth factor is it has to be agile. The question of today is not necessarily the question of tomorrow. How do we make sure that we can continue to have a data strategy that allows us to answer questions as the world is changing around us. So if we think about it it's
[02:02] from a timeliness, it's from a trust trust trusted, has to be repeatable or scalable, affordable, and then agile. So those are sort of our minimum requirements around our data strategy. So if we know that the gap is from getting from question to answer, we have to think about how what are the factors
[02:19] that go into closing that gap. The first thing is your data. Right? We know data is messy. We know that data is unstructured and structured, it exists across multiple silos, it's buried in legacy systems. There's net new data that's coming into into place. Most of us have multiple tools across our
[02:35] ecosystem that we're looking to deliver with and those tools overlap. We have storage that's being underutilized so we're not leveraging our budgets effectively. And when someone on your team says, "Why can't I just get that number?" Who's ever gotten that question? You guys have to participate a little bit with me. Who Who has gotten that question from
[02:50] somebody in the business like, "Why can't I just give me the number?" It's because it requires us to actually aggregate across multiple ecosystems, multiple tools, multiple silos, actually collate it, bring in four analysts, and by the way one of them is retired on Pebble Beach right now, right? You just
[03:06] cannot get the information pulled together to give them the number. And that's not a failure on the part of of you or your teams being able to deliver against that data strategy, it's the reality of the world that we live in. So here's the opportunity that people miss. You We have to think about it as data is an asset asset. You know, all
[03:23] too often we hear about data where you've made these investments already on collecting it, on storing it. And the question isn't data is valuable, we know the data is valuable, but it's how quickly can you unlock it? How quickly can you give it access across
[03:38] the organization? And that's why most agencies are actually seeing funding of net new data sets being attributed within the organization creating more silos across is because they're just trying to unlock the data to do their job and serve the mission.
[03:55] That's a solvable problem, right? Lots of organizations have actually made significant progress here. They've worked on consolidation, worked on migration. It's all about getting it all together in a single place, right? The single source of truth that we're trying to build out for the organization.
[04:11] The problem with that today is that it's not scalable, right? And as the world continues to evolve around us, as technology continues to evolve, that becomes really difficult. A single source of truth becomes really difficult to maintain. It's not like the old days of data warehousing where we just threw
[04:26] up another SQL server, right? To be able to accommodate it, but it's about how do we work with the data where it exists today? How do we mitigate this having to move data into our single source of truth to have it be accessible across the organization? The other thing we need to think about
[04:41] is people. I mean I think often times the Databricks sellers hate me because I love to talk about people. I think that's your biggest asset. That's really where you're going to unlock the power of your data strategy. Today your analysts are probably spending most of their time just trying to wrangle the data, right? Get it into a usable format. Institutional knowledge of how
[04:58] those data sets sit together actually sits in their head, it doesn't sit in the technology or the processes that exist. There's a real opportunity for us today to think about how do we flip that ratio? How do we give the team access to the data that they need so that you can preserve the knowledge and make it
[05:14] accessible multiplying the capacity that your team can actually deliver. So democratization isn't about efficiency, it's about scale. How do we scale that expertise across the agency? So here's the path forward as we continue to mitigate some of this complexity as it grows, we can think
[05:30] about how do we invest in our data and our data strategy as much as we've invested in the technology to make sure it's ready. The technology is ready, the patterns are here. We're going to hear some great examples from two organizations and so there's a real opportunity. So let me just give a little bit. If we
[05:45] think about you know, mitigating some of that gap, we talked about the what, now let's talk about the how. So if we think about optimizing your data strategy, it's really rethinking how you see data. We've all taught heard them talk about data as an asset, data is the oil, data is gold, whatever
[06:01] analogy you sort of want to use. Really I'd love for you to think about data is your mission, right? It's not separate from what you do, it's an actual record of what's being done to achieve the mission. Thus it's part of the mission. And so if we think about that sort of
[06:17] context as we think about our data and our data strategy, we don't want it to sit in a silo. We don't want it to be segregated from the things that we're doing. We don't want it just to become another cost center that we don't actually get value out of. So how do we think about a platform approach that actually allows us to take that dormant
[06:34] records of the activities that are being done across the agency and turn it into mission velocity? So instead of weeks to get a response or an answer to those questions that we have, we can mitigate it down to hours. We can actually start to make decisions based in evidence and data instead of
[06:49] instinct, which is often times how we operate. Your people have done a lot of hard work I'm sure over the last couple of years thinking about doing that and so how do we now get that insight out of the systems to be able to do that? So how do we actually unlock that?
[07:05] So the first instinct that most organizations or agencies have is consolidate, right? Continue this consolidation motion. How do we break down these silos? How do we get rid of the fragmentation? We have scattered systems across. Let's start with a clean slate and we'll do a migration and we'll get it all into one single source. And
[07:21] that makes sense in a document or a whiteboard, but the reality is for most organizations that takes you 18 months to plan, takes you two years to execute if you're lucky and you don't run into any migration issues. Takes you six plus months to actually transform the culture
[07:36] of the organization to use the new data and platforms that you set up. And by the way the data models and everything else that you set up three years ago don't apply anymore because everything's changed. The mission has moved, the world has moved and now the mission has to keep up with that that change
[07:52] that has happened. So you didn't actually eliminate the silo by pulling all these things together, you've just created a bigger silo that existed before. And the problem is it becomes almost dated or legacy by the time you actually get it stood up. And so what's the alternative? We think the alternative is not moving the data but
[08:08] thinking about unifying the data where exists today. And this is what we think the power of Unity Catalog offers you. How do you leave it in place but you still have the ability to govern and track manage the access, lineage, quality and compliance in a single place without actually having to physically
[08:24] move the data into a different source of truth. So the goal isn't actually relocation, it's cohesion. How do you actually allow your entire platform to operate as a single platform without having to rebuild anything that's already outdated by the time that you launch it.
[08:42] So let's also talk about people. Like I said, I think you're they're your most valuable asset, it's your best line item to actually get scale. We often think about scale from a technology perspective, but technology isn't there for just technology's sake. Most of the time if I see an agency or an organization think focused only on the
[08:59] technology and not thinking about the workforce and the people, it becomes obsolete before you can actually ever get it to execute and you've built a great field of dreams issue where you have an amazing platform and tools and no one ever uses it to achieve the mission. Now how do we think about actually
[09:15] scaling your workforce? Most of us are being asked to do more with the same or more with less. How many sort of have to think about that these days? I know I do. I just got asked last week like how am I going to do more with the same number of resources. And so how do we think about scaling people not just
[09:30] compute or your architecture? And so for most organizations right now data access is the actual bottleneck, right? An analyst needs information, they submit a request. Your data engineer who already has is super over overloaded has a phenomenal backlog that they're already trying to work through.
[09:46] Eventually we'll get to it in weeks or months. Meanwhile the analyst has either already made the decision without the data or they've stood something up and they tried to work around the fact that they couldn't get the access that they wanted to. Now we have our senior data engineers who are actually working on things like building reports or writing simple SQL
[10:03] queries. All of the things that really if we gave democratization or access to the analysts in the first place, we could help break that cycle. And we think a uniform platform unified platform can help you do this. How do we make sure that analysts can actually access the information directly? How can they
[10:19] answer their own questions? How can we now free up our most senior technical resources to focus on our most technical problems that we're trying to solve? And that allows us to actually start to think about the aggregation of headcount across missions. So, instead of having an engineer assigned to just one
[10:34] program, could we actually have that engineer, if they're not having to solve every analyst issue, work across three programs? How do we now start to get efficiencies of scale of our most technical resource? For most government entities, you are not going to be be able to compete with the private sector when it comes to
[10:50] acquiring new talent. So, how do you get the most value out of your existing capabilities and talent that you have on top of the team? And so, once you have your people ready to move, we think that the Databricks data intelligence platform is the best platform to be able to help them move across that. We think about it from the
[11:06] data unification as opposed to a data consolidation point of view. How do we make sure that you don't have to move everything into one place? That you actually have the ability to make it discoverable and governed and usable while it stays in place where it is. And in this multi-year migration cycle that
[11:24] we've all been stuck in for decades. Every organization wants to be data AI organization. It wouldn't be a pitch if I didn't talk to you about AI. I don't think you go to any data event without hearing about AI. But what I will tell you is that good AI depends on good data, right? And
[11:40] so, how do we think about solving that with the platform? By giving access it we need to make sure that the data quality is of that that AI actually becomes practical, right? That the organization can use it. And that you can set that governance, the security, the guardrails in a unified way across
[11:56] your data assets and your AI assets, so that it's already built in. It's not a bolt-on that you're trying to add to the platform later. It gives you the capability to be able to deliver AI with all of the structure that you've worked so hard on your data platform to be able to do. And multi-cloud is definitely something
[12:11] that we think about is how do we make sure that you're portable? So, as you think about cost savings and efficiencies and all of those aggregations that you're looking to make, we want to make sure that we have the right foundations in place for you to be able to deliver against those AI initiatives. Not add more tools to your toolbox, but how do you best leverage
[12:28] the tools that you have today? And so, I'm sure you're tired of hearing me talk by now. So, what I would love to do is invite out a couple of customers who have actually leveraged the data intelligence platform for their data strategies. They're going to share a little bit with you about the good, the bad, and the ugly as they've gone
[12:44] through that process. So, let's go ahead and start with Freddy Diaz from the USDA. Please join me in welcoming Freddy to the stage. Okay. Good morning, everyone. How you guys doing?
[13:00] So, I was told this was a breakout room. So, in my mind when practicing this, I thought this was like yay big, not yay big. All right. So, my name is Freddy. I'm the Deputy Chief Data Officer for USDA. And I'm going to just tell you a story about our very early and nascent
[13:15] journey of using Databricks. Robin talked about empowering people. So, this is going to be a story about empowerment. Throw in some infrastructure there, but I think that's really going to be the story I'm going to tell today. So, a little bit about um my journey. I've been at USDA about 3
[13:31] years. And I learn, feel like every day, uh what USDA actually does. Functions like a bank. We are a bank. We also give out public benefits like SNAP, WIC, school lunch, school breakfast program, as well as things like the Forest
[13:46] Service. I didn't know that Smokey Bear was an employee of USDA before joining. So, there's so many other functions that I've learned during my short time, relatively short time at USDA, that you have to think about that when developing your strategy and making sure that it fits all things for all people.
[14:09] I'm going to tell a story about not any of those things in particular, but about environmental permitting. That sounds a little boring and you're probably like I I need to go to lunch at this point. But really, I think it's a small but very powerful example of the type of challenges that we have that I think are relevant to other agencies as well. So,
[14:25] just like with anything, we have a process where we have to go ahead and have farmers and ranchers and those that need a permit to file some paperwork. Ideally, in a data-driven world, this would be a like Domino's Pizza Tracker type of
[14:42] thing. A lot of data unified and a decision made in minutes, potentially hours, but not the weeks or months that it can take for this process. Right? So, I think we're falling short and that's something that we've looked at and we're looking to improve. Well, this is not just a story about one
[14:59] specific process, but rather this is actually an example of what happens across the department. Whether we're sharing data within our agencies, we have data silos like I'm sure many organizations have. Also, when sharing data outside of the department, there's issues. And also,
[15:15] when it comes to unifying data for other purposes, this is where we just have problems within our agency. But at the end of the day, it's the mission that suffers. It's not just that it takes somebody more time, it's that ultimately we are not living up to the realization of what we can be as a
[15:31] USDA with some of the data silos that we have. I'm going to give a graphical representation of what this might look like. So, doesn't matter what what your cloud of
[15:48] choice is, AWS, Azure, Google, we have on-premises, we still have mainframes at USDA. But that's not the point. It's more about infrastructure connectivity or the lack thereof. The lack of that infrastructure connectivity is really is really what's hurting us. Right? So, not
[16:05] having an interoperable standard between our different data silos and being able to move data in between, this is where we have data copying come up. This is where we have many issues happen because we don't have a standard across the board, right? It's just a lot more
[16:20] difficult to unite the data in any way, shape, form, or fashion if we're not able to unify things all in one spot. Unfortunately, that means our analysts, our data scientists, our practitioners are spending more time wrangling the infrastructure and trying to get that aligned than they are doing the actual
[16:37] mission, the actual work itself. And by the time you get to the mission, you're already tired. You spent weeks or even months in some cases just getting to the point where you can actually do the work that you were there to do. The thing that you raised your hand, especially those of you as federal employees, and took an oath of service,
[16:53] it's not to wrangle this infrastructure, it's to do the mission of your organization, right? So, that's where we have our issues and that's where we're trying to stitch together through infrastructure connectivity.
[17:11] So, in 2024, our Chief Data Officer put out a USDA data strategy that really reinvigorated, reaffirmed our direction. So, this was a update to our inaugural data strategy where we were focusing on unified governance, modernize analytics, open data architecture, and just making
[17:27] investments in those areas, including our workforce. However, if our foundation is a house built on sand, that's where we can have some issues and some problems. Right? So, how good can our strategy be in execution if we have some issues like
[17:44] custom integrations and dependencies that we have to work on that are still prevalent to this day. So, that's something that, you know, just we have this strategy, we need to make sure it was not dead on arrival by putting it out. We had to make sure that we were able to execute and make sure we
[18:00] realized the vision of the Chief Data Officer and the organization as much as we could. So, back to that environmental permitting example. So, earlier this year, uh the president signed an executive
[18:17] order that, to paraphrase, said, "Use AI and technology to expedite the process for environmental permits. Turn this from what it is into something better. Make it faster. Make it better." So, we had really a choice to make as an organization. We can either continue our
[18:34] current path of how we're doing business or we could try something else. We could try a different approach. And this time, we chose to make the infrastructure and the data almost invisible. We already had that taken care of with a small investment in
[18:51] our AI lab, which is our AWS commercial Databricks instance, and said, "What if we already provided all of this?" And as an experiment, went out to the workforce and said, "What if you can redesign this entire process to get to
[19:06] an environmental permitting decision in 30 days?" So, 30 days for them to go from the inception to actually developing it. So, we did an internal hackathon with our employees. Some of those employees are here today, which is great.
[19:21] So, and what we what we had as a result of that is some amazing work. So, we had, and I think the more important thing before I talk about the work is the types of people that I think raised their hand to be part of this challenge. These were not just IT specialists.
[19:39] These were not just folks in environmental permitting office. These were folks that felt the need to come to this challenge and provide a solution and took some of their time over the summer, over a 30-day cycle, to actually contribute across our various mission areas of USDA to all work on one problem
[19:57] for that 30-day cycle. So, as a result, we had something that instead of taking weeks or months to make a decision, we now had prototypes that had some viability to make decisions in minutes, worst case hours
[20:12] in some point. So, a drastic difference from what we had before. The best part is is that we have a path to production on what we're doing. Since we're using code and we're leveraging the Databricks environment, now we have a path to take that and actually put that into production so that that farmer
[20:28] and that rancher doesn't quite have that Domino's Pizza Tracker for everything, but at least the decision is not something that will take weeks or months. But, the question is what changed from our approach? And really it's a a just a shift. It's not just pre-staging data
[20:45] and making Databricks available. It was making sure that we had all of that kind of pre-figured out and ready to go and empowered our workforce to be able to do those things.
[21:01] So, we took a step back and really thought about what that meant. And to make sure this was not a fluke. And what we thought about is our infrastructure should not be front and center. Instead, it should be almost invisible and be like a USB port where it works, it's
[21:16] standard, it works across the different cloud service providers. And it's only when it doesn't work that it becomes an issue, right? So, you usually plug in something in your USB port to do something, charge your phone, transfer data, do some function,
[21:32] right? It is not front and center, but it provides a very essential function. I think my phone's on like 20% right now. So, like we have to make sure that it is not stealing the show, if you will. So, if this works effectively, we can get to the mission, which is really
[21:48] the star of the show. How do we fix the things that are preventing us from getting to mission outcomes? So, our approach is really thinking about how infrastructure can connect across clouds and get to the point where it is no longer a barrier and now we're getting
[22:03] to the right questions, which is not whether I'm using cloud A or cloud B or tool A or tool B, but rather what are the pain points that the business is having and how do we get to answers quicker and faster.
[22:19] So, what's next for us? And I think this is a very important part of our journey right now is that we're at an inflection point. What we don't want is that that hackathon that happened over the summer to be a fluke and to be a one-time a one-off event. So, what we're doing now is with the help of our partners, we are
[22:35] scaling our AWS Databricks footprint across the entire department to make this available to as many practitioners as possible through our enterprise data platform. That's the first thing. Second thing is going to be to focus on key focus areas
[22:51] that are important for the department. Things like loan monetization and fraud, waste, and abuse detection to commodity grading and some of the other critical functions that we have at the USDA. So, now we can actually get to providing those services in a better, faster, and
[23:06] even cheaper way than we have in the past. That way we can empower the workforce to focus on those things and make our data strategy come to life. So, I want to just say thank you for for listening. I do want to say one
[23:23] thing about our journey is that we are still early in our journey, but I think it is a very good representation of where we plan on on going over the next 12 to 18 months, really empowering our workforce using tools like Databricks, but really putting the tools and the data in their hands versus spending lots of time
[23:40] worrying about that messy spaghetti map that I showed earlier, making sure that we can get past that and start focusing on the mission at hand. So, just want to thank you and if there's any questions, feel free to get me during the conference. Thank you.
[24:02] I love that story only because I think it's a great example of how the USDA not only thought about it from a technical perspective, but from a business perspective. How do you do a hackathon across your entire agency where you're actually bringing people that are having to do that type of approval work into the solutioning. And I think we see
[24:17] amazing progress across across multiple agencies when they bring those types of groups together on those types of hackathons. With that, I would love to introduce now to Matt Sokol, who's the Chief Data Officer of Office of the CTO or OCTO of City of
[24:32] DC. Matt? All right. Thank you everybody. Good morning and here to talk about DC's data monetization journey with Databricks. So, again, I'm I'm Matt Sokol, Chief
[24:48] Data Officer amount of DC's Office of the Chief Technology Officer. I just quickly what I want to talk about today. I want to talk about just who OCTO is really quickly. I want to talk about where we were, again, this is part of the journey. What were some of the goals of the journey we had and then
[25:05] where are we now? Where are we going? I don't have time for questions, but I'm happy to stick around and some of my team is here as well. So, what is OCTO? We're we're the Office of Technology, right? We're the centralized service provider for IT in the city, right? We do data, we do apps, networking, security infrastructure, all
[25:21] your common IT operations and digital services. We serve lots of people, lots of agencies here in the district. We actually do provide services to the federal government, also nonprofits and ultimately our goal is to better serve the residents of the city.
[25:37] We focus on some things like procurement, provide enterprise solutions and we talk about technology policies and standards. So, that's that's us in a nutshell. But, let's talk about in the past, right? How do we How do we start with all this and kind of get to where we are today? So, back in 2017, the city made a
[25:54] capital investment in in big data tech, right? We had money, we went out and bought this really cool bare metal hardware that we threw in the data center and we bought some proprietary, you know, Hadoop software that we used some open source with, too. So, we tried
[26:09] to give everybody a little bit of something, you know, in the big data stack, right? We had a a need or a start to thinking about how we did IoT in the city, right? Lots of sensors being deployed and trying to figure out what was the best way to get information from all over the
[26:25] city about various things, whether it was air quality or how people were accessing spaces or how vehicles were even moving about. Um and all of that resulted, right, in a lot of unstructured data, you know, where we are today, that's it's not just traditional relational databases. We've
[26:41] got semi-structured, unstructured. We've got video, we've got pictures, we've got lots of streaming data, right, constant data flowing in. Um and we also had an emerging data science community, which really was interested in things like parallel processing and how did they take all
[26:57] this data and do something really cool with it, right? But, what what happened? Um lot of overhead, right? We had staff that had to be sort of working on this full-time to keep it up and running, making sure we gave people the tools they needed. It was a complex learning
[27:12] curve, even some of the data scientists couldn't figure it out, even some of the the folks who were the security, you know, we had MIT Kerberos in our back end, right, serving as our security platform. So, some of these things were complicated. But, I want to talk a little bit about
[27:27] the journey goals and think about some of the things you've heard today about Unity Catalog, Delta Sharing, AI/BI Genie and you'll see how these things those products and services from Databricks tie into these goals right here. So, we wanted to deliver an enterprise service across the district, right? We
[27:43] needed to make something that we could deploy, we could manage, and we could give ultimately services that everybody can invest in and make use of. And of course, breaking down silos, nobody's ever heard of that in government before, right? There's no silos, it's not a real thing. Doesn't
[27:58] exist. Self-service, right? We need people of different walks of life in the district to be able to access this and from various places, whether you log into Databricks or you're building a dashboard in Tableau and Power BI or utilizing tools like Box.
[28:14] We need those people to be able to do that themselves, right? We can't with a with a staff of a couple, we can't hold hand handhold everybody through every single step all the time. And they don't probably want us to do that either.
[28:29] So, integration with other DC enterprise platforms, right? Databases, file transfers, flat files, other cloud providers, right? So, having that ability to have native connectors all over the place to connect to the other enterprise services that we offer
[28:46] today was really key for us. Um overarching governance and security. So, as OCTO, where, you know, security lives, our role is to really provide security guidance platforms. As the data guy, I have to talk about governance and how
[29:02] that works and how we apply it across the city. So, again, having this come out of OCTO really provides us the overarching governance and security on a platform like Databricks that has Unity Catalog and some other things around role-based access and attribute based access. So,
[29:18] those things are really key for us and how you saw earlier about limiting what people can see based on their roles and their their knowledge and and what they're supposed to have access to. But, we really want to have across the city this single pane of glass, right? We want to know what data's out there, we want to know how it's being used, we
[29:35] want to know the metadata, we want to know the lineage, and we want other people to be able to find that data to a degree, right, know it exists and be able to talk to other agencies about using it and understand how they can use it and how good it is. Um and then supporting the DC data workforce, right? It It's a wide range
[29:52] of people. We have business analysts, right? Who have very specific functions in their agencies. We have data analysts who look at data across the spectrum. Um some of them are GIS people, some of them are regular database people, some of them you know, they they have all different skill sets. We have data
[30:08] engineers. Um they build ETLs. They build pipelines, right? We have data scientists and they love to not click on buttons. They love to type code. They love Python. They love R. They love complex analysis between data sets and
[30:24] they like their raw uncurated data to do things with as opposed to other business analysts, data analysts that love that kind of well-curated gold layer. So, we have a lot of people we need to support. So, let's talk about where we are now,
[30:41] right? So, in 2024, we started planning on our migration to Databricks. Right? So, how did we go about this? We evaluated the platform's capabilities. So, again, some of the things we've been talking about, the governance and the security and how well it integrates with other things, but we also sort of
[30:57] combine that a little bit with user needs, right? So, everyone in the district is in a little bit of a different place with data, right? Some people are looking at data warehousing, right? Some people are kind of in the fundamentals of of data. They don't have a lot of staff, but they need to kind of unify all this. They need to do
[31:13] reporting. Uh there's other agencies that are much further along and they're building ETL processes and they're building pipelines to other and they're using APIs, right? But what we thought about this platform is it gives a little bit of something to everybody.
[31:30] So, it didn't really matter where you were in your journey uh overall within your agency, but if you want to do this, you can do this. If you want to do that, you can do that, right? So, we're giving everybody a little bit of something. Um so, we're not really excluding anybody and you can find value in the platform and move through your own
[31:45] journey within an agency using Databricks. So, again, like why did we do this? This had a lot of technology that we thought was was beneficial, right? Um again, it had AI. Uh it's got connections to business intelligence. Um all that security based thing we
[32:01] talked about. So, we we went here because we thought this provided a good kind of, you know, ecosystem for us to work in. And it was it's really much lower overhead than what we were experiencing in our in our previous data lake and big data adventure. Um a lot
[32:17] less, you know, we can get and we can provide some overarching policies and then we can let the folks who need to do the work do the work. Uh we we don't need to give them permission to use Python or R or build notebooks or build pipelines. They can do that if they need to do it. Um and we don't have to like sit there
[32:33] and monitor them. Um and we kind of have this hybridized governance and security model. So, what happened? Ultimately, we went with Azure Databricks. We're we're an Azure shop, so that's where we are. Uh we built a process through which we
[32:48] could automate the deployments of Databricks to agencies. So, we used tools like Terraform to build uh models that we could actually automate the deployment for people, um set them up with policies, set them up with governance, come up with uh a
[33:04] structure for how they would fit, whether they were an administrator or an analyst or an engineer, right? So, they get certain roles and and um things like that. So, that's where we are, you know, right now. And additionally, you know, within the
[33:21] last, I'd say, 9 months, we've onboarded lots of agencies. So, OCTO's in it, Department of Transportation, Deputy Mayor for Education. I could read them all. Homeland Security, Department of Buildings, and we're still expanding. So, we are talking to the agencies about
[33:36] the benefit. They are coming to us asking how they join because they are seeing the successes of other agencies using Databricks. Some project successes. Uh probably our marquee one uh is with the Deputy Mayor of Education. So, there is a program called
[33:53] Education Through Employment Pathways. Um for those who familiar with P20W systems, they are longitudinal data systems that look at education and workforce data. And so, the idea is that over time, they
[34:08] can see how education in the district impacts economic mobility, impacts how they are in the workforce, right? And so, we are actually building a record linkage system on top of Databricks. It's pretty cool. Um OCTO is doing some things. We provide
[34:25] call centers through Amazon Connect. There is a lot of data that comes through those call centers. We are pulling all of that call center data for each agency into Databricks. That allows them to examine metrics and KPIs on how their call centers
[34:40] performing. But not only that, but we also now have a single unified view of the AWS Connect data for the city. So, we can provide city leadership dashboards, metrics, SLAs, KPIs on how the city is performing overall for their call centers.
[34:58] Uh there's other ones with DOB that's partnering with the with um DC Water and they're looking at billing analysis uh for water bills between buildings that the district, you know, provides and housing um and and the water company. So, we're also looking with uh Department of
[35:13] Licensing and Consumer Consumer Protection. Uh unification of disparate data source systems, right? They have a lot of systems that require, in order to get a license for a business in the city, that you have to go through and, you know, be in a good state um to get
[35:29] that license. And so, they are looking to speed up the business licensing validation product uh process by unifying all that data within Databricks and more quickly being able to process that information. So, I just I'm not on much time left, but I just want to talk about the future
[35:45] and like where we're going and what's on the road map for us coming up. Um we're going to continue to onboard agencies. Like I said, we have many other agencies asking us to to come on this platform and then begin the process of easily sharing data through things like Delta Sharing.
[36:02] We're looking at a full replacement of our citywide data warehouse in a more modern technology using things like lakehouse and using things like the lakehouse. Um we're actually investing uh a proof of concept in a security lakehouse. So, we have tons of security data from tools
[36:18] like Splunk. And so, how do we look at that information and to do predictive analytics on the security side. So, how do we how do we be more proactive than reactive, right? Uh and then there's a lot of talk about AI. Uh starting to figure out agent
[36:34] bricks and figuring out how as a city we can leverage reusable agents. Maybe they have different supervisory agents, but we have these underlying agents that could be used across various agencies. So, that agentic AI concept. Um so, we're excited, you know, we
[36:50] there's a lot of a lot of things going on here and honestly, if you've followed Databricks on LinkedIn and other places, there's like new stuff like every day. It's kind of hard to follow. Um so, we're excited about what the platform can do for us and where it's going to go and whether or not we provided some base fundamentals versus some very cool
[37:07] technology, we're excited about where it's taking us on our modernization journey. So, um as I mentioned, I will be around for a while. I have folks from my team that are here that are doing the work actually. Um and there's some other folks from DC agencies around if you have questions about what they're doing.
[37:23] Um so, I I'm happy to uh thank you for having me today. Uh really appreciate. Hope you enjoyed the the presentation and you enjoy the rest of the conference. Discovery and sharing, I think we've been talking about that for decades. So,
[37:39] it's super exciting to see some progress being made. As uh please join me in one more time in thanking both Matt and Freddy for sharing their stories. As they both mentioned and whether they agree or not, I'm committing them both to stay at least through lunch. So, they'll be available if you do have questions since we didn't have time
[37:55] during the session. Feel free to catch them. This is the lunch time. It is on the other side of the expo. Uh so, feel free to grab some lunch. Please ask Freddy and Matt or their teams any questions you have about the presentation, but thank you very much for joining us.
Learn more about the Databricks Data and AI platform.
The information provided herein is for general informational purposes only and may not reflect the most current product capabilities or configurations.