Enterprise Data Governance at Scale: Mesh and AI Readiness
Summary
- General Motors transformed data governance from a centralized bottleneck into a scalable capability by adopting a domain-driven data mesh with Terraform automation, three simplified personas, and governance embedded directly into Databricks with Unity Catalog.
- The unified data readiness pipeline includes automated classification, attribute-based access control, quality monitoring with Anomalo, and lineage tracking, with Genie workspaces enabling self-service access for teams across the organization.
- In 12 months, GM doubled classified data assets to 8 million, achieved faster onboarding, reduced regulatory risk, and positioned all governed data as AI-ready, demonstrating that scalable governance requires both technical automation and cultural enablement.
Enterprise Data Governance at Scale: Mesh and AI Readiness

Enterprise data governance struggles to keep pace with organizational scale and AI adoption. Traditional centralized models fail when teams operate across domains, and inconsistent processes make it impossible to maintain data quality, compliance, and trust. Sherri Adame, Head of Enterprise Data Governance at General Motors, transformed governance from a bottleneck into an automated, scalable capability by embedding governance into Databricks and adopting a data mesh architecture.
Learn how GM implemented a domain-driven governance model with Terraform automation, simplified personas (domain owners, data engineers, stewards), and a unified data readiness pipeline including automated classification, access control, quality monitoring with Anomalo, and lineage tracking. Discover how team enablement, self-service tools, and measurable outcomes drive adoption. In 12 months, GM doubled classified data assets to 8 million, achieved faster onboarding, reduced regulatory risk, and positioned all governed data as AI-ready using Unity Catalog and Genie workspaces.
🤝
Chapters
00:00Introduction and General Motors03:26Scale, Challenges, and AI Governance Needs06:27Three Failures and Data Mesh Architecture11:15Implementation: Domain Accountability and Control Plane14:31Team Organization: Capabilities and Governance Engineering16:40Democratized Access, Enablement, and Culture21:30Three Actors and Simplified Personas23:57Data Readiness: Classification, Access, Quality, Catalog29:56Progress: 12-Month Results and Measurable Impact33:44Business Outcomes and Future Priorities
FAQs
How did General Motors transform data governance from a bottleneck into a scalable capability?
General Motors moved from a centralized governance model that could not scale to a domain-driven data mesh architecture with accountability distributed to domain owners, data engineers, and data stewards. Terraform automation handles provisioning and configuration, while Unity Catalog enforces governance policies automatically, removing manual bottlenecks and enabling teams to self-serve within guardrails.
What is the data readiness pipeline GM built on Databricks?
GM's data readiness pipeline is a standardized workflow that every data asset passes through before being considered enterprise-ready. It includes automated classification of data sensitivity, attribute-based access control configuration, data quality monitoring using Anomalo for anomaly detection, and lineage tracking through Unity Catalog—ensuring that all governed data is also positioned as AI-ready.
How did GM measure the success of their governance program after 12 months?
After 12 months, GM doubled their classified data assets to 8 million, achieved faster onboarding for new teams and data sources, and reduced regulatory risk across the organization. These measurable outcomes helped demonstrate the business value of investing in governance infrastructure at enterprise scale.
How does GM use Databricks Genie within their governance framework?
Genie workspaces provide GM employees with self-service access to governed data through natural language queries, enabling teams to explore data assets without requiring engineering support. This capability is enabled precisely because the governance infrastructure—classification, access control, quality monitoring—ensures the data Genie surfaces is trustworthy and appropriately permissioned.
Full transcript
[00:08] Um, we're right on time. It's 4:10 right now. I do want to take a quick poll because I'm super curious. Well, one, I think you guys are all super cool because you are in a governance topic and this is by far I think the most important topic that that we could possibly have at the data and AI summit.
[00:24] But how long um who's been doing data governance the last year? What about the last two years? How about five years ago where you're doing data? Oh, couple hands. Yeah.
[00:39] Keep going. Keep going. 10 years. 10 years. 15. 15. Awesome. I've been doing data governance about 15 years as well. So, um it's definitely a passion of mine. It started when I worked at a large banking
[00:56] organization called HSBC. I supported an an um account servicing platform. We built standard interfaces um across all the different systems that we had in HSBC. And after we had released um I don't
[01:13] know 50 or so services I realized that we hadn't collected the data that we needed to collect in the right way or shared the data we needed to share in the right way and that was about our customer interaction. So that was around the time CRM was a big deal. Um so that
[01:31] was really when the data bug hit me. From there I moved to a healthc care organization. So I went from a regulated to what I call a highly regulated organization. Uh I focused on data governance in that um healthc care organization as well. I was there when
[01:48] we did HL7 implementation which is the interoperability for the uh Medicare information. And if for folks that are not familiar essentially there was an executive order that said you had to give to somebody
[02:04] who had Medicare all their document medical documentation including the handwritten notes from the doctors. You had to make that transparent and available to them in a data package. you had to you they had to be able to download. Um I also worked at a
[02:22] billion-doll distribution organization. They distributed electronic components. Um I rolled out information security policies there because I tried to stand up data governance at that organization but they didn't have ins
[02:38] policies there yet. So we did the first version of that and then it did a fast follow with a business glossery as well as um rolled out standard enterprise level reports and we took the generation
[02:53] of those reports down from six months to inside of four weeks. So like I've been living and breathing data governance for a while now um about it's been three years since I've been at General Motors. So about three years ago, I got bored
[03:10] with the healthc care stuff that I was doing and I was like, I want to try something new. And I'm a firm believer that data is data is data. Um I've been across many industries. I've been in highly regulated kind of regulated organizations.
[03:26] So um I saw this role open up for um software defined vehicle data governance lead and I was like I'm going to try it. So that's when I made the jump to General Motors and what I'm going to talk to you today about is the data
[03:43] governance that we've stood up. So this is real story. We've built it into a capability. Our primary uh data platform is data bricks. We're primarily an Azure shop although we are multicloud. So in
[03:59] the last six months we've sto stood up a AWS deployment of data bricks and we're in the process right now of standing up a GCP deployment of data bricks as well. the data governance tools that I'm going to talk about in the framework that we have is what we're going to be deploying
[04:15] across all those clouds and I'll tell you about what's working now um and then what the next set of priorities are to realize um AI governance.
[04:30] So, um, how many folks are in organizations that are pushing forward with AI tools like they've unleashed them within the organization? Like everyone, it's terrifying, isn't it? And it's so
[04:47] hard to keep up. Um, so, uh, if you think about your grandma's data governance, right? It was like a centralized queue and you had these people that were constantly chasing you say give me the definitions I need the metadata or um tell me what the data
[05:03] quality rules are we've got to document them and reflect them out right like you just cannot operate like that anymore now you have AI at scale it has to be fully automated so when there's a piece of data that lands into data bricks you
[05:18] need to be able to say it lands in a governed state So um we've pivoted from the typical like data life cycle to the life cycle that you see on screen which is our govern by design data um platform. So the first
[05:37] thing is we define purpose and use for data. Then we build the data. Now we're starting where you see the two stars on there. Those are the new things that we're deploying right now and are moving right now. So we deploy on top of our metadata. So we have business, technical
[05:53] and operational metadata. We're now deploying context on top of those things as well. And then we have trust and I'll talk more about the approach to data quality that we're using. Um that's where your eval come for AI governance now. And then we have share and um our
[06:11] ongoing operations where we're uh monitoring cost and use. So what are the three failures that we saw when we started to scale? Um so when
[06:27] I started at General Motors within 60 days of being there we identified the data governance process that we wanted to have. We wanted to collect metadata upfront. We wanted to classify data immediately. We wanted to apply data
[06:44] policies based upon our classification. And then we wanted to have um data quality run automatically. The three things as we started to scale on the data bricks platform that we saw was that the typical definition of ownership
[07:00] that you would have in on-prem environments was you had someone who someone or a couple of teams that would focus on one or two on-prem assets and then fully control that little thief. Right? So when we went into data bricks
[07:19] now you've got a cloud and for me conceptually it's just like one big container right like it's now we're all in the backyard together we all need to play together so you can't have this rigid thing of ownership you need to have a flexible approach to ownership
[07:36] and we're pivoting from saying that this team owns this data or this team owns this data to General Motors owns this data and the data is dri um divided up by domains. So um the first thing that
[07:52] we said was we're going to adopt a data mesh architecture. Data mesh is different than data fabric in that data fabric would allow you to connect your data and organize it um in a certain way but data mesh extends on
[08:07] that by supporting it within operating uh organizational structure. So our teams are divided up by domains. So typical domains that you would think sales, marketing, customer, uh we have
[08:23] an after sales. Uh we have the data that comes off the vehicle. We call that connected vehicle. Um so we have all of those domains. We typically have one leader who's in charge of that domain and then they bring all of that data together. they're responsible to land it
[08:39] in bronze, bring it forward to silver, and in some cases deploy gold assets. In other cases, we just allow business teams to crowdsource on the bronze data that we put we put there. So, we flexed immediately on ownership and said we
[08:56] need to have this new approach. We made everybody migrate from their current setups to the um new domain setup and that's what helped us organize our data and now it's allowing us to understand who's responsible for what as we continue to mature. The other problem
[09:14] that we notice is inconsistent autonomy. So if we hadn't set up the domain structure and uh backed it with Terraform code that deployed the domains through a consistent onboarding process, things would have really been a mess,
[09:31] right? So uh we addressed inconsistent autonomy by putting in the right architectural components to onboard people to these domains in a very consistent way. So it doesn't prevent them from working. It doesn't slow them down from working. It just says here's
[09:48] your box. You can play in it and you can read from all the other boxes or all the other domains, but we really want you just to write to your own domain. And um that's been working. And then the last thing that I've just noticed in the last
[10:03] six months is like AI is really amplifying the risk that we have in these spaces. um because everybody's deploying these AI tools and using data in ways that we had never imagined, it feels like these risks are compounding, right? So, we really need to continue to stay on top
[10:20] of them so that um we're staying ahead of the the risks and we're not seen as blocker, right? So, long time ago, well, probably not that long ago, three years ago, everybody was like, I don't do data governance. It kind of slows me down.
[10:36] It's not a capability that can scale or move as fast as me. And uh we just can't behave that way now. We have to be just as fast and when things are released to production, we have to meet them in that production state and have everything done.
[10:59] So um how have we addressed those failure modes? I talked a little bit about domain accountability. So within our domain structures, we have teams of engineers that build those data products within those domain spaces. Um they'll make an onboarding request to be
[11:15] onboarded to a particular domain. I have somebody who works with an entity model to um understand the data set and then map them to the appropriate domain. So there is some rationalization that happens. We have the governance control
[11:31] plane. So um terraform code stands up the domain. We assign a works an account in a workspace for individuals. There's multiple tenants within the context of a domain. And now we have the governance control plane. So as soon as data is
[11:48] landed in datab bricks and I'll talk more about this. It's classified then we go through and we'll apply tagging to it. um we'll apply quality to it and um we'll list it within our enterprise data catalog and I'll talk about the tools
[12:04] that we use. It's not all data brick centric and then we had to do platform acceleration. So if we're going to onboard um to the tune of 200 teams across a large complex organization you have to do infrastructure as code and you have to
[12:21] have a seamless process to allow people to on board. So um in order to meet all of those things we organized our team across four capabilities. So this is uh I have people that uh work in these
[12:38] different areas. So on the left side is really I call the enablement side right literacy and domain enablement. Literacy is the training and the content. Um, everything that we give to people to
[12:53] meet where they're at. Uh, we use uh the Slack hole. I mean, Slack. Does anybody else use Slack? We have Slack. We have thousands and thousands of Slack channels. I call it the Slack hole. Um, which there's a couple people that don't like that I
[13:09] call it that, but but I do. So, um, but we have a help data Slack channel. So, uh, people can go in and submit a request and ask for onboarding, ask for help, ask for more compute, ask to help
[13:25] get something resolved. Uh, this team is very like they're amazing. We used to take um up to 10 days to service about 300 uh 300 requests a
[13:40] month. We're now close to a thousand requests a month. And we've pushed down the average time to service to about four days. And the ones that take the longest are the weird ones, right? Somebody who comes and you're like, "Well, I'm not sure we want to do that with an Excel thing. I'm not sure that
[13:57] we want to put your data there. Um, no, you can't just move it over the way you want to. You're going to have to make some changes to improve or optimize it." So, that's where we get slowed down now. But we keep trying to optimize that with more and more automated workflows. And
[14:16] um we really the the vision there is to get to the natural language interface. Somebody asks for something and we just make it happen seamlessly. Uh so that's like data literacy is the training part. Domain enablement is that
[14:31] help side uh where we enable people to be successful within our domain and data mesh architecture. On the other half, we of course have the enterprise governance team. Uh for the folks that raised your hand, you guys are probably the folks that would be on that team. They think
[14:47] really deeply about how we manage and care for our data. Um the role is different than a data steward because we have to define how the data will be cared for and governed and then we have to implement that and then we have to
[15:02] scale it. So um for us it's about how do we find the practical processes to actually execute governance in a way that it's uh low touch or no touch by this enterprise governance team and when people put their data in the right
[15:18] domain it's already governed and then the last piece is governance engineering I have worked so long and so hard to have an engineering team uh I feel like now if you're going to be successful
[15:33] uccessful and build a capability around data and AI governance, you need an engineering team. That's the team that's building the right interfaces, connecting the different platforms that you have that are executing pieces of governance in different ways. And
[15:50] they're they're your tech team that make all of this work. Without that team, it's just like it used to be. you're going out to people saying, "Can you come into the enterprise data catalog and put all the definitions in and can you do this, can you do that?" And you
[16:06] just like at at GM and probably like you guys, you can't scale in that way. So, you need an engineering team that's building super smart technical solutions.
[16:24] So, um, I'm going to talk a little bit more about literacy and enablement. I'm right on time, too. We've got a timer. It's great. Um, so, uh, democratized access. So, when I started at GM, they were like, we need data democra democratization, and that was
[16:40] coming from the analytics team. But to them, that just meant, I just want access to everything, and I want to do whatever I want when I want to, right? It's just like, no, it doesn't mean you just get free access because we still can't break the laws. And we have a
[16:58] thing. I work for a guy called John Leech and um our motto is let's keep John out of jail, right? So, we do have to have some things in place, some policies, processes, restrictions, all of that. So, we control and at least
[17:14] meet our regulatory and compliance obligations. It's not just free access. Um, we also focus the team on enablement. So, we're not allowed to um come back at someone. So, like our enablement team is there to serve. So,
[17:32] no matter how crazy the question is, no matter how insulting the other party is, no matter uh what crazy thing they're trying to do, we're going to meet them where they're at in a very pleasant way. try to understand the situation that
[17:47] they're going through and then steer them to do the right things. Very rarely do we have to escalate, but when we do, my VP has my back. So, I've never had a situation I'm very fortunate. I think top down is really important to build
[18:03] these capabilities and be successful because we're not just building the capability. We want to build the culture, right? You want to build the culture where the product managers are saying, "Do we have governance in place? Do we do everything that we're supposed
[18:18] to? You have engineers that care, that are looking into the data catalog to make sure that the definitions, the classifications, everything are appropriate. And then you have people that care for the nature of the data and are checking that the anomaly detection
[18:34] is set up appropriately. So, uh, you can't just put out, uh, generic training. You have to focus on enablement, meeting teams where they're at and, um, giving them the right tools, um, pleasantly, right?
[18:52] Because we want to always be on everybody's side. And then self-service also helps reduce friction, but not responsibility. So, we put checklists out there. We put all kinds of things out there that people can use. Um, and
[19:08] we tell people that they're accountable for these things and just because we've enabled a lot of self-service things doesn't mean they can skip those things. So, we ingrain that through our enablement processes and the training that we have.
[19:33] So, what does that look like? um when we like what makes that model work. Um so it's really the three things together right it's the data mesh architecture and setup the domain ownership uh pieces because now people know where to work
[19:49] and then the governance piece tells people how to work right so we cover those those three things. So the domain owners they own approve and oversee what's happening with respect to their domain area. If we're onboarding a new
[20:05] tenant into a domain, we send them to the domain owner to have the conversation of uh are you okay if we add this information in your domain? Uh the domain owner has the responsibility to make sure we're not doing uh
[20:22] sending in duplicate data sets or we're not setting up something that doesn't make sense for our organization. Uh duplication is still a problem. like we haven't totally figured that out but the intent is there and I think as we mature we'll get better in that uh and then you
[20:39] have platform and data engineering teams those teams are really focused on the building parts so the platform team our platform team manages data bricks they also have fiverr for hyper ingestion of
[20:54] uh raw data sets into the bronze area very quickly so that we can set up our raw data sets super quick and then our data engineering teams um go in at that data build the pipelines and then build across the data bricks medallion
[21:11] architecture and then that last piece is governance and enablement. So they um enable they help define and then they align the organization.
[21:30] So what does that mean for people? We have three actors and actresses, right? Um and this is all we need. So if your data governance organization has 16 personas and 12 different kinds of stewards and
[21:46] you're still trying to explain who the custodian is versus the technical steward or the business steward and who does what like I think maybe rethink that pivot just a little and make it very simple. So uh your domain owner is
[22:01] the person who's in charge of that domain. They have a responsibility to make sure that what's happening in that domain is appropriate for your organization and they have some accountability, right? Uh the data engineer is the person who's building
[22:17] the data and is accountable to work within the guard rails or on the superighway that we gave them to deliver data out to a production environment in the way that we've described and steward. So this um our team spent a lot of time
[22:34] talking about stewards. We did have 16 personas. I think we still have eight. Um not for long if I can help it. But uh to me with AI, this is the other thing that's changing. The nature of ownership
[22:50] and responsibility is changing. It's not just IT teams that are building these technical solutions. Now with some of the AI tools, you now have business teams that are building very technical solutions to solve for uh very specific
[23:08] business questions. So I don't care, and I say this all the time, I don't care if the steward's on the business side or the IT side. I don't care what team they're on or what they're doing, but if they volunteer and say they're the
[23:24] steward, they're going to validate the classification that we applied. They're going to take a look at the quality things that we've deployed. That's fine with me. So, um there's still a lot of hangup within our organization about we have to have a business steward, we have
[23:39] to have a technical steward, and who's the custodian taking care of the data? I just think it just doesn't make sense anymore. We just need someone who's going to be accountable and someone willing to um oversee, care for, grow, nurture the data.
[23:57] So this is my favorite slide. I've used this slide like a thousand times, but these are our data readiness capabilities. So um the journey that I'm talking about happened really over the last 12 months. So we have data bricks as our data platform.
[24:13] We did a homegrown data classification tool. Um, it's an automated process. As soon as the data lands in data bricks, we run it through a tool of course called Chavevel because we work for a car company. Uh, it's super cool. It
[24:29] goes through, it classifies the data, and then it stores a reference set of what we've classified. then that gets reflected out into our data catalog where a steward can verify that classification.
[24:46] Um, if the classification changes, it comes back to Chvel and those reference data sets. So the next time we run classification, we're that much smarter about how we're classifying our data. We use that classification to drive our
[25:02] data access. So um we pivoted our data access from being down at table level to being at domain level. We worked with the legal team and we said hey if people are building within a domain or they
[25:20] have to answer business questions within the context of a domain. We should just give them access to all the data in the domain not some of the data. And they agreed with that. Um so we have uh when somebody requests access they get access
[25:36] to a domain set of data and they uh can't read the PI or like the PII or the SPI or if there's restricted or secret data that's masked or like uh obuscated or omitted. Um and then they can request
[25:52] for elevated privileges and we use Ammuda for that. And then we have data quality. Uh this is the coolest tool ever. We love Anomalo. Um I I have only two people. So data access uh policies set up and roll out to the organization
[26:09] which is massive right we have um over eight million assets is two people two people support Ammuda and the roll out to the data governors that we have across the domain. Um data quality is the same. We have two people, one
[26:25] technical engineer that supports the platform or solution in an ongoing way and then I have one data quality lead that understands intimately what the anomal platform does and they've built self-service tools. So now teams when
[26:41] they when their data gets into data bricks they can selfserve register their tables in anomal. um a IML capable so it learns from the data over time. It detects anomalies and drift and it allows for very simple
[26:58] rules as well as very complex rules and we give that to all the domain teams so that they can apply quality. uh this is another space that I think with AI governance is pivoting right because AI governance for high-risk applications
[27:15] requires you um if you read any of the legislation or the NIST uh AI RMF standards they talk about having uh quality applied consistently on high-risk use cases so this allows us to
[27:32] do that team was amazingly smart and they've defined find like 20ome essential core quality checks that any team can essentially subscribe to for their um data set. So they don't even have to think about the core basic
[27:48] checks. They're already out there. They can just set those things up. And then the last piece is uh I would call it data catalog. I've referred to it as data catalog. We're now changing the language. It's the context layer, right? So it has all the metadata, the
[28:04] technical metadata, the operational metadata. It has lineage. Um it has our classifications. It's our user interface for whoever the people are that steward who has to ver verify the classification. Uh that's the UI that we
[28:19] give to the organization to discover, find, uh and find data. And then we leaned in on Atlin and Ammuda so that if somebody discovers data in Atl's
[28:36] button right there they can just click to request access and we'll um issue that request out. If um there's no secretator or anything, we'll automatically approve. If there's something of concern, then we route it to a data governor who will review that
[28:52] request and uh either approve it or not. And then in terms of architecture, I'm just kind of buzz through this. I'm slightly couple minutes behind, but that's okay.
[29:08] Um, so in terms of architecture, I talked about this is the data data readiness. Everything that I talked about is what we do to make data AI ready. So any data that goes through this process, it's classified as quality. We know what it is. We have an
[29:24] owner assigned. Um it's in the right structure as part of our data mesh. Uh that data is AI ready. Um we took it a step further. Um and I'll talk about that in a minute because there are some
[29:40] things that this is missing. So this is essentially source, ingest, curate, publish and then we have consumers off that data. So in 12 months what was our measurable pro progress. So we set up 16 domains
[29:56] that were deployed with Terraform infrastructure as code. Uh we had 230 teams that migrated to our data mesh. our active catalog users went up over by a thousand% and we moved from having assets that
[30:14] were classified from about 1 million to just over 8 million. So um we are operating at scale now in terms of trust and transparency we drove up the lineage piece. So the lineage wasn't automatic. we wanted to
[30:31] really focus that on the core or essential um critical data sets that we had. So we focused on those and I also have responsibility for data transfers. So we register all data that ingresses
[30:46] into GM and data that egresses out. Um so that's a pretty manual process we're looking at. Is anybody familiar with data contracts? Okay. Yeah. So we're using we're testing
[31:02] data contracts now for automated pipeline builds. We think if we can dep automatically generate that specification, we can then automate the whole generation of a pipeline. But I also see that playing a place uh within this data transfers.
[31:19] So before you ingress data in if you define what that contract is now you have a systematically a systematic way to gate uh what's happening and version and control the uh definition of the data that's ingressing in and then the
[31:35] same thing for anything that's going out to put that contract in place we have like sterling I see you in the back we have sterling gateways and we want to use the contract within the that context um and Then the last piece, boy, it's
[31:50] hard to read that one. AIML grade data quality. I talked about the anomaly detection that we have. Um, how we scaled the medium cost per table per day is about 8 cents. So that comes to about
[32:06] $25 a year to protect our table. We had to show that number. There are so many people that love to build those really sophisticated data quality checks by hand against their tables and we had to
[32:23] really sell that and and we also had to divert. So the other thing about enablement because we're always serviced with a smile. If we find a team that doesn't want to budge, it's fine. There are loads of opportunities with other teams that are ready for us and will
[32:39] meet us like where we want to meet them and uh we'll step away from the teams that are like not willing to change to divert our energy AC to the teams that are willing to change and come back to that team later. So, we had one team
[32:55] that just loved building their quality and it took probably close to a year before we could convince them to try the Anomalo platform and now we're pushing them in that direction. So, it feels like a win finally. But yeah, and the other cool thing about this and I didn't
[33:11] make up this saying uh one of my colleagues Janine did. It says we scale data quality with compute not headcount. So every time we want to do more quality, I don't have to say I need six more people to build this blah blah
[33:28] blah, right? I just say, hey, we need a little bit more compute because we're going to deploy some other sophisticated observability or checks. So what's our next set of priorities now? Or okay, actually we'll talk about
[33:44] business outcomes first and then the next set of priorities. So um the business outcomes that we had of course are faster onboarding. Some some folks they place place the requests for onboarding and they can be onboarded within a day. Um we also have reduced
[34:00] regulatory risk. So that whole process that I talked about produces evidence right? So I come from a highly regulated background. So I was like when we do these things we need to produce evidence. um which was atypical for this manufacturing organization because
[34:16] they're not as regulated. Um but everything that we do produces something. So like in five minutes I can see who has access to what data from what country. So if there's an incident I can reflect that. Um all of the
[34:33] quality gives us visualizations and information that we can reuse. The catalog is exposed. Everything is transparent and we keep that list of who owns what. So we always have a network of people to talk to should something go
[34:49] wrong or we have questions. So it's uh definitely reduced our regulatory risk. Uh we also automated enforcement. So uh I don't have to tell teams now that you've
[35:05] landed your data please apply these data quality things. Right? So if you land your data in a catalog that's part of the mesh, we automatically deploy this and all the data data governance that I talked about is it's just happening. Um
[35:22] I had a lot of I reported to the VP of data engineering. I have director peers that own the domains and they were constantly asking me what does my team have to do to govern our data? What's the next thing that we have to do? Are we
[35:38] governed? And I'd be like, "Are you in the right catalog?" They're like, "Yes." I'm like, "Your data is governed like that. That's it. You don't you don't have to do anything." And um so it it was like it was a confusing thing for
[35:53] them because they always thought like, "Okay, now I'm going to have to dedicate time, resource, and energy to make sure that we're governed." And it's just not like that anymore. If you deploy in the right way, then your data is governed.
[36:08] And then we have measurable AI readiness. So for all the data that's governed, we're ready like it's ready for AI. It has all the eyes dotted, all the tees crossed for what you need in order to be um in a governance state for
[36:25] AI. So our next set of priorities um so we want faster AI accuracy and assurance. So the testing piece, the eval, the judges, those kind of things.
[36:42] This is like atypical for our teams. I'm actually starting to wonder like did they ever test these data products that they were putting out because there's some very simple things where they they release some like agent or something out there and it's completely heinous and
[36:59] it's not working and not all the right people have access and I'm like did you do testing and um I got a lot I get a lot of weird questions back but I think we need to really train people on this like your typical data scientist I think it's kind
[37:15] um uh embedded within their psyche to do some testing or to test how the uh model is working. But with these agent things, it's it doesn't seem to be embedded that way. Like the production of genie rooms,
[37:31] we have over 2,000 genie rooms and like I don't I very uncomfortable. I don't think anybody's really tested those things. they just put them out there and they're like, "Okay, this is my data side and I programmed all this SQL and
[37:47] have at it, right?" So, um, this is a this is an area where, uh, over the next six months, I'm going to do a lot of work. I spent a lot of time. I did the agent training here. I understand the evals and the judges that are now available in the platform and I'm going to start pushing that out with our
[38:04] enablement team and start demonstrating to people what it looks like when it's appropriately done. Um, we also need automated risk assignment. So, I talked a little bit about the NIS standard. That's really what we're going to embed ourselves in for General Motors in terms of
[38:21] evaluating AI risk level. Right now, we have a manual form. You fill out 14 questions. If it says because of those answers that you typed in, we think it's high risk. Then they give you more questions that you have to answer so
[38:37] they can assign to you the right controls to deploy on your data, your model, your AI application or in your AI use case and um we just can't operate like that. So we're trying to figure out how do we do automated risk uh based on
[38:54] these use cases in a really quick way. uh then we also need to align the enforcement to the risk. So I know pretty much how to govern data when it lands in the data platform and make it AI ready. Now we're the next step of
[39:10] maturity is once I've assigned the risk, how do I make sure that all the right controls are in place in an automated way and how do I reflect that evidence out and have something measurable that I can uh show to people to demonstrate
[39:25] and that's it. So uh we're right at 40 minutes. I really talked. Sorry guys. Uh so uh the govern by default full life cycle is like we go through the purpose, the
[39:40] build, the context, the trust, the share and the operate. And how would you start? If you were going to start, I would say establish ownership, build the minimum viable engine, make enablement part of the
[39:56] onboarding, and then measure and iterate. And that's it.
Learn more about the Databricks Data and AI platform.
The information provided herein is for general informational purposes only and may not reflect the most current product capabilities or configurations.