Skip to main content

AI-Ready Data at Scale: Building Master Data, Ontologies, and Agentic Operations with Databricks

Summary

  • Baker Hughes, a $25 billion energy technology company, built AI-ready data foundations on the Databricks Data and AI platform by implementing master data governance across 250-plus systems, creating ontologies that connect jobs to invoices, and establishing Unity Catalog as the centralized governance layer.
  • Semantic alignment and ontologies enable Baker Hughes to reduce revenue leakage, improve days sales outstanding in order-to-cash processes, and control hallucination in AI agents by grounding responses in clean, connected master data.
  • Genie spaces provide natural language analytics on top of the master data foundation, and cross-functional team alignment around data domains is the organizational prerequisite for seamless AI deployment at scale.

AI-Ready Data at Scale: Building Master Data, Ontologies, and Agentic Operations with Databricks

Watch: AI-Ready Data at Scale: Building Master Data, Ontologies, and Agentic Operations with Databricks
Enterprise AI success depends on clean master data, semantic alignment, and process integration. Baker Hughes, a 25 billion dollar energy technology company, demonstrates how to build these foundations on Databricks to unlock agentic operations and accelerate order-to-cash processes while eliminating revenue leakage.
Learn how Baker Hughes implemented master data governance across 250+ systems, built ontologies connecting jobs and invoices, and deployed Genie for natural language analytics. Includes real examples of reducing days sales outstanding, controlling hallucination in agents, and aligning cross-functional teams around data domains for seamless AI deployment.
🤝

Chapters

FAQs

What is Baker Hughes and what data challenges did they face?

Baker Hughes is a $25 billion energy technology company that manages data across more than 250 systems globally, creating significant master data fragmentation challenges. The company's AI potential was outpacing its data foundations, requiring a systematic investment in master data governance, semantic alignment, and ontologies before deploying agentic AI at scale.

How does Baker Hughes use ontologies to improve AI accuracy?

Baker Hughes built ontologies that connect related business entities such as jobs and invoices, establishing semantic relationships that give AI agents a grounded understanding of business context. This ontological foundation is described as critical for controlling hallucination in AI agents, ensuring that generated answers are anchored in verified data relationships rather than approximations.

How does Databricks Genie fit into Baker Hughes's data strategy?

After establishing the master data foundation and ontologies, Baker Hughes deployed Genie spaces to enable natural language analytics across their data domains, replacing static dashboards with conversational interfaces. This video describes Genie as the consumption layer that allows business users to interact with AI-ready data without requiring SQL expertise.

What role does Unity Catalog play in Baker Hughes's governance model?

Unity Catalog serves as the centralized governance layer for Baker Hughes's data estate, providing lineage, access controls, and a unified catalog across their diverse system landscape. The lakehouse capabilities of the Databricks Data and AI platform, instantiated through Unity Catalog, are described as enabling a qualitatively different data experience for the organization.

Full transcript

[00:09] Awesome. Good morning, everybody. Hope you're having a great day. I feel like this is an earnings call, so the forward-looking statement. I think everybody knows what this is for. Um Cool. Complete your surveys so if you could at the end, would love some feedback. So would love to understand how Dan does in his presentation. Don't judge me.
[00:25] Um Uh so for today, um you can't say kissy face. Um So for today, we're going to talk about enabling AI-ready data. So I think it's one of the topics discussed over the course of the last 24 or so hours. Um happy to be here with uh my friend
[00:41] and and colleague Dan McPartlin. I'll have him introduce himself. Uh for those soccer fans, uh Portugal the tied 1-1. Yesterday was a big day. So a couple goals by some really good players. The goat had three, so that's awesome. So uh you know, we'll keep you updated as it relates to the score. Um for context,
[00:57] the conversation that we're going to talk about here is You know, I think everybody understands the potential of of AI, right? And and the foundation of data. We're going to talk about how it's been built and how that's enabled specific business outcomes for Baker Hughes. You You okay?
[01:12] I'm good. It said kissy face again. It It's That's the fifth time it said kissy face on the screen. We're not saying it, I promise. I don't know what it is. nothing to do with the conference. That's right. Yeah. Okay, good. Just making sure everyone's clear about the what we're here for today.
[01:29] So I think similar to the themes of of Data and AI Summit, this isn't about model scores and benchmarks. It's really about how do we enable the infrastructure for AI and what that ultimately means from data. So what Dan will touch on is how you know, lakehouse capabilities were
[01:45] instantiated, Unity Catalog, so on and so forth and how that enabled a different experience for for Baker Hughes, which I I think is really important for the crowd here. So before we jump into the content, I'd like to ask Mr. McPartlin to introduce himself and maybe a little bit of background on Baker. Hi, my name is Dan McPartland. I have
[02:01] the pleasure of running the data and AI office for Baker Hughes. We are an energy technology company. And I just want to start off by saying thanks for having us. This is great. Um typically in the energy world, we start off meetings with health and safety moments. So, we're going to do a very
[02:17] quick one here just to stay compliant. Just for everyone's opinion on the matter, um let's not look at our phones when we're walking through very full hallways and stop and look around. End of health and safety moment. Thank you, Matt. Friends don't let friends uh text and
[02:33] walk. So, please just be careful. Um Awesome. My name is Matt Oulano. Uh I help to lead the global data and AI team for Genpact. Um I also lead our Data Bricks business group. I was fortunate enough to be acquired into Genpact. I was co-founder of a company um that Genpact acquired in June of last year.
[02:49] Data Bricks Ventures was an active investor in my company. We were their first successful exit. So, we were very close to the Data Bricks ecosystem um helping them consult around the product, drive product development, and ultimately the enablement of the technology with clients like Baker Hughes. So, thank you for the opportunity, and I really look forward
[03:05] to to the discussion today. So, what we'll cover, we'll talk a little bit about the inflection point, the problem that we see within data, uh the partnership that we've developed both between Baker Hughes Genpact as well as Data Bricks, the layers of the architecture, how we went from platform
[03:22] to connected data to aligned semantics, and ultimately the experiencing experience leveraging Genie. Um we'll talk about Genie in action. I think some of the work that Baker has done is is really paramount to their success, and then open it up for questions. I think most companies are are stuck between
[03:38] kind of the the problem and the solution, and um they they try to resolve or revolve too much around kind of the tool, where tools are kind of mean, you know, the means to the end in helping to enable. But we'll talk about kind of the intersection of people, process, tech, and how some of the
[03:54] leading edge technology of of Databricks helped to unlock kind of that that data problem. So, the the potential of AI, I I think is is out pacing data foundations, unstructured data, semi-structured, you know, data is moving at different speeds. I think you've seen that through
[04:09] yesterday and Lake Flow and the new advancements, but everyone has decided that AI matters. I think few and and maybe those in the room have, but few executives have really decided the same about the data, right? And it it's hard and Dan will talk about to justify the
[04:25] investment into the foundation, but I think people are starting to realize that the foundation is critically important um to enablement. I think as you think about the evolution of business, for centuries we've kind of relied on this single operating model. So, most work was done by humans, validated by humans. We're now moving to
[04:43] um machine-driven, human-validated environment, and because of that, right, we need ways in which the humans and machines can interact with data. So, you see thematically things like markdowns and how data is ultimately going to be uh put into memory and the context
[04:59] behind that. So, all of this I think kind of um leaves us to the fact that we need our data structured, we need it semantically aligned, we need it connected, and that allows us to kind of have different experiences of data. And those data, again, those experiences are both for
[05:15] humans and machines. So, it's really, really important. Um we recently conducted some research with uh HFS, and we talked about the debts. So, the debts organizationally, we view this as about an $18 trillion opportunity, and these aren't new, but it's four things, right? It's process debt. We talk about agents
[05:32] and how agents are going to help us to accelerate different types of business outcomes, but relying on legacy processes built on tacit knowledge doesn't actually allow an agent to operate most efficiently. So, the process debt, the data debt, ultimately, how do we make our data more actionable,
[05:48] uh the technology debt, so the enterprise architecture and the relationships, and finally the talent debt. So, those are the things as we think about Baker Hughes, as we think about our client base and bringing this all together, it's truly that intersection. So, Dan, I'll start with
[06:03] you've been on this journey for 18 months, 24 months now. Sumish, some of that together. When you started this, what was the the driver behind it, right? Was it we just knew there was a problem of connecting data? Was it truly a business outcome? How did you kind of get the organization rallied
[06:19] around the idea of modernization? Um very interestingly enough, our somebody one of our senior vice presidents um came to us and wanted to talk about how he runs his business reviews. All right. And ultimately, he said, "People typically show up with a
[06:35] PowerPoint presentation and a screen grab of Excel spreadsheet that says, that's my third-party spend." And he described it as a way of driving a car where you couldn't see out the windshield. There was no gauges.
[06:52] So, you didn't know how fast you were going, and that's what it was like running the business. Another way of saying it is looking out the window of an airplane, making sure you're going to the correct destination. And we ultimately formed a group and Sorry about that.
[07:07] We ultimately formed a group um and we took a very software type approach to how we delivered data products. Um so, think very much CICD, um proper code management, and things of that nature. So, um again, we were anti-low code. We were very much like
[07:23] SQL is the universal language of data. Therefore, everyone in the department writes SQL. Um and when we first launched our first we call them zero-touch solutions, which was primarily a Power BI report. You know, looking at the third-party
[07:38] spend, automated in very real-time with a high-level visibility for a business review, but the click down to actually show PO level detail. Right? And it was very interesting when we started to see some of the data sets because
[07:53] obviously things were scrubbed before. Hence the the screen grab of the Excel. And once they got that taste of the third-party spend kind of a feel, then they wanted everything that came along with it like a a digital P&L, head count reporting, visibility into
[08:09] travel and expense, sales revenue, so on and so forth. And that was great because we were able to build and visualize a nice beautiful data product for them. Very quickly it became
[08:24] a master data problem. Like how come I'm seeing three different copies of the vendor NOV? And it's like, "Oh, because that's the as-is data in the source systems." And that's when our data quality master data
[08:40] uh you know, concept and approach started to take take place quite quickly. And all these things came together into "Whoa, I just watched this I listened to this podcast and Anthropic said they're doing this. We want to do this. And we are
[08:56] going We're not ready to do that. There's a ton of foundational capabilities that we just don't have yet that we need to build. And that has been the Look, we need to invest in a bunch of things that you're going to see very little return on investment immediately, but
[09:14] when the cool stuff starts working, that then you're going to realize the investment and that's a little bit of the challenge we're running into right now. It's interesting. One thing you mentioned I'd like to ask about is is data products. Right? I think people have talked about the importance of data products. If you listen to Dario or
[09:30] Jensen Wein, they talk about the foundation. Data bricks will be a critical enabler right of agents and unlocking generative processes in in agents. But, what is the definition of a data product to Baker Hughes and how does that apply to the business? I would say the definition of a data
[09:46] product to Baker Hughes, we think of it like what problem are we solving? Right, so I want visibility into such and such thing and they really naturally fall into these these buckets. So, I would just say answer to a question would be what I would refer to as a data product. That's great cuz I think it becomes a
[10:03] religious conversation sometimes. At some point, yeah. Is it Is it data? Is it code? Is it a BI report? Is it a machine learning model? So on and so forth. I think to your point as long as it's answering a problem, right? That's That's what's most critically important. Is it searchable? Is it understandable?
[10:18] Can it be used by other people? You know, data products I think are driving reuse and that reuse hopefully drives down the marginal cost to answering the next problem. Cool. So, master data as we think about the distribution of data and the fragmentation,
[10:35] some of the problem statements that we saw was around revenue leakage. Baker Hughes is a a massive organization, 25 billion. So, 5% is pretty meaningful. Absolutely. I think that's not a small number and then as it relates to kind of the global deactivations, that's not just a data
[10:50] problem but a controls problem. Dan, I'd like to ask for for the crowd. Most of these conversations start with we need a data foundation. We as data practitioners know the problems for the business that needs to manifest in some type of business case and it's hard to put a business case together sometimes
[11:06] around foundation. We know it's brittle, we know it's cracking. How do you resolve it? You need to kind of layer on the different business opportunities. So, Yeah. how did you actually do that for Baker? How did you get your sponsorship of the executive team around making the type of investment that you needed?
[11:21] Again, it it gets back to what problem are we trying to solve? Okay, so what everybody wants in the organization is I'm about to go see Saudi Aramco. Generate an executive briefing to tell me what to talk to them about. Are they happy? Are they sad? Are they mad? Uh
[11:38] has something happened in their company recently like, you know, change of leadership and things of that nature? And when we took that problem statement and say could we do that we ran into a major problem of master data again. Sorry. I'm like I like master data but then I don't know what
[11:54] that Um where we what we ultimately ran into was we had a master data solution for customers that really only knew about things that were in your CRM and things that were in your ERPs, right? Um so things that are not in your CRM or in your ERP, uh cash
[12:12] disputes for example, uh service delivery, are they happy with the jobs, customer incidents and such. So if somebody came to Genie per se and said, "I want to do that." I'll be able to tell you about open orders cash
[12:28] opportunity win loss ratios but nothing really to that next level. So we definitely came up with uh why does master data only know about a handful of systems when there's over 250 other systems that actually are the real stuff that people care about for data.
[12:44] Uh and this is where um the 360 degree view very marketing term, sorry. Uh it became a real thing. So the way I was able to sell the executive buy-in was ultimately uh when you go to see a customer, we want to tell you who was there recently.
[13:00] What did they do? Uh is the customer happy? Do we have upcoming jobs? What's the probability that they're going to pay their bills before quarter close? And a lot of those things. And when we started thinking about that, the investment in master data was a fraction of the value that we're going to be able
[13:17] to provide the company, which was great. That's great. One thing I've learned from you and and your Baker team is resolving master data in the platform is important, downstream distribution, you get consistency. One thing you guys did really well is kind of solve the problem up front in authoring. So, how did you
[13:33] change the paradigm of actually how customers are inputted, getting consistency up front, so you're not always reconciling things downstream? We had a lot of fun implementing some data governance protocols. Uh so, for example, what what we what happened before is we had a what's
[13:50] called know your customer process. So, you met somebody at a conference like like this. And you would make sure they're not affiliated with organized crime. They have They're not having free cash flow issues and such. Um and then when we were able to
[14:05] um connect those things together, uh we very quickly realized that wow, we have we have some challenges, Matt. We have some big-time things we needed to do. And when the when we started talking about we don't want you to go create a
[14:21] customer directly in an ERP, because you're bypassing the compliance, the governance processes. Uh it became pretty clear that we needed to centralize onboarding into um a platform that people are familiar with. So, we we selected uh
[14:36] salesforce.com in this case. So, now um everybody, whether it's a ship to, a sold to, whether you're modifying a customer's hierarchy, is all managed through the centralized onboarding. And then all of the downstream systems have been locked down, so you can't like
[14:51] modify them. Sure. Yeah. That's great. Yeah, I think it's important because many times as data practitioners, we're solving it in platform, right? But there's a business process change that it would actually make it easier than, you know, writing writing code to solve the problem. I think you guys did that really well. Um
[15:07] the the partnership, so, you know, as I mentioned, um we were fortunate enough to partner with Baker over this 18-month journey. So, you know, Baker brought kind of the urgency in the domain, we brought the build capabilities, and um Databricks brought the platform. But what's really important is the sequence
[15:23] of activities. So, making sure that you're building a sustainable platform built on Databricks leveraging the capabilities of Unity classification, um our back a back, um enabling the different experience and we'll talk about that from a genie both a code
[15:38] perspective as well as genie spaces, the ontologies and we'll get into that in a second, but I think the connection of data becomes critically important. We've been talking about ontologies for about 5 to 10 years now. I think people struggle to kind of understand the difference between that and semantics
[15:54] and we'll talk about that, but above the the ontology layer is metrics and then obviously agentic AI. So, Dan I guess question for you shameless plug for for us. Why did you choose to partner versus to build this internally cuz you have a very capable engineering team?
[16:10] I think when we started, we needed we needed some industry guidance. I'll put it that way. And um technically we started with exponential. Yes. There you go. Plug. Perfect. Um and what we greatly appreciated about that partnership was uh you were small, you had about a
[16:26] couple of niche players who like really knew their their stuff like nailed it. So, not so much uh we didn't use you in a uh air quote body shop consultant. We we used it in more of a consultative manner and we we came to you with very specific problems to say look at we're trying to
[16:43] implement CICD building data products. We want to move real fast. We want to go from um we deploy once a quarter to we deploy several times a week. And I think that your your practice back then had the promise to show us that you you had the
[16:59] industry expertise and you came in and you put your money where your mouth was and you delivered. So, that was great. Great. Thank you. The commercial's over. Okay. Yeah, good. Yeah, there we go. All right. So, I I think Dan would love for you to kind of talk to the audience about the
[17:14] foundation and and how building this master data foundation in Databricks really unlocked some some business opportunities for Baker. Yeah. So so once we were able to centralize the onboarding for customer master, uh so we were technically it's pairing Informatica and Databricks together.
[17:30] Uh a couple things that we realized, one was um if you were to read a Gartner article about master data management, you're going to hear the words hub and spoke model, you're going to hear system of record and the things of that nature, and we ultimately have the opposite of that working. So what we run into is
[17:48] we would create something in a master data platform, and then we would just ship it out to an application, and then we would just pretend that it it worked, like everything was great. And what we weren't really paying attention to was other systems would then connect to
[18:05] those systems and basically get their master data, and then systems would connect to those systems to get their master data, and we get this whole spider web approach, where we're we're actually enabling it's much more of like a bidirectional synchronization of data sets.
[18:20] So for example, if you were in salesforce.com and you went to sell one of our wireline services, for example, um we would know that here's the 10 systems that need to execute the order to cash process. Like we know that
[18:37] there's a that use this one SAP platform for quoting, they use this other one for the service delivery, and so on. And once we um connect those data sets together, what gets really fun and enjoyable, I guess what I'm trying to say is we do
[18:53] a bunch of unsexy stuff to make the beautiful sexy stuff work at a later point in time, hence the initial investment. Um and the bidirectional sync is great because now we know what systems are using that particular customer and for what reasons. And then we can bring a lot of that into
[19:09] Databricks and start connecting things. We're going to get into semantics in a little bit. And now you can start to really know your customer. Really understand are they happy, are they sad, are they mad, are they glad? And we can even start, you know, configuring LLMs to determine why, right? And And for me,
[19:27] you know, master data, if we don't get this right, I have no idea how we're going to stop hallucination. Like I have absolutely zero idea how we'll do that. What do you think, Matt? Yeah, I think that's right. When So we do a bunch of work in the agentic world as a business process company now
[19:44] into agentic operations. We're deploying agents both for our operations that we're, you know, helping clients with supply chain, accounts payable, accounts receivable. The number one thing we struggle with with clients is cleanliness and harmonization. So if you don't have harmonized customer, vendor,
[19:59] pricing, product, the agents just don't work, right? The agents aren't smart enough to kind of understand how to route back to different CRM systems, different ERP systems. So getting that reconciled, getting that single view, and then having Databricks be that point of transaction from agent back to agent,
[20:16] and then writing into the ERP systems to post into a ledger, right? Or reconcile an invoice. All of that right back becomes really, really important. So I think that the two things for us are this isn't really a technology problem, right? It is a business problem. There's procedural process, there's procedural
[20:32] aspects that need to be changed to make agents work. And the two areas of data that we should be focused on within agentic operations is cleaning and harmonizing, right? If you're not doing those two steps, the agents will just fail. 100%.
[20:47] Cool. Um So big topic of conversation again about a decade ago, the ontologies became all the rage. Palantir talks about the ontological models. Yes, they do. I think this is something that people are kind of figuring out within the semantic construct. So I think in
[21:04] simplest terms, just for the audience, like ontologies are what exists and how they relate. Semantics of the definitions, right? So, we need to define things in business terms. We as data practitioners think of SQL, right? We understand SAP tables. The business
[21:19] needs understand data, but how things relate. So, Dennis, you looked at business problems, right? And let's take customer and and disputes, right? How did you build the relationship and kind of where did that ontology manifest within Baker? Um where it manifested
[21:36] was pain points. And I know that's a little bit the opposite of what you're probably expecting me to say. Uh so, what we would we would focus on something that really bothers the bottom line. Okay, so for example, uh days sales outstanding is something
[21:52] that really bothers our bottom line. We want that number to be less and less and less as we grow as a business. And but to actually make that work, we really had to understand how these transactional processes interacted with one another. Like what were the levers and pulleys to make this
[22:08] thing work? And if we take a very simple example of job to invoice, right? We there's a lot of uh triggers and events that we can listen for that actually can indicate the process has moved forward, but maybe the data in the systems hasn't. Sure.
[22:24] Okay, so probably a good example would be a a rig which drills a well, um can really only be at one well site at a time. For example. So, if we've if we detected that that rig has moved because there's been a logistic which moved the rig, we
[22:41] go, that job must be completed because such and such thing. Uh so, again, we can detect, hey listen, uh dear Matt, service delivery coordinator, is job 1 2 3 4 done? It we we think it is, but the system says it's not. And hopefully this is a
[22:58] lovely chat interface with Genie or some type capabilities that say, hey, do you want me to go into that job planning thing? And go and finish the job for you?" And you're like, "Yeah, that would be phenomenal because that's one less thing that a very, very busy very busy service delivery coordinator
[23:13] has to do. And if we if to break that down a little bit more, it's like don't treat invoicing like we treat our travel and expense data. Okay? So, travel and expense data, you get a copy of your hotel receipt, and
[23:29] then 20 days later Concur tells you, "Hey, come on, go file your expenses because we don't want to pay the APR on that credit card." Uh into like if we treat our invoicing information like that, like part of that ontology, then we have a real problem
[23:44] because we're waiting 25 days to recognize revenue. We get the the ticker started on our payment terms. Um and so again, when we looked at how were we going to structure these things, it's usually well, how do we take a major pain out of somebody's side that has a
[24:00] major financial impact? And that was like the the birth of how we structured our ontologies. So, prior to you would be stitching tables, now we have the relationship of jobs have rigs, rigs have operators, right? Once we understand a rig move, now we have the ability to close out that job, process invoice, get cash in
[24:16] faster. So, it's the relationship that you help to manage. Absolutely. And as this evolves, uh so this is a very market to cash uh type of business process, right? But, you know, and then imagine there's another like a supply chain based uh set of
[24:32] processes. Like we sell something, then we have to buy the raw material to make it, then we have to make it, and then we have to ship it, and then So, all of those little sub processes, understanding how those stitch together in a larger ontology is a adventure that
[24:47] we have yet to start. Um but I think that's where we're going to be truly game changing in the way we do our business. Wonderful. That's great. So, as you think about the ontology layer, which is the connection points. People talk about metrics layers, supervisory agents, and then on top of
[25:03] that, obviously agents need to talk to common language and then the connection. So, how do you guys drive semantics, semantic alignment? Where'd that get captured? How did you leverage Unity for it? I know we've heard Unity Gateway here over the course of the last 24 hours around model routing, MCP
[25:20] management. I think that's the next evolution, but how did you actually get alignment from the business? How did you store it and how did actually help you operate that that kind of semantic alignment across I would love to say that we're 100% semantically aligned.
[25:36] Okay. I think we have we have some areas where we're growing in that space for like for sure. Um but but one thing that we found like how do we get the correct data structured to actually solve the problem? And
[25:52] that was and and there's sort of been an evolution in our group where we used to be sort of order takers. Someone said, "I want you to build a dashboard that does this." Yep. And that was the culture for a while and we're fixing that. Uh and now it's much
[26:08] more okay, could you could you give me a little bit more information as to why you're looking for this particular? And usually it comes down to a sad, mad, glad data point, right? Going back to the uh jobs and invoicing, um it makes me very mad
[26:26] when we wait 25 days to invoice a customer after we finish a job. It makes me very glad when we do it immediately, right? So, cool. So, now we understand a data point and this drives our critical data elements, our governance, our quality, those things
[26:41] into okay, so why does that make you mad? That makes me mad because that affects our billing timeliness. That affects our day sales outstanding. That affects our operating income, right? So, in our in our world, what
[26:56] we're gearing up for is document hunt Let's say I'll go with a thousand sad, mad, glad data points. So, that you could walk to Genie and say, "Do you have any recommendations to decrease day sales outstanding in North
[27:13] America land, for example?" And we're hoping that combination of supervisory agents and Genie 1 and wait um will crawl through all of those and provide like actual real solid recommendations to the people. And um we've really only run a couple of very small pilots using
[27:29] like metric views, but I think that this is going to take off and it's going to be a really big thing for us moving forward. That's wonderful. Could you touch on I think one thing you guys have done really well is a lot of organizations align themselves to capabilities and domains. Um domain being customer,
[27:45] capability being data engineering. Um I think what you guys have done is the value orientation. You talked about market to cash. So, how did you structure the teams, right, and the type of skill sets from uh process operator to data engineer to AI engineer? And how do you think that actually helps you look end-to-end in business processes to
[28:02] better align semantics, ontologies, and relationships? Uh so, one thing became abundantly clear. Uh and I probably needed to provide just a little bit of background about the organization itself. Uh so, in October of last year, uh we we were operating a data and
[28:19] analytics group in like a division, oil field services. And then there was this other data group which was more of like a central function, data And then there was a different group that was running RPA, generative AI, and those types of nature. And and then in October, we all consolidated into one gigantic
[28:36] organization. So, the data and AI office uh has our scope is unbelievable. I'm going to put it that way. Uh we have everything from data management like master data reference data data governance integration between
[28:51] systems fantastic an API program a generative AI capability robotics process automation and we have data products genie and so on and so forth so we have a very large scope let me put it that way and
[29:07] one thing that was super clear was the is that we weren't aligning with the business like we had an intake process that's like Matt in such and such division wants you to build a dashboard and then you do it right and it's just it
[29:24] they weren't they didn't know what we were capable of right and so what we ended up doing is um taking a functional point of contact for like sales and commercial for example and said that is the business's single point of contact for all things data day
[29:41] I regardless if it's management of data and so on and we took a very agile type approach to the world where we meet on a monthly basis with like the like you know our CFO from the finance world and we say this is what we're working on top-down here's what
[29:57] everyone's asking us to do and boom now we have this great alignment with the business which is super fantastic but the secondary part of it is um instead of having data engineering sprinkled in all these individual functions for example we put that under one delivery
[30:13] organization and we we kind of have a couple of must-haves in the world of data so for example we want each product to be built using the same methodology the same type of approach to software so that we have consistency
[30:29] across our delivery and then that way it creates a really cohesive structure because somebody who was primarily doing just data engineering work could say hey I saw that there was like a more generative AI capability feature in our
[30:44] backlog. Um is it cool if I work on that and I get some help from somebody else?" And then I think we're we're we're developing this um this talent pool, which is unbelievable because you're learning RPA, you're learning AI,
[31:00] you're doing some data engineering, you're writing APIs, and I think that um the business is pretty happy because they're engaged with us, which is they haven't been before. They know the art of the possible. And our functional leaders
[31:15] are constantly pushing the envelope. So, I said, "If you go to build a dashboard for them, build a dashboard, they didn't ask you to build a generative AI chat, build that, and then do something absolutely mind-blowingly crazy so that they know that how good we are." So, closer to the business, more
[31:31] responsive, and then upskilling your team. 100%. That's great. So, as we think about, you know, I think the work you guys would did with Genie spaces was was pretty amazing, and I don't think dashboard we could debate if dashboards are going to live on and for how long. I don't think dashboard is the
[31:47] end state, right? The end state here is people stopping people from dropping data into Excel, pulling it out of systems. But, how did you see Genie change the experience for people going from first I look at a dashboard to then I ask questions to first I ask questions and then maybe render something
[32:03] visually? Yeah. So, every time we deployed a product, uh number one reaction is that can't be right. There's no way that that data set's right. And once we have actually validated it with, again, the business,
[32:18] so we're pushing it out, um that's solved a huge problem for us. But, we still have the where's the export to Excel button, right? And one thing that we're planning to do is take all of our reporting infrastructure and basically disable exporting data. It was going to be very disruptive.
[32:34] People are going to hate us, but it's kind of what we need to do, right? Um and so it's great because dashboards show you things, phenomenal, but they actually don't answer like the business critical questions that people are asking like um which customers have the worst DSO? For example, and can you maybe lay out some
[32:52] reasons why? And what's nice about that is it's it's something you truly can't the best dashboard in the world wouldn't be able to tell you those reasons why. And this is where the sad, mad, glad data points plus the ontology coming together gives
[33:07] a different look at data and now it has the business asking us things like how do I start taking action? So they start describing agents to us even though they don't really know they're describing it. Like um what you know, how come such and such
[33:24] person keeps expediting payment for a customer or for a vendor? And what why are they doing this? And how would you even visualize the random questions that people come up with? And truly I think we get the best requirements by looking at the questions that people ask Genie because those
[33:40] aren't the requirements that we get when we're asked to build things. That's awesome. Yeah. So I think you know, we've talked about disconnected data. I think everybody understands that the better you govern data, the better described, the better it's connected, um the easier it is for
[33:56] both machines and humans to kind of operate around that, right? And and that's been something we've discussed for a long time. I think the uh the opportunity and the investment has always lagged. You guys have done a good job of kind of rallying the business around you. So what's the next domain? We talked about customer, you gave some use cases, DSO, kind of how
[34:13] that bad data's manifested. What's the next set of domains that you're going to onboard, build some ontological relationships, and and get semantic alignment around? I would probably say assets. Okay. Yeah. Uh and probably vendor in parallel, quite frankly, because there's
[34:29] a um there's we're about to acquire uh a large company. And one of the things Forward-looking statements, please. Yeah. So, remind remind everybody. And and one of the first questions that people are asking us is where how can we
[34:45] find some synergies within our supply chain organizations? Who who are the vendors that we we both do business with? Who has better pricing? How can we leverage the the the new size of our organization? And so, vendor is definitely there. Um, and then our asset management. Uh,
[35:02] so sometimes assets just sit and they could be available for auction. And or we could use them differently. Um, but of true visibility into like uh life-based maintenance of tools and things of that nature are all going to be uh packed in that asset 360 stack.
[35:18] Do you see the reconciliation of data changing a little bit? 18 months ago, you know, Genie code wasn't there, Genie 1, some of the capabilities. But now kind of using the probabilistic approach and we've done some of that, right? But it's hardened around reconciling across, you know, uh a new company that's coming
[35:35] in versus your legacy data set. Do you see that being a an accelerator around getting to kind of a deterministic set of rules? Sort of. Okay. Yeah. Uh so, for us every time we've stood up a new domain in master data, um we we end up with like
[35:52] match rules that are perfect for good data and then all the garbage that's in between. Right? Um, and so typically, let's just say you had a million customers um and then maybe 200,000 of them are like need to be reviewed by a human, which
[36:08] we never really get through that queue. I I personally see Genie or some other AI tech tech uh fixing that problem for us. Like going through like um I don't think that AI is going to like clean all of our data. I think we need to do that. That's our job. Um, but I do think that that gray
[36:25] area, how are we going to reconcile and harmonize? AI has a major role to play there. The stewards have relief. Yeah. Finally. Thank god the poor stewardship community. Awesome. Um so, thank you everybody. Again, I think we've talked about enablement. I'll leave you with a a tagline from Genpact
[36:42] is um data intelligence and process intelligence really equals artificial intelligence. And and I think what we talked about here is getting data a little bit more intelligent, enabling the governance around it, marrying that with process, market to cash,
[36:58] um it could be procure to pay, whatever those are, but the more you can bring the data to the process, and then I'm seamlessly unify it, the better opportunity you have to actually agentify, and I hate that word. Um but I had to use it in this conversation. So, thank you. We have 3 minutes left, happy to stay after, but any questions from
[37:15] the crowd?

Learn more about the Databricks Data and AI platform.

The information provided herein is for general informational purposes only and may not reflect the most current product capabilities or configurations.