Skip to main content

Accelerating Maintenance with Conversational AI: How BP Uses Genie on Databricks

Summary

  • BP reduced the time to answer equipment maintenance questions from weeks to approximately 30 seconds by building a conversational analytics solution on Databricks using Genie on top of Unity Catalog, consolidating 60 source systems including 47 SAP instances.
  • Julie Nguyen shares a seven-step AI readiness framework that guides organizations from raw data consolidation through silver and gold data layers, metadata governance, and data quality controls before deploying conversational AI to business users.
  • BP is scaling the Genie deployment from 400 to 9,000 users, requiring deliberate cost management strategies and robust data modeling so that engineers across refineries, offshore platforms, and pipelines can access maintenance insights without SQL expertise.

Accelerating Maintenance with Conversational AI: How BP Uses Genie on Databricks

Watch: Accelerating Maintenance with Conversational AI: How BP Uses Genie on Databricks
Energy companies like BP manage complex operational environments across refineries, platforms, and pipelines where maintenance decisions must be made quickly and with confidence. Yet accessing critical maintenance data traditionally requires deep technical expertise, weeks of coordination between teams, and queues of requests to data engineers. This bottleneck slows decision-making and limits business users' ability to act independently.
Learn how BP built a seven-step AI readiness framework on Databricks to empower business users with conversational analytics powered by Genie. Julie Nguyen shares how consolidating 60 source systems into a unified silver and gold layer architecture, combined with robust metadata governance and data quality controls through Unity Catalog, enabled BP to accelerate insights from weeks to seconds. See the live demo of maintenance data queries, discuss scaling from 400 to 9,000 users, and explore cost management strategies for production deployments.
🤝

Chapters

FAQs

How does BP use Databricks Genie for equipment maintenance analytics?

BP deployed Genie on top of a unified silver and gold data architecture built on Unity Catalog to let business users ask natural-language questions about equipment status, failure risk, and maintenance schedules. This replaced a bottlenecked, queue-based request process through data analysts, reducing insight time from weeks to approximately 30 seconds.

What is BP's seven-step AI readiness framework for Genie?

BP's AI readiness framework begins with consolidating 60 source systems into a unified layer, then progresses through building silver layer data models, creating gold layer curated datasets, applying Unity Catalog governance and metadata, establishing data quality controls, configuring Genie, and finally scaling to a broad user base. The framework ensures data is trustworthy and well-structured before conversational AI is deployed on top.

Why did BP need to consolidate 60 source systems before deploying Genie?

BP's maintenance data spans multiple business units — offshore production, refining, terminals, pipelines, and logistics — each with separate systems of record including 47 SAP instances. Without consolidating these sources into a unified data layer, data analysts remained the manual bottleneck for every domain engineer request, making timely maintenance insights impossible.

How is BP managing costs as it scales Genie from 400 to 9,000 users?

This video addresses cost management as a critical consideration when expanding conversational analytics to a large enterprise user base, covering strategies to maintain performance without runaway query spend. BP's approach relies on careful data modeling in the silver and gold layers and governance through Unity Catalog to ensure efficient query execution as the user base scales significantly.

Full transcript

[00:08] So, I am I called my prep private chat GPT yesterday, okay, and asked how should I start today? So, my husband said, "Whatever you do, okay, do not start with a joke, okay?" So, here it is, okay?
[00:23] So, and I need guys to to participate with me, okay? Knock, knock. Who's there? Genie. Genie who? Genie. How long it take you to actually
[00:42] find out that your equipment is going to fail is have a potential to fail in the next 30 days in your organization? Oh, thank you for laughing. I need to tell my husband this, okay? Thank you. Thank you. Yes. Yeah.
[00:59] So, how long? A day? 30 minutes? 30 seconds? In BP, it used to take us weeks, weeks,
[01:15] right? Because we have 60 source system that only take care about maintenance, right? This is you talking about SAP, you talking about multiple 47 instances of SAP data, right? And then you have
[01:32] maximal, you have upstream, you have downstream, you have midstream, you have refining, right? And then you talking about logistic, right? You talking about activity integration. So, it's a whole lot of system of record. And
[01:48] you know, we have everything is like on the shoulder of data engineers and data analysts, right? Where we don't have that much domain knowledge, right? And our engineers sit out there is waiting for us to give them the data
[02:04] so they can have the answer to that specific question. Okay. So, today I'm going to walk you through how BP using Genie, right? On top of our Unity catalog, right? Answer that question now in 30
[02:21] seconds, right? So, let's go. Let's do it. All right. So, BP is as complex as you can see, right? We have so many different business
[02:38] units, right? Like I said before, we have offshore production, we have refining, we have terminal and pipelines, and all of them, right? All of them need safety, right? Like we we really care about our safety,
[02:53] need logistics, and we need the availability of our equipment, right? This is not a data problem, it's the nature of our business. Okay. So, the real challenge is how do you how that complexity now is show up in our
[03:10] data organization, okay? Where data engineers and analysts was working directly with the raw data set and sometime curated data set, right? But the business, they don't care about data bricks, they don't really care about
[03:27] sequel, they they they just want us to give them data in the form that they want it. Okay. So, with that, right? You can see in the black box, the analyst the data analyst is become
[03:43] the bottleneck, right? It's we're going to go through the requirements. This is what I want you to do, right? That requirements might take weeks, right? And then we take that back, right? We map the data, right? We map the data. And then we come up with sort of a
[03:59] solution back to the to the maintenance engineer. And they was like, "Yeah, this is not what we we really want, right?" And that going back and forth with for quite some time. And sometime it's there, but they were like,
[04:14] "Yeah, but we want something more, right? We want something more." It's always something more. Okay? So, uh Not only that, right? You can see the no AI readiness right now, right? All of us
[04:30] try AI, right? But underneath that, the data is not very consistent, right? That's 60 sources that we have, right? Everyone talk their own language, right? I didn't know there is 11 ways for us to
[04:45] spell Gulf of Mexico or Gulf of America now, right? That's many ways. Yes. Yes. Okay? So, you can imagine, right? If your data is dirty, is your AI good? No. No, it's not, right? So, how do we do it?
[05:01] This is seven step on how we do it, right? And I truly believe all of us here, right? We is somewhere in the seventh step, yeah? So, step number one, yeah? Data need to
[05:17] be together in one place, yeah? That place we choose is Databricks with our storage, right? Databricks on top of our storage. So, for BP, that storage could be anything. It could be ADLS Gen2 in
[05:33] Azure, or it could be S3 buckets. We don't actually force this at the moment because the Unity layer, right? The Unity catalog kind of solve that problem for us. So, we pull all that system of record together. Yeah?
[05:48] Step number two, we're going to build what we call the silver layer. Yeah? So, remember AI data readiness, and we will talk about this a little bit more. Yeah? So, silver layer, this is where we require there has to be a data model
[06:04] with proper PK and FK, right? It has to, right? For the AI later on, for our AI data readiness, and for us to put Genie, for us to put other agents on top of that.
[06:23] And then, in this in the silver layer, we stop speaking the SAP language. We stop speaking the Maximo language, right? We have our data where we actually name the table sort of, right? So, it's no longer it's going to be ABCD. It's no longer it's going to be
[06:40] that SAP PRO does, you know, XYZ, right? It will be equipment. It will be work order. It will be notification. It's something that all of us could understand, sort of. Yeah, so that is
[06:56] our standard. Yeah? And then, we move on and and look at the gold layer. In the gold layer, we now start talking a little bit more on the business, right? So, we work with our domain expert. How do you How do How do you want your data?
[07:13] What language do you speak, right? So, when they come to us, they know they're asking like, "Do you have uh Do you have work order data, right? Do you have object data? Do you have logistic data for this specific region for
[07:29] this specific asset, right? That's how they're going to talk." So, we're going to model the gold layer according to what our domain expert is going to say. Yeah? And for each layer, raw layer, what we do and just capture metadata, right? For
[07:47] silver layer, what we do our data model, capture meta metadata. The gold layer, same thing, capture metadata. In our metadata, we have to know, number one, who owns this data set, right? What is your data classification, right? And
[08:05] your data quality definition. You have to have it. So, for us, for later on, remember Genie will sit on top, our agents will sit on top of that. Um the next layer you have is Unity Catalog, right? So, one is is in Unity
[08:22] Catalog now, right? Your domain expert not going to look at it, but you do. Your data analyst will. Okay? So, we need to proper check this. So, later on when we put the what we call the AI data readiness scoring model, right? It's
[08:39] going to go through that scoring model and come out on the other side with the score of four or five. That's mean it's ready for AI at this time. Um Genie layer, right? Once you you're done
[08:55] with the the the AI data readiness, which I will talk a little bit more later, okay? So, the Genie layer is become quite easy. It take us probably 10 minutes to set that up, okay? And I will go through that process as well,
[09:10] okay? So, when the Genie is set up, the vision and action become quite natural. There is multiple way for you to actually deploy this Genie to your domain expert, right? For your engineering as a platform. And and we,
[09:26] you know, MVP, you name a technology, I'm pretty sure we use it somewhere. So, it's same thing at this, right? We currently do not have a standard on how we actually deploy our Genie our agent we deployed it through team to an API
[09:41] call to our own platform to AI AI solution to BI solution you name it we do it we do what our data expert or our business ask us to do at the moment.
[10:00] All right, so when you talking about AI data readiness, what are we talking about? So there are five things that we're looking at at BP. So number one is our data quality. It's very very important to us because as we see if we roll out something to the business
[10:15] to use and they look at it and they were like what is it and that's it. Once you lost the trust, it's really really really hard to get that back. So for us data quality is number one. Yeah. In our model it score 30% at the moment. Right?
[10:33] How do you define your day data quality? How do you actually do it? Right? Because data quality is not something that you just score one and it's always like that. Tomorrow the source system decide that they're going to add a new column. They're going to change the data type. Your spreadsheet that you pull
[10:49] pull in today is look a little bit different than yesterday and then your data quality is out of the window. Right? So how do you keep that? Is is is the is the ongoing things. Right? Is is the is the thing that you always have to care
[11:05] about. Right? That's is where data quality agents come in. Yeah. We will we will demo that in a little bit. Yeah. From a structure and standardization we work with engineers so
[11:21] engineering standardization is quite important to us. Right? We follow the industry standard. We were talking to our shell friends here. Right? Like in BP beside the engineering center, we also do like OSDU standard. We will put that in also. So, that is built into our
[11:39] data model at the moment. Yeah. And then, uh in the uh Sorry. In the metadata and lineage, like I mentioned before, right? In the data metadata and lineage, we always capture
[11:56] metadata and track the traceability. Now, with data brick, you can trace all the way back to where the data source come from, right? We do also capture the data source. So, we know the data source data quality from the origin. And then,
[12:11] from there, we track it all the way into the golden um uh the golden data product that we have. And then, with that, the the the Genie sit on top, right? But again, Genie sit on top. Genie right now is have
[12:27] Sometimes it's have a mind of its own, right? How do you actually make it think the way that we want it to think, right? Make it respond to the way that we want or our business want it to respond. Yeah. Uh the next one is accessibility and
[12:43] usability. If you working in the main data maintenance, you will see the equipment the functional location, the notification is not only used by the maintenance engineers, right? It's It's the most popular data set in EDP at the
[13:00] moment. So, we do share that with quite a few number of other areas, right? So, so that whole data set is fit right into the digital twins area. We use that the same data set will be fit into the production uh management work. It will
[13:16] go into the um uh the risk management, for example. So, the usability and the share of our gold data product become quite important, okay? So, when you enter go, you go through the AI data assessment,
[13:32] it will have the certification certification for each of your gold data product come out. So, you know that if this data product is able is certified to share with other area for them to use before you put in any
[13:47] anything on top of the gold layers. Okay. And last and not least is the the government, right? Um It's a data the data that we have, right? Has to
[14:02] have the data owner, the data steward, right? The organization structure, okay? Sometimes BP gone through quite a bit of reorganization as you see in the news, right? So, before we have P&O, now we're
[14:18] going to have upstream and downstream, right? So, the data product also go through that transition, right? Go go through that transition. And today, this data product could belong to P&O, tomorrow we're going to have to break down the maintenance now, right? so we're going to have upstream maintenance
[14:33] data, we're going to have a downstream maintenance data. How do you manage that? We manage through the data government process, right? Uh and a lot of data mappings, I guess, right? Right. All right. Um Next is the is the genie itself, right?
[14:52] The genie itself is actually the the easiest part here, right? When you have your data ready for AI, genie become the uh the cherry on top of the cake, per se. Okay?
[15:07] Uh At this moment, before we go through deeper, I will do a demo. Okay?
[15:26] All right. So, oh I will need to move.
[15:53] I think we should duplicate the screen, right? Oh, you need it there? No, I need it here. Or else it's going to be hard for me to All right. The way we have it, they they're not going to be able to see it. Okay. All right. All right. I'm going to walk it over here.
[16:11] Yeah. So, this is our equipment uh maintenance management. So, as you see with the first start, you see the list of general instruction there, yeah? So, that is a skill that we put in for our maintenance model. So, first we're going to teach
[16:29] we're going to talk to Genie and ask Genie tell Genie that we build this maintenance um uh management, right? Using what engineering standard. Yeah? So, we talk to it a little bit. Even we capture that already in our metadata, right? We
[16:46] capture that in the government. When we build that, we put that into the building of the data model, but we tell Genie anyway, yeah? Right? And then after that, we give them some we give Genie some uh uh instruction, right? Mean time between
[17:03] failure, right? So, we we we we teach that what it what that mean, yeah? Uh later on, I will show you the data model and what is news with Data Bricks in term of the data model. But at the moment, right? I think this week they have a new release in term of the data
[17:19] model and how data model will work with uh will work with Genie, but at the moment you still have to tell Genie the relationship. Yeah, you still have to tell Genie that. That is how you build your data model, how is your work order and your equipment, right? Uh uh uh have
[17:37] the relationship with your functional location, how is your notification have relationship with uh uh with your order, for example, right? You So, you still have to teach Genie that part. Yeah? Uh
[17:52] And then, you can see over the here, right? We will need to give definition because this is what our engineering will use out of the platform, right? Out of the the the the wells that they they operate in. So, this is what they're going to see. And then,
[18:09] in here,
[18:35] I need to duplicate it, right? Hold on, let me duplicate my screen.
[19:06] It's already duplicate, so can I duplicate here and there? One and three. Maybe one and three.
[19:29] One and three. Yes. Yes. Thank you. Oops.
[19:54] It's not responding very well. Yes. Oh. So, that is here. It's not here. Yeah.
[20:14] So, currently the limitation for Genie is we can only use from 25 to 30 um uh 30 tables or views that you can use uh uh but after this week, that will be uh that will be different, right? That will be upgrade. You will have more um
[20:29] uh you will you can work with more uh uh table or queries or uh um or views that you can. But right now with the limitation that when you build when we build the maintenance, you stay with that 30 tables that 30 view 30 days
[20:45] taste tables that you have, right? So, if you need a bigger model, you can keep building Genie 1, Genie 2, and Genie 3, right? On top of that, you can use the uh the agent supervisor, right? So, when you have the agent supervisor, they will
[21:02] work with all the they will pull all your Genies together, and then uh and then you can have a bigger or more data to work with. Yeah? Uh I guess.
[21:20] All right. Huh? All right. So, uh my attempt to share with you the demo is is is not uh Yeah. All right. Hold on. Let me see.
[21:43] Yes, that's good. All right. We're back in business. Yes. All right. So, uh you can see down on the bottom, right? It's very I have to actually
[21:58] tell it's scale, right? This morning in the in the keynote, they're talking about scale. So, this is just an example of that, right? So, when calculating cost, always sum from cost line table. So, we make sure it stay in line, right? Genie, it's just like my daughter, right? Don't
[22:15] touch it. Like that, right? So, now we put that back in line. Yeah? So,
[22:31] This is the data area, right? So, in the data area, you can see that we pull together, right? And this is at 29 tables. So, I'm also playing around with limitation too, right? 29 table. Uh you have the instruction. You have
[22:46] the about, which is telling them this is the metadata that I'm talking about. And after this week, this will go to Unity Catalog, right? In the in the session where you have all the agents together. You have all the genie together. So, the
[23:02] more you tell Unity Catalog about your agents, your uh genie, your data, the easiest it's going to be for you to manage. Right? So, uh back to
[23:19] back to the catalog itself. So, now we're talking about raw, silver, and gold, right? And then we talk about data AI data readiness to prepare for our agent. So, how's what's this that look
[23:34] like? Get? So, we start with our schema, right? So, this is the silver area that I'm talking that that we were working through earlier. So, in the silver area
[23:50] when you do the data model, right? Uh currently, Databricks cannot show you or they not be able to show you the entire data model itself, right? The diagram itself. But after this week, they can, right? Thanks to BP, actually. We are
[24:07] the one who suggest that we should be able to see the whole data model. But for now, you can see, right? You can see your PK right here, and then you can view the relationship, right? So, you can see that. Right? So, now, if tomorrow you're going to feed the data
[24:23] in here, even it's not enforceable, right? But your PK is, right? Your PK is. So, if you cannot put the data into this table if you have a primary if you have the data and the primary key column that you have is is a notable column.
[24:40] So, some has been improved from a Databricks standpoint. Thank you. Databricks. Yes. Right. Yeah. Yes. Yes. So, you can after this week, you will be able to see a bit of your data model that you built
[24:58] in here is going to take a lot of it's it's going to start taking that step over to the Genie area. So, you will not have to set up the relationship in Genie anymore, but that will be a step-by-step, right? We're working the baby step. Yeah?
[25:19] Uh you can see here that we have the column definition, right? So, the column definition quite important here because this is how Genie is going to find your data, right? How Genie going to find your data. And it's make a big difference in the cost model as well. So, what we did is we test we test a
[25:36] root exactly the same table without definition without key and then we take the same set of table where we set up the proper definition and proper key and then we let's Genie take care of it. We're asking the exact same question on
[25:52] how uh uh what is our top 10 most fell equipment data for example and then we will see that Genie, right? So, the root that the Genie that with proper data model, proper metadata, proper definition can give you the answer
[26:11] four time quicker than the one that Genie has to go and find the data itself, right? Without definition so they had to actually scan the the the content much more in order to find the answer for you.
[26:31] And here is the metadata that we're talking about, right? So, these will be fit into our AI data readiness scoring. And then when you see you can see that, right? From the data owner to data classification to
[26:46] uh frequency to the retention period all that, right? Into the data tag and then the contact support. When we see other stuff in term of the metadata, it would score you it depends on it depends on how complete this it is
[27:04] when you capture the metadata, right? So this is the one that when we when we build our silver, we capture the metadata and then we have a automatic workflow to flow and push all that data
[27:19] into data brick, right? So this will live into your system catalogs in the in the information schema. So you will see that and our AI readiness scoring model would tap into the system table
[27:34] as information schema and get the data and turn into the score for us to know if this is ready for us to put agent on top or to put Genie on top.
[27:58] You can see here is the lineage and here is the data quality. So we testing the data quality that provided to us by by data brick is right now is a POC in the POC period we still have our own data quality capture where we define
[28:14] what is mean by good, right? And then we also feed that into data brick and data brick actually runs a workflow and then spit out the the score for us in term of the the data quality score. But we are testing it with data brick
[28:29] right now to see instead of using our own tool to do that to capture the metadata to capture the data quality if we also can do that in data bricks.
[28:45] All right. Back to Genie. Okay. So, we're going to test this together. So, once we set up the Genie and it's if you have your data ready, right? AI data ready already, it's going to take you
[29:02] about 10 minutes to set it up and that's it. Okay? And now you can run. Which vendor have the best performing rating? Okay? So, we ask Genie.
[29:17] Right? Start thinking. It's thinking, right? Yeah. Okay. This is my uh my test box, right? So, don't have a lot of of of cap juice in there. Okay?
[29:34] So, we're waiting. While we're waiting for Genie to think, right? Do you guys have any questions? Yes. In your designing you mentioned you used a lot of silver layer. Yes. Any experiences using silver as a core
[29:52] for your Genie? Um we actually when I when I did this for the presentation today, right? Of course, I could not use BP real data. So, everything that I build right here, I use Genie card, right? And uh and a
[30:09] little bit of class, actually. Okay? So, the way that we There are two way for you to build your silver bottles, get your data your silver layer. One is the way that you already have the data. Okay? You already build something. Someone already curated your 60 sources
[30:27] or your 100 sources already. Okay? So, now what you have to do is you will need to ask You can You can use a Genie card for that, right? Ask them to go through your data and do what we call a data profiling. Okay? So, that's mean look
[30:44] into my Look into my data objects. Look into my table, right? And give me the relationship, right? Give me the the data profiling for this. They will do that. And usually usually it take probably about a few try for for it to
[31:01] get it to what we you want it, right? The other way is you start fresh, right? So, yeah, now you have a new project. Okay, you know your data source. You got to have to go and build your silver layer, right? But you don't want to start by go into some sort of uh data
[31:17] modeling tool, right? And start doing entity modeling. Start to do That's going to take you a quite some time, right? You can ask Genie code to do that. Maintaining management, right? SAP and logistics, those are the industry
[31:34] standard system, right? There are open source out there. You can actually cloud code it, right? Or you can use Genie code in this situation to build you the model. The model will not be exactly 100% of what you want it to be, but it
[31:51] will give you pretty much if you if your organization is follow the industry standard, it give you the 75% of that already. And then from there you twist it according to your uh your subject matter expert in the company. So, that
[32:06] is what the process that we currently use right now. Uh the result is back, right? Wonderful. Right? You can see that in here, right? Right? They give us Right? This is again
[32:24] my mock data, right? It's just not real data. So, if your company name show up in there, uh blame uh Genie code. Yeah. So, you can see it give us, right? It give us the uh the result. It's give us the specific vendor and it's give us the
[32:41] performance rating as well. Okay? You can continue this conversation, right? You can say you can ask more, right? What does that mean by 4.5, right? They will explain the process to you. Okay? You can start
[32:57] what do you recommendation that you have for us? How do we do better, right? How do we do better? If you ask what that should do to do with say I I I can't give you that answer. But if you ask what your recommendation, Genie is
[33:13] always is going to give you the recommendation based on the data that you currently have, right? And if you wanted to say go out in the industry, look into Shell, see Sorry. I'm picking on you, right? And see what
[33:28] the industry out there do, right? They will also give you that perspective as well, right? So, I see we rolled this out MVP, we rolled that out to 400 users, right? Our business users. We have a very a good feedback back and all we
[33:46] also have quite um can you do more? Can you check into this? Can you do that? So, this morning I actually called Srini here and say Srini, we're ready to roll that out for 9,000 people, right? And with the new um
[34:03] with the new cost model that um that Databricks have is is worrying me a little bit, right? Because the Genie model now is you have $10 free per user per month, right? And that $10 free is equivalent to 150 DBU and it's
[34:22] equivalent to about 80 to 100 question. Yeah. And it's could go pretty fast. You can see that, right? And then after 100 question now you pay as you go. Right? And then if you give this to a maintenance engineer that's sitting
[34:39] out there and they on three weeks on three week off shift, this could go pretty fast. Right? This could go So cost management is another thing that we will have to pay attention with the new cost model. It's It's a new cost
[34:55] model is good if you use in the moderate amount. Just try anything else, but if you start, you know, asking a tons of question just like our user currently is, and then that is something for you to
[35:11] implement the new operational model that in the in the keynote that we have this morning, right? With the new with the new governor, and you actually can put control specifically cost control on any genie, any agent that you
[35:27] have, so you can tailor your cost down a little bit. Okay? Uh I will be available down there if you have any more question, but I think this time is the time for we get. Thank you.

Learn more about the Databricks Data and AI platform.

The information provided herein is for general informational purposes only and may not reflect the most current product capabilities or configurations.