Multi-Agent AI for Upstream Intelligence: PETRONAS Case Study
Summary
- PETRONAS, a Malaysia-based oil and gas company, built PetroVisor on Databricks to unify upstream intelligence across 1.8 million engineering documents and multiple systems including OSDU, PI, and SQL, where 90% of data was previously unstructured.
- The multi-agent architecture routes queries through a supervisor orchestrator to three specialized domain agents — Subsurface, Surface, and Sensor — cutting data discovery time from 3–5 days to 30–45 seconds with over 90% query accuracy.
- Governance and quality assurance using LLM-as-judge evaluation with MLflow monitoring ensure the system is model-agnostic and maintainable as PETRONAS plans to expand to additional domains including P&ID diagram reasoning.
Multi-Agent AI for Upstream Intelligence: PETRONAS Case Study

PETRONAS faced a critical upstream challenge: 90% unstructured data fragmented across multiple systems (OSDU, PI, SQL, and 1.8 million engineering documents). Engineers spent 3-5 days just searching for data. The company built PetroVisor, a multi-agent assistant on Databricks that routes queries across specialized agents (Subsurface, Surface, Sensor) and a supervisor orchestrator, reducing data discovery from days to 30-45 seconds with 90%+ query accuracy.
this video showcases how consolidating structured and unstructured data into Unity Catalog, building a semantic layer, and deploying purpose-built domain-specific agents enables enterprise reasoning at scale. Learn the architecture, governance patterns, monitoring with MLflow, and how multi-agent systems solve the data fragmentation problem that fragmented AI systems cannot.
Chapters
00:00PETRONAS Upstream Intelligence Challenge and PetroVisor Solution02:37Vision: From Barrel to Dollars Faster with Trusted Data Foundation04:14Data Landscape Complexity and Interconnected Domain Systems06:42Real Cost of Data Fragmentation: 3-5 Days Per Discovery Task08:05Strong Foundation Required: Why Generic AI Is Not Enough09:09Architecture: System of Records, Unity Catalog, Semantic Layer15:25Multi-Agent Architecture: Supervisor with Domain-Specific Agents17:02PetroVisor: Live Demo with Document and Data Agents22:10Vector Search, RAG, Hybrid Search, and Genie Integration24:02Quality Assurance: LLM as Judge, Ground Truth Validation, MLflow Monitoring25:06Key Lessons: Architecture Over LLM, Governance, Model Agnostic Design26:09Future Roadmap: Scaling to Surface Domain and P&ID Diagram Reasoning
FAQs
What problem did PETRONAS solve with the PetroVisor multi-agent system?
PETRONAS engineers were spending 3–5 days searching for data scattered across OSDU, PI, SQL systems, and 1.8 million engineering documents where 90% of the content was unstructured. PetroVisor, built on Databricks, consolidates this data and returns answers in 30–45 seconds with over 90% query accuracy.
How does the PetroVisor multi-agent architecture work?
PetroVisor uses a supervisor orchestrator that routes incoming queries to three specialized domain agents — Subsurface, Surface, and Sensor — each purpose-built for its domain. The agents reason across structured and unstructured data consolidated in Unity Catalog through a semantic layer, using Vector Search, RAG, hybrid search, and Genie integration.
How does PETRONAS validate and monitor PetroVisor's accuracy?
The team uses LLM-as-judge evaluation against ground truth datasets and monitors model performance continuously with MLflow. This combination of automated evaluation and operational monitoring ensures the system maintains accuracy and allows the team to detect regressions as new data or model versions are introduced.
Why did PETRONAS choose a multi-agent architecture over a single AI system?
Generic AI systems lack the depth of domain knowledge required across PETRONAS's complex upstream data landscape, which spans geology, surface operations, and sensor data simultaneously. Purpose-built domain-specific agents allow each to be optimized for its area, while the supervisor orchestrator coordinates across them to answer cross-domain questions at enterprise scale.
Full transcript
[00:08] Uh thanks everyone uh for joining uh today's session. Uh In today's session, we are going to talk about how Petronas, which is Malaysia based, one of the biggest uh oil and gas company, is using generative AI to unify intelligence in the upstream business.
[00:25] Uh quick introduction about myself. Uh my name is Prashant Singh. I'm a senior solution architect at Databricks. And I'd like to introduce Musrin to take over the session from here. Okay, good morning. Um very pleased to be here in this morning and um
[00:42] um we try to uh talk about this and uh I can see that we can have some uh discussion towards the end of the session. Okay, it's going to be I'm going to I'm going to I'm going to make this a bit brief at the beginning. I think we have more interesting thing towards the end. We we're going to show
[00:58] you some demo. And then take it from there. Okay, um yeah, uh I'm Musrin from Petronas. Again, coming all the way from Malaysia. So, it took me um 2 days ago uh 42 hours to get in. Missed some connecting
[01:14] flights because of the luggage and so on and so forth. Long story short, I'm here. I'm happy to see you guys. Uh I see if you jet lag. Okay, um let me start with something that probably uh some of us can, you know, relate to. Uh some statistics. Uh
[01:31] so, is is this an open secret? Uh McKinsey reported 1% at least spend um 1 day every week just to search and gather all those information. And Gartner reported that um
[01:47] a company um um spend I mean losses around 20 12.9 million uh uh, because on this uh, uh poor data quality per year. So, for us it's not just a
[02:02] technical data problem. For us is a real productivity and decision-making problem. So, that's why we are really, um, focusing on this as as as big in organization in Petronas, 1% of change
[02:19] can translate to millions. So, that's what we are doing and and and and, uh, today we share how we change this by, uh, unifying upstream intelligence with GNI. Of course, partnering with, uh, Databricks.
[02:37] Okay, um, yeah, we started with one vision. Always, so that we know where our North Star is. Um, for us, we're going to, you know, have this target from barrel to dollars
[02:55] faster. You know, so it's it's always it's always it's always interesting where, you know, the top guys in the company want always something faster and faster, you know. Uh, but it's not we must we must be very careful. Yeah, we can get faster over
[03:11] time. We can get faster and faster over time. We want to make sure also that we can depend on those decision that we make. Right? So, um, as I said, it's not for us it's not just to build another data platform or any AI tool.
[03:28] It is just it is about, um, to help the team to get decision to make decision better and uh, faster from subsurface until production, maintenance, HSC, and so and
[03:43] so forth until business value. So, that is the goal. For us, for us speed matters, but only when the data trusted. So, that's important. And to understand that vision, we first need to just I I just
[03:58] uh put it in high level uh the data landscape of what we are dealing with. So, this is very level because I don't even intend to go to a you know every aspect of it just to give some some some context and I I'm sure everyone is aware of it. And um
[04:14] the main message here is upstream data is very broad, complex, and highly connected. Um So, as I said, we covers all those domains, you know, subsurface um has its own complexity, and coming
[04:32] into production, and maintenance is the biggest chunk of of what we call this operation cost, where the money we spend every year, the highest money we spend. And uh HSE and um until the um uh decision is made.
[04:49] So, in upstream, it's really the answer is not just sitting in one system, uh but rather it is highly dependent on a few uh uh complex system, and the chain of of a
[05:06] few things. Okay, and because the landscape is so broad, um now the challenge is very very very very clear to us. The data exists, but the answers are not. It's very hard to reach. This This is the real
[05:22] challenge for us. we are dealing with it's not much to say that millions of documents um from every domains.
[05:39] So, we are talking about first from daily operation, all those reports, all those maintenance, all the drawings. Um as we speak, we have drawings alone, we are dealing with around 1.8 million drawings.
[05:55] So, that's how complicated it is and then we know which one we should focus on to ensure that they uh decision is is um is uh I mean where we are going to focus decision on. So, that's very huge.
[06:11] And because of this because of this um um um very broad uh data and documents and drawings this create delays. Um repeated work um and loss of the
[06:26] context. So, this is where the challenge is. This is the start of everything. Okay, and uh for us this is not just an inconvenience um when the data is fragmented
[06:42] there is it causes the um the real cost to the organization. Um yeah, the the the team spend around in Petronas, what I'm saying is
[06:59] on average 3 to 5 days just to look for data, gather it and do some analysis on it. That is very conventional. So as the world's going to more challenging space, I think we cannot live with that
[07:16] and the margin the the margin becomes smaller, I think we we need to do faster about that. So um this fragmentation slows down the technical decision. Uh that is the dangerous thing and uh
[07:32] business respond is um compromised. Okay, and a few days may sound small uh but for us but for especially for in this uh in this complex upstream um operation
[07:49] all that matters a lot to us. So, we are trying to reduce that from days to minutes. Okay, now the message here is the AI is powerful. But the more important thing is the
[08:05] foundation must be very very strong. That is the real key. So, we found out in our few a few years journey in AI uh it works in the beginning in the first three to six months, but it slows down after that because the foundation is not
[08:21] stable and uh then it becomes just another investment that do not return any value to us. So, we try to avoid that. So, that's why I said in the beginning it's not just another tool. We are really uh focusing on resolving something in in in the
[08:37] organization. So, we need a platform that connects all these structured and unstructured data um to support multiple use cases. So, we are very um uh fortunate here because we have been
[08:53] doing quite a few things with Databricks. Um that I see this this is one of the uh latest thing that we collaborated together. And um we we did not start with chatbot. We started with the foundation behind
[09:09] the chatbots. That that's our that's our thing. Um the um Yeah. So, so the question becomes how do we solve this properly? So, always the
[09:26] answer we start with some architecture. I think this you always see this this is a typical architecture, but for us this um the foundation is to bring together system of records at the at the very bottom layer and then um governance through the UC of course
[09:42] Unity Catalog and semantic layer that can make sense of those data coming in I mean going in and but again the the the the architecture alone is not enough and the the
[09:59] the the the underneath the foundation underneath must be very strong as well. Okay okay I think it's not working. So now I'll leave it to Prashant
[10:15] and to proceed with certain thing before you show you this some demo. I'll take it back to see it. Thanks Muslim. Thanks for setting up the context. What Petronas was facing was a immense challenge. You heard about the
[10:31] vastness of the amount of data that Petronas is holding and those are not just one year of data those are the data that is coming from 20 years so you can think about how the PDF files were looking like from 15 20 years back and how the drawings will look like. So the challenge what we had was how to help
[10:48] Petronas to solve this challenge where engineers geoscientist everyone is spending four to five days just to find that data and getting insight out of it and then making the business decision was somewhere goes more than one to two
[11:04] weeks. So what we did was and the Petronas was actually working on couple of iterations of using AI how to use this AI to solve this problem. And what happened was as Muslim was talking about AI was everywhere. Each of these upstream application whether you take
[11:20] OSDU or you take some other internal application every application somehow in different form or other they have the AI. But eventually what happened was you have 20 applications. Each application is having their own embedded AI. But
[11:36] still you're Now, you lead from the problem of data fragmentation to the AI fragmentation. Your information is not in a single place. So, that was not solving the problem. And what it turned out to be that AI is not just the answer. The answer what
[11:53] lies is in building a very strong data foundation. And this is where we started building this data architecture where we are bringing all these system of records, whether it's a seismic log, uh well drills, uh drilling activity logs, production data, both the
[12:10] structured, unstructured data in a single place. And on top of it, what we are doing is Unity Catalog helped a lot because now we have the structured, unstructured data in a single place. Unity Catalog gives that metadata, the lineage, and also the context on this. And on top
[12:28] of that, then we started using some of the very best capabilities of Databricks to run the use cases. Those use cases are not only just AI use cases. It is also addressing some of the very high business value use cases like finding emulsion in oil tanker,
[12:44] which is another very high value use case because if there are oil if there are water droplets in the oil, then those oil tankers are getting rejected. But it also helped us to basically run our next generation of AI use cases.
[13:00] So, what we solved with this architecture is two fundamental problem. One is consolidating all the data. So, all the structured, unstructured data we consolidated. The second one is having a very strong metadata layer, lineage, auditability, and governance because
[13:17] this solution what we are rolling out is going to be rolled out for 3,000 engineers and geoscientists across Petronas. So, we need to have a very strong governance on top of it. And what Unity Catalog helped us, and you might
[13:33] have heard from Ali yesterday as well, AI is as good as your semantic layer and context. And the problem with fragmented AI, or we call it as AI sprawl, is when you have AI embedded in many applications, there is no unified
[13:48] context. And what Unity Catalog helped us to do is build this unified semantic layer and a context layer on top of it, we were able to build our purpose-built agents.
[14:03] So, as I was talking about that it always used to come so when whenever I used to have this conversation within Petronas, it always used to come this natural question that why can't we just plug in all these AI together? And there are reasons to it. The first is
[14:21] all these agents which is available within these applications are very much vendor locked in. Data residency is another big problem. So, you have to move data from one place to another place, and that leads to subsequent challenges of latency, concurrency issues, and
[14:37] a lot more. Uh another challenge what we have found is there are many use cases where you want to do reasoning between the different data what you have. So, you have some unstructured data, you have some structured data, and you want to reason on top of these data using AI. That was
[14:53] not possible. Observability was a huge gap, and if you don't observe something, you cannot improve the quality of that. And the knowledge was very static. So, when I say knowledge was very static, that means the data was not getting updated. You don't have a single place,
[15:09] so you don't know which data is uh how stale and which is fresh. So, these were some of the fundamental challenges we had. Uh so, how the agent that was built on Databricks is solving the problem? So, when we build all these
[15:25] uh system of record in a single place, we get the end-to-end observability. So, we know how our agents are behaving, who is accessing that. The second is continuous improvement because now we have got all the uh agent telemetry is all the agents
[15:41] that are stitched together to answer the queries about the user, we can improve these agents over the iteration. Data is in single place, so we can do batch real-time injection. Uh all the sensor data coming from different fields, different uh
[15:57] uh Pi systems are in real time now. Uh semantic accuracy, this is this is going this is one of the area which has helped us to improve our accuracy, the agent which are built within Databricks itself. We can do multi uh system reasoning.
[16:13] Now, we can do reason on different data sets. And then the data stays local. We don't have to uh shuffle around lot of data. So, this is what high-level uh looks like. What we did was using all these
[16:30] four fun fo- three fundamental pillars, we were able to build very purpose domain-specific agents. So, you can see there is a subsurface agent, there is a surface agent, there is a sensor agent, and then we have a supervisor agent which is orchestrating around all these
[16:46] agents according to the intent of the user query. And which resulted in a project we call it as a Petro Visor. So, what exactly the Petro Visor is? Uh I think this is working.
[17:02] Uh so, what exactly Petro Visor is is a multi-agent assistant which help uh upstream engineers and geoscientists to get the data on their fingertip. Uh so, there are two basic type of agents we have within
[17:18] PetroVisor. One is your document agent which help you to answer queries on unstructured data, and then you have a data agent which we are using is Genie behind the scene. And then I have a uh supervisor agent which is supervising
[17:34] around all these agents together. And doing this, what we have found is the challenges we had was getting the answer from 3 to 5 days it turned out to be that these engineers and geoscientists now can get the answer within few seconds.
[17:52] We did a very extensive evaluation uh in term of We did a very extensive evaluation and it turned out to be more than 90 plus in accuracy. We did a
[18:07] survey also within Petronas to understand how the user satisfaction look like. And in our MVP and uh the user satisfaction was 4.25. Uh number of active users that is going to be rolled out is is going to be 3,000, and we
[18:24] wanted to have a full telemetry and auditability to look at how the AI is being used uh is not being getting abused within the organization. So, all these things are available. So, having said that, the next what I
[18:39] want to do is I'll walk you through the demo how the PetroVisor look like. Okay. So, this is not exactly the PetroVisor
[18:55] UI. This is one I write coded two days before using Genie code as an uh Databricks app, but the idea is this agent is behind the scene behind this uh UI. And what you can see here is
[19:11] I have this UI where I can go and ask questions. And what happens is when I will ask this question. So, I have this question that what are the latest drilling activity happened for uh well F12 and
[19:27] how it is impacting my production in 2014. So, what you will see here is still thinking, but I want to get your attention in this agent activity. So, you can see here there is a supervisor agent. Then I have this rag agent, which is
[19:44] querying the unstructured data. I have a genie agent, which is looking at the structured data to get the answer. And then it is trying to correlate this insights together. What is happening here is the supervisor agent is
[20:00] trying to understand the intent of the question and then decides whether what will be So, it is trying to understand the intent of the question and with the intent of the question it decides that do I need to call both the agents or I
[20:17] need to just call one agent. So, in this question it has called both the agents together. Get the information from these two agents, correlated these answers, and then it presents answer to you. So, you'll see here there is some drilling activity has happened and it says that
[20:33] in October this was the overall production. Uh production has gone down by 70% then in December it was somewhere around zero because of some did something happened. And this data is not present in one single place. So, there are two
[20:49] different systems. One is your unstructured data and there's a structured data and it has correlated this. And to get these kind of informations without AI and and without having this multi-agentic system, it is very challenging because then you have
[21:06] to do everything correlation everything manually. I'll also show you the next question. So the next question I have is Okay, show me just the oil production for F12 in 2014. You will see here is using this intent of the question
[21:23] the supervisor agent is just calling my data agent which is Genie to get the answer which this is present in our one of the Delta tables within Unity catalog. Yeah. So
[21:38] this is the result. So now switching back, this was just a small demo I wanted to showcase that this is how the Petro Visor look like. That's not exact UI that is being used in Petronas. Petronas has embedded this supervisor agent in their in-house
[21:55] application which is being used by the upstream engineers and geoscientists. Okay. So this is what I was talking about that I showcased you that how
[22:10] these two agents are being called together. Now what exactly how this is working behind the scene is we have this vector search and rag built for unstructured data and Genie space is for answering the
[22:26] queries from your structured data using the natural language. One thing I wanted to tell here is we have not just used vector search for the similarity searching, but we have done uh some improvisation on top of it. We use
[22:42] hybrid search and what exactly hybrid search does is it uses the similarity search along with the lexical keyword matching of B BM25. And it actually helped a lot because
[22:58] when you have these millions of document, you want to be very accurate. And what we have found that with hybrid search, our accuracy went up by 20% compared to when we started building this PetroVisor. The second thing because we were dealing
[23:14] with the vast domain of the data, uh unstructured data basically, well logs, and then you have maintenance logs, then you have uh the sand logs uh for petrophysics and all. So, we use metadata filtering. That also helped us to not only uh
[23:31] filter the data faster, but also help us to get the result faster from our LLMs because now you have very small uh token size to input token and output token as well. So, that's what runs under under the
[23:46] hood of uh PetroVisor. And not only that, because uh we built this, and the next thing is if your AI is not giving you a right or quality answers, better you don't use it because you don't want a wrong information.
[24:02] So, we started with uh LLM as a judge, which is within uh Agent Bricks and Databricks platform. And every response is scored and stored within the inference table in Databricks for later uses.
[24:18] We also built a ground truth uh Q&A. Uh we worked with domain experts and evaluated each and every result when we were building this uh uh project. And we extensively use MLflow monitoring to actually capture each and every response
[24:35] and look at uh the model degradation, the LLM model degradation, whether it is degrading before the users start complaining. So, we monitor everything.
[24:50] So, what exactly we have learned uh while building this project, and it took uh us almost 7 to 8 months to build this from the day we started uh this project. The first is generic AI GenAI is not enough for building these kind of complex
[25:06] system. You need a very specific domain specific agents to work together. Just throwing an LLM is not going to solve the problem. Second thing is the most important lesson that we learned was architecture
[25:21] matters more than just an LLM because this is the architecture which helped us not only just call one LLM agent, but it helped us to chain the different agents and tools together. And governance and flexibility. Now, when we are dealing with the vast domain
[25:37] of the data and different types of the data, the governance become the core fundamental that we wanted to put and we started with small and what what it turned out to be the if you don't put a very strong governance in the front of the users, then there could be
[25:54] scenarios where the data leak can happen. Some different team will be able to look at some other data and there's lot of challenges will will be there. So, where where we are headed? This is not the end for Petro Vision. What we are headed is
[26:09] this architecture what we build is actually helping us to scale this Petro Vision for other domains as well. So, we have started with subsurface. We are expanding this to surface because of this modular design where we can is
[26:26] very much like a Lego design where you can hook different agents together and you can have different supervisor agents as well. This modular architecture has helped us a lot. The next phase what we are doing is we are going with all the P&ID diagrams within engineering documents
[26:43] and then see that how AI can help us to answer the queries on those P&ID ID diagrams. Which is a harder one is is is quite challenging, but we love challenges. We love to solve challenges in Data Bricks. The next is
[27:00] this architecture is model agnostic. So, every week or month you see there is a new model LLM model is launched which is having a better accuracy than the previous one. So, this architecture is
[27:15] very much model agnostic. So, you can switch your models according to your choice or what has happened with Fable today today, it might happen with some other models as well. So, you always need to have your architecture which is model agnostic so you that don't run
[27:30] into these kind of scenarios. And also when there is a new model which is cheaper in the market, then you can switch the models. The next one is governance as I talked about you should have a very strong governance because lot of people also put
[27:48] PIIs and send companies confidential information into the AI. So, you need to have something like AI gateway even Ali was talking about yesterday. So, that kind of interface you need to filter out those informations and have the guardrails.
[28:06] Before ending, I would like to call Mushtaq again to share what he looks how the upstream is going to be going forward and using GenAI. So, I'm going to close it with one simple thought. Um
[28:22] for us every decision matters and every delay has a cost. So, that is very key to us and this Gen agent based GenAI has helped us very much moving from searching and gathering those data to
[28:38] acting or acting based on that trusted intelligence. So, that's we want to be and I truly believe um that we are going to make better decision faster so we can see
[28:53] that in the next 6 months or so. Right? So with that, thank you very much and we have around 9 minutes or 8 minutes if we we have some questions from the floor. If you're not then
[29:13] Okay? Thank you.
Learn more about the Databricks Data and AI platform.
The information provided herein is for general informational purposes only and may not reflect the most current product capabilities or configurations.