Skip to main content

Scaling Enterprise Data: Lakebase for Legacy Systems and Operational Distribution

Summary

  • Quantum Capital Group uses Databricks Lakebase as a PostgreSQL operational data store to deliver ODBC connectivity to legacy desktop tools used by geologists, economists, and deal teams, bridging the gap between lakehouse analytics and operational workflows.
  • The firm ingests 1.5 billion records across energy markets into a medallion architecture, applies geological modeling to tier well assets by internal rate of return, and uses safe Lakebase branching to let basin champions test data governance rules without risking production data.
  • Quantum is extending its platform with agentic AI to automate individual well diligence decisions at scale, incorporating human-in-the-loop validation to maintain institutional governance as the volume of deals processed increases.

Scaling Enterprise Data: Lakebase for Legacy Systems and Operational Distribution

Watch: Scaling Enterprise Data: Lakebase for Legacy Systems and Operational Distribution
Private equity deal evaluation demands integrating diverse data sources: public records, vendor datasets, and proprietary models. Quantum Capital Group ingests 1.5 billion records across energy markets but faced a persistent challenge: transforming analytical gold datasets into operational data accessible to legacy systems, geologists, economists, and deal teams using desktop applications.
Learn how Databricks Lakebase acts as a PostgreSQL operational data store, enabling ODBC connectivity to legacy tools while maintaining the scalability of the medallion lakehouse. Discover how Quantum built a custom master data management tool with Lakebase Apps, democratized data governance to domain experts and basin champions, and automated individual well evaluation using agentic AI. See how safe branching for data testing and human-in-the-loop workflows enable institutional AI without governance risk.
🤝

Chapters

FAQs

What is Databricks Lakebase and how does Quantum Capital Group use it?

Databricks Lakebase acts as a managed PostgreSQL operational data store that sits alongside the analytical lakehouse, providing ODBC connectivity for legacy applications and desktop tools. Quantum Capital Group uses it to deliver curated gold-layer datasets from their Databricks lakehouse directly to geologists and economists working in traditional desktop applications, without requiring those users to interact with Spark or notebooks.

How does Quantum Capital Group evaluate oil and gas assets using data?

Quantum ingests 1.5 billion records from multiple energy data vendors, applies geological modeling and economic curves to characterize individual wells, and tiers assets into red, amber, and green categories based on internal rate of return. This medallion architecture-based workflow enables the firm to screen large numbers of assets and focus deal diligence on the most promising investment opportunities.

What is Lakebase branching and how does Quantum use it for safe data testing?

Lakebase branching allows teams to create isolated copies of operational data for testing new governance rules or transformation logic without affecting the production dataset. Quantum Capital Group uses this capability to let basin champions and domain experts safely develop and validate data quality rules before those rules are applied to production data.

How is Quantum Capital Group applying agentic AI to private equity deal evaluation?

Quantum is building AI agents to automate individual well diligence decisions, addressing time pressure on smaller investments where manual diligence is impractical at scale. The agent workflow incorporates human-in-the-loop validation to ensure institutional AI governance is maintained while increasing the volume of deals that can be evaluated.

Full transcript

[00:07] All right, let's go ahead and get get started. So, um first of all, uh thanks everyone uh for coming. There is a forward-looking statement here. So, just give that a quick read. Um I have been asked to say uh you know there there is a survey. So, we we do welcome your feedback on on this presentation and all the the other presentations at the
[00:24] summit today. Um by way of introduction, my name is Ian Brown. I'm head of digital engineering at Quantum Capital Group. For those of you that don't know who Quantum Capital Group is, we are a private equity firm uh based in Houston, Texas. Uh we focus exclusively on energy
[00:40] investing um right across the spectrum. Uh predominantly upstream oil and gas, energy infrastructure, renewables, decarbonization, and thermal power generation. Um the workflow that we do, the workflow we're going to talk about today, scaling deal evaluation is pretty
[00:55] much the lifeblood of the firm, right? A private equity firm runs on deal flow and deal diligence. And so if you can't diligence deals, you don't get the opportunity to uh make an investment decision. And likewise, if you do a poor job of diligencing lots of deals, you're
[01:10] likely to make a sub-optimal investment decision. So, so really the holy grail is we're trying to get as much deal flow through the company as possible and diligent diligence those deals in a robust manner. The outcome should be a win-win where we should get to make a decision on the best deals in the market
[01:27] and get the best risk adjusted return to our fund and that is really the goal of what we're going to talk about today. Um I am going to set a little bit of context. Um energy is a very interesting space right now especially if you're you're in that market. Um and especially with AI and data centers and all of
[01:43] that. So we're going to touch a little bit on on on energy market demand. We're going to touch a bit on data centers and AI. Uh and then we're going to touch on how that's changing market dynamics. Uh what that means as a private equity investor and private capital in that space and then we're going to talk a little bit about uh how we actually use
[01:59] data bricks to diligence deals or as part of the deal diligence process. Um and then uh lakebase um so what we've done with lakebase in the last 12 months to scale that even further um and and some of the real advantages of using lakebase that I think some of you in the room may appreciate and might not have
[02:15] considered when you're talking about your lakebased strategies. And then finally a little bit of agentic AI. I didn't want to be the only one at the conference without anything agentic. So the last slide is talks a little bit about how we're going to scale our diligence process even further potentially using agents. So with that
[02:32] um let's start with three charts. Right? This is just some energy by the numbers. Right? On the top left here you've got past, present and future or past, present and present going up uh or past and present I should say. On the top right there you've got the US uh current
[02:47] energy mix and then at the bottom here you have got uh US energy uh forecast in green and the rest of the world forecast in blue. On the top left uh just want to spend a minute there. Uh what you can see there is um around 2005 uh something
[03:03] changed in the energy markets right you can see it starts going up very very dramatically and that was the start of the digital revolution the digital economy. The cloud was born. Steve Jobs invented the iPhone. Um, Internet of Things came along. SAS uh came along. Uh, and big data, machine learning, and
[03:20] uh, EVs. Elon Musk invented Tesla. And all of that, all of that's a big draw on power. And we've become a very power- hungry economy. All of our devices are switched on all the time. We've got sensors in our houses that are always on. Even our cars are are beaming things to and from from the cloud these days. So, uh, we are becoming a very energy
[03:37] uh, um, hungry uh, economy. And it's not just us, it is the rest of the world. As the uh the chart at the bottom there shows, the chart on the top right there, this pie chart is actually quite interesting. Despite all of our efforts, uh to decarbonize here in the US, we are still uh 70% fossil fuels, uh 10%
[03:54] renewables, 10% nuclear, and 10% coal. Natural gas has been obviously, if you look at the top left there, natural gas has been replacing coal over the past 20 or 30 years. But ultimately we use natural gas uh for for electricity and we use uh oil for transportation and uh
[04:12] that's not looking like it's going to change in the the near future. So let's bring data centers into the mix. So the numbers that I showed you do not include recent data center buildouts or or or or demand. And this this chart here is from the EIA and what it tells
[04:30] us or that they're forecasting is that uh demand from data centers is going to triple uh in the next the next decade. Um and more than half of all US electricity demand growth will come from from from data centers. And I think the most
[04:45] important thing here on the left is again in the purple in the bar chart, natural gas is going to be the uh fuel that powers data centers. Why is that? Um there's three reasons for it. Uh first of all, the US is in a very very privileged position. Natural gas is
[05:01] abundant in the United States. Uh the second uh and perhaps the most important is it is very very cheap. Uh once you've built a data center, there's a huge capex cost there, but then you've got to power it and that's obviously your second biggest cost. And uh you want the cheapest fuel uh that you can have and
[05:18] that is uh usually natural gas, especially here in the US. And then third, it's a reliable fuel source. A data center has to be running 24/7, 365. Uh if you power your data center with renewables, you've got some challenges keeping a steady load there. So NAT gas
[05:34] is a reliable base load fuel under just about every political scenario and under every uh weather uh scenario as well. So, so NAT gas it probably is going to be a major major um player in our energy mix in the future and it is certainly going to be the backbone for the
[05:49] majority of of data centers uh as to to power them. So, going forward a little bit further, what you're starting to see is a shift in energy markets and um when it comes to to to to energy consumption and and where data centers are being built. Um,
[06:05] not to labor too much on this slide. There's quite a lot on there and you guys can read read some of the details. But where you're you're seeing more and more data centers being planned and being constructed is largely where natural gas is cheap and in abundant supply. It's also the states where the
[06:20] the amount of natural gas that can be produced out far outstrips the takeaway capacity. That what does that mean? That means it's difficult to pipe or store all that natural gas. it's difficult to get it to point of sale on one of the coasts for LNG export. Therefore, it's kind of locked in. It's trapped in into
[06:37] uh that geographic area. And so what's happening is if you go back to the early days of data centers, data centers would be built much like buildings, right? Uh you'd construct them, you'd arrange a commercial power uh agreement with Con Edison or whoever and you'd you'd run your data center at commercial rates.
[06:53] Those days are long gone because of the amount of power consumption that that these things need. So what we're doing now the industry is moving towards is uh basically power purchase agreements and basin supply agreements and and very very briefly what that means is we're starting to collapse the stack. We're starting to bring data centers to where
[07:10] the energy is not bring energy to the data centers. So a hyperscaler will build a data center or propose one. They will arrange a power purchasing agreement with an energy provider a utility company. that utility company will build a power a power generation
[07:25] asset for that data center or collection of data centers and they will arrange a basin supply agreement with an oil and gas producer who produces natural gas. And what that does is it takes the grid entirely out of the equation. You've got a direct line from the from the wellhead through into the uh the power generation
[07:42] through into the data center. Everything's locked in on a long-term agreement. You've now got reliable power for for your data center. So keep an eye on this slide. Look at Texas. Look at the other dark blue states. That is where the vast majority of these data centers, at least the ones that have been announced recently, are going to be
[07:59] uh constructed. And now take a look at this map. This is a map of the US obviously. And what you can see here in the red is the major oil and gas producing basins in the United States. And you'll see up on the east coast there, um there's a huge basin with a lot of natural gas. And down in Texas,
[08:15] there's a huge basin with a lot of oil and natural gas. And so you can see that those buildouts of the of of the of the data centers are going right where the cheapest natural gas and the most abundant source of natural gas is. Uh what does this mean as a private equity firm? We invest in the energy markets.
[08:32] We invest in oil and gas. We invest in power generation uh and and energy infrastructure. So for us it changes uh the value and it changes the attractiveness of of of our portfolio uh our investments. And it also means for us that there's a lot of mer mergers and acquisitions. Um, and with that comes
[08:48] devestature. So, it's an opportunity for us to maybe exit assets and maybe pick up assets. And, uh, maybe just to touch on on on on, you know, private capital and oil and gas a little bit more. We've traditionally been an innovator in oil and gas. We we will we will get into assets that might be unattractive to
[09:04] others. We will unlock new technology that make previously unattractive assets much more economical. will drive process efficiencies and usually we will seek to grow those assets, make them very valuable to a a larger publicly traded company and perhaps exit that asset
[09:20] hopefully with an appreciation on the capital that that that we invested. Um, with all this activity, uh, it's a great time to be in in the business. Uh there's a lot of deal flow and that goes back to this presentation which is there's record deal flow in the in in in the market not just in upstream oil and
[09:37] gas assets in midstream and in in in power generation. How do you process all this? How do you diligence all this and how do you get the right uh the the right risk adjusted return get the right portfolio and uh and and and keep doing the right thing by by your fund. So,
[09:54] talking about natural gas, because I I think you you kind of were kind of honing in on natural gas as the base load for for for energy uh for data centers. Um when you think of a natural gas or an oil and gas asset, don't think of it as like a building or or or an office space. It's land. It's a
[10:09] geographic area of land where somebody has the rights to drill oil or natural gas wells. And assets usually come as a package of land with producing wells on them and with what we'll call inventory, which is land that could in theory be be drilled on. And when you come to try to
[10:26] work out what one of these assets is worth, you've got to work out the current wells that are producing, what are they doing, why are they doing that? Could they be produced more efficiently? Could they produce more? And then what is that the actual remaining inventory? And then how do you drill out that inventory? What sort of well spacing?
[10:42] Where exactly are you going to drill those wells? because not all of the earth is equal. So you have to understand exactly what's under the ground and somehow you got to put a price tag on that and then you've got to think about how you would develop that asset over a 5 to sevenyear period and then perhaps what the exit strategy for that asset would be. You times that by
[10:57] hundreds and hundreds of deals and you've got a bottleneck a real headache on on on trying to do all this valuation so that you can you can deploy capital and again the diligence process of all of this is robust. You're not just looking at insurance papers looking at the building structure signing the deal and away you go. You literally there's a
[11:13] whole subsurface geological element to this. There's an economics, there's an engineering model that sits on top, an economics model that sits on top and all of it has to be pieced together. So the diligence process really is quite complex is the message that I'm I'm trying to get across here. It's not that um simple. Now
[11:31] we have exactly the same problem that every single presentation you've heard today has, right? Which is we ingest a lot of data into data bricks. It's about 1.5 billion source records. Here's the problem. The United States doesn't have a central oil and gas authority. It has federal lands, but ultimately each state
[11:47] has their own oil and gas rules. Each has their own oil and gas commission which records all of the data and most of them rely on the oil and gas companies to self-report that data. Therefore, the data is uh not reliable and there's no one source of record where you can get all of the data. This
[12:03] is very very important because if you're looking at areas of the US where you'd like to invest, you don't have any private data necessarily and you have to rely on public information to look at what might be attractive or where there might be be opportunity. Quantum spends far too much money, frankly, on buying these data sets only to have to go and
[12:19] clean all of those those data sets up. And so we buy six different versions of the truth across six different vendors who aggregate all of this data from public data sources. We then basically go on this workflow that I've shown here all done in data bricks to cleanse that data using 125 data validation rules uh
[12:35] survivorship rules etc. We then put our own algorithms on top to enrich that data. Basically we put our own subsurface model across the United States and then work out where these wells are really drilled and how they're going to perform. And then what we do is we uh we have basically a model of the whole of the US uh for which we can then
[12:53] lay deal specific data sets on top of and evaluate what the what the seller of the asset is uh is is is basically saying and what what our models say and there only then do we start the actual diligence process. That's not the end of the diligence process. That's actually
[13:08] the start. So digging into uh again this is our um current data bricks process. This is being done with uh the lakehouse using uh Spark uh SQL. Um and we're not at the lakebased part of this yet. That's coming a little bit later in this
[13:24] presentation. So we we over the past um I guess three years, but really the current incarnation of this is about uh 12 to 18 months old. Uh we go through once we've got from the the slide previous, we then go through a huge
[13:39] amount more of enrichment. We start looking at ge geologically similar areas. Not to get into too much detail, but we're using unstructured machine learning to work out which rock is the same. And gradually over time, when does that rock become different rocks or when does it get more gassy or when does the oil dry out and become water or where is
[13:55] there just no no hydrocarbons at all? And we start to learn the subsurface much like you forecast the weather, we're kind of trying to forecast the subsurface. And from there, we look at all the wells that have been uh drilled into that that rock. We look at historically how they've been producing and we build prox approximate curves.
[14:11] They're they're curved based on a lot of different parameters which say how an oil well or a gas well might produce in that geologically similar area. So if you drill a well into that geologically similar area you would expect to see this type of production profile from from the well. Uh and from there we put
[14:27] an economics model over the top and the economics model is capex, opex and of course crude and natural gas prices. And we do a a bunch of scenario modeling and ultimately what we do at the end of it is we boil down to this this picture on
[14:42] the right which is a red amber green um model of where we think there's attractive oil and gas opportunities and where there's less attractive oil and gas opportunities. Tier one, tier two, tier three. And that tiering is based on internal rate of return IRR, right? And
[14:58] so we're looking obviously at tier one. Everybody wants the tier one stuff, right? Everyone like the best wells with the best rock. Yeah, we all want that. So tier two and tier three become attractive, but only if you understand exactly what the what the opportunity is in those spaces. So as you can imagine, we do this for all of the basins in the
[15:15] United States. And you've now got this map, right, to start with. And this map gives you a ready reckoner for any single deal that you want to look at. And then you can actually place the real world data over the top. Um, very very powerful to be perfectly honest because
[15:32] a lot of I think you've probably seen other presentations most of the time is trusting the data, getting the data, working out how to format that data, getting it into systems and so on so forth. So by taking care of this, you get 50 to 60% head start on your on your competition. And again, if you're seeing
[15:48] hundreds of deals a year, that's more than, you know, one or two in any given week. There's a lot of deal flow going through the firm. The value of this alone is is pretty immense to our petrochemical diligence team. So, as I said, data bricks runs all of this. Um, we follow uh the medallion
[16:06] architecture the in the lakehouse uh the bronze, silver, gold. I recently learned just 10 minutes ago there was platinum. Um, we we haven't got that far yet. I'm not quite sure what platinum is above gold, but uh one day hopefully we'll we'll find out. And we we we put all that in in into a data set in our gold
[16:22] which largely is a consumer focused data set and it's focused purely on a data set which will work for our engineering team. Remember I said there's a subsurface model, an engineering model, an economic model. This focuses purely on the engineering aspect of of it which is where are the wells, how are they
[16:37] placed, how are they going to be drilled, how will they how will they produce and all of the different scenarios around that. And that's largely consumed by Spotfire. And again the graphic there hopefully you can see it's all using data data bricks jobs uh in in data bricks it's all being pieced together uh by by our team and we we
[16:54] have our our our our models and everything running in in in there and it's worked very very well for us. So let me give you an example of value. Um when you work in private equity and certainly at Quantum Capital Group you're probably as proud of the deals that you don't do as the deals that you
[17:11] you do right. So when you're looking at let's say a 100 deals, you're probably maybe do one or two of them, right? So a lot of these deals are not going to get done for for for a lot of reasons. So this is an example here of where the seller came to us with an asset and the seller said, you know, hey, we think,
[17:28] you know, this asset is is awesome. It's got uh what is it? We think it's got two 2,000, sorry, 3,516 uh uh remaining locations, and we think the wells are going to produce X, Y, and Z. and therefore we think the assets worth X. What we did is we overlaid our
[17:45] golden data set including the subsurface data and we said well actually when you bring all of the faults and all of the the understanding of the rock in there we actually believe that uh half of those um sites that you say we can drill on the inventory is uneconomical and other parts we just can't drill on. So,
[18:01] we actually think that it's worth uh you know, there's about 1,800 locations there. And then when we derisk it for the underwriting case, we actually think, you know, there's 1,700. So, in this case, if we were blindly using public data and blindly using what the seller was telling us and trying to
[18:16] piece public data and what the uh the seller was saying, we could have overestimated this asset by 100%. Right? Uh we didn't do this deal. Somebody did. Good luck. Right? We we we chose not to. Uh and that's the value of of this data. It's not just to to to to win. It's to
[18:32] it's to not not lose. Now, again, we talked a little bit about scaling uh deal flow. Uh here is some dollar value of deal flow that we've been doing over the past uh how many years. What's that? 2023 to 2026. So, the old process 23 and
[18:50] 24, you'll see quarter by quarter the amount of deal flow in dollar value going going through the business. I should clarify this isn't this isn't capital we've deployed. That's values of deals that have gone through the the deal diligence process. 2024 we released this new data bricks powered uh
[19:06] automation workflow that does what I said previously and you can see that the amount of deal flow going through the business rose substantially. We didn't hire many more people into our technical business. We had a better data set. We utilized it better. That freed up our
[19:22] technical team to change their process processes and focus on different things and streamline their workflow. and we got a very different response and a very different business outcome. 2026, this is Q126 on this slide. Look at the numbers for Q126. We've almost met the
[19:37] same deal flow as the whole of 2025, right? So that's the power of this workflow. So that being said, if that's so wonderful, why do you care about Lakebase? That's a good question. Um question that we had
[19:52] to justify into the business. We came back from the tech uh the AI summit last year and went, "Oh my goodness, Lake Base is amazing." And everyone said, "Well, what are you talking about? We've just built a new system. It works great for us, right?" But here's the reality. Um it doesn't solve all of our problems. As I mentioned, uh bronze, silver, gold.
[20:09] We produce this data set and the data set goes to Spotfire. It's great when it's in Spotfire. It works great with Excel. It works great for the engineering component of the uh the workflow, which is where frankly a lot of the bottleneck was, which is why why we solved that problem. But there's a lot more people that take part in the
[20:25] diligence workflow. There's our subsurface, our geologists, our prophysicists. There is our econ economists and then there's our deal team. And all of them do not use the latest greatest cloud stuff. They certainly don't diligence things using PowerBI dashboards. They don't use little cloudy things that have got
[20:42] little trend lines that go up and bar charts and pie charts. They're using prophysical modeling applications and those are desktop-based applications. They don't interact with anything except a 1990s database and none of them talk to each other unless it's through CSV exports. I'm sure many of your
[20:57] businesses have the same problem. And so really we just pushed the bottleneck a bit further down. So a huge amount of gain to be had by by by what we've done, but also a lot left to go. And so we asked the business for some feedback. And this is a slide that uh that the business gave us. I I had to sanitize it
[21:14] and make it a little more polite because the original one was a little more blunt. But um basically three things they were they were basically saying look great Ian this data set's awesome but we're still having to manually manipulate it to get it into our systems. A lot of reverse ETL was happen
[21:30] happening further down the pipeline. A lot of people shleing a lot of data still uh and then 10% of the actual diligence team are happy and going home at 3:00 because they've got the gold data set and they're they think everything is wonderful. But really they were saying it's integration integration integration. None of it talks to each
[21:47] other. It's not that valuable. Help us out. Let's get this sorted. So, we came back from the AI summit last year and we had no idea that data bricks was going to release this this lakebased thing which turns out to be Postgress. And the great thing about Postgress is
[22:03] it's a legacy relational database and that we can use this now as an operational data store. It has ODBC traditional ODBC available to it. It will connect to anything literally anything. And that solves a lot of problems for us. The other thing that we did is we took our gold data set which
[22:19] again we we created as more of an analytics ready product and we put it back into more of a third normal form relational product. We did that because these tools and applications love that sort of thing. When you've got a desktop app that needs to connect with an adapter to a database, they kind of expect primary key, foreign key, even
[22:35] spotfire kind of expects that sort of stuff, right? So having a more traditional relational stack is a huge advantage. And the other thing is the two can sit side by side. The gold product in the lakehouse and the lake base can sit absolutely together. Right? So no one has to lose in all of this.
[22:51] Right? We just create yet another avenue for distributing our data. And so what this has done, it's turned what was once just an analytics data set uh into what is now a workflow ready data set and brings a lot more participants from the deal diligence process into the business
[23:08] and and it's live. This is the greatest thing. They're all looking at the same data at the same time. And we in the data engineering department don't have to do a thing. They everyone knows how to connect to databases. Everyone knows how to hook up their their legacy systems. And it's all pretty much self-service. You're going to say, "Right, yeah, but
[23:23] Ian, okay, you know, Simba driver, OBC can work with with with the gold data on on on on the lakehouse." Yes and no. Uh yes, and that it's it's not really true OBDC. And number one, these legacy systems really don't have any concept of personal access tokens, cloud-based
[23:40] security, and so on so forth. So really, they don't um not not in the way that you would you would like like them to. And so Postgress and and Lakebase really solve challenge number one, which is turning it into an operational data store and distributing live data across across the enterprise. The second
[23:58] problem that Lakebase solved for us was we combined it with Lakebase apps. the there's a lot of talk around master data management and governance. I mean, usually from a a security stat standpoint, but but MDM is a very interesting subject and it gets more
[24:13] interesting the more you get away from consumer data sets and more into niche data sets, maybe biioarma, oil and gas because there's not a lot of great off-the-shelf products that are going to help you with some of those industry specific master data management. So we built our own master data management tool using lakebase using data bricks
[24:31] apps live link to the uh the data set and um we did use cloud code which helped a lot because that was quite quite a daunting task and so now what we've done is we've got our gold silver bronze in the lakehouse we can master it using a lake base with a with an MDM attached to it and feed the final
[24:46] product back into gold and then distribute that out to both the gold data set and to to the lakebase data. um that's brought a huge amount of satisfaction to the business. When you're looking at uh one and a half uh
[25:02] billion records of production, oil and gas wells, there's about seven or eight million wells in the United States. U Canada's got its own challenges. They they they name have naming conventions and and and different numbering systems for wells. When you've got all that going on and you're trying to work out
[25:18] what the gold record is, no data engineer can do that. Ultimately, you have to push some of that back to the domain experts. And what we have in quantum is we have basin champions that look at each one of the oil and gas basins. We, you know, they've worked at operators for many, many years in their careers. They know these basins. They
[25:33] know what these wells look like. They know what these wells should be. And and and they know when something's not right. And so, we're still bringing the human element to data quality. We're not automating data quality end to end. Because when you when you try to automate it blindly, what happens is someone gets it wrong and we
[25:50] overestimate production by 25 50% and the whole thing after two weeks of doing diligence, someone says, "Hang on a minute. This none of this looks right." And then we have to scrap the whole thing and you know, you basically don't get a chance at at at making an investment decision on the deal. So what this has done to us is democratized the
[26:06] master data management. It's democratized data quality to a large degree. And it's made the data owners and the basin champions a lot happier because in the old workflow we would have to go into the Python scripts and go if basin equals this and the well is this and it's a Tuesday you know change
[26:22] this rule because there's a lot a lot of niche stuff that actually goes on as I'm sure many of you have probably seen in in in your own business that the 80% is great the 20%'s impossible to govern correctly. We we also got into a state with our our current version or or the older version now where you change one
[26:39] rule and it breaks something further downstream. So you satisfy one problem, you create another one, you've got that kind of whack-a-ole going on on through your through your gold your gold data set. So this goes um a long way to alleviating that. And most importantly
[26:55] from my selfish perspective, very very very few of those issues come back to my team. they're now in the hands of the uh the actual protr techchnical team that are doing the diligence workflow. Where this goes from here, we haven't
[27:11] actually implemented this yet, but this is what we're going to do right now. When someone remasters the data using our master data management tool, what happens is it blindly pushes all the data into gold and they can review the data set once it's been pushed to gold. Lakebased as many of you know has
[27:28] branching and there's really no reason why you can't branch the data set because it's it it you know it's it's it doesn't take take a copy of the data. So you can just branch that data set, rerun your new rules, test your new rules, make sure they look right without ever having to go back to the production live
[27:43] database. And so this gives this will give our users a and and our own internal team a lot more confidence to try out different rules and see what effect it'll have on the data. uh try overwriting data with our own in-house data so on so forth without ever ever bothering a live deal because again if
[27:59] it's if it's being mastered back into the gold data it's actually going back into every other deal that's also being diligenced and could screw up a deal that's that's in progress and so this gives a safe space for the uh the MDM and the data management team to to govern that data without any fear of
[28:15] wrecking a $500 million investment decision or or or what have you. So, as you can imagine, it's something that uh that we think is going to be pretty valuable in the future. So, a little bit of a lesson learned.
[28:31] I'm old and one of the other people in my team's older than me. And when we heard about Lakebase, we were like, "Oh my goodness, relational databases were so excited because we just been through a decade of big data and big data and unstructured data and don't worry about it. It's just a data lake. It'll take care of itself." The reality is I think
[28:48] we've all been through there. We created a lot of data swamps and things like that. So when when we heard about when we heard about lakebased we were like oh we could use a lake base for bronze, a lake base for silver, a lake base for gold. We could have just relational databases everywhere. It'll be heaven. Let's get back to the '9s and be database administrators, you know. So we
[29:06] did that and it was a bad idea. Um, what happened is we slowed the whole system down to the point where it was unusable and we got on the phone to Data Bricks and and and and I'll say this, our data bricks account team's been amazing with us with this. We we adopted this pretty much the day it came out and then started misusing it right from the start
[29:22] and uh and they they they beared with us, let her make our own mistakes. We we found a few bugs and had some lessons learned along the way, but we honestly couldn't have got this to the right place without them. But the the key thing here is if you do what's on the left, it's going to become three times more expensive, three times slower, and you're always going to run into these
[29:38] weird issues where you've got cold start data that's in cold storage. It all seems fast, then it all seems slow. Just don't do it is is is the advice here. Use the lakebase for what it is, large analytical workloads, use the uh the sorry, use a lakehouse for large
[29:53] analytical workloads and use use the lake base for for for the last mile delivery. And that's where I think we're seeing the most value is that last mile delivery into into the all of the the different applications because again the value is not realized unless people are actually using that data. One other
[30:09] thing we adopted it's not really lakebased but I think it's worth touching on this presentation. All of our Python models that we built were built in regular Python on someone's desktop trained on a data set or whatever and then just shoved into a into Unity catalog or into sorry you
[30:24] know data bricks job and said go run it. Um, now we've adopted MLflow, but again in the tradition of quantum not using data bricks technology exactly as it's meant to be used, we're putting a a lot of non-machine learning models in there. So, we've got one or two machine learning models that I described earlier
[30:39] in the presentation, but we're also putting a lot of statistical models in there, you know, curve fitting and and things like that. And believe it or not, MLFlow does do non-machine learning models and manages those pretty much as well as it does for the for the uh the machine learning models. So yet again
[30:55] another technology we're playing with and not using it exactly as as as as designed. So anyway they they did all that to say on the right yes on the right do that on the left don't. Uh basically um I think the last thing I want to touch on in this presentation is
[31:12] we've talked about buying entire assets and diligencing assets. Well, we actually have a business, a very a very uh growing a rapidly growing business in in quantum capital where we're actually investing in individual wells. Uh we're taking working interest in wells. What that means in in summary is that what
[31:29] we're doing is an operator of that well say, "Hey, do you want to take a share of the expenses for a share of the revenue?" And we're like, "Oh, look, that makes sense, but it only makes sense if the well's going to be super productive and we think the costs are going to be super low. Otherwise, it doesn't make sense." So now what we're
[31:44] going from is diliging an asset where you can make approximations. Remember I talked about aggregating curves and geological similar areas. You can do all that and you can you know over the long run of a thousand wells you're you're probably going to make statistically a good a good decision. When you're in looking at individual wells you've got
[32:00] to be a lot more correct because you get one well wrong that well is going to be uneconomical and so on so forth. You get more of them wrong than right and you you know you know the outcome. So the other problem with this workflow is it's time pressured. Um and what what what that means is we've got to get a
[32:16] turnaround or an electing it's called election decision electing indecision usually within 3 to five business days. And so we're seeing 10 to 50 of these AFS they're called AFS or authorization for expenses. We've got to make a decision as to whether we want to participate and take that expense for a
[32:31] share of future future revenues or or or we don't. Um, and so obviously if we can't process these, going back to what I said at the start, if we can't even diligence them, the answer is always going to be no. If we just blindly look at them and say, "Yeah, well, the one next door looked looked okay, so let's do this one as well." Then we're
[32:47] probably going to make not the best investment decision. We're not doing as robust diligence as what we need to do. Enter the agent. So there on the slide is the timelines. It takes us to diligence one. Well, uh, my math's not great. I can't add all that up, but it's basically a working day. And so if you
[33:03] got 50 of them, you've got 50 working days. And obviously that doesn't scale well when you've got a team of maybe five or 10 people looking at not just these deals, but also our all of our other other um investments that we're that we're looking at. So agents are a perfect task. And this will leverage our
[33:18] gold data. It will leverage Lakebase because what we propose to do is obviously put each agent into lakebase, spin up a branch for every AF, process the AF, produce a report, the same report as what the human would produce and then allow the petrochemical expert
[33:34] to review that report doing the diligence and then send that report to the investment team with their own commentary attached and then perhaps make an investment decision. So now we've got the uh agent doing the leg work and the human doing the diligence.
[33:49] So we believe this is a perfect use case for agents. Um we're not we're just starting to look at how it's going to be achieved and obviously there's a lot of risk management around this. I think like you know we've all talked about in the conference you know you can't just blindly trust agents to go and do things and say that that that that'll do us. So
[34:05] so a lot a lot of diligence. This will I'm sure run in parallel for many many months before we we adopt it. But this is a great use case for Lakebase and for Aentic AI. And if it works, which we see no reason why it won't, we think this will uh really game change our our non-operated uh oil and gas well
[34:21] business. This is a visual of it. Again, it's agents do the work, humans do the diligence, deals go in, agent spins up on a branch, each agent starts make doing its own work and makes a decision. Ultimately, every morning we should get all of all of the AFS that have came
[34:38] into to Quantum. They come in automatically already. Run them through the agent overnight and then have a a decision recommendation ready for all of the AFS first first thing in the morning. That's that's the goal here. That's the output. Last slide. I wanted to leave you with a a few thoughts on on AI at quantum and
[34:56] also AI and energy generally. Um I think starting on the left efficiency versus opportunity AI. think like a lot of people here, we're using AI for individual productivity and I think everyone agrees, you know, some of these tools are are phenomenal. Uh, opportunity AI, I think it is opening up
[35:13] um a lot of opportunity for us to do things differently and we're kind of feeling that out as to what that actually means because when you do one thing differently, something else all of a sudden can be done differently. I think individual versus um institutional AI I think the golden data set and this and this deal diligence workflow that
[35:29] that we've demonstrated is kind of the first step of institutional you know um not just AI but data governance data processing and if we can put the agent element on top I think that's kind of really institutionalizing um some AI in the firm and I think there's a lot of other opportunities not just in our deal
[35:45] diligence side of the house but also in the fund administration and accounting side of of of the firm as well over here on the right couple of thoughts here too you know we have markets for everything, gold, silver, oil, gas, uh, agriculture. Maybe a market develops for tokens,
[36:01] right? Maybe that's something that happens in the future. I I don't know, but there's, you know, why not? We've made a market out of just about everything else that we want to consume. Uh, and then lastly, sovereign AI. You know, there's, you know, my thoughts and on this are, you know, if you have the cheapest source of energy and the most
[36:17] abundant source of energy and you have data measures, uh, sorry, data centers and, uh, and and and you have access to data. Chances are you're going to be a pretty powerful nation in the future. Think of Saudi Arabia with oil. You know, um, they can they can move oil
[36:32] markets and they can they can they can move influence geopolitical decisions. Imagine what the AI revolution looks like if you control the data controlled data centers and have the cheapest access to energy as well. So that's just uh maybe maybe a a more wider thought uh for to end this presentation. So that's
[36:49] it from me. I think I just wanted to kind of conclude with hopefully this presentation has shown you perhaps how quantum capital group is adopting data bricks in part to do a deal diligence workflow which actually provides private capital to the energy markets which eventually ultimately powers data bricks
[37:05] because data bricks runs in those data centers right so we're actually using data bricks to power data bricks in a way so with that I'll leave it there thank you very much and if there's any questions happy happy happy to take

Learn more about the Databricks Data and AI platform.

The information provided herein is for general informational purposes only and may not reflect the most current product capabilities or configurations.