Skip to main content

Multi-Cloud Databricks at BP: 43% Cost Reduction and Unified Data Strategy

Summary

  • BP achieved a 43% cloud cost reduction and accelerated job runtimes from multiple days to under 8 hours by standardizing on Databricks as a unified platform across AWS and Azure.
  • Unity Catalog and Delta Sharing enabled BP to eliminate the 40–50% SAP financial data duplication that had accumulated across clouds by providing zero-copy virtualization and governed cross-cloud data access without data movement.
  • BP's multi-cloud consolidation unified governance — including lineage, audit, and data stewardship — across systems that had become fragmented through 117 years of mergers and acquisitions spanning multiple business units and cloud providers.

Multi-Cloud Databricks at BP: 43% Cost Reduction and Unified Data Strategy

Watch: Multi-Cloud Databricks at BP: 43% Cost Reduction and Unified Data Strategy
BP's 117-year history spans multiple business units, data sources, and cloud platforms: production workloads on Azure, customer data on AWS, plus legacy systems. This fragmentation created data duplication (40-50% of SAP financial data copied across clouds), slow queries spanning days, governance gaps, and no single source of truth.
this video shows how BP standardized on Databricks as a unified platform across AWS and Azure, using Unity Catalog for governed multi-cloud data access and Delta Sharing for zero-copy virtualization. Learn BP's results: 43% cloud cost reduction, job acceleration from days to under 8 hours, elimination of data duplication, cross-cloud collaboration without data movement, and enterprise-scale governance including lineage, audit, and data stewardship.
🤝

Chapters

FAQs

How did BP reduce cloud costs by 43% with Databricks?

BP achieved a 43% cloud cost reduction by standardizing on the Databricks Data and AI platform across AWS and Azure, eliminating duplicate workloads and redundant data copies. Delta Sharing's zero-copy virtualization removed the 40–50% of SAP financial data that had been duplicated across clouds, directly reducing storage and compute spend.

What is Delta Sharing and how does BP use it for zero-copy data access?

Delta Sharing enables direct access to live data across clouds without creating copies or running ETL pipelines, keeping data in place while granting governed read access to consumers. BP used Delta Sharing alongside Unity Catalog to provide governed multi-cloud data access and eliminate data duplication across their AWS and Azure environments.

Why did BP end up with data spread across multiple clouds?

BP's 117-year history and multiple business units resulted in data spread across different clouds — production and operations on Azure, customer and commercial data on AWS — largely due to mergers, acquisitions, and different teams choosing different tools. This fragmentation caused data duplication, slow cross-cloud queries spanning days, and governance gaps with no single source of truth.

How does Unity Catalog support multi-cloud data governance at BP?

Unity Catalog provides centralized governance including lineage tracking, audit capabilities, and data stewardship controls across multiple cloud environments. For BP, it served as the governance foundation that unified data access, compliance, and discoverability across AWS and Azure within the Databricks Data and AI platform.

Full transcript

[00:10] Let's start. Um, a few months ago, BP roll out a co-pilot, office co-pilot. So, I will wonder what is the most common question that go through my I don't know how many years of emails that I have, right? So the common term is
[00:27] Julie where is the data that's is the most common question that people ask me where is the data not you know what what uh what we going to do about the report not the requirements that what we
[00:43] benefit we're going to bring to the business no where is the data okay so how many of us here that your organization have data that in more than one cloud.
[01:04] A lot of us AWS, Azure, GCP, right? How many of us that actually enjoy reconciling a same dashboard in three different clouds because that's the requirement for your business. We're not right. No. Right. No. But most of us are not intentionally
[01:23] have data or analytics in so many different places. But it's kind of is because through merger and acquisition, right? Because different business want to do different thing in different cloud because of the tools that we choose to
[01:39] use that force us to go in in one cloud versus the other. So we are here today right now. Imagine if where you store the data doesn't matter right where you uh where you want
[01:55] to do analytics doesn't matter the matter is you have your data in one place right and you can do report dashboard you can have AI you can have journey on top of that so that is what we are here to tell you our journey
[02:12] right tell you our story my name is Julian wind and I I am with BP as the principal ible data engineer. Hi everyone, my name is Shini Chandulu and I uh I'm working as a lead data
[02:27] platform engineer back in BP. Welcome to our session today. Oh, too fast. All right, so let's start with the challenge, right? I should have
[02:43] I shouldn't say the problem, the challenge, right? Why one cloud is not enough for us, right? Actually, we don't know. We think it should be enough, right? AWS or or Azure or GCP, we just pick and choose.
[03:00] But in our BP with 117 years of history, we are end up with you name it, we have it. That's kind of organization. Okay. So, our production and operations is on Azure, right? So we have everything in
[03:17] our our Azure stack, right? Our CNP is AWS. So we have S3 storage, different tooling, different different team doing uh different things, right? And we don't have a single catalog. We don't have a
[03:32] single access model, right? We don't have a a single government for that, right? All of us we doing different thing. is depends on when we where we are with that. We copy our data everywhere.
[03:48] The the upstream and downstream the the production and or uh and operation also need SAP data. The finance of course need SAP data. The retail business, the midstream business also need SAP data. So the same SAP data got copied into
[04:06] multiple places, right? And then we wonder why we don't have a single source of true. Okay. And with that right because of the because of data is everywhere right we
[04:22] have analytics we have report and everyone say my version is the right version. But if you have you know two version that from the same data you don't have the true at all. Right. There's a lot of lot a lot of issue with
[04:39] the business right in PN that is where I am working right now we have seven pabyte of data sit in one link right we have jobs that run multiple days it's with that amount of data in our old in
[04:56] the the old setup that we have right we have report that the analyst under the the oil and gas platform, right? The analysts that in the refinery want that data, but we would like, sorry, it's going to take us
[05:12] a little bit longer because SAP is a lot of data there and we couldn't clean it up fast enough. We couldn't make sure that what we give you and what finance is going going to get your finance team is have to match, right? Because if what I we give you and what finance give you
[05:28] is not match and then we're going to have a problem. So sometime it take days for us just to resolve the data quality issue that we have.
[05:44] So two path one platform right. So how are we end up today? So in production and operation we start with very simple right we have staging right and then we have operational data store we have a
[05:59] EDW enterprise data warehouse and on top of that we build many many data mods right and then that's not enough right that's still a lot of problem so we move to IBM at niza right so in there we can
[06:15] now consolidate all that uh all that data in one place but every time someone want to change even one column we have to rule everybody together and then make sure that you know my change is not going to impact your change so after
[06:32] sometime that's not working so we move to the cloud right we use the Hadoop right we use HDI it still take us days to actually get the maintenant management report together right it the compute is not there at that time for
[06:49] us. Okay. And we're using of course we move our stuff to Azure right our journey on Azure we use uh uh we use starting with uh sign up right and then we move to fabric and now we in data
[07:05] brick shini going to tell you a little bit about journey for our CNB uh business. Yeah. So when it comes to the customers and product part of our company so like the same way like Julie was talking about we had the concept of staging operational data store enterprise data
[07:22] warehouse and again we have data marks in the era of initial one then historically we moved to cloud data lakeink and then eventually we decided that we have to go back to the AWS native version of the data processing back here using AWS red shift. So
[07:39] irrespective of whatever the technology data is not in one place that is a fundamental problem what we have realized in this whole world. So what we decided is basically in the recent past is that the entire CNP workload will be migrated into the datab bricks on also
[07:58] that means that we are going to have the datab bricks on both cloud providers AWS and Azure you have data in one place and in the further slides I'm going to talk about what are the advantages what we got because of that and the CNP journey is in the middle of it we did not
[08:13] complete that journey we are uh we have completed almost 30 to 50% of that workload off of whatever we have in Amazon red shift back to data bricks on AWS and if you see in the uh if you see here
[08:29] go back yeah please yeah if you see on the right hand side we are talking about we have one datab bricks ecosystem which talks about both data sitting on the ADLS gen 2 and also AWS S3 irrespective of the storage we
[08:45] are going to have one compute engine which is running back on data bricks. Okay. Yeah. So this this picture talks about the what are all the grand advantages of what we got as a part of this ecosystem. If you see this on the left hand side in
[09:02] the Microsoft Azure part of it, we have the upstream which is like a production and operations data which is sitting in the Azure part of the world and on the right hand side we are talking about the customers and products and customers and markets kind of data sitting in the AWS part of it because BP went with a dual
[09:18] cloud strategy historically to start with. So if you see in the middle you you talk about the capabilities of data bricks lakehouse what we are taking the advantage because of this consolidation what we are doing at this moment we're talking about we're getting the capabilities of the SQL warehouse where
[09:34] people can actually consume the data and again when it comes to the governance the unity catalog is the big picture where we are trying to get the capabilities of the governance and uh again data sharing is a very important piece which we are using where we are
[09:49] talking about if There is a data sitting in one part of the ecosystem. It is much easier to share that data into the other part of the ecosystem because in the past this was leading to lot of data duplication altogether which can be avoided through this data sharing
[10:05] capabilities what we have in picture. So this can be data sharing in talk in terms of data within the data bricks ecosystem and if you want to share that data outside also this delta sharing capability will give you the advantage of sharing the data across
[10:20] and we I'm going to talk about a little bit more in the further slides about the lineage and the audit capabilities but if you see the middle part of the slide is where we are trying to get the advantage of having single ecosystem across the clouds.
[10:37] Yeah. Yeah. And in this particular slide, we are particularly talking about how does that flow work, right? So we like a typical data injection pipeline using the medallion architecture, we are trying to use the native capabilities of the data bricks to ingest the data into the silver or gold layer like a typical
[10:54] data lakeink architecture and then on top of it we are using the SQL warehouse capabilities so that consumers in the downstream can actually consume that data. In our ecosystem, we have consumers like PowerBI. We have data being consumed by using some kind of
[11:10] JDBC interfaces. And we also have the Palunteer ecosystem also where we are actually virtually sharing the data between Palunteer and databicks ecosystem through this new capabilities what we got as a part of this uh Unity
[11:25] catalog. Okay. And in the right hand side we are particularly talking about we are using the capabilities of the serverless and we are also using the JDBC and OBBC interfaces when it comes to PowerBI consumption and uh Julie would you mind adding a little bit more
[11:40] about the genie capability what we just uh started. Yes. So at UC um we follow the medallion architecture. So we have our raw data in just in asis and then on top of that we build a silver layer with the proper
[11:56] data model okay and proper metadata capture and we build a goal what we call is the data product layer where we speak the the business language right so we we take the data from SAP and maximo we
[12:11] convert that into the equipment model uh as a silver so we have a data model where you have your equipment and your location and your material connect and joy together. And then you're going to build your data products where your uh uh uh where you're going to say uh my um
[12:30] uh uh my health uh mater uh equipment health check for example and we make all that data available for Genie to use or for AI to use. Okay. So that is what the data brick assistant the jinny and the
[12:46] jinny code layer that build on top of our SQL warehouse and behind the scene if you enter AWS you stay in a AWS ecosystem we don't make you we don't force you to migrate we don't force you to come over to Azure if you in Azure
[13:02] you stay in the Azure world but both of us right operate with data brick on top on top of our cloud and then in between is sit the unity catalog right where I can see the the the the data formed at
[13:18] AWS if I'm stay on Azure and if I'm stay on AWS I I can see the data over on Azure right so no copy zero copy that's what we have over to you yep so can you go back to the next one please
[13:38] yeah so when it comes to the governance aspects of uh data in BP right In this particular slide, I want to emphasize a little bit more about the access control and the security aspects what we are trying to get the advantage of the unity catalog. So when it comes to the BP in the Azure part of it, we actually
[13:53] upgraded to UC in the past 2 three years back we are already seeing the advantages the UC upgrade. So one of the example in the top is basically the column level security or role level security right we can actually apply the control of the data at that RLS or a CLS
[14:11] feature which can be enforced across the consumers in some cases we are talking about the consumer part of the datab bricks ecosystem and sometimes it can be the outside the ecosystem like palenteer as well so this kind of control really made lot of improvements in terms of
[14:26] what kind of controls can be applied at the data Actually and again in the past we had business units some parts of the business units wanted data or they went with the duplication of the data because these kind of controls cannot be enforced across the ecosystem because
[14:44] now we are getting all the business units are going to come back into the data bricks on unity catalog. We can actually apply whatever the controls what are needed across the ecosystem. That is a big thing which we are a achieving as a part of this change actually and again there are many again
[15:01] many companies would have the same requirement where we are talking about what kind of additional fine grain controls what we can apply for kind of secret data or PII data or whatever right we're actively working with datab bricks also to figure out what can we do for that kind of fine grain permissions
[15:16] like attribute attribute based access control or any kind of role based access controls which is possible because of this unification of the platforms what we did so far okay and when I talk about the lineage aspects of it this is again a big thing
[15:32] when it comes to lineage because now we have the end to end lineage of how the data is being consumed across the ecosystem in the past what happened is basically you have data upstream data which is sitting in one cloud provider using a different tool and in the
[15:47] downstream as an example it's a different cloud and different tool altogether it was always a issue where operation ally if we talk about how does this flow work end to end nobody knows and everybody's talking to each other or talking between the teams to figure out what do I do so in this case uh this end
[16:06] to end lineage feature really is showing us the picture in terms of what is the upstream and downstream uh capabilities when it comes to observability aspects of it and again when it comes to the audit logging also this is again very important feature for kind of sensitive data or a regular data because we get we
[16:22] get audits all over the place in terms of who used what, what happened at what time. That kind of capabilities are something what we can get across the ecosystem because of this uh uh unification what we are trying to do and the column level lineage and the table
[16:37] level lineage is also possible only because we are using the same uh tool across the ecosystem and that is the advantage what we are trying to take and we are also talking about the capabilities of because right now we always get the audit request from KPMG
[16:52] or whoever that is and then we can actually export the what are all the audit controls what are followed or not followed. In the past it was a several weeks process but now it is much easier because you have everything in one place. Now we are trying to get that to a near real time also so that we don't
[17:08] need to spend the time on audits in general as a platform team because if you know about the audit it's a long process which we went in the past but now it is much easier because of the capabilities what we have at this point. Julie, would you mind talking about the
[17:25] discovery uh from a data discoverability and the stewardship standpoint? Now you have the data everywhere together not from a storage standpoint but in the unity in the data brick environments. So what we do is we have a um uh we develop a quick
[17:44] uh application where we pull all that data together for uh our business users. Right? So this is not designed for the uh for data engineer or data analyst because for us we just go straight to the uh unity catalog and we use that but
[17:59] our business is need a little bit more right it is need to speak in that term had to have the organization data on top of that as a meta data it have to have drilled uh down it have to have hierarchy built in so we call that the discoverability right uh disco the data
[18:17] discoverability in there You can go in and you can search for any data across the organization giving that data is a structural data. It's a table uh it's a schema table that sit in a unity catalog or you can have the unstructured data
[18:33] all the engineering documents that we have right all that is a streaming the pi data the sensor data that we have is all available for the business to search in their term right in their language. So that and from that discoverability
[18:49] you can uh if that is the data that you need right that exact you just click a button and that would say request access right so that request access will bring you back to the cal uh to the unity where the data actually physically sit
[19:05] right and then they they can start working uh with the data set that they have or they can access the jinny on top of that data set and start talking in uh uh talk to that genie in the in the English right in plain English
[19:21] from the stewardship standpoint. This is quite important because this is the uh every single data set that we have we put in a governance uh process to make sure that we have a owner right we know who's the owner of this data set is we
[19:36] know the data classification we know how long we have to keep the data set around in term of uh the data retention the data archive strategy so all that stuff is is is uh is taken care of right now by our data steward and that is fit
[19:52] right back to the unity catalog and then uh and and then into our data discoverability. Okay. Yeah. So this is a good example what Julie was talking about earlier earlier
[20:08] in the presentation where there are some jobs where we waited for two days or up to days for the data availability in general but now the availability of the data is much faster in hours because right now we don't have this ecosystem of different clouds or uh sorry
[20:24] different tools all together where everything is a batch process or whatever. Now the data is sitting in the datab bricks ecosystem as an example within the uh one of the we call them as business nodes or whatever and anybody who want to take the data out into the other ecosystem is much straightforward
[20:41] and that's the reason why we are able to save that time and this zero copies is a very big thing where we are not moving the data physically which we did that in the past and we all almost know that entire financial data for example the
[20:56] financial data which is sitting in SAP P in our ecosystem 40 to 50% of the data is getting duplicated because people are actually loading the data into the other ecosystem and doing their processing and the reconciliation also is a problem but now we are introducing the zero copy
[21:12] where we are actually doing the we are using the capabilities of the delta sharing so that we can actually virtually access the data and avoiding the duplication of the data all together. Duplication of the data not only solves the coh cost problem but it actually the availability of the data
[21:28] itself is faster because it has to go to multiple ecosystems in the past. Now it is much faster because you are actually taking it from the source. Any change in the source is automatically used in the other ecosystem and again it's not only within the databris ecosystem we have
[21:43] palunteer also which can take the advantage of this virtualization capabilities what we have in place. This is not only the cost thing, the performance problem and again the timing problem also can be solved using this approach what we are taking and uh the unity catalog governance layer like what
[22:00] Julie was touching earlier. We can actually apply all that controls whatever we want to apply at this level versus talking about doing that at multiple levels which we did that in the past. So this governance capabilities are evolving day by day. So we are
[22:15] taking the advantage of all that capabilities as we go. So as of now we are using all the capabilities what we have but we are also actively working with data bricks to say that we want this further features because that will solve our governance problem in general. Julie yes so from the government at scale
[22:33] right as you can see how much already we can save in this process before every all the data need to copy from one cloud to another. Okay. So not only we store the data in cloud one we store the data in cloud two but the process the pilot
[22:51] the job that we have to replicate the data from one another and the cost that we have to reconcile right the cost that we have to keep and changing the code over here oops I need to change it over there right so together in that right we calculate that number is quite
[23:07] impressive right but is that enough for our finance and our boss probably not right uh In our case, no. Right. So, what we need to do next is we need to now AI, right? We start working with Genie, we start working with AI and we
[23:24] see that the way that we actually control the cost or scale it up, right? Scale it up is we go vertical now, right? So for every single data domain for every single data uh uh business unit or what we call for uh portfolio
[23:41] and subportfolio we give them their own data brick workspace right what do you spend in there you pay for it the horizontal layer is where all the data will stay together right the the platform will pay for that layer but if
[23:58] you sit right here you're going to curate it your data you're going to create your silver you're going to create your golden you pay for that if you're going to use Jinny and you going to build your data set proper so Jinny will don't have to go and search and do
[24:13] all that contacts by itself and then you will save much more money your cost will see go down right but if you not following the standard and you're not doing doing what you supposed to do in term of build in the data model build a context into your data and then you will
[24:31] see your cost starting going up, right? But now because you own that cost, we say the be we see the behavior change, right? And with that, it's easier for us to scale actually, right? Now, if you
[24:46] have a business case, if your business support for what you to do, they will actually pay for that. Right? And then you from an extra small you can go up to the extra large in term of capacity if you have a business case to do so right
[25:03] if you have the amount of data if you have the the the the nonfunctional requirement for you to be that. So it's actually significant lesser time and really really easy for scale when we
[25:18] have a stable platform and everybody else just built on top of that. Okay. uh and that is also easy for you to do the self-service as well right uh we automate everything right through shiny uh platform we automate if you need a
[25:34] new workspace all you need to do is put in a request and boom you have a new workspace right and then but we monitor behind the scene right we handle the operational and the support aspect of it if you we monitor and if you see that you don't need an a medium you actually
[25:52] need a small, right? We will we will scale that down, right? We will scale that down for you, right? So with that, right, the cross the the crossbu collaboration is actually quite well. All if I need finance data, all I need
[26:08] to do is go to the portal and say I need finance data. Here is my use case. Why I need the finance data is go to the data owner. The data owner evaluate my case approved. Boom. Magically I see my finance data that actually sit in AWS
[26:26] now is is my in my Azure data brick workspace and I can work with that right and and um uh shi the lineage. The lineage is a wonderful thing. I can see my data is actually from SAP go to um uh
[26:43] AWS uh uh uh and now is available for me in Azure for me to use to to use so I can see all that lineage and now the the the conversation with the business about where your data come from right is
[26:59] significant easier actually with all that data context I hope in another three months right when I search my co-pilot again and where is your data is not going to be the top number one question that I see.
[27:15] All right. Um we actually um have a quick demo that we want to show you. Um hopefully can you see it? Yes. All right. So um as you can see right
[27:31] here this is all our uh data across organization. So it's it's an awful a lot of uh catalog that you can see is right here and then sometime you don't even need to ingest your data you can
[27:46] federated right so if you have a SQL server or you have something uh uh uh uh uh something that data brick actually support it right no need for you to actually ingest the data in you can federate it right just like you see that
[28:02] uh our naming standard FC Right? So that's federate catalog. So now data is there for you to use without even move anything. Right? So no cost for jobs, no cost for workflow, no cost for operational and support and and the
[28:19] storage as well. So it's quite convenient. We we use that quite a bit. Okay. And imagine that data is just like that, right? It's available in real time for you to use. uh but it's also in the same is subject to the same standard and
[28:36] the same uh uh uh metadata uh uh requirements that we have. Okay. And from here right from here I am in the azure environment right now. You can actually do anything. You see the
[28:54] AWS data and as long as you have permission to use that data you can you can do as you please right quite convenient. One point there Julia there were several cases in the past when we have to exchange the data between the ecosystem like what Julie was talking about when
[29:10] it comes to Azure or AWS. It was it was a huge deal where we have to go through several steps of request tickets whatever whatever now like what Julie was talking about it is pretty straightforward process as a user if somebody actually wants something it's a pretty straightforward process where
[29:25] Julie the owner as a steward will approve and they'll get the access to the catalog and again the permission can be at the schema level table level row level column level whatever that is all that controls are applicable which is a huge deal in the past so BP also have Palunteer before. What
[29:43] we have to do is we have to take the data in our legs, right? And we copy that data to Palunteer or even worse, right? We copy that data into a storage area because we don't want Palunteer to tap into our internal storage. So, we
[30:00] had to create an external storage, put the data in there and then Palunteer take that Marit agent to tap into that external storage. And now you can see the cost, right? and you could see the operational nightmare that we have. But now a day with the UDT catalog, you no
[30:17] longer has to do that. You just virtualize your data. So whatever you have in your legs either in any cloud, you can immediately see that in Palunteer and they can use that data just and they are subject to the same governance that we have in um in in in
[30:33] data brick or in our AWS or Azure cloud that we have. Okay. So um the way that we do it is we call our platform is unifi data platform right. So in here you can go and you can discover your
[30:49] data you can go to access control. So in when you want data to um uh when you want access to the data all you have to do is click on access control and you're going to say I am uh uh whoever you are right I am Julie and I am coming from an
[31:06] Azure environment right and I need the data that sit in AWS right and this data and then they're going to tap into the unity catalog right who are that and then there are multiple drop down and you just saying that I need finance data
[31:23] and that finance data is SAP PRO and the table is B set right so you can say that next to it you're going to see the owner who own this data set you state your case why do you need to see this data set right I why you need to access to
[31:38] that data set boom it's a automatic workflow it's go directly to the owner the owner evaluate your case say yes and then you immediately see that appear in your data brick workspace and then you and work on that, right? And then from the um uh from the for the data share,
[31:56] if you want to share the data with Palunteer, you do the same thing. You will go to Pal go to the data share and you say I want to share my data with Palunteer, right? And this is the landing zone in Palunteer and your data will automatically land there. Okay. Uh
[32:13] from the observability is the same thing, right? You can go in there and then you can say I am working in uh finance or I am working in wells and subsurface and here is my workspace. Boom is will give you all the data that we pull from the system catalog. Right?
[32:30] And then you can see your job what uh uh what the hiccup you have off the hiccup right now is automatically sent to you anyway. Right? You wake up with uh the hiccup there or the h with the recommendation of what you should fix about that. Yeah.
[32:46] uh for the um you can see here is we have the data quality and this is the journey right so for the data quality you will have to go and define right I have this uh I have this data set right I have pressure data pressure will go from negative one to one if anything
[33:03] other than negative between negative one and one uh that's not a good data set right so now what you going to do about it right are you gonna go back to the system of record and fix it so there is process that we have in term of what what we're going to do with the data
[33:19] quality or with the with all the hiccups that you have with your data uh with your data. Okay. uh you can see here is is the master data uh we have the master data management right in BP because we have so many different business units
[33:35] and so many different uh cloud different tool different system different application and each one of us speak the different language right so we have so many different way of dispelling the United States of America for example right or the country the region the
[33:51] government of Mexico have 11 ways of spelling that word right so now is That's is where it's come to our reference data and master data for us to um uh how we going to uh uh uh map the data to that right so so it's is is in
[34:07] that module okay and then we also have a way to actually what we call is uh dark data intelligent right so for dark data intelligent uh we give it a definition and we saying that right we're going to
[34:22] go and evaluate your data and see If anyone use it, right? If no one use it, can we do something about it? There is a retention policy, right? Even no one use it, by definition, you still have to keep it, right? Do you want to keep it
[34:37] in the expensive storage or can we move that into the coal or into the cool, right? Or into a cheaper storage for us to save the data for you. So, we don't just say here's the report. We're going to come to you directly. We're gonna say Julie, you are the this is your
[34:54] workflow. No one touched that data. You are the data owner, right? Do something about it. Give us a decision on can we move this data into the cool or call or archive it. And we also give you the um we also give you how much cost, right?
[35:11] The organization and and how you going to go and do something about that. So it's it's right there and submit a ticket and we will move that for you right. So this is the effort of using what is available from data bricks right
[35:26] in the system catalog uh that that help us save the cost and this is across the organization right so this is the part of why in the uh the title of today talk that we um uh we say with everyone
[35:43] moving to data brick right bp of the winner right um you guys have any questions

Learn more about the Databricks Data and AI platform.

The information provided herein is for general informational purposes only and may not reflect the most current product capabilities or configurations.