Skip to main content

Zero-Copy Multi-Cloud Data Sharing: Apache Iceberg and Databricks

Summary

  • McKesson Compile provides analytics-ready patient claims and provider reference data to healthcare organizations, and replaced a provisioning-first customer onboarding model with zero-copy data delivery using Apache Iceberg, Delta Sharing, and Databricks Unity Catalog — eliminating the need to duplicate data for each customer.
  • The architecture publishes data once from a Databricks source so customers running different platforms including Databricks and Snowflake can access the same governed data without moving any bytes, with schema evolution managed centrally from a single source of truth.
  • This video includes a live cross-platform demo showing a Databricks source sharing data to a Snowflake consumer through standard protocols, with governance and trust enforced by Unity Catalog so operational costs remain constant as new customers and platforms are added.

Zero-Copy Multi-Cloud Data Sharing: Apache Iceberg and Databricks

Watch: Zero-Copy Multi-Cloud Data Sharing: Apache Iceberg and Databricks
Scaling healthcare data products across multiple customer platforms creates complexity, cost, and confusion. Traditional approaches require duplicating data for each customer onboarding, leading to maintenance overhead, inconsistent refresh cycles, and expensive operational support. McKesson and Databricks solved this vendor lock-in problem through open standards: Apache Iceberg for unified data format, Delta Sharing for platform-agnostic access, and Databricks Unity Catalog for governance enabling zero-copy delivery to customers running Databricks, Snowflake, or other platforms without moving a byte of raw data.
This talk deconstructs the engineering journey from a provisioning-first model to an operating model centered on trust and governance. You'll see a live cross-platform handshake between a Databricks source and Snowflake consumer, learn how schema evolution works across systems without syncing, and discover how to architect data products that scale seamlessly as you add new customers and platforms while keeping operational costs constant and data governance centralized.
🤝

Chapters

FAQs

What is zero-copy data sharing and how does McKesson use it?

Zero-copy data sharing means granting customers access to data without physically duplicating the underlying datasets for each recipient. McKesson uses Apache Iceberg and Delta Sharing so that customers on different platforms can query the same governed patient claims and provider reference data without McKesson maintaining separate copies or pipelines per customer.

How does McKesson's architecture support customers on both Databricks and Snowflake?

Apache Iceberg provides a unified open table format compatible across platforms, while Delta Sharing enables platform-agnostic data access through standard protocols. Customers on Databricks, Snowflake, or other compatible platforms can all access the same governed data product without McKesson maintaining separate ingestion pipelines for each.

How does schema evolution work across platforms in this architecture?

Because data is published from a single Databricks source with Unity Catalog as the governance layer, schema changes are managed centrally without needing to synchronize updates across multiple copies. Consumers on different platforms receive the updated schema definition automatically through the sharing protocol.

What challenges did McKesson face before adopting this zero-copy architecture?

McKesson's previous provisioning-first model required duplicating data for each new customer onboarding, leading to maintenance overhead, inconsistent refresh cycles, and expensive operational support as the customer base grew. The open standards approach with Apache Iceberg and Delta Sharing allowed operational costs to remain constant regardless of how many customers or platforms were added.

Full transcript

[00:07] Good evening everyone. Thanks for joining in. I'm not sure why I always get the slot just before the happy hour start, but yeah, that's how it has always been. Uh Uh before we start uh let me quickly introduce what we do. We are representing here McKesson compile. Uh
[00:23] we are a company which provides uh analytics ready patient ready uh patient claims and provider reference data to healthcare organizations. Um and with those data our customers get the uh insightful uh insightful uh decision-making power
[00:41] uh because of uh both for the research and the commercialization purposes. Uh And before I start uh before I talk about uh uh open passport, before I talk about
[00:57] Databricks, before I uh uh talk about uh Apache Iceberg or data sharing, I would like to uh start with uh with one incident which recently happened. On this Sunday, I know we I mean many of us are flew have flew in from different parts of the world. Uh have anybody visited Golden Gate Bridge
[01:13] recently? Quick show of hands. Can you please tell me what's the color of San Francisco Golden Gate Bridge by any chance? It's a trivia.
[01:34] What brown? Go ahead. International orange. That's right. I always thought this is the orange, but uh that's the international orange. And apart from that, I went and read the history about San Francisco Bridge that how big an engineering marvel it is. It was made 90
[01:49] years uh 90 plus years back and it's still serving. I was not sure whether I'll be able to take a walk at the bridge or not. But I was able to um there was a lane for cyclist, there was a lane for pedestrians, there was for commercial vehicles with the speed signs, with the highway patrol,
[02:05] everything. And I wished that maybe our data ecosystems should have been like that. That you build an infrastructure once, it lasts for generations, and people just need to reuse it just by following the protocols, following the
[02:22] rules, and if something happens, there is a patrolling or governance. So, that would summarize the session that we would have in the next 40 minutes. Um Uh before I start, uh we'll just add that this theatrical would have five
[02:38] facts. I'll start with the problem. Uh then I'll quickly cover the architecture. And then I'll tell that how we get there, what was the problem, how we were solving, and like any story there should be some hiccups, and then how we overcame that. And then I'll hand it over to my partner in crime, Vinya,
[02:55] he will uh take you through the live demo and the takeaways from this entire POC that we did uh with the Apache Iceberg and Delta Sharing with Databricks. And uh four key terms which which I might repeatedly use during the entire
[03:11] session, so just wanted to introduce them so that they are out of the way. Apache Iceberg, it's nothing but a a common uh table contract. Delta Sharing, nothing but a common access uh contract. Unity Catalog, nothing but
[03:27] a certified uh governance lineage and one audit uh or control. And lastly, the Flux, it's an internal one. Uh so, we have our own product provisioning uh tool that we developed, and how the journey has been will be sharing about
[03:43] that. With that, uh I'm sure uh most of you uh uh might have faced uh this challenge where um every new customer onboarding on a data product looks like a small project.
[04:00] Uh which should not be. And that is uh the problem that we faced. Uh In In reality, a customer doesn't want a new pipeline or a customer doesn't want uh data to be in a certain way.
[04:15] What they say is, uh put the data in my platform. Actually, what they are trying to say is that we have evaluated our data and I'm coming from um our data product experience. Uh what they mean is that we trust your data, we have done our due diligence. Um we are ready to pay the price. Now,
[04:32] what we just want is get the data in our platform so our team can start using it. So, basically, it's an access problem. Uh but nine out of 10 times for all the people who have been this journey, uh what we see is that it becomes a project. Um
[04:50] there would be a request, ticket, approval, pipeline creation, some customizations, and then the person get the data. And today's world when speed could be a moat, that becomes that that becomes a real challenge.
[05:09] So, again, taking the analogy from uh the Golden Gate Bridge, what how would it feel that uh if you drive a car over there and after every few meters, we need to uh on unboard the car, inspect it, and then fix it again, and then again go.
[05:25] The data provisioning uh in the older world looks like this. Every time you do it, every step is there for a reason, by the way. But in today's world, it seems like an over-engineering because uh we are into we are we are solving the access
[05:40] problem, not actually the sharing problem. And that is the complexity part. The next part is the timeline. And as I said, in today's world speed is the mood. And with the approach the old
[05:56] approach, it takes a lot of time than required. Now comes to the cost. Every storage creates a cost. And it is not just a storage part.
[06:12] The real challenge is what comes after that. Like for compute, validation, support, operations, the cost multi-folds and believe me, the value that the customer gets doesn't extrapolate like that.
[06:34] So, we talked about complexity, we talked about cost, but the real challenge as a product manager, and I'm sure many of you who are from the engineering fraternity will realize that the real challenge comes from confusion. We hate as a data product team when the conversations are about the
[06:49] sanctity of data rather than how to use the data. Because when you create multiple copies, it creates confusion and all the steps are necessary sometimes. Sometimes some metadata changes, sometimes refresh changes, and every product evolves
[07:05] because there's a need for that. But because of these copies, unnecessary confusion is there. So, when we sat together as a product and engineering team and decided that how to change that,
[07:20] we decided that this would be our true north question. Uh can you just deliver the data directly into our platform? How to solve that? Because somebody wants in Databricks. Somebody wants in some something which
[07:36] which is SF. Can we say that? Okay. So, yeah. Um So, that was the true north for us and with that we begin. Uh that brings me to the the act two,
[07:52] which is about the architecture. So, we have defined the problem. Now, we are going towards that what are the building blocks of the architecture to achieve what we aim to. So, that question really changed how we look
[08:08] at our product architecture. Earlier, we were thinking we need to provision the data, how fast we can do that. But now, the question is not about the speed, it's about how securely we can get just give the access.
[08:28] So, the old model where every new business integration, every new business unit integration within a customer, or if there is a just a new AI tool or a BI tool integration from the customer side,
[08:44] it costs us to copy the data with the with the privileged access to the certain folks. But with the with the new model, which which I'll go into the details in the upcoming slides, it is just becomes the access part.
[09:04] So, the three open standard, which are the backbone of this architecture, at the bottom you can see the data layer, which is the Apache Iceberg. As I said, it is the the the open data format. Uh basically, uh we want the customers to read from
[09:20] anywhere. The second, Delta Sharing is access layer is a common access contract where uh we can share with our customers from any uh to everywhere. And lastly, uh the trust layer and coming from a um
[09:36] regularized industry like health care and I'm sure other industries like us as well. Trust is the heart of our product or governance is the heart of our product. So that we we try to bring in with the Unity Catalog of Databricks.
[09:52] So with these three building blocks we could achieve with the problem which we wanted to solve. And uh all three are necessary. If you miss one, then you won't get the desired result.
[10:08] So it's like uh it's like a PDF. So when you have a PDF, PDF is not necessarily a data format, but it's a data contract. Chrome can read it, preview can read it, Adobe can read it. So once you have this, so there's a shared understanding that what a
[10:25] document is. Similar to once you have Apache Iceberg that set that standard that what the table means and whether it's an analytics engine or an AI application or or a research or the future future future customizations, they all be able
[10:41] to understand and understand the data in the same way in which it is intended to. As I said that same goes for another data warehouse and the platforms as
[10:58] well. Same Apache Iceberg provides that privilege. Now, how to understanding the data is one thing, but how to access that and that's where data sharing comes in. Uh you can see here uh
[11:13] the customer can discover, request access, authorize query, and receive the results. One part is missing in this channel and rightly so, which is the copying the data. So what what gets shared is the metadata, policies, and query requests,
[11:29] but not uh the actual data or the bulk copying of the data. Uh that's the power of um data sharing. And lastly, for the governance, Unity Catalog, one policy, one audit, one lineage, and you
[11:45] don't need to replicate or or be worried about uh the multiple access controls. So, the data will be delivered with the people who wants that right access, and it is shared across multiple platforms seamlessly.
[12:06] So, again, taking the analogy from the Golden Gate Bridge, so your data layer becomes your road or the infrastructure. The access layer uh becomes the traffic rules. And the governance layer becomes the highway patrol. So, with these three, you can you can
[12:23] build the infrastructure first and can reuse it without any worry. And lastly, on top, the experience layer, which I'm going to talk in my upcoming slides, is that how to make it a user experience great user experience because the customer need not worry
[12:40] about all these um these these backbones. They just need to share I mean, get the data and go with that. So, that brings uh me to the act three, which is our
[12:58] engineering journey or the experiences we had while solving this problem. And as I mentioned that like um any good story, we need to have that what we faced as a hurdles and how we came over them. So, when we started solving this, we
[13:14] thought that um the idea is to increase the speed and the user experience or or enhance the user experience. So, we focused our energies on how to provision our data in a best user experience way. How can we can make it
[13:31] self-serviceable as self-serviceable as possible. But that didn't work out well. So, this was the the case that whoever is using the our product for the provisioning, they can choose the data set, they can they can go for the entire
[13:47] data set or the subset of it. Uh they can choose the destination where they want it, whether it's uh data bricks or SF or uh flat files. And we can provision the access of the light levels, and it is done.
[14:03] But one thing was missing over here. Uh we didn't uh solve for the the the open format and the data sharing, and hence the copying was the problem with with each share. So, we simplified the request. Part of the problem is definitely solved. So,
[14:20] from the customer intent part, it was great. Customer was able to articulate what they need very clearly, and we were able to provision that as well. But the problems came later. Got more requests.
[14:38] At the same time, we got more copies of the data. And as I explained earlier, with each each copy, it creates cost, complexity, and confusion. And again, this slide just just strengthen the previous argument that
[14:55] that how the journey goes when you create more copies uh both in terms of storage, in terms of transfer validation, monitoring, support, and the operational coordination.
[15:17] And with every schema change it becomes an event. Because then every chain, people start questioning the sanctity of the data, which creates further confusion. And it becomes uh at a beyond a point, it becomes like we are in the business of maintenance.
[15:39] But with Apache Iceberg and the uh and the data sharing, uh it becomes one format and we can share it with the many platforms without copying. So, the copying is uh taken out of the equation with with the help of these two building blocks. And with the governance, we can control the access as well.
[15:54] We don't need to create uh the multiple policies again and again. It's a one audit, one lineage. So. So, this is to summarize. Uh this has been the journey where we came from uh
[16:10] as a provisioning tool. In fact, our our tool's name was a provisioning tool when we started. But now we can safely say that it has become an operating model for us. So, we we came from the provisioning, from data engineering, from from a platform or product thinking.
[16:25] With that, I'll I'll bring uh Vinay to the stage who will will uh give us the proof of the demo. I'm sure you'll like it and also explain that what are the takeaways from this experience.
[16:52] Uh thanks. Thanks, Ravi. Um Yeah, it has been an interesting uh journey and uh and uh as you see it, right? Uh uh when you actually go into the into the handshake motion, basically, uh we are basically replacing what we think
[17:08] today is is working for us is actually not the real thing. And uh what we really want to get into is how do we publish once and and then leave with that kind of thing, right? Wherein your refreshes you don't have to
[17:23] go and refresh for every customer. Wherein you don't have to go and uh and again fight for it in the sense go and figure out, okay, for this customer did the data go right in the right manner? Do you go have to go validate that data again? So, you you
[17:39] keep repeating again and again the same process for various customers. Also, how do we avoid it? Uh kind of That's where we started with publish once and and then once you grant the access, we go into the what we call as query. We can
[17:55] see some of these things. I would put it like I don't have the entire demo set here, but I have a small bit that I can walk over here. So, what do we do to prove that this architecture really works?
[18:11] You you you basically go create a share, give the access, connect with the customer wherein you share the metadata uh and then you then the customer can run his queries kind of thing. And then we repeat this process, right? Wherein you refresh uh
[18:27] From a traditional model, what changes here is we are we are not going and refreshing it for individual guys, but it happens for once and everybody gets the same data kind Right? And this is this is what we are trying to
[18:43] achieve by by moving towards more of a open standard flows. And to start with, how do I go create a share? Right? Once once the share is created,
[18:58] basically you can share it with any of the things, right? In fact, you use some of these things and the engine that we have built across you just do a couple of clicks and it goes and shares. So, you don't need uh expertise there
[19:13] uh to say, "Okay, I need uh I need to go into a particular uh a database or a a particular infrastructure." You can You can just use your engine uh and it it does the things. Uh what we are not doing uh here is uh what we call as the copy or the
[19:30] export of the data. Uh lot of times when when we actually get a request from the customers, "Okay, I need it in a particular platform." And we are not supporting Right? We end up giving them something like a flat file on Azure or uh or S3 kind of thing. And uh what we
[19:46] want to avoid uh is some of these things, right? Where we can support it if you're if you're actually on a open uh protocol uh flows. And any of these, even when you are going into a futuristic purpose, also this really helps. And uh as Ravi pointed the
[20:02] analogy of Golden Gate Bridge still standing uh uh kind of thing, right? So, that's what we are also looking in terms of a uh software product here. Okay. Uh so, now I created the uh the the share. Uh how do I uh
[20:19] then get this uh get this into the customer's hand, right? Uh so, we basically uh get the metadata, uh share it with the customer. Uh there's access end point that you give to them. And then there is an authentication that you share over a secure channel. And uh
[20:36] yeah, they are then the cus- customer just connects it. Uh and he's good to go uh kind of thing. And if you see in in the entire flow uh right, we are not really moving any data. It's only the the trust and the metadata that moves across. And we are
[20:52] not really uh touching the any of the data. Data is still a single copy. And this really helps us to manage it better. There is not uh much of the confusion that one can add there because you have a single source and it really helps uh for you to move
[21:08] forward kind of thing. Cool. And uh And and this is what you have when when you have uh This is an example of the Snowflake one here where the data is shared from a
[21:24] Data Bricks into Snowflake. I'll walk over the live example in some time. And how do we do it, right? We basically use the Iceberg REST and the Delta sharing part of it and then Yeah, you give the data out. So you're at this point what we are
[21:40] putting across is is basically the the Delta sharing part of it. And this really helps us to connect to I mean, here is an example of Snowflake, but we can connect to S3, we can connect to blob
[21:56] and then move forward kind of thing. And We can even do some other uh databases also without even moving any of the data that we are talking. I mean, uh I think at some point we are basically looking at not really a technology, but
[22:12] what is a governing principle, how we can drive this. And lot of architectures which stands tall is because of this governing principles that that we have. And I'll talk more about that as we move forward. And And yeah, next is basically
[22:30] how how customer accesses it and you still query the uh data that has been shared. Advantage of being on Apache Iceberg. And you see all the information that is there, right? Uh
[22:51] So so what what we bring here is basically you update once and and everybody gets the same data kind of thing, right? So so whenever there is a new publish that happens, Iceberg actually takes care of your schema evolution. There is
[23:06] also a version history that is kept and also you have the snapshot management. So, you have one copy and and that copy itself works for everybody and you can share it across different ecosystem like the
[23:22] you you talk about the AI applications or the future engines that we have and you can you can actually use any of that flow share. There is there is this what do you call has one update and and and and what consumers get is is
[23:40] actually already synchronized. You don't have to tell them that okay, now tomorrow you are getting tomorrow this and day after tomorrow one more consumer. It happens on the fly and everybody is is already synced up kind of thing. So, you don't really go and communicate with everybody and
[23:57] that's a that's a good advantage that we bring with Iceberg. And as I said, right, the freshness is without synchronization. So, you don't really go and keep updating everybody and
[24:14] because you have updated you have published the data and the contract gets updated, everybody gets to see the data on the on the fly kind of thing and it's all the ground up source. So, you don't really have to
[24:29] what do you call has put lot of resources there to manage it over a period of time. So, if you if you were to look at it, creating data pipelines, creating the ETL jobs for each consumers, that takes lot of time and also there are lot of these
[24:47] what do you call has the maintenance work that comes into picture where things don't just fit in there, right? Because you keep doing it again and again and it becomes a repetitive process. I think one of the things that we avoid with this is having lesser
[25:03] people to manage it also. And what we have done in this entire process is we have not moved the data, but we have moved
[25:19] the the trust part of it or the the governance part of it along with with when the data travels to the customers end, right? So so we have we are basically talking about how how we can actually move data without without actually creating a copy
[25:35] of it, but still have the same access across different different customers kind of thing, right? So that is that that's one of the things that really helps with the governance part of it and this is where we we are using the Unity
[25:51] catalog that comes into picture. Um, let me just uh
[26:08] Okay. Uh so what I did is a small prototype here and what we have done here is basically what we call as data products, which is which is slightly different from the way we are looking at data because right now we are we are
[26:24] basically saying that we have a data product which is a supply chain issue and I have these these products in my warehouse and I want to ship it kind of thing. All right. So that's how you you can see this and and I'm just giving one example here.
[26:51] So I would like to share this data and I would like to maybe use Snowflake as an example here. I've already populated with some of these things. Uh
[27:08] We've been like testing for some time now, so So you you can see here uh what we have done basically is uh a simple click and then the data uh there's no there is no uh what do you call copy that gets created, uh but it is just the data that gets shared across and and
[27:26] everything works uh seamlessly there. Uh we can actually Yeah.
[27:42] Uh so this is this is how it eventually uh uh comes into like uh you you can see that it has been shared across and and and this is one small example that I had here, uh but uh you you can actually replicate it across uh different products, too.
[28:07] Uh coming back right. Uh I mean uh so one of one other challenge that we actually generally face is a lot of times there is uh this access issues or once you give data to your customer, it's not that it always works, uh right? Uh there are always issues that comes back and uh and lot of times one of the
[28:24] common issues that we actually face is I don't have access or it is not working for me. And uh with with this open uh way of uh doing things, uh all we have to fix is mostly the access issues rather than the data issues.
[28:39] All right, if you have creating like multiple copies of uh data, then uh then the the challenge will be you have to fix the data as well as access and uh with with with a single copy you don't really worry about the copy but you go and fix the access if at all for some
[28:55] reason it got uh removed part of it. Right? So, you're uh you're debugging your uh becomes also easier and your uh you don't really spend a lot of time uh fixing uh issues looking at uh data but uh you are actually looking at the trust uh part of
[29:12] it uh which uh which which really helps to move forward quickly and uh your innovation uh takes the front seat there. Okay. Uh so, so uh what we discussed till now is uh uh basically like there are multiple
[29:28] uh actions that we want to do uh but everything uh is resolved uh is is centered around one thing which is the trust uh part of it. That is that is what we call as our operating model. Uh that is that what's also the basis of this uh architectural change that uh we
[29:45] are uh discussing now here in this session, right? And uh what what happens uh here is we have been talking till now around data uh but uh it it is not about uh anything around moving of the data but it it is mostly
[30:02] around uh how do we build that trust uh without actually moving the data right? Uh and uh and that's what uh we have been uh putting across uh and uh I I think uh uh one of the things here uh that also comes in like what Ravi pointed we spend
[30:19] a lot of time on onboarding a customer and getting them the access part of it. And uh with this flow I I think we also get into more of a self-service mode for the customers kind of thing. Right? I mean, if if I can go click a button uh I can actually give that uh uh
[30:35] give that to the customer for him to uh get the data itself kind of thing. So, uh coming to the the the last act, right? So, we basically went through one flow
[30:51] wherein showed like how the data can uh uh where where we where where the data lies in one place and but the but but shared to multiple customers across the different databases. And
[31:08] what what we understand from this is the the shift that we are looking at is basically not not from the uh uh data perspective, but we are looking at it from a more of a trust perspective, right? How do you How do you ensure that
[31:24] there is a good governance around a single copy of data which is shared with multiple people, right? How do you How do you get into the principles where your your architecture model has changed, but you you still stick to
[31:40] what we have been doing it in the sense of delivering data to our customers. Uh And and how do you decide when to copy the data and when not to copy the data? All right. And also ensure that when you get into new
[31:56] uh futuristic engines, how does it work, right? And and and I think as we move forward sticking onto the open protocols really helps us to uh keep it more open uh rather than getting
[32:12] fixed to one platform or multiple uh what I would say is like multiple platforms exist and each one of them have their own advantages and disadvantages, but sticking with sticking to open protocol, you can
[32:27] actually shift across based on how the market moves, based on what the new tech comes in part of it. Right? Uh so so the entire uh architecture that uh we are uh putting across uh is is based on uh
[32:45] the underneath uh part of these principles, right? Which is uh open standard. Uh we talked about uh open formats, which is the iceberg part of it, the open protocol, and how it really is future compatible. Uh there's also the governance part of it that we
[33:02] talked about the Unity catalog part of it. And the final thing around not really data, but it is more around data products. Uh it's more around uh uh I would I would put it at we we really did not create a pipeline or anything like that, although I showed you in the
[33:18] nice uh flow of data how it happens kind of thing. Uh but but all these principles eventually gets down to your trust protocol, right? How do you ensure that uh the data that you're giving to your customers has uh is has the right
[33:34] access. And how do you ensure that uh there are no multiple copies of the same uh uh data. So so uh overall, I I I think uh when you when you look at architecture uh part of it, right? Uh we we are basically looking at these principles, uh which governs our
[33:51] architecture rather than uh uh rather than saying that we built a like great uh architecture and and nothing works in the end kind of thing. So so if you're around this uh principles, I think it takes us a long way uh on on how to move forward.
[34:12] And uh here here is uh the various things that we uh discussed about, right? Uh I mean, we've always been uh discussing about uh new customer uh comes in and our our growth uh is basically add new ETLs, add new uh pipelines, and then we keep
[34:28] moving, but that's not a scalable system. Uh right? Uh the more uh pipelines you add, the more you maintain, and the more you maintain, that means that you invest more time, more people on it as time goes by kind of thing. And again, you end up with a more tech debt,
[34:45] and then you spend It's a It's a vicious cycle part of it, and that's where the the trust protocol what we have built across really helps us to move forward in the right direction. Um
[35:01] So so you don't have copies. You don't get stuck into one platform, and and you have a ecosystem which helps you to scale and scale for the future part of it, right? And when when you scale and
[35:17] your operational cost remains constant, that's when you're actually having a right growth. All right, a lot of times as I pointed, right, your you add new pipelines, your customer is coming, your operational cost also grows. So so how do you how do you
[35:33] ensure that your cost does not grow, but you you still have have new customer added, and that's where this uh scale comes into picture, and the trust protocol and the and operating model that we are discussing about comes into
[35:49] picture. So I think if you were to look at the transformations that has happened over a period of multiple years also, all right, a lot of times we started with
[36:06] the data delivery where we said, "Okay, for this customer we'll give manually like run the query, run the copy commands" kind of thing. And then, okay, then when once we started growing, we started doing the shared platform where we said, "Some from data we will give to multiple people" kind of thing.
[36:22] And uh now we are at a stage of data products where we have uh we have the data sets and then then you can do a self-service part of it. Uh and and then the way we are going towards is is where we want customer to enable,
[36:41] which means that uh it's a faster onboarding for him and and and and there is less friction uh part of it, right? So, and as we move forward to the next stage, it becomes uh we are getting into more of a ecosystem scale part of it, uh where we can scale
[36:58] without actually increasing the operational cost and the consistency remains across all the uh layers that Ravi actually put across kind of thing, right? Uh so, uh so, this is more from a product uh part
[37:15] of it, right? Where where we are talking not really the architecture, but but the operating model uh to take this forward uh and if you were to do it, how do you go about doing it? I think this is the
[37:31] blueprint uh that that you would uh suggest to have in the sense that get into the iceberg format, then get into uh delta sharing, uh have unity catalog and then there is we call it as a flux, which is our
[37:47] internal uh tool, but yeah, it's a self-service layer which can really help to uh push push a lot of things without having to get into uh the infrastructure, without having to get into a lot of permissions uh part of it
[38:03] or the approvals flow of it and uh it's a self-service tool in the end kind of thing. All right. Uh with that, I think what we achieve is not not really like copying uh the same data to multiple customers, but
[38:19] we have like one operating model that you use across everybody and uh helps in really scaling this to the next level kind of thing.
[38:37] So just summarizing what we discussed till now. Uh it's it's basically around uh one data and you have different consumer platforms and you have different consumers unlimited consumers kind of thing and and it works across because we are sticking to the same
[38:53] operating model which which talks about governance and sharing part of it uh uh part of it, right? So so uh
[39:11] Just to summarize, right? So so we are basically talking about uh one product many consumers and one trust model and this is all built around our operating model which is basically what Ravi pointed out in his
[39:27] Golden Gate analogy, right? Wherein you you still have it and it has the governance part of it. You can take care of the standards and your customers get the data in a easier manner, right? And and they don't really have to worry
[39:43] about okay, when is the next refresh happening kind of Okay. Uh With that I would like to thank all of you for making it here and listening to us. Thank you.

Learn more about the Databricks Data and AI platform.

The information provided herein is for general informational purposes only and may not reflect the most current product capabilities or configurations.