Skip to main content

Iceberg meets Databricks: OpenSharing for cross-platform data collaboration

Summary

  • Databricks OpenSharing now provides first-class support for Apache Iceberg tables, enabling live, zero-copy data federation across Snowflake, Trino, and other Iceberg-compatible systems while keeping governance centralized in Unity Catalog.
  • Foot Locker eliminated terabytes of data duplication and dozens of SFTP export pipelines by federating Iceberg tables from AWS Glue and sharing Delta tables to Snowflake through OpenSharing.
  • Unity Catalog maintains centralized governance over all cross-platform data sharing, providing auditability and access control across multi-cloud, multi-region, and multi-business-unit architectures.

Iceberg meets Databricks: OpenSharing for cross-platform data collaboration

Watch: Iceberg meets Databricks: OpenSharing for cross-platform data collaboration
Organizations increasingly use multiple data platforms across clouds and regions, but data sharing traditionally required expensive copying, complex pipelines, and scattered governance. Databricks OpenSharing now provides first-class support for Apache Iceberg tables, enabling live, zero-copy data federation across Snowflake, Trino, and other Iceberg-compatible systems while maintaining unified governance through Unity Catalog.
Learn how Foot Locker eliminated terabytes of data duplication and dozens of SFTP export pipelines by federating Iceberg tables from AWS Glue and sharing Delta tables to Snowflake through OpenSharing. Discover the architecture enabling real-time data sharing to multiple vendors and partners, how governance remains centralized through Unity Catalog, and the implementation playbook for designing multi-platform data strategies.
🤝

Chapters

FAQs

What is Databricks OpenSharing and how does it differ from Delta Sharing?

OpenSharing is an evolution of the Delta Sharing protocol, extended to provide first-class support for Apache Iceberg tables alongside Delta tables. While Delta Sharing enabled zero-copy data sharing with downstream recipients, OpenSharing adds bidirectional exchange with Iceberg-compatible platforms such as Snowflake and Trino.

How did Foot Locker benefit from using Databricks OpenSharing?

Foot Locker used OpenSharing to eliminate terabytes of data duplication by federating Iceberg tables from AWS Glue and sharing Delta tables directly to Snowflake without copying data. This removed the need for dozens of SFTP export pipelines, significantly reducing operational complexity.

How does governance work when sharing data across multiple platforms with OpenSharing?

Governance remains centralized through Unity Catalog even when data is shared across different platforms and clouds. Unity Catalog provides auditability, access control, and lineage tracking so data owners maintain visibility and control regardless of where data is consumed.

What are the main use cases for cross-platform data sharing with Databricks OpenSharing?

OpenSharing supports peer-to-peer sharing between individuals, marketplace distribution for notebooks and datasets at scale, sharing across business units in large multi-cloud organizations, and integration with SaaS platforms such as CRM systems to centralize or distribute data across different systems of record.

Full transcript

[00:08] Good afternoon, everybody. Thank you for joining us for the the final breakout session of the day. Very thankful that you're all still with us and still engaged and present after a long day. I know it's uh I know it's been a lot. Um we have an awesome talk for you today on on Iceberg interop- interoperability.
[00:25] Make me say that five times fast and I won't be able to do it. Um please give a huge round of applause to uh to Ti Chuang and Balaji uh Swaminathan.
[00:42] Well, thanks for being here. I know it's literally the end of the day, so I hope you guys are in the right room talking about open sharing, Iceberg interoperability. We're almost done, so bear with us. Okay, so let's talk a little bit about collaboration, right? And I imagine you guys are here because you probably need
[00:58] to do some sort of external data sharing or AI asset sharing or some sort of sharing in as part of your organization. Um and obviously this is like business critical, right? Just breaking it down to a couple of use cases. Um maybe you'll relate to some, maybe some is more forward-looking. So the first one
[01:15] is peer-to-peer sharing, so this is you're sharing from one person to another. Um maybe this is, "Hey, I have a data set that I want to share with you." or "I have a view that I only want you to see." right? So it's point-to-point sharing. The second one is maybe you're also sharing some sort
[01:30] of third-party asset, right? So we also have the marketplace which is powered by open sharing. And let's say you're listing uh a notebook that you want to share or a data set that you want to share, um but it's actually a one-to-many type of distribution model that's also powered
[01:45] by built open sharing. Third one is maybe your organization is massive, right? And you have a lot of sprawl again uh across different regions, across different clouds, um and also across different business units. Um so BU sharing is also something that we see a lot. And then finally, maybe you
[02:02] have different systems of records across different SaaS applications. Um, maybe you need to share your data to different CRM platforms, or maybe you want to get data from say Salesforce and centralize it. Um, that's obviously also a key use case that we hear about a lot.
[02:18] Um, and with this, right? I we also very exciting announcement actually as of a few days ago. So, we were always known as Delta Sharing, and Delta Sharing was founded back in 2011. Um, so it's been a couple of years, but outside of, you know, the existing stuff that I
[02:34] just talked about, so supporting those four use cases, you can share uh, data assets, unstructured assets. Um, we have now officially also graduated to open sharing, and that enables for a couple of new capabilities. The first one is, this is
[02:50] really the open protocol that we uh, have also donated to the Linux Foundation because it's set for the agentic era. And this includes AI skills, um, you can also share AI models, and then you also are able to now share genie agents. Um, additionally, you can
[03:06] also put data controls in top of that. We actually just gave a talk today earlier about this. Um, it's probably recording somewhere if you guys want to hear more, uh, we can definitely get that. And then another thing is we also now support sharing to IRC clients, right? And
[03:21] that's kind of the point of why you guys are all here today is talking about how Iceberg is now officially a first-class citizen in open sharing. And everything, the kind of general thesis or consensus that we have now is the data format of whatever tables are sitting in really
[03:37] doesn't matter, and it shouldn't matter. Um, everything should just be able to work, no matter where you're sharing from, where you're sharing to, and what you're sharing. And the last thing that we also added is instead of, on top of just being able to share to and from the three major clouds, we also support a
[03:53] bunch of different on-prem sources as well. Cool. And this is These are just kind of um a flash of the logos of who has been as part of our like uh open sharing ecosystem, right? So, you see a lot of brands um across different industries.
[04:09] There might be big brand names you recognize. There might be smaller boutique names as well. Um but it's really this is a industry agnostic um protocol, right? And all of this is we're growing day by day. We want to increase kind of these like data sharing
[04:25] hubs that we're calling it um day by day as well. Cool. Now, let's talk a little bit more about Iceberg now. Um and I set the background a little bit, right? So, the lakehouse, what does that kind of encompass? Obviously, you have all sorts
[04:40] of data sitting on it. But we found that actually there's a pretty even split almost across Iceberg as well as Delta tables. Um and under the hood, right? Traditionally, they actually operate off of very different APIs.
[04:56] Their metadata are not compatible with one another. And oftentimes customer feel like, "Hey, I I feel like I'm locked into one of the formats, right? If I've already standardized on say Iceberg or on Delta, I feel like I'm I have to do that forever." Um and if I want to convert to one or the other, I
[05:13] have to make so many different data copies. I don't even want to think about migration. It's just so complicated. Um so, like we see that you're forced into this like false choice um or like a yeah, false trap basically. Um but we're super excited to announce
[05:30] and actually this is kind of a continuation um of an effort that has been over almost the past year already. Um we've really tried to invest in kind of this like open community open ecosystem um as well as open connectors. So, you can see like
[05:46] some of the brands here, like you might already have been using them in production. But on the left side, so I wanted to start with some of the foreign catalogs that we support. And what this means is that hey, if I have data sitting in Snowflake for example, and they're in Iceberg format, but maybe I'm running
[06:01] some AI analytics on Databricks. Like I have two different work streams. Both are fine, but like they're just for two different purposes, but I want to be able to centralize them, right? So now you can actually federate those foreign Iceberg tables into UC. It's all
[06:17] governed into one centralized source, so you know exactly who's accessing it, you know exactly what assets they are, and everything is tracked and logged as well through Unity Catalog. And then also on the other side, right? So using Open Sharing, you can now also share to
[06:33] all types of recipients. Whether this is the more traditional say are like Databricks to Databricks sharing. So of course you can share to another recipient on Databricks, but you can also share to the Iceberg RC ecosystem as well. Probably 95% of the
[06:49] use cases that we hear is like you're sharing to Snowflake, right? But also support like Trino, Flink, Spark, all that as long as they have like the RC basically endpoint implemented, that will work. And on top of that, so this is also just adding more to our connectors, you can also share to Open
[07:05] Delta clients. You might have say like BI tools, so like Power BI or Tableau that you want to share to, or if you want to share say to SAP or to Adobe or to Oracle, all of that is also supported. Great. Now you might be wondering like
[07:21] okay, I have a sense of what's happening now, but what is actually happening under the hood, right? Like what how are you getting this data? Like is it even live? Like what is happening? So let's just start from the first scenario. We were talking about sharing a foreign Iceberg table. So let's say it's sitting in
[07:40] yeah, like a Snowflake or a Hive metastore or like a AWS Glue metastore. That also works. You actually first register as an external connection. And I'll show the show a demo later as well, so you guys will see this in action. But you're basically just like pack
[07:57] making the path of hey, where is my metadata living? But there is no copies being made. Um And before like pre-open sharing, you're actually making a copy every single time, right? But today through cloud tokens, you actually get access
[08:13] to the source catalog that is external, that is foreign. And then you see Unity catalog vends these like short-lived read-only credentials so that it scopes specifically only to the tables that you want to read. And then you would be able to get access. Um and then metadata also refreshes in
[08:30] the background, so everything you get is updated. Great. And one little note, so ignore the public preview. Um we actually are GA. So that's something that we're super excited to announce as well.
[08:46] Great. Now moving on to the other direction, right? Let's say my data is living in Databricks. It's living in UC and I want to share it to somewhere else. I want to share it to an Iceberg um IRC client, for example. Then how does that happen? Um so through our D2O protocol or Databricks to open sharing, you actually
[09:02] create a credential file, which includes a couple of different things. Includes like the credential to which you authenticate against, so like the bearer token. Um and then it also includes like hey, which specific tables or which specific schemas do I want my recipient
[09:19] to have access to? And once again, these are short-lived tokens. Um and then the alternative, which is something that is brand new that we are launching this week, is OIDC authentication as well. Which means you can actually use your own recipient's IDP, so their identity provider, um to create like a short-lived scope token
[09:35] that is authenticated by Databricks as well. Um so a couple of options there for you to choose from. And all in all by the way, all of this on the recipient side is read only. So that's how you ensure that like this is actually a live non-copyable share. But of course, if you want to do any
[09:52] like life downstream like queries or if you want to do any manipulations or put it in a top pipeline, you can make a copy there as well. Cool. So yeah, so outside of Iceberg, right? Like you might be wondering, okay, is that kind of like a standalone product?
[10:09] Like I get that there's kind of this web of things. Are they all compatible with one another? And the answer is yes, right? A couple of things that I wanted to highlight is for view sharing, we hear from a lot of customers like actually sometimes I don't only want to share the entire table. Maybe I only
[10:25] want to share some subset, right? So that's compatible in both directions. And then we talked about OIDC just now. And then coming soon as well and and global distribution and secure connect we are announcing this week. Global distribution is we hear a lot about hey,
[10:40] maybe I'm doing like cross region sharing. But every time make that share, I have to pay egress. It's expensive, right? I have to say I'm sharing to five different recipients in a different region. I have to pay egress five times. But with global distribution, you actually replicate whatever you're
[10:55] sharing to that external region once and then from every then on out, you're not paying egress anymore. So this really makes that process a lot cheaper. And then similar with secure connect, we want to make sharing cross cloud, cross region a lot easier. And this is basically you set up your network
[11:11] configuration once and then that flows down, propagates to all the different recipients you might have. Great. So now that we've like talked a bunch about what open sharing supports and how it kind of works on a technical level,
[11:26] I want to hand it off to Balaji to actually ground this in an actual business use case at Foot Locker. Thank you, Tia. Um So, how many of you have uh shopped at Foot Locker? Okay. More than I expected.
[11:43] Okay. So, we we are a global sneaker athletic retailer. We operate 2,400 more stores across 20 plus countries. So, what that means, we have a lot of data to deal with. Right? Um So, before I hop into the architecture of the data platform landscape, I want
[12:00] to give you a a sneak preview of what we did, right? So, like like 7 years before, we were operating completely in on-premise cloud. So, we want to migrate to cloud. So, Databricks was our primary partner where we engaged and we decided, "Okay, let's use Databricks for our analytics analytical processing, right?"
[12:16] So, today Databricks is our primary AI, gen AI, and analytics engineering platform in which we run our workloads. Plus, they're also our open sharing hub that we leverage them. So, we do have So, if you look at the horizontal right? So, we are Azure workshop. So, the
[12:33] horizontal layer is our data lake where is where all the copies of data resides, a single layer in open format, right? So, we have two platforms, Databricks and Snowflake. They coexist with us. And we have And the consumption layer is Power BI where business users users use
[12:50] for dashboards and insights. We have around 5,000 plus users who use our dashboards and take decisions on it. On top of it, right? So, if Databricks is not just for analytics engineering, we also use it for sharing our data to enable business, right? So, coming to internal or external vendors,
[13:07] right? Our internal partners, we leverage them. Uh for example, a customer or identity resolution, right? Uh the pricing and promo optimization side of things, assortment allocation intelligence, AB testing, these are all not experiments. These are truly happening today in production, right?
[13:24] So, we adopted a strategy uh data strategy, right? Bring your platform where your data is and not the other way around. Don't bring your data to the platform. So, this is how we protect our investment and we stay flexible. Okay? So, I'll
[13:39] walk you through the next slide. What is the problem What was the problem we had before, right? Before I did that, right? How many of you in this crowd sharing the data through SFTP or some type of file transfer here? Okay, see?
[13:54] I see. And how many of you had a pipeline failure in in 2:00 in the morning where you have to fix it and ensure that okay, lots of engineering work to make sure that sync is happening so you don't have a problem when you share the data to the recipients.
[14:10] Okay? Still there, right? This is exact world where I was living in couple of years back, right? That's why I'm here to showcase what we can do better there, right? So, we had the same problem. If you could see, I had the data produced in Snowflake just for the purpose of
[14:25] sharing because we build two types of products, analytics as a product and data as a product. So, when data as a product, we need the data to be present where we can share to the external vendors, right? I have to extract the data, copy the data, the customer pricing or allocation some of the examples here. And I have to sync the pipeline to keep
[14:42] the data fresh. I have to export the data to the external vendor so they get the data from me. So, now I any transformation I apply when I move the data, I need to sync it back to the source so I have a same copy of data present in both the system. So, what this means, right? Is a It's not
[14:59] just engineering cost to it, right? There is a massive governance, right? I have a customer data sitting in two places so I need to tokenize them. I need to encrypt them, right? I need to have access control sitting on top of it. So, more problem, right? So, but if you look at this particular
[15:15] problem, what you see here, right? Each and every arrow points there is a job that can fail. Uh there is a QA overhead, right? Our governance is going to be missing, right? Too much work. Uh, but what is the intent, right? We want
[15:30] to be sharing the data to external, right? So, in the hindsight, we were discussing what what should be our data strategy should look like, right? So, we should have a consistent one copy of the data. We don't want two copies of the data just for sharing, right? And one copy of the data resides in one
[15:47] place. One consistent schema. But we can bring whoever wants to consume. Self-serve, zero touch. That was the strategy we adopted, right? Um,
[16:03] right? So, with that particular strategy, right? Uh, this is the fine This is the solution, right? This This is the architectural solution that is currently in production. This is not a road map slide. This is where we operate, right? So, here the left side the Snowflake is
[16:18] our uh, prime uh, is the source of platinum data exchange hub we create, from which we federate the data. Databricks also produces the data for AML use cases. They produce Both write the data into the same Gen2 data lake in open format, uh, in Iceberg format or
[16:34] uniform format, right? So, now the data is sitting in data lake. Now, if I want to share the data, I'm I'm federating the metadata into Databricks Unity Catalog. So, no here, no copies, no replication, no
[16:50] sync, no governance overhead, right? The same data is what you see in the source platform that is generating is what you could see in the Unity Catalog. Now, right?
[17:05] In Databricks, I'll be having the governed layer on top of it, right? It governs the data. From there, I enable intelligent enterprise use cases. Um, sharing to our CDP partner, sharing to our intelligent uh identity uh resolution provider. Uh so, everybody we
[17:20] use all the three protocols uh based upon where the provider is, what type of share we have to provide to them. All right? So, we do have prescriptive and ML analytics provider who we who consume data in the Iceberg native format. So, we share with our parent company Dicks with both ways. We share some data in Delta, some data in Iceberg. So,
[17:37] interoperable. All right? So, write once, federate the metadata, and govern from one single place. Right? So, one physical layer data layer, zero replication, there is no platform to platform that has to happen, right? Um so, I'm going to pass it on to
[17:55] Tia where she's going to show us how simple it is to set this up. Once you have that particular Platinum data governed, set it up the platform. Yeah, go ahead. I have one more slide.
[18:10] Great. Okay, so let's see this in action. Um and I know earlier we talked about there's like two kind of directions of sharing. Where did my mouse go? If I can find my mouse, there we go. Okay. Um So, let's skip over to this is
[18:26] Databricks and we are in Unity Catalog. And right now, let's say I want to bring in a foreign catalog, right? So, I have Iceberg tables living in say a AWS Glue catalog and I want to I want to actually be able to federate that into
[18:41] Databricks, then I can worry about sharing. But, I'm not even doing that step yet. So, let's say I click add, I want to create a catalog. And for demo purposes, let's just say Dice Iceberg. Foreign catalog, and then you can see here that you actually need to configure
[18:57] a UC connection under the hood, right? And I've done this already. Um but you do need like some credentials and you do need to do the setup so that um is something that I would recommend doing beforehand. And then, authorize path, I'll choose this, and then the
[19:13] storage location. Basically, what these two are is they store different metadata. One stores the Delta metadata, and one stores Iceberg metadata. Um and then I click create. And then what that happens, what happens here is you can actually see that I have created a foreign catalog that came from
[19:30] this AWS Glue connection. Right? And we can come in here and actually see that this is a foreign catalog, and it is in It should show up after a bit.
[19:45] You'll be able to see that it's Iceberg format as well. Um so, if we come, and let's say then I want to downstream be able to share this to let's say a Databricks recipient, right? We can just do a D 2 D share for now. Um come here.
[20:06] It's actually a Glue database, right? That's part of wherever I foreign uh federated from. And there's actually an employees table that was part of it. We can see what's in this table as well. There's some sample data if it if we want to share it if it would load, I mean.
[20:21] Um but here, let's just say I want to share it. So, I click share. Okay, great. I actually loaded it. So, you can see this is a foreign table. The data source is Iceberg. So, this is a nice foreign Iceberg table. And then then I want to actually share it. So, I click on share. I see, "Okay, great. I can share via open sharing." And I want
[20:38] to create a brand new share. So, I want to name this Iceberg share for Dice. And for my recipient, in this case, I've already um added a recipient, but we'll just set we'll just share with this uh sandbox environment, for example.
[20:55] So, I click share. And then I can even go and and look into like what is part of the share, right? And I can track, "Hey, if I want to revoke this at any time, I can." Because once again, this is a live share. So, once you revoke it, the recipient will lose access. And then you can also go in
[21:11] here and look at um you can even set say like row level restriction. So, like ABAC controls on that you can also put it um on all these tables that you're sharing to any recipients. Great. So, now let's go over to the recipient side. So, this other
[21:28] um you can see this is kind of where the recipient is sitting and let's see if I actually got that share. Um so, if I go back here and my provider will provide me with their uh sharing identifier which is also the same thing as your metastore ID. Um so, if I just go to that just to know
[21:51] if it's actually here. So, I click on um my provider to see oh, great. I actually have this iceberg share dice that we just shared to the recipient. So, I mount it to my catalog and I can just create let's say Tia catalog.
[22:08] It already exists. Great. And now I'll just mount it to an existing catalog. Or we'll just call it something else. Um then you actually have this foreign iceberg table living in a recipient's Databricks catalog that is now consumable and you can use this in any
[22:24] downstream use cases you might want, right? So, you can start writing a query on it. You can put it in a notebook. Um you can put the data and you can even say like put it in a genie space for example um and ask questions about it. So, at that point the data is yours and
[22:40] it's living in your environment now. So, that's the first direction, right? And then now let's switch to the second direction of maybe my data is already living in Unity Catalog and I actually share it want to share it out to say Snowflake, right? So, how does that work? Um so, let's
[22:55] jump back to the other environment that we have. And let's just look into what I want to share. So, in this case, I wanted to share a managed Delta table. So, very simple. All right, a managed Delta table, but I actually want to consume it as an Iceberg table. So, what happens
[23:12] here is Let's say I want to share this CRM data um with say like my Snowflake counterpart who who lives in Snowflake. So, you can see here it's a managed Delta table. Um I want to actually be able to share
[23:28] this. It once again this is probably something similar that you've seen now. Um I want to create a new share and this is CRM. And in this case, I actually want to share to an open recipient. So, you
[23:43] wouldn't choose a recipient from this drop-down list cuz these are only if you're a part of Databricks already. So, I'll choose a recipient later. I'm going to make the share first. And I can actually go to the share and see, "Okay, I have this table as part of the share, but I have no recipients yet, which is totally expected." So, now
[24:00] let's add a recipient. And I want to create a new recipient. And you can actually see, "Okay, this recipient, I want to make it an open recipient." And like what we talked about earlier, you have two kind of ways of authentication that you can choose from. You can either choose to do this
[24:16] through a bearer token or you can choose to do this um via an OIDC uh kind of credential. So, in this case, let's just go with the bearer token. And then I want the token to be valid for 30 days. And as the provider, you can also
[24:31] come back and rotate this token token whenever you want. So, it's truly very secure. And then I'm going to create and add the recipient. Great. So, I'm going to add the recipient to the share. And I actually see that, "Oh, okay, I have this activation link." Like what exactly is
[24:46] this activation link? And it leads to the credential file that I was talking about earlier. So, here what usually would happen is let's say I'm a provider um and I have this activation link now. So, I'm going to send it to my recipient. Uh Um, so let's just open a new tab.
[25:03] And let's paste it in. So, here you'll see this is our like kind of uh Delta open like client page. Um, and first let's go and download the credential file so that I have it in my local machine. And then I see here, okay, great. I actually do want to
[25:19] consume it um using the Snowflake open connector. So, I click on that and then I choose the credential file that I just downloaded. Uh, and then here I can see, okay, yeah, the share that I want to import is the CRM table. And then I can actually very easily
[25:35] point and click generate a SQL command that will get a get me um this exact table once I just run this command in Snowflake. So, if we want to look a little bit together what is part of this command, it's saying that, "Hey, I'm creating a catalog integration." All right, so I'm making a new catalog in
[25:53] Snowflake. And I'm calling it I can I can rename this to whatever I want. And the source is actually coming from this Iceberg rest um catalog with the table format as Iceberg. And then here is the I'm using vendor credentials, so the
[26:08] bearer token that is contained as part of this credential file to authenticate against the source catalog to make sure that, "Hey, I do have the right access. I do have the right tokens to actually be able to read this table." Um, and yeah, bearer token is here. And the refresh interval is like the metadata um
[26:26] refresh basically to make sure that like the data you are getting is fresh. And then creating a linked database, all that does is if you later on down the line decide, "Hey, actually as part of the schema that I've shared, I want to add five more tables." What you needed to do before is you need to go back and
[26:42] do this five times, but with the linked database everything should just link automatically. Um, cool. So, I've generated the SQL. I want to go ahead and copy it, actually. And now let's go to my Databricks environment
[26:57] and I just paste it in. Um and let's say I call this my iceberg recipient.
[27:15] catalog And then let's see if there's any other placeholders. Yeah, also the database name. So let's just call this dice.
[27:30] And then catalog we'll call it dice. Oh, yeah. It's the same thing. So, the same thing. Cool.
[27:56] Iceberg format and the credentials. I have the bearer tokens and I'm creating a database. But let's maybe pretend we're in an ideal environment where every single
[28:12] demo out there works. Theoretically, what will happen is that you'll see a preview of the iceberg table of the delta table that was living in Databricks show up here in Snowflake. Great. Let's just pretend that worked.
[28:28] That's the curse of the demos. Yes. All right, let me go back to the slide. And I forgot to mention one more thing also is you need to enable uniform, which that might be the reason, now that
[28:43] I think back, why that didn't work. So, what uniform does is it stands for Universal Format um and it's if I'm sharing a Delta table, it actually creates also an Iceberg metadata layer so that whatever like client I'm trying to like read it from, they'll be able to
[28:59] read the Iceberg metadata so that I can actually that's how I can basically process as an Iceberg table. Okay, back to biology. Okay, thanks thanks to you. Okay. This is where it gets real, right?
[29:16] Architecture looks great on whiteboard, what I showed you before, right? But what did we actually enable by doing that architecture, right? So, these are the six plus use cases that's active in production today for business, right? The first one
[29:32] is the CDP platform, right? It's a Delta to Databricks to Databricks sharing, it's a native format. We're able to share the customer data without copying the data, right? So, no PA, no governance overhead. They're able to consume, no business they're able to activate marketing orchestration,
[29:47] campaign activation, segmentation, what not, right? The entire CDP is powered through this share, right? The second is IDR resolution, right? So, we get cross channel, right? We have Omni, we have we have both retail and stores, we get different type of information clients
[30:04] customer information. So, we share the data to our leading identity to our leading identity resolution provider. So, now they resolve the identities and they give the data back to us. We integrate that back into CDP platform to power CDP, right? So, that's all that is a cross platform, so they accept in open
[30:21] data format. That is activated. So, third is so we our parent company Index, right? We exchange data in both ways, so we share data, they share data back back to us. So, this way we can power our own analytics within the organization, right? Both organization.
[30:36] So, the fourth one is a huge win, right? The that's the markdown optimization one. If you look at that, so that's a big win for us. So, we are sharing that uh our data to a leading prescriptive analytics vendor. They are sharing us back the
[30:51] real-time pricing information uh through which now we are integrating the pricing information back into stores and digital, right? So, previously we used to have the same price across all the stores. So, now we are pricing based upon the clusters of stores and by size. So, this is a huge win. So, this will
[31:07] this actually helps us lift our margin up. So, the that's twice buck sharing. So, next is allocation and size, right? Here again, this is the same uh type of share. We leverage Iceberg for it. Um so, here we share the size curve
[31:24] information. So, they give us back the DC quantities and placements at PO level. So, through which we'll be able to uh reduce the overstock and have and we have a better sell-through, right? Having the right uh at the right time, right place, at the right size. drives retail. This is how
[31:41] we enable it, right? Um So, now we are expanding this across multiple different strategic retail partners. We have the same pattern. Leveraging it, we are sharing this particular data uh completely governed. So, vendor is able to get our data back to them, right? So, so so we have six plus recipients.
[32:00] Three different protocols shared through one governed copy. No no copies of data. So, we share terabytes of data between our vendors. So, we started our uh data sharing journey in 2025. We started adapting Iceberg this year and it's going on well.
[32:15] So, by building this real use cases, right? What did we learn, right? I want to share that with you all, right? So, previously we used to duplicate the data. So, we are we no longer do that, right? So, now we need to measure the impact that we create to the org, right? So, first, the storage savings. Tera
[32:33] Tera terabytes of data is no longer have to be copied and stored and maintained, right? Then uh we retain the old extracts. So, we have a lot of extract jobs running in and keeping it sync. We don't need We no longer need to run the compute run the compute, right? Once again. So, now third, right? In terms of
[32:49] business impact, right? What did we do? The pricing impact, right? So, we are able to do real-time in-season price optimizations for the SKUs. And when it comes to allocation, right? We are able to improve our stock, reduce the overstocking a particular store with the same with quantities. And we are
[33:06] able to sell through the stock better. That's the real business impact we are able to create. So, before if you want to create an extract and share the information to external vendors or anybody, right? It takes weeks of engineering work, QA testing, deploying the pipelines, making it work, monitor it. We no longer no
[33:22] longer need to do that. It just Yeah, showed us it's a few click of a button the same data living in Amazon or Snowflake, right? Or the Glue Catalog is getting easily shared across. Right? So, that So, now we have a standard uh onboarding model that we can adapt and
[33:37] scale faster. Okay. So, if you are thinking about building uh this solution, something like this, right? I have a playbook for you, right? So, first, right? The first is the design phase. So, when you design,
[33:53] right? You need to make sure that you separate the data plane from the sharing plane, right? The The data plane and the sharing plane are two different things. Okay, data is originated in a different systems of truth, but we are able to share in the same governed way, right? Second, we need to identify your
[34:09] recipients. What type of share they are able to achieve, right? Are they Databricks? Are they going to be doing open share? Are they going to be Iceberg native, right? Identify the recipients. Uh because each of the steps are a little bit different when it comes to activating them, right? The you have to design it pre-made, right? So, use one
[34:25] storage layer. We use Azure because we are Azure workshop, but depend upon your cloud provider, I would your own storage layer. One storage layer, no duplication of data, right? So, define define your what workloads should run in Snowflake, what workloads should run in Databricks, right? Don't uh let it emerge
[34:43] organically because that is where we ended up having duplicates of data in both places. Uh that's something I learned. If somebody would have told me before, we would have not done that, right? So, now you have designed a solution. Now you have you're going to build it. So, uh when you build it, right? There are certain limitations at every, right? So,
[35:00] a limitation last month would not be a limitation this week, right? For example, I was not able to create views on foreign table, that's no longer a limitation. So, keep uh knowing what is your limitation because a partition files handles differently, um the CDF has its own quirks when you share it. Uh
[35:16] know the limitation. So, uh make sure that you validate the ADLS endpoint compatibility because the DFS versus blob endpoints uh will silently block you, right? I almost lost a week to it when setting this up, right? I want to share that with you all. So, now measure.
[35:33] So, you designed it, you have built the solution, now you need to really measure it. How do you measure? Uh before you implement, take a baseline. What is the storage size? How many terabytes of data? How many pipelines you're going to retire? So, that will give you the cost of pipeline. What is the cost that we are going to optimize, right? And
[35:51] the second, measure the impact in terms of business value, okay? Because because leadership would like to hear what is what did we enable, right? Not about how many tables I reduced in the end of the day. That's not what they want to hear, right? Measure in terms of business, right? And report your business-facing ROI metrics
[36:06] in their language. Speak the same language what your uh leadership talks to us, right? So, what is a net takeaway, right? So, we started with SFTP exports. We had duplicate copies. We had data sitting everywhere. Governance is a nightmare. We had terabytes of data. So,
[36:23] the the the componenting of the problem at that scale. So, now uh we have six plus recipients, three hubs, right? And we are able to share the data seamlessly because of the right strategy, only because we adapted open format and open sharing.
[36:39] Thank you. We'll take any questions.

Learn more about the Databricks Data and AI platform.

The information provided herein is for general informational purposes only and may not reflect the most current product capabilities or configurations.