Skip to main content

Delta Sharing SecureConnect: Simplified Networking for Private Data Collaboration

Summary

  • Delta Sharing, rebranded as Open Sharing to reflect its expansion beyond structured tables to volumes, AI models, and notebooks, grew over 100% in the last year and launched several major new capabilities at the Data and AI Summit.
  • SecureConnect addresses the enterprise networking burden of data sharing by introducing a Databricks-managed network proxy that eliminates per-recipient IP allowlists and direct storage exposure, replacing complex per-recipient configuration with a one-time setup.
  • Amadeus, a global travel technology company, uses SecureConnect and Open Sharing for real-time data collaboration at scale, demonstrating how private connectivity via network endpoints keeps storage behind firewalls while enabling zero-copy sharing across AWS, Azure, and GCP.

Delta Sharing SecureConnect: Simplified Networking for Private Data Collaboration

Watch: Delta Sharing SecureConnect: Simplified Networking for Private Data Collaboration
Scaling data sharing across multiple recipients creates operational burden. Managing per-recipient IP allowlists, exposing storage to the public internet, and coordinating with network teams can stall deployment. For enterprises with multi-cloud footprints, the complexity compounds. Delta Sharing SecureConnect addresses this by introducing a Databricks-managed network proxy that eliminates the need for direct storage access.
Learn how SecureConnect simplifies networking configuration with one-time setup instead of per-recipient changes. Discover how private connectivity via network endpoints keeps storage behind firewalls while enabling zero-copy sharing across AWS, Azure, and GCP. See how Amadeus implemented this for real-time data collaboration at scale. Understand the setup flows, access patterns, and roadmap for cross-region and cross-cloud deployments.
🤝

Chapters

FAQs

What is Databricks Open Sharing and how does it differ from Delta Sharing?

Open Sharing is the evolved form of Delta Sharing, rebranded to reflect its expansion beyond sharing Delta tables to include volumes for unstructured data, AI models, and notebooks. Delta Sharing was launched approximately five years ago and grew over 100% in the last year, with Open Sharing representing the next phase of the product's vision to allow sharing of any asset with anyone anywhere.

What problem does SecureConnect solve for enterprise data sharing?

SecureConnect solves the networking burden of managing per-recipient IP allowlists, which requires coordination with network teams and exposes storage endpoints each time a new sharing recipient is added. By introducing a Databricks-managed network proxy, SecureConnect reduces this to a one-time setup that keeps storage behind firewalls for all current and future recipients.

How does Amadeus use Open Sharing and SecureConnect?

Amadeus, a global travel technology company with a large-scale big data platform, uses Open Sharing and SecureConnect for real-time data collaboration across its operations. Lead architect Pedro Pova describes how the platform's data strategy and sharing principles around security and scalability align with the private connectivity model that SecureConnect provides.

Does SecureConnect support multi-cloud data sharing?

Yes, SecureConnect is designed to support private connectivity via network endpoints across AWS, Azure, and GCP, and the roadmap includes cross-region and cross-cloud deployments. This video covers setup flows, access patterns, and the planned expansion of SecureConnect's multi-cloud capabilities.

Full transcript

[00:07] Um, hello everyone. Thank you for being here. Um, just a quick round of introduction. My name is Huey. I'm the product manager for uh, Open Sharing, formerly Delta Sharing, and I'll let you guys introduce yourself. Hello. Thank you for being here today.
[00:23] My name is Pedro Pova. I work in Amadeos. I'm based in Spain. I'm an lead architect of the Amadeos big data platform. And uh hello everybody. My name is Eli. I'm specialist solutions architect at databicks. I'm based in Paris. Uh yeah,
[00:39] and happy to be with you all today. Awesome. Yeah, thank you guys for being here on the last day. And I see uh Bruno, we just met yesterday at the at the exact meeting. Um okay, so thank you guys for being here. Um
[00:56] so quick overview and agenda of what we're gonna cover uh today. Um we're gonna first talk about sort of a high level overview of open sharing and and delta sharing and then we're going to talk about the problem that a lot of customer run into when it comes to
[01:12] networking and then we'll talk about this new feature uh we actually launched um two three days ago called secure connect uh intend to uh solve the problem and then I'll pass it on to Pedro to talk about how Amad is using the features and then Algalli will do a
[01:27] demo and we'll do some Q&As's. Awesome. So let's talk about an overview of uh open sharing. Um as many of you guys know uh we launched uh delta sharing about uh five years ago and just curious
[01:45] like how many of you guys are using delta sharing today. Okay, so most of you guys and um and in Delta sharing you know we've seen tremendous amount of growth over the last five years um many partners and you know in the last year we grow over 100%. Um and we built a
[02:04] large ecosystem with delta sharing but over time we also added a lot of new capabilities with delta sharing um that went beyond the original intent of sharing delta tables. For example, we added volumes uh so ability to share unstructured data but also AI models,
[02:20] notebooks to a point where we think it's now a new a time to introduce a new projects um which is why we introduced open sharing uh that sort of uh I think reflects better what we're trying to do uh which is allow you to share any
[02:37] assets um with anyone uh anywhere uh that's not just sort of structured table sharing. And for you specifically, uh we're launching uh quite a few things at this data AI summit. So just want to call
[02:52] that out and there's sort of three pillars of features that you know uh the new things that we're building. The first is format interoperabilities specifically with iceberg. So uh two things uh either you know GA already or
[03:08] being GA very soon is the ability to share to every uh open iceberg client that speak IRC. So that's one. The second is the ability to share uh foreign iceberg table. So these are iceberg table managed not by unity catalog
[03:24] um but uh uh let's say external catalog like glue um that you can federate into unity catalog and be able to share. So that's the two sort of features for for iceberg in ter of operability. The second pillar is AI. Uh and for that
[03:40] um we are also launching in beta the ability to share a genie agent formerly known as Genie space. And what that means now is that you can as a data provider offer your recipients a agentic experience on your on your data. But
[03:58] that sounds a little high level. What does that mean? So previously when you share a data sets many of you probably share fairly complex data sets with you know a lot of columns pretty complex schema some nuances with your tables and what the recipient when they get this table what they need to do is that they
[04:14] need to do like select star limit five try to explore do some describe try to understand it uh your schema and it takes a while to get used to your table and be able to get value out of it with genie sharing uh what you can do is that you can offer a curated genie agent or genie case that have instruction on how
[04:31] your table works to AI to genie and then when recipient get that genie agent they can just query it directly without first having to learn the nuances with your table and sort of how your table works. So that's the second pillar which is genie sharing.
[04:47] The third is crosscloud sharing which is what we're actually going to cover uh today. And for that we're launching sort of two features. Uh one is called secure connect meant to address networking challenges that a lot of customer run into which we'll we'll dive into and
[05:03] then second is called uh global distribution. It should be global distribution instead of replication. Um but what it does is that a replicase of data uh from provider region to recipient regions automatically managed by data bricks to minimize uh query egress costs and also um increase sort
[05:21] of query performance for um uh uh for recipients because it's local uh in region queries. So just want to give you guys a sense of on a high level what we're launching at this data AI summit and our customers already love open
[05:37] sharing and we have you know leading AI companies like open AAI um but also traditional enterprises like uh SAP. So let's talk about networking. Most of you guys today here you know
[05:54] already use open sharing delta sharing. So this is probably already familiar with you. Um but just to rehash um one of the key uh selling point and key benefit of open sharing is this sort of zero copy sharing approach where you know you don't need to duplicate your
[06:09] data and recipient access storage directly and then the way it works is that you know data providers uh you know you have your tables in your storage whether S3 ADLS and uh and and these are either managed
[06:24] table or foreign table uh in unity catalog and data Data provider will create a share, add a recipient to the share. The recipient can either be data bricks on data bricks, what we call data bricks to data bricks sharing or they can be off data bricks called data bricks to open sharing for data bricks
[06:40] to open sharing. Uh you can create a credential file to distribute to them so that you know wherever they are they can query the uh the shared data from a recipient point of view. Uh if they're on data bricks they will mount the sharing unity catalog once they receive it. If they're open and they're not on
[06:56] data bricks, they can use a credential file you give them uh to query uh the share and most importantly what happens at the request when they start querying the uh the shared data. What happens is that the client would first talk to the delta sharing or open sharing server
[07:12] saying hey here's who I am can I get access to the shared data the server will you know look up permissions try to figure out if they can get access if they do we then direct storage credential so these are S3 ADS credentials presign URL or uh SD SCS
[07:27] token SS token then the client then access the storage directly and there's a lot of benefits of this approach of bring your own storage. Um so minimal cost since you have one copy of storage uh to manage operational
[07:43] efficiency one piece of storage to manage also and then single source of truths um because one piece of storage to manage so they always recipient get the freshest and and latest data.
[07:58] But from a networking perspective, this zero copy um you know bringing your own storage approach can also introduce problem and specifically for a lot of enterprises you know your resources typically sit behind firewalls. So if you're on ADLS it's probably your ADLS
[08:14] is not open to the internet. It's behind some sort of storage firewalls. Your compute uh is also probably sitting behind some sort of VPC that have pretty strict egress and ingress control. Um so what that means is that in this paradigm of recipient needs to access the storage
[08:32] directly. Uh these firewalls both on the storage side and on the compute side need to be opened up and create some sort of uh you know punching a hole so to speak so that they can talk to each other. And for a lot of you guys where you know you're sharing maybe between
[08:48] different business unit within a company or sharing externally with different uh uh recipients punching this hole can introduce a lot of operational overhead. Um because you know this you know the usually what we see is that there's a data team that's responsible for creating the share and
[09:05] sharing the data and curating the data. Then you have a platform team that's responsible for managing the networking policies and and storage firewalls. And so you have those two teams on your side and then you have two other team on the recipient sides. So all those team need to talk to start passing you know IPs
[09:22] passing storage endpoints emailing and slacking to make sure this get to get this thing working. So this introduced a lot of sort of operational overhead for a share uh to work and this gets worse if you do let's say crosscloud sharing where you doing from S3 to you know
[09:39] Azure um where you know things can behave a little differently and things might not work super out of the box. So what are we building to help solve this networking problem?
[09:57] Open sh so we're we recently launched this feature called open sharing uh not open sharing secure connect uh which is meant to simplify networking and and really does a few things. uh first is that uh the the the crux of this feature is I introduce a data bricks manage storage proxy that sit in front of your
[10:13] storage so that you don't punch this one to one uh hole so to speak with your recipient uh instead you on board to this proxy once and from your perspective this proxy actually just sit in data bricks u serverless data plane
[10:29] so if you use serverless today and granted data bricks serverless access to your storage you're kind of done you don't need to do any anything extra from a networking perspective. You enable this one toggle uh in the UI or via API which Algi will show you guys a little later. Uh and then from a provider point
[10:47] point of view, you're kind of done. Any new recipient you on board, they would just talk to this proxy. So you no longer need to do these sort of one by one networking configurations. Um and from a recipient point of view, this also gets easier. um instead of
[11:02] getting your storage you know asking you guys like a provider to get the storage endpoint and you know allow listed in their VPC from their point of view secure connect is also just a piece of uh data data bricks uh infrastructure so they only need to allow list if it if
[11:18] they're on serverless there's nothing they need to do uh it would just work out of the box because we manage the networking behind the scene if they're using classic cluster or if they're open recipient they just need to allow this data bricks IP specifically ally data brick control plane ips uh and those can
[11:34] be self-s served they could just go to data bricks documentation follow existing documentation so it's also a lot simpler for a recipient as well last but not least is private connectivity uh which is um you know very important for a lot of uh enterprise customer you want
[11:50] your storage to sit behind uh a private link so for that we also offer uh private connectivity via NCC or network connectivity configurations which offer from this proxy to your storage. Um, uh, this private link connectivity um and,
[12:08] um, and a lot of enterprise customers are, uh, very interested in this as well, especially if you're operating on Azure and your storage behind uh, behind uh, uh, uh, Azure Private Link today.
[12:25] So, how does it work? Specifically, the way that you set up is and and we've covered this already a little bit on the last slide. And so there's a setup flow, there's access flow. The setup flow is provider enable this one time. You toggle this thing on in your UI. And then if you already uh and and then you just need to allow sort of data bricks
[12:42] uh serverless data plane access to your storage. If you use serverless today, you probably already done that. If you haven't then haven't, you can follow the data bricks documentation to do that. And then from a recipient point of view, if you're on serverless, nothing needs to do. If they're on classic clusters or open recipients where they manage the the networking, they just uh follow data
[13:00] bricks documentation to allow list the IPs during access time. Uh what happens is that client same as before they will request uh access to a shared access by talking to the open sharing server. The server will authenticate and authorize
[13:15] and it will give back the the storage credential. The difference here now the storage credential is not a direct storage credential. So not S3 or ADLS credentials instead is what we call it's a data bricks specific storage credential what we call datab bricks
[13:31] token or data bricks uh urls and then the client will then use those credentials to talk to uh this proxy uh secure connect secure connect then will then fetch the bites from the uh from your storage over private link if you have that configured and then returns
[13:48] the bites to the clients. So the difference here really is sort of three step three and four in the access flow. Now with this um you kind of get the best of both world where you still have the zero copy sharing uh approach where it's minimal cost one copy of your data
[14:07] operational efficiency one piece of infrastructure to manage single source of truth and the data is always fresh but then with this uh secure connect this sort of storage proxy you also get simplified networking and private connectivity uh even in cross cloud.
[14:28] So one last thing before I pass on to Pedro is the road map ahead. So this feature right now is in public preview. Um and so you know you guys can try it out yourself. Um and uh in the future what we want to do is sort of spend two pillars. The first is expand asset
[14:45] coverage. Uh today um the assets that we cover is uh structure and tabular data. uh we want to support every other assets that uh you know open sharing supports which means volumes notebooks AI models and probably you know future AI assets
[15:02] as well and the second pillar is um because now we manage the networking so from serverless for example to this proxy is we manage networking we want to make that even better so what does that mean um our traffic team actually has a lot of exciting initiative to build
[15:18] cross region crosscloud data bricks managed backbone um that u you know traffic is completely off internet and and we manage that and um so uh if that whenever that's ready we'll start migrating off this traffic
[15:34] onto the new backbone and what that means for you is that the traffic will be more private and it maybe more performant as well and lower cost and for you also the the thing is that because we manage networking whenever these are ready we could just manage it
[15:50] we could just migrate it on our and there's nothing uh you need to do and I think just get better. So with that said, I'll pass it to Pedro to talk about Amadeus. Thank you. Way. So let me start by introducing Amadeos
[16:06] for the ones that are not familiar with us. Amadeos is a technology company dedicated to the travel industry that operates at a scale. So we are present in more than 190 markets. And to share with you some numbers, last year our
[16:21] reservation system process more than 485 million bookings and more than uh 2,200 passengers boarded, okay, by by the airlines.
[16:38] Our solutions are designed around the traveler journey. So from the inspiration phase through the shopping booking and all the trip stages pre during and postrip okay understanding travel behavior is actually one of our
[16:53] main priorities. We do business at a scale and we operate technology at a scale too. We have a global multiloud footprint with present in more than 15 cloud regions where we are running more than 200 Kubernetes
[17:10] clusters. At peak times our systems manage to process around 150,000 transactions per second. Okay, that's an scale really comparable with the Google search itself
[17:26] and that reaching 99.95% of availability and actually Madeos has implemented one of the largest serviceoriented architecture in the industry with more than 600 applications and more than 10,000 microservices
[17:41] running there. And to illustrate that with an example that I'm pretty sure that you are all familiar with. Let's see the fly search puzzle. So every day amo system receives around three billion fly search requests
[17:58] that are exploring billions of combinations as well to return from 50 to 1,000 offers as average. And that's done thank you to the uh machine learning. We are running around two billion machine learning
[18:14] inferences every day that are able to provide our customers an efficiency gain between 20 50%. And that's how we make the travel ecosystem better through five main
[18:29] levers. Personalization, providing the right content at the right time. Analytics, turning data into actionable insights. This is an area where data bricks is actively helping us business automation
[18:46] events reacted on real time and integration between the different uh actors of the of the industry and that vision it's implemented in our open platform which is a set of travel
[19:03] capabilities deployed on the public cloud that rely on trusted data that is allowing us to switch from a product oriented company to a solutions oriented company.
[19:23] To operate that open platform, we need a standards and one of the standards is the open data. The open data it's a common data dictionary an interoperability standard that we are introducing by design in the platform which allows all the different
[19:40] solutions to talk the same language. It is governed in the data mess. We have implemented a data mesh architecture where we are enforcing go the the governance of at the different domains
[19:58] locally and with this federated governance approach each domain is ensuring data confidentiality and privacy on the on their side. Today this data mess is composed by 10
[20:14] different data domains across the entire company. We have more than 290 application work spaces. It means project teams working and collaborating in our data platform with our data mess.
[20:31] And we have today 38 pabytes of data stored in this data mess which is divided into 1,700 data sets instantiated in around 36,000 containers in the cloud
[20:52] and how datab bricks is helping us to unlock value from this data mess with our data platform. Of course, datab bricks is actively helping us on the data processing side. On the data engineering, our teams are actively using data bricks to extract
[21:09] actionable insights from this row open data. But we want to go beyond. We want to enable our open platform vision by being able to connect with our customers, to connect our data with them.
[21:26] And for that we have two guiding principles. Security and scalability. Security because we have we want to be able to share data in a secure manner. We know
[21:43] that the data is secure at rest. The data is protected following our data governance framework. But when it comes to exposure, we want to be able to expose the data without exposing the underlying infrastructure as we was presenting before. And of course, scalability because at
[21:59] the at the volumes and at the scale that uh that we that we operate, we must ensure that this deployments will be efficient and automated. Of course, there was historically a way
[22:15] to do so in the traditional BI era. That was the batch file exports. Okay, I'm pretty sure that uh you all experienced that. Uh you extract the data into flat file, you export to a customer storage.
[22:33] That was okay in the past and that's okay for some use cases, but uh that's not valid anymore. Why? Because the data is refreshed and the insights are available on real time in our data
[22:48] platform and customers wants to access this data in real time as well. So we want to activate this sharing with zero copy. We want an in place data sharing. So how can we enable this near real time
[23:05] or real time data sharing in place with zero copy with full security and in a scalable manner. Well, this is where secure connect comes into the picture. And to summarize,
[23:24] Amados is operating at a scale business and technology. The open platform vision allows us to unlock new opportunities thanks to data. And for that we need enablers that allow
[23:39] us to activate this data sharing in a secure and a scalable manner. And now I pass the floor to GI who will demo how it is working. Thank you. Thank you Pedro. So now we will go through directly the demo and uh
[23:55] for that we will do directly also a live demo. So I don't know who's watching right now the world cup uh but you might know right now that you have lot of people coming passing through the airport taking flights etc etc. So today
[24:12] we will imagine that an airport is trying also to share data to other airlines multiple airlines for example San Francisco airports. Uh so basically they get operational data they want to share that to make sure that they can
[24:29] predict for example the price of the tickets or predict the customer journey and make it personalizable. So in that case uh let's assume that here I have my workspace as a data provider. In that case uh I'm able to access my workspace
[24:49] and basically just have a look at this table. So in that case I have a table here called flights uh and I'll be able just to get some uh information related to my flights. This uh data is on Asia.
[25:05] So in that case uh it means that the underlying table is stored in a storage account and this storage account we will assume that this storage account is basically private. So in that case what I mean by private it means that it has no public network access um and
[25:21] basically no one can access that uh without having setting up a private endpoint. So if I see here for example in my table flight I go if I go into details I will see that this is the storage account underlying storage
[25:37] account in which uh I store my own data this storage account is private and in that case I can access it as data provider because I have in place private endpoint but what what's happening is for example I want to do sharing and in
[25:54] that case for example one of my uh airline that I want to share data to is on AWS. First option is to basically uh get the address IP of this uh I mean uh um airline uh from its own compute and
[26:12] then white list IPs into my storage account. The issue is that it's super super uh difficult and gets some frictions once I have to manage uh hundreds or thousands of airlines at the
[26:27] same time and need to whitelist all these IPs even more if the this IP changes. So in that case what I will do is that I can enable secure connect. So uh when I enable secure connect
[26:42] basically what I have to do is that I need to go into the metas store of my workspace as data provider. So in that case my metas store is called world cup flight operations. I go into this meta store and in that case whenever I enable
[27:00] a open sharing uh uh I mean with third party organizations in that case I get the ability to specify here a network configuration connectivity. This network configuration connectivity is basically
[27:18] the network that is associated to a serverless compute that will be used by the manage proxy. So basically secure connect in order to access the data and in that case what I need to do is just to go into this network configuration
[27:33] connectivity create some private endpoints associated for to it for example into my own storage. So in that case the storage is the underlying storage of my data
[27:49] and that's it for the networking part I have to do do this once and then second part is that once I want to share this data what I need to do is that I need to create a delta share open share in that
[28:04] case so I create a share specify my recipient my recipient here for example is a workspace that is based in EU West one region uh and just share my data. In that case
[28:21] if I do not enable open sharing here in that case the compute that I will use will not be able to access my underlying table. So if I go into table code flight
[28:36] start my compute in that case it will take a little bit of time it will try to fetch the underlying data but after 30 seconds 40 seconds it will time out because the underlying compute of my data recipients won't be able to access the data unless I have a private
[28:53] endpoint associated to it. So let it just launch. In that case here I get this error called this storage account firewall is not authorized because my IP are not authorized. So in that case just go back into my recipient
[29:11] again as data provider. Enable secure connect. I confirm that the recipient is allowed to use secure connect that I have allowed secure connect to access my storage because I have set up the underlying configuration and then if I
[29:28] enable that go back into the recipient level refresh the page then I'll be able to access the underlying table associated to it without any friction. So in that case at the recipient level I don't have to involve any devops any network
[29:45] security uh people because in that case everything's happen at the provider level. The provider needs to set up once the configuration for secure connect to basically access the data and then the recipient will be able to access the
[30:01] underlying table without having to worry about the networking part. So again open sharing allow you to access the Android table in real time securely and without friction. This is what we
[30:16] wanted to basically prove you. Cool. Cool. Awesome.

Learn more about the Databricks Data and AI platform.

The information provided herein is for general informational purposes only and may not reflect the most current product capabilities or configurations.