Skip to main content

OpenSharing SecureConnect: Multi-Cloud Data Sharing Without Network Complexity

Summary

  • OpenSharing SecureConnect is a new Databricks-managed network proxy that enables secure cross-cloud data sharing across AWS, Azure, and GCP without requiring providers to expose storage to public networks or maintain individual IP allowlists for each recipient.
  • Amadeus implemented SecureConnect at enterprise scale to simplify network security for their data mesh platform, achieving secure data sharing without the operational burden previously required for each new sharing relationship.
  • SecureConnect supports private link connectivity for truly private cross-cloud sharing, and extends OpenSharing (formerly Delta Sharing), which has grown to 300+ partners and 100%+ year-over-year growth over five years.

OpenSharing SecureConnect: Multi-Cloud Data Sharing Without Network Complexity

Watch: OpenSharing SecureConnect: Multi-Cloud Data Sharing Without Network Complexity
Enterprise data sharing often requires complex network configurations to maintain security. Providers typically must manage individual IP allowlists for each recipient and expose storage to public networks, creating operational burden and security risk. OpenSharing SecureConnect introduces a Databricks-managed network proxy to simplify this problem, enabling secure multi-cloud data sharing without requiring direct storage exposure.
this video explains SecureConnect architecture, covering how it routes data across AWS, Azure, and GCP through a managed intermediary. You'll learn the setup flow for providers, how recipients access shared data through the proxy, and how private link connectivity enables truly private cross-cloud sharing. Amadeus shares their enterprise implementation at scale, demonstrating how zero-copy sharing patterns with SecureConnect deliver both security and operational simplicity.
🤝

Chapters

FAQs

What is Databricks OpenSharing SecureConnect?

SecureConnect is a Databricks-managed network proxy that routes data sharing traffic between providers and recipients across AWS, Azure, and GCP without requiring direct storage exposure. It was launched as a new capability within OpenSharing, the renamed and expanded version of Delta Sharing.

What network problems does SecureConnect solve for data providers?

Without SecureConnect, data providers must expose storage endpoints to public networks and manage individual IP allowlists for each recipient organization, creating operational overhead and security risk. SecureConnect routes traffic through a Databricks-managed intermediary, eliminating both requirements.

How has OpenSharing grown since the original Delta Sharing launch?

Delta Sharing was launched approximately five years before this video as the first open-source, vendor-neutral enterprise data sharing protocol, and has since grown to more than 300 partners with over 100% year-over-year growth. Over time the product expanded to include volume sharing for unstructured data, AI model sharing, and notebook sharing, leading to the rename from Delta Sharing to OpenSharing.

How does Amadeus use OpenSharing SecureConnect?

Amadeus is a large-scale travel technology company that implemented SecureConnect within their data mesh platform to enable secure data sharing with partners and customers at enterprise scale. Their use case demonstrated that zero-copy sharing patterns with SecureConnect deliver both the security guarantees and the operational simplicity needed at scale.

Full transcript

[00:08] Thank you and welcome to the second breakout of your D&I Summit journey. This session is about open sharing Secure Connect simplifying the network configuration for external data collaboration. And with us we have a couple Databricks folks. So Elgally and Huey from
[00:26] Databricks and then we also have Pedro from Amadeus to speak to us on this topic so I'll turn over to them. Uh can you guys hear me? Okay, perfect. Uh just quick introduction. I'm Huey. Uh I'm the product manager for uh Delta Sharing now
[00:42] known as open sharing and I'll quickly let you guys also introduce yourself. Yeah. Uh nice to meet you all. Elgally. I'm specialist solutions architect at Databricks. I'm based in Paris and happy to be with you today. And hello. Thank you for being here. I am Pedro Poveda. I'm lead architect on
[00:58] the Amadeus big data platform and I'm based in Madrid. Awesome. Well, thank you guys for coming and thanks to you for your time. So let's uh dive in.
[01:14] Just a quick sort of a overview of what we're going to cover today. Uh first we're going to give a overview of you know open sharing formerly known as Delta Sharing. Uh then we're going to talk about uh network complexity sort of the problem we're trying to address. Uh then we're going to talk about this new feature called Secure Connect that
[01:30] we actually you know recently launched yesterday uh to simplify this networking problem you're running into today. And then I'll pass it to Amadeus to talk about how uh Amadeus is using Secure Connect and how it's solving the networking security challenge. And then Elgally will wrap up with a demo and
[01:47] then we'll give you guys some time for a Q&A if we have time. So let's do an overview of open sharing and Delta Sharing first. Um as many of you guys might have know, we launched uh Delta Sharing about 5 years ago at also at Data AI Summit as
[02:05] the first sort of open source vendor neutral protocol for enterprise data sharing and collaborations. And in that 5 years, we experienced a lot of growth. Uh we now have, you know, 300 plus partners, uh over 100% year-over-year growth, uh and a wide ecosystem of partners and uh data
[02:22] providers and clients that are using Delta Sharing. But over time, uh we also start adding a lot more capabilities to Delta Sharing. Uh we introduced volume, for example, um for unstructured data sharing. We introduced uh AI notebooks, uh AI
[02:39] models, and notebook sharing. Uh so, over time, sort of the name Delta Sharing uh kind of just doesn't do us justice anymore. And that's why we last week introduced this new open source project called Open Sharing, uh which is a new uh Linux Foundation
[02:55] projects, um sort of build on top of Delta Sharing and evolving it further for this new agentic era that that we're in right now. But what does that mean for you specifically? For Open Sharing on Databricks, uh we're
[03:10] introducing a suite of capabilities um this few days at Data AI Summit. The first is sort of Iceberg interoperability. As you might have guessed, part of the reason we renamed uh Delta Sharing to Open Sharing is because we have extensive support for Iceberg. Um so, at this Data AI Summit,
[03:28] we're sort of announcing two uh capability for Iceberg. One is uh be able to share to open Iceberg clients that speak IRC. And the second is uh be able to share uh foreign Iceberg tables that's not uh managed by Unity Catalog, but managed by
[03:45] external uh Iceberg Catalog, let's say uh Glue or Snowflake. The second capability we're also introducing uh is a Genie sharing. So now you can share Genie agent, formerly
[04:00] known as Genie space, and offer your collaborators a sort of agentic experience on your data. And as part of that, we're also going to introduce very soon the ability to add data control as part of this sort of Genie sharing. So
[04:16] you can allow you to give control like I want my recipient only do 10-20 queries per day or I want my recipients to only do export 200 rows. And and this allows what we think is very
[04:31] exciting cuz it allows sort of unlocks new business model for data providers. And last but not least, which is the focus of this topic, is multi-cloud and cross-region cross-cloud collaboration. So two feature we're launching, one is
[04:47] secure connect, which we'll focus on today, to simplify networking configuration, and the second is is called global distributions, which minimize egress and and improve query performance by replicating the data from provider
[05:02] region to recipient regions. And there's a lot of session on these individual topic. Just want to give a quick overview of what we're launching and if you guys want to dive into any of these feature more, go to those sessions. And our customers already love open sharing since we launched. We have, you
[05:17] know, leading AI company like OpenAI using open sharing, but also traditional enterprises like SAP. So let's dive into the networking and and network complexity.
[05:34] So before we dive into the specific networking challenges, let's just do a quick refresh of how open sharing and how Delta sharing works today. And I just out of curiosity, how many of you guys are using Delta sharing today or uh so a good chunk. Um so as you many of
[05:50] you uh might already know, uh Delta Sharing or Open Sharing uh has this sort of a zero copy sharing model that allows you to bring your own storage. And the way that you work is I I, you know, you register uh your uh tables uh or your
[06:05] storage with Unity Catalog. Then as data provider, you can create a share. Uh you add a recipient to a share. Uh and then um and then sort of that's it. That you're done. And then uh part of the secret uh sauce where the the the perks of Open Sharing is that it allows you to
[06:21] share to uh recipient not on Databricks. So what we call Databricks to Open Sharing. So as part of that, you can also create a uh credential file that you can distribute to this um recipient not on Databricks and they can use that to access your shared data as well.
[06:37] And for data recipient, if they're on Databricks, once you get this share, they can quote-unquote mount the share in their Unity Catalog, which creates this um assets that they can query against. Uh and then they could just query. And for open recipients, they can use that credential file we share uh to
[06:53] access the shared data. And during uh this share process, what happened behind the scenes is that the client will first talk to the Open Sharing server uh or Delta Sharing server to say, "Hey, here's who I am. Can I get access to this uh shared data assets?"
[07:08] If they have the right permission, uh the the the server will actually give them direct uh storage credentials for the client to access the storage directly. As opposed to uh you know, a replica of this uh of this data. So that's where
[07:24] sort of the zero copy sharing come into picture. And there's a lot of benefits of this sort of zero copy uh direct storage access model. First is minimal cost. So you only keep one copy of your uh shared data. You don't need to upload it to another uh place and have a second
[07:42] copy. And the flip side of that is also operational efficiency, which means that if you have one copy to maintain, um, you have left less infrastructure overhead, less storage to manage, uh, less pipelines. And last but not least is you have single source of truth. So, recipient
[07:58] always get the freshest data cuz you have sort of one copy, um, and there's no sort of lag or or freshness issues with your data. But, uh, the sort of with the zero copy model, um, we also
[08:13] see a lot of customer running into networking challenge. What does that mean? So, for a lot of enterprises, um, your storage or your resources typically don't just don't live in the open. Uh, typically is behind sort of enterprise firewalls. Uh, for your storage might be behind firewalls or your compute
[08:30] resources is behind VPC. What does that mean for sharing and this sort of zero copy direct storage access model is that the provider and the recipient need to punch holes, uh, so to speak, so that the recipient can get access, uh, to the storage.
[08:45] And for a lot of large enterprises, this means many sort of teams and stakeholders are in involved. You typically have a data product, uh, team that's responsible for curating and producing the data. You have a storage admin that's responsible for managing
[09:00] the storage firewall. You have a networking admin that, uh, responsible for VPC policies. And and provider has three of those stakeholders and recipient also have three of those stakeholders. So, all of these people need to sort of coordinate and, uh, uh, pass information along,
[09:17] update policies, uh, before this sort of plumbing, uh, can work correctly. So, that's the problem we're set out to solve. And how does the feature secure connect, uh, address this problem for you?
[09:33] On a high level, what secure connect does is that it introduce a Databricks managed proxy. So, instead of this direct storage access where the client you know get storage credential and access storage directly, now it will talk through this Databricks managed proxy
[09:49] and this is Databricks managed proxy will fetch bytes and and stream bytes to back to the the recipients. And what does that mean for simplifications is provider only need to onboard to this proxy once now instead of doing this one one by one
[10:06] relations with recipient that scales sort of all then you only onboard to secure connect this proxy once and thereafter regardless how many recipients you you have, they will just all talk to this proxy. And
[10:21] and and from your perspective, this secure connect this proxy is sort of just Databricks data plane. So, if you're using serverless today, you might have already granted serverless access to your storage and there's sort of nothing you
[10:36] need to do. From a recipient point of view, this proxy is just Databricks control plane. So, if they you know they might have already allowed list of those control planes in their networking. If not, they just need to go to Databricks documentations, grab those
[10:52] IPs and allow list in their VPC policy. And if there's a recipient on serverless because we manage the networking, we actually can get it just work out of the box. So, recipient don't need to do anything from a networking perspective as long as there's a right Unity Catalog
[11:08] permissions that give them access, networking just sort of work out of the box. And last but not least is private connectivity and this is very important that we hear from a lot of enterprise customers is that I want my storage sit behind private link. So, with this feature also,
[11:23] you'll get private link connectivity between this proxy and your storage via network connectivity configurations um, is this uh, sort of private link for Databricks serverless. Um, and uh, and and this works even for cross
[11:38] cloud, uh, which if you've done, uh, you know, cross cloud networking, you know, private connectivity, uh, for cross cloud is extremely difficult to do. So, how does this work? Um, I'll break it down to sort of the first, the setup flow, and then a access
[11:55] flow. So, uh, the setup flow is uh, the provider uh, enable this feature, um, uh, once, uh, and Ali will show you guys a demo, but this is basically a a metastore level setting that you can go to Delta Sharing, uh, to enable.
[12:11] Uh, once that's enabled, um, by default, uh, all the new recipients and and new sharing you you you try to do will, uh, get proxied through this, uh, secure connect. Um, but if you want sort of more granular control to either migrate
[12:27] existing shares or for certain new shares you don't want them to route through secure connect, there's a recipient level control that you can toggle on and off, uh, to say whether I want this to, uh, route through secure connect or not. Which Ali will also show you guys.
[12:43] Um, and then so that's the setup for you. Once you you do that once, uh, you kind of done. Regardless how many more recipient you on board, uh, you don't need to do any sort of extra uh, special things. And then from a recipient point of view, um, if they're serverless, there's sort of
[12:59] no setup needed, um, because we handle the networking behind the scene. If they're classic, they're using classic cluster on Databricks, or if they're open recipient that are not on Databricks, because we don't control the networking, uh, if they're if they have firewalls or VPCs, um,
[13:15] sitting in front of their compute, they just need to allow this the Databricks IP. Which they can grab from, uh, uh, the the documentations. And once that's set up, uh, during access, what happens is that um you know, the client same as before will
[13:31] request access to a shared assets from the open sharing server. Um we will then, you know, authenticate and authorize for to figure out if this guy can actually have access. If it does, we'll give them the credentials. But the credential we give them now is not no longer the direct storage
[13:47] credential that you use to call the storage directly, but rather is a storage uh is a credential you call to call the secure connect, this proxy. The client now talks to the proxy with the credentials, and then the proxy then will go fetch uh the bytes from cloud storage, um
[14:03] get it back and stream it back to the client. Now, if you have a private link configured for your cloud storage, this leg, number four here, between the proxy and the storage, uh will be over private link as well. And then the recipient will, you know, query as normal. Uh so, there's really
[14:18] sort of no change from a uh sort of recipient point of view on the query experience. With this, what that What What does that mean? Is that you sort of get the best of both world with the zero copy sharing. Uh previously we talked about some of the benefits is that minimal
[14:35] cost, operational efficiency, single source of truth, um because you have the single zero copy uh data. Now, with secure connect, this storage proxy, you also get the simplified networking and also the private connectivity uh that you can use
[14:50] to um for your storage. Next, uh I'll Before I wrap my section, just want to talk about the road map. Um so, there's sort of two uh pillars of uh things we're working on.
[15:05] The first is expanding more uh coverage of different asset type. Um so, right now we support uh mostly tabular assets like tables and views. Um we'll also expand the assets to more, sort of let's say, volume, notebook, AI models. Um so,
[15:21] that's sort of the first pillar. The second pillar is uh now the networking because it's Databricks managed now. Uh let's say you're from serverless classics, you're uh you're talking to this proxy this via this Databricks managed proxy. We're going to try to make that uh even more private and and better for you
[15:37] guys. Um so we're actually working on our own um quote unquote backbone for uh cross-region cross-cloud connectivity that's more private uh that's you know more performance. So over time we'll migrate that over to our backbone as
[15:52] well. Um so that you know without you guys doing anything need to do anything because we manage that you just get um sort of better networking experience. So with that said, I'll pass it on to uh Pedro to talk about how Amadeus is using uh Secure Connect.
[16:09] Thank you, Wade. So let me start introducing Amadeus for the ones that are not familiar with us. Amadeus is a technology company dedicated to the travel industry that operates at a scale. We are present in more than 190 markets.
[16:25] And to give you some numbers, last year our reservation system processed more than 485 million bookings and more than 2,000 million passengers boarded.
[16:43] We design our solutions around the traveler journey from the inspiration phase to through shopping, booking to the all trip phases including pre during and post trip as understanding travel behavior is one of our main priorities.
[17:02] We operate business at a scale and we implement technology at a scale as well. We have a global uh multi-multi-cloud footprint with a presence in more than 15 cloud regions where we are running more than 200 Kubernetes clusters. At peak times, our systems can process
[17:20] around 150,000 transactions per second. So, numbers that are really comparable with Google search itself. And that uh reaching 99.95% of availability. Yeah, actually Amadeus has implemented one of the largest
[17:35] service-oriented architectures in the industry with more than 600 applications running more than 10,000 microservices. To illustrate this technology scale with an example that I'm sure that you are
[17:50] all familiar with. Let's see example of the flight search. Every day, Amadeus receives around 3 billion search flight search queries on our systems that are exploring billions of combinations as well.
[18:07] Who at the end provide uh from 50 to 1,000 different offers. And that's thanks to the machine learning models that we are running every day around 2 billion of machine learning inferences that uh allow to to get an efficiency gain on
[18:24] this uh process from 20 to 50%. And that's how we make the travel ecosystem better through five levers. Personalization to provide the right content at the
[18:39] right time. Analytics, turning data into actionable insights. This is where Databricks is helping us. Business automation, events, reacting on real time to business events, and integration between the different
[18:56] partners of the travel ecosystem. And that vision is implemented in our open platform. It's a set of travel capabilities deployed on the public cloud that combined with curated and trusted
[19:11] data allow us to switch from a product oriented company to a solutions oriented company. And to operate that open platform, we need a standards. We need a common language and that's the
[19:27] open data. This is an interoperability standard that we are enforcing by design in our platform and allows our different solutions to talk the same language.
[19:43] And that's implemented in our data mesh. So, we have implemented a data mesh architecture in Amadeus where the governance is federated by domain and each domain ensures that the data is protected with the right confidentiality and access authorization at the
[19:59] the right side. This data mesh today in Amadeus is composed by 10 different data domains where more than 290 application work space. It means project teams are working and collaborating together.
[20:16] And it's storing 38 petabytes of information instantiated in 1,700 data sets. That represents around 36,000 containers.
[20:34] And how Databricks is helping us to unlock value from from our data mesh. Well, of course on the data engineering side, but Databricks is helping us converting the raw data, the raw open data into actionable insights in our golden layer. But we want to go beyond that. We want
[20:49] to implement our open platform vision by connecting our data with our customers. And here we have two challenges that we presented before. One is security
[21:06] and the other one is scalability. Is security because the data is protected at the rest. The data is governed all of we know our data governance framework on the data mesh. But when it comes to exposure, we have
[21:21] to make sure that we are able to expose the data to our customers without exposing the underlying storage. And scalability because on the the volumes and the scale that we have seen that that we manage in Amadeus, we want to be able to implement that vision
[21:38] without having to repeat the manual steps. Of course, there is a classic way to implement that. That was the the traditional BI way, which is to export batch files. Okay? That's a way to expose data securely because you are
[21:54] not exposing the underlying storage. But that's not valid anymore because customer wants to access the data on near real time because the data is refreshed on the source on real real or near real time. So, we want to implement in place data sharing, zero
[22:10] copy data sharing. So, for that Delta Sharing uh is a is a very suitable component, but respecting the security and scalability principles. And that's how Secure Connect can allow us to share the data
[22:26] without replicating, but respecting security and scalability. So, to summarize Amadeus is operating business and technology at a scale. We are unlocking new opportunities thanks to our open platform and thanks
[22:41] to all the data uh capabilities that we are offering. And for that, Databricks is helping us thanks to Delta Sharing and now with Secure Connect to do it respecting security and scalability. And now, I'll
[22:56] pass the floor to Gali, who will show us the live demo on how does Thank you. Thank you, Pedro. Uh For now, we will go ahead with the demo. And basically, for this live demo, what we will do is that basically, we will go
[23:14] directly into my environment. So, let's assume that right now, I'm an airport. And I want to get data, flight data, to be sent directly to an airline because
[23:29] of the World Cup today. Basically, I want to share that in a scalable way without any friction and in a secure way. So, basically, as a Databricks user, I want to use Delta Sharing and Open Sharing. The issue is that
[23:45] basically, if I want to do that for each airline that I operate with, that case, it might be introducing some frictions. Why? Because I don't want to expose my storage, but I have in that case to align with each airline in order to make sure that I
[24:02] trust their compute in order to access the storage in a secure way. That's how we introduce right now Open Sharing with Secure Connect in order to streamline this process. So, let's assume right now that I'm the
[24:17] data provider, so I'm the airport. Basically, I have right now a catalog, World Cup flights, in which I also store some information related to operations of my uh airport. In that case, my tables are
[24:34] right now stored in a storage account on Azure called World Cup Flight Operations DLS. If I go into this storage account, basically, we will see that the storage account is uh in that case, private, so public
[24:51] network access is disabled and only accessible by my own compute and also with other kind of compute, which is the serverless compute. And why am I basically allowing serverless compute to access my storage? Because in that case, I will use the
[25:08] underlying storage, I mean, compute of serverless in order to introduce secure connect. So, basically right now, my storage is is I mean, private. I want to share the table, so I still have the capability to share this table
[25:24] to another user. In that case, I'll be here in my workspace as data consumer. I have done the sharing to my consumer. The consumer does not have yet done the setup with me in order to create its
[25:39] own private endpoint. So, if I refresh my page, won't to access to the table, I will have access to the metadata, but after 20-30 seconds, it will then time out and tell me that basically I cannot access
[25:56] the underlying table because I don't have access to the underlying storage of this table. Now, let's set up secure connect. So, in order to set up secure connect, I need to go directly into the metastore of my
[26:12] workspace. So, in that case, I can go into this metastore. In that case, once I go into this metastore, I first enable open sharing with parties out outside of my organization. And second step is that I can configure here
[26:27] an NCC. So, maybe I can zoom it a little bit, but I can configure an NCC, network configuration connectivity for serverless, in order then to use that as the backbone in order to access the underlying storage. So, in that case, I
[26:42] will use a dedicated NCC with the specific private endpoint in order to access the storage. And it will be this NCC that will be leveraged each time I want to access my storage using secure connect and that will be set at the provider
[26:58] side. Once I set up that, I have the possibility to set up this secure connect by default at the secure at by default the secure connect capability by default for each new
[27:14] recipient or I can also specify that at the recipient level. So, if I go back to my workspace, go into Sorry, here. Go into my recipient. In that case, my recipient will be AWS workspace in US one region. If I
[27:32] enable secure connect, in that case for my specific recipient, I make sure that I want to use secure connect and I make sure also that the underlying network configuration connectivity have access to the underlying asset. I enable it automatically.
[27:48] And then, I can just tell to my recipient that's all good. Everything is all good. Everything is set up. You don't have to involve your network team. Basically, I go back as consumer. I refresh the page again.
[28:04] And the magic happened. In that case, I access the storage in a secure way without any friction and without having to set up any kind of other uh configuration in my own storage. And this is the way to
[28:21] basically access securely to table in real time without any friction and in and at scale because in that case, I can set up that once at the metastore level. And then, I just have to set up my recipient, allow secure
[28:38] connect, and then all my recipient will inherit these properties. And there it is. Uh I think that's it's all good for the demo. And it's all good also for this session. Thank you so much. Please fill out your surveys and have a great summit.
[28:54] Yeah. All right. Thank you for your time. Thank you.

Learn more about the Databricks Data and AI platform.

The information provided herein is for general informational purposes only and may not reflect the most current product capabilities or configurations.