Multi-Tenant Databricks Architecture: Hub-and-Spoke at Enterprise Scale
Summary
- P&G consolidated their European operations from 34 workspaces to one shared platform supporting 30+ teams and 80+ pipelines while lowering costs and enabling faster development.
- The hub-and-spoke architecture uses three control layers — Terraform for infrastructure management, Databricks Asset Bundles for pipeline deployment standardization, and Unity Catalog for data sharing and governance — to balance centralized oversight with team autonomy.
- With over 8,000 monthly users, P&G's multi-tenant design preserves working team autonomy while eliminating workspace sprawl, duplicate governance, and the administrative overhead that had made their original approach unsustainable.
Multi-Tenant Databricks Architecture: Hub-and-Spoke at Enterprise Scale

Scaling Databricks at enterprise scale creates challenges: multiple workspaces lead to fragmentation, duplicate governance, admin overhead, and inconsistent data. At Procter and Gamble's scale (over 8,000 monthly users), the traditional approach of adding workspaces became unsustainable, blocking cross-team collaboration and slowing adoption.
this video shows how P&G designed a multi-tenant hub-and-spoke architecture using Databricks as a multi-tenant platform. You'll learn how to consolidate workspaces while preserving team autonomy through three control layers: infrastructure management with Terraform, pipeline deployment standardization with Databricks Asset Bundle, and data sharing through Unity Catalog. P&G reduced their European operations from 34 workspaces to one shared platform supporting 30+ teams and 80+ pipelines, while lowering costs and enabling faster development.
🤝
Chapters
00:00Introduction and the Workspace Proliferation Challenge02:51PNG's Data Evolution: From EDW to Databricks Lakehouse03:56Enterprise Challenges: Fragmentation and Admin Overhead07:28Multi-Tenant Architecture Framework09:56Hub-and-Spoke Design for Autonomy and Governance13:45Infrastructure Management with Terraform18:39Standardizing Pipeline Deployment with DAB21:22Data Sharing and Governance with Unity Catalog23:16Case Study: Consolidating European Operations24:23Key Learnings: Ownership, Self-Service, and Data Governance
FAQs
What is a hub-and-spoke architecture in Databricks?
A hub-and-spoke architecture consolidates multiple workspaces into a single shared platform where a central hub provides governance and infrastructure while individual spoke teams maintain autonomy. P&G used this model to reduce 34 European workspaces to one platform supporting over 30 teams and 80+ pipelines on the Databricks Data and AI platform.
How did P&G reduce workspace sprawl on Databricks?
P&G implemented a multi-tenant hub-and-spoke architecture with three control layers: Terraform for infrastructure management, Databricks Asset Bundles for pipeline deployment standardization, and Unity Catalog for data sharing and governance. This consolidated their European operations from 34 workspaces to a single shared platform, lowering costs while enabling faster development.
What are Databricks Asset Bundles and how are they used for pipeline standardization?
Databricks Asset Bundles (DABs) standardize pipeline deployment across teams in a multi-tenant environment, ensuring consistent and repeatable workflows. P&G leveraged DABs as one of three control layers in their hub-and-spoke architecture to manage 80+ pipelines across 30+ teams without fragmented deployment practices.
How does Unity Catalog enable data sharing in a multi-tenant Databricks environment?
Unity Catalog provides centralized governance and data sharing so that multiple teams within a shared Databricks platform can access data while maintaining proper access controls. In P&G's architecture, Unity Catalog served as the data sharing and governance layer supporting 8,000+ monthly users across their consolidated European operations.
Full transcript
[00:09] Hello everyone, welcome to the multi-tenency architecture session. My name is Carol Ro. I am the lead data and AI architect in PNG. Together we have Larry John who is the senior solution architect in data platform. We are going to co-present a showcase on how to sc
[00:26] drive massive scale with a management administrative overhap synergize data analytical and AI use case and the wonderful part is still get the working team the autonomy they want.
[00:41] Let me start a small quiz. How many workspace do you have in your company? Raise your hand if you have more than 10 workspace. Good. How about 50? Good. How about 100?
[00:57] How about 500? We we end up with many many workspace in PNG in our journey. So today we will use workspace as a simplified example to
[01:13] share why we end up with so many workspace. what's the challenge or the scalability and the solution to handle this challenge. So here is the agenda. Before we go to the multi-tenency architecture, I would like to share the
[01:30] call from Shash, our CEO. He was saying that to leading the consumer relevant business strategy can and only be delivered by leverage super superior data technology and capability.
[01:47] It is found in this strategy and his executive sponsorship motivate and mand to invest on scalable governant data and data platform. Let's go through a little bit journey of
[02:04] PNG uh PNG data journey into layhouse. Starting from 1996, we have been introducing enterprise data warehouse where mostly invest in canonical data model.
[02:19] From 2013 when big data start to born, we adopt big data to process uh machine learning user case. Our journey with the data brick start from 2018 where we utilize data brick to process both
[02:35] structured data and unstructured data and starting from 2023 we had entered the layhouse journey where you can see the famous medallion architecture and the multi-tenency architecture. Let's step back a little bit back to
[02:51] 2018 when PNG first introduced data brick to our company. Data brick offer a super wonderful scalable data processing platform. So it unleash a lot of potential of data use case. A lot of team are knocking our door. Hey we want
[03:07] to leverage data brick to build our solution on top. We end up with supporting various uh use case on data platform including reusable data pipeline analytical use case and some team are using data to process the last
[03:23] mile ETL before data feeding do solution and BI solution. If you think about that at that time many of the famous data brick future are not available yet. Unity catalog is not yet born. End to end lineage is not
[03:41] available. So we have to copy our data from a reusable data pad to analytical product. And we also have the no direct query capability etc. So to tackle the demand
[03:56] the team end up with a prioritize infrastructure provision to get things done. So like I need another workspace to get my job done. Here you are. This is the code we have been doing. So it was a fast and quick solution at the
[04:12] very beginning but it quickly end up become a barrier at enterprise scale. The first challenge we encounter is about fragmentation as I just mentioned right. So the platform handle many many
[04:27] of the use case reusable data analytical data a lot of analytical data use case are use case basis supporting a particular business question supporting some part of last mile ETL. So I mean you can easily figure out some of the
[04:44] analytical use case they will have repeatable repeating compute same data is being copied and works worst the KPI is inconsistent across across those digital product
[05:02] there is also sizable duplicate if on workspace administration this including feature governance CI/CD pilot and also disaster recovery Thinking about hundreds of workspace the administrative effort is really really sizable. Consequently the new feature
[05:18] adoption is getting delayed. Legacy feature sunset is also become a problem. Consider you need to align with hundreds of people workspace ad. Hey would you adopt this feature? Hi would you sum high meta stop because you see it's there. There is tremendous effort from
[05:36] from administration uh perspective. It also come with a governance overhead especially when feature require additional info security vetting and we are hidding and already hit some other data brick limit
[05:52] on the frick side data democratization continue to become a priority for PNG. So at PNG scale we support multi- category multi- region. So do you can we expect a central team to do all the data
[06:07] work for the entire company? That's impossible. So we have been utilized federated the data management oppose since many many years ago with the AIA agent the demand data and meta data and business metadata and
[06:23] sematical alignment is is also getting up expecting the central team to do the work for the entire company getting less and less possible. We also have many many team knocking the door continually knocking our door to on
[06:38] board on a weekly basis. Our user base has been grow uh around 800 uh 8,000 monthly user. This user uh spread across data engineer, AI engineer, data scientist and data asset manager. At the same time, the expectation to data
[06:55] engineer is also kind of expecting they to play a bigger role. It's not just ETL. It's building a end to end data analytical and AI product from a data engineer perspective with the other democratization. Is it
[07:12] the data governance going to be stringed down? Of course, no. Quite contrast with the all the ever above the data governance demand is getting up. So this amplify the urge for us to evolve from the prior model to a new model. So I'm
[07:28] going to introduce the multi-tenants architecture model here. So it composes three different layer. The first layer is the central platform team. This is a central team in charge of resource provisioning including storage account
[07:45] unit catalog catalog workspace firewalk network etc. It also in charge of feature qualification. Hey data brick release a bunch of feature let's make sure that we qualify them from technical perspective from a security perspective
[08:03] it define a lot of standard including using naming standard data model standard meta data center etc to our company it also provide the platform governance and monitor in the back end the second layer of the of the
[08:19] architecture is the multi-tenant implementation across PNG we anticipate Okay, by our scale we will end up with around 50 to 100 multi-tenant uh multi-tenant implement uh implementation. This will be mostly organized by data domain first plus our
[08:36] category and market operation to support our category and market specific needs. Number three, each of the multi-tenant implementation will have h and spoke model. In our internal team we call it mothership and satellite and this will
[08:52] be the core focus of today's presentation by Valaria. The result with a s uh with the proven success is of course move through and continue input the data democratization you will have all the data in the central platform and then you can find
[09:09] all the data and reuse all the data as much as possible. And the second one is about productivity with all the standardization governance. We expect free eggs faster developer at the same time lower the cost especially lower the
[09:25] cost for administrative effort and we also expect the time to publish data to the data lake is going to be shortened thanks to the multi-tenance implementation. With that, I will invite Valaria to show how exactly we designed
[09:41] solution to to achieve this result. Thanks, Carol. Yeah, we've come a long way. Um, now let me pick up from here and open up the how behind it. So, how how do we scale? Um, the secret actually
[09:56] lies in a very simple mindset shift. When we ask ours this question, how do we scale? Instead of focusing on the technology itself, instead of focusing on how do we scale data bricks adoption, the real question we should ask ourself is how do we scale autonomy, governance
[10:14] and collaboration. Those will be the three keywords that you will hear repeatedly in the next uh 10 15 minutes. And with those three things in mind, we came up with hub and spoke multi-tendency architecture. What's cool about this model is it decentralize the
[10:29] domain ownership but at the same time centralize the um infrastructure and governance. So let me walk you through it conceptually first and then I'll go into the implementation detail. Let's look at hub first. So hub it serves as a
[10:46] control plan. So when I say control it does not mean it control every single business process. It contains a database workspace in the hub. Um but it's for it's provide a shared foundation for your uh standards, guardrails, security
[11:02] development patterns and reusable components. And we have a hub team and again this hub team they don't build any pipelines or they don't build any models. They focusing on enforcing engineering standards and maintain those reusable libraries and models and
[11:17] monitoring cost and usages. That's what the hub team do. So you can think about it as a management office in a well-managed building. So a management office won't tell their unit what to do. It won't tell the law firm calls to consult laws or software company how to
[11:32] build an app. They make sure the building works. They make sure you know the the elevator works, the maintenance is on track, the utility works and that's also the right way to think about a hub. So a hub is basically a common operating environment, a place to build and we make sure it's durable, well
[11:49] managed and fit for the enterprise use. Okay. And moving outwards we have spokes. So each spoke is a its own resource group and serves as an execution domain. So um it usually aligns with a domain owned workload or
[12:05] individual data products. Those are the people who knows their data the most. They know what exact use case their data should suit. They know what their transformation logic their quality rules their delivery cadence should be. So in that way a bulk is designed to be
[12:22] decentralized. It's although it's a resource group you should not think about it as um a technical container. It's more the domain ownership. It's more of a ownership boundary. So you can think about them as tenants living in this
[12:37] building. They absolutely own their office. They decide do I want to be a law firm or software company. They decide their arc structure. Do I need two people in my team, 20 people in my team? They recruit their own clients and fully accountable for their own business outcomes. So a spoke is not just a passive consumer of a hub. They are the
[12:55] tenants working on the hub's shared data bricks but owns their own workload delivery and business value. So that's a spoke. Okay, if that sounds clear to you conceptually, let's uh think about implementation because hoke is not the most um
[13:11] innovative IT PNG created. Um it you might have already seen it in the industry but what we've been asked a lot and what we observed our peers find difficult moving towards multi-tendency was the question how do you govern centrally while still maintaining um team autonomy well with data bricks
[13:29] technology data bricks is designed to be a multi-purpose multi-tendency platform it depends on whether you use it in the right way um as a technical foundation enabler and I broke down our secret for separating centralized governance from decentralized execution in into three
[13:45] control layers. Infrastructure management, uh, pipeline deployment and data sharing. Let me go over it one by one with you. So, infrastructure management uh infrastructure is a backbone that glue your your tendants together. Why should you we care about infrastructure?
[14:00] Because we will be able to centralize your governance standard but at the same time uplift the administrative work your tenants are doing to maintain the workspace. So your tenant can focus on the actual data work the value generated work and we use data bricks terform provider to do that. I think have anyone
[14:17] here have not used data bricks terform provider. Okay I think most of you know that that that's that's great because uh I'll be focusing on how we actually implement it instead of the conceptually what it is. Um and and this is what we do on a hub level. We don't manually maintain any
[14:34] configurations in UI. We codify the platform in terraform. Right now we have five reusable models. We have access model. It uh manages the service principle, the users, the workspace access. We have compute model. It manages the jobs cluster, the instant
[14:49] post, the opose cluster, cluster policy, service execution model and this permissions. We have directory model. It manages the workspace folder structure and folder permissions. Schema manages the schema and grants and secret manages the secret scope and ACL. And it it can
[15:05] be extended. Let's say data bricks now announced the unified AI gateway. It can we can add another module once it's available. Um so for the new AI tracking and I want to show you an example of how it looks like um in the Terraform repo.
[15:22] So this is an example of how a hub define their job compute cluster. You can see that we have a list of u items we defined and then the pipeline would uh when they create their cluster they will choose from it. So here we have VM types, we have runtime versions, we have
[15:38] security mode, we have maximum workers in this example set maximum 10 auto termination maximum 120 and we have a list of tags they must apply so that we can do cost tracking and monitoring. Well, this is the guardrails in those
[15:54] shared modules. The tenants absolutely own their own configuration as code. So they again they have their own resource group and they own their work level decisions. So use compute is my favorite example. Show you another example how they manage their own work level decisions. So those are the two job
[16:10] defined by uh the tenants. You can see that they are under the same cluster policy but uh they use different node types because um they have different workloads. Okay. So this standardize the foundation for infrastructure. But if you have a large number of teams and
[16:27] pipelines, you need a seamless operation model between your hops and spoke. So you don't want to trade speed for governance. You don't want your central team become the bottleneck. Your tenant wouldn't be happy. So operation is very important during this operation process. We are trying to make it as automative
[16:42] and as self-s serve as possible and we use git flow. So the pro the process look like this. Let me walk you through it. Again the bulk they own their configuration as co as code. they own it and they added their configuration file. If they want to make changes, they would propose change through a pull request.
[16:59] Once it's it's proposed, once it's um opened in GitHub, the pipeline would validate the change and generate a preview. And then a hub admin would come in and review it and approve it. And once it's merged, the um deployment workflow would detect the change and
[17:15] automatically de develop it into the environment. So this gives a bulk team a self-service operating model without bypassing central control. But you might ask what if you know we have thousands of updates, we have many many teams, thousands of updates. Um the central team can review it one by one. So this
[17:32] can be further automated by enforcing naming convention and approval rules. So um for example, if I'm a pamper analyst, Pamper is one of our most famous brand. So you are under baby care tenant uh this category this tendency um all your
[17:50] asset will have the prefix BC underscore as a naming convention and then your approval rule in this case in your code owner fail would be you know this user will be able to addit this asset as long as they have prefix BC underscore so in that case as long as your change is
[18:07] confined to the honors to your designated area and met the pre-approval approval rules it will be automat automatically approved and deployed to your target environment. Um, but I do want to remind you be careful of this. Only use it when you try to stream the approval process for a large number of
[18:23] configuration and updates because we don't want to blendly just approve everything otherwise you lose monitoring and heat workspace limit very soon. So use it in certain scenarios. Okay. So the next area after uh infrastructure management is pipeline
[18:39] deployment. We also find this is very valuable to standardize because once you ensure your consistency across the pipeline, you'll be able to circulate developer resources and reduce your pipeline on boarding time. Um and we particularly find three area that's value adding uh the ripple structure,
[18:56] CI/CD stages and branch in targeting and we use data bricks asset bundle. I think the new name is declarative automation bundle to facilitate this. Um so let's look at repos structure first. This is what we designed for our company. So you
[19:11] can see we have GitHub folder for um CSED uh related content. We have DDL folder for schema related content, ETL folder for our transformation logic. Um test folder for integration test and unit test. So by initiating your pipeline
[19:28] using DAP you automatically get the same core layout uh for every single pipeline and but the the specific business um content like the workflow the YAML the ETL the schema can still differ under the same uh logical organization. So
[19:44] they will just be the Python file or SQL file your engineer created under the same um logical organization. So that's a repo structure and another area the bundle help us standardize is CSD and promotion path and this is what we do. So our CI includes a company standard
[20:01] stages and checks. Um so we have SQL fluff if you are using SQL code we have we use pest if you are using Python code and we have sonar cube is our in-house uh built code scanning tool where we want to apply in every single pipeline and cd uses the data brick CLI to
[20:18] validate the bundle run schema migration and deploy the workflow uh the environment target is defined in your database.yammo YAML if you use DAB you you know this is file your pipeline need to put the specific environment variables like your workspace URL and your credentials. So the CI/CD pipeline
[20:35] would deploy the um to your right environment. It could be prod u you might have more or less. So this creates a repeatable release pattern where quality, security and code standards is already part of your CI integrated path
[20:50] and it save their developer um time to redefine the stages and creates a predictable promotion path from merge to deployment. But do remember your engineers still have full control of um their work level decisions. If they want to add more CI test they can just add
[21:06] the YAML. If you want to add more stages than deploy and publish which I showed here, they can added the file. They still have full control. like it just create a baseline for them to have a jump start and under the same structure as an enterprise will be able to circulate resources.
[21:22] Okay, last but probably most important for your collaboration is your data sharing strategy and storage flexibility. So we use uni car lock as a key enabler as Carol mentioned multi-tendency was not possible at the time we adapted data bricks because at that time it was still
[21:38] have metasur you see it wasn't there and in have metasur there's no way you can isolate access in one workspace but with unic log uh one amazing thing we can do is separate infrastructure ownership from data product ownership that means your governance can be attached to data
[21:54] but but detached from infrastructure and how do we do it on an enterprise level we just define three things we define the catalog structure the naming convention and access pattern and all the rest are live to the spokes so the spokes they own their storage account if
[22:10] you remember my conceptual diagram they have their own uh resource groups and they own their storage account they manages the life cycle of their data products they publish data define views you know manage permission at the schema uh table and view level they have full control over that and under this model
[22:27] it means the permissions s date with data product. So one spoke can publish to multiple catalog. U multiple spokes whether they share the same workspace or not sharing the workspace doesn't matter can publish to the same catalog. One spoke can
[22:43] subscribe from multiple cataloges and this spoke that already subscribes to this data does not mean another spoke who share the same uh workspace will automatically have access. So they still need to subscribe separately if they want the access. So this ensure the
[23:00] access isolation and uh data security and as a result the use of this reuse and discoverability facilitate cross domain analytics feature reuse and model training at our enterprise scale. Okay. So we went through those three control layers that make this happen. I want to
[23:16] show you an example of how it looks like in practice. So this is a real implementation example of our European market operation team. At peak they have 34 um workspaces. 12 of them were for data pipelines and 22 of them were for uh analytics product like dashboard and
[23:32] modals. And at that time we needed more because we have emerging AI needs. We have u um like agents we want to build. But before we add more workspace we're like wait a second let's see if we can make it more efficient. And by adopting
[23:48] hub and spoke multi-tendency architecture, we are able to consolidate them into one single shared workspace while still maintaining team autonomy. So in this new multi-tendency instance, we have 30 plus distributed developer teams operating on one platform. We have 80 plus pipelines managed through a
[24:05] common governance model. We develop much faster. The shared analytics and AI model becomes reusable cross tenant domain. But at the same time, we observed a reduced maintenance overhead and platform cost. Okay, I want to end our presentation here. But I do want to end with this. So
[24:23] if you go out of the rooms and you go back to your hotel, you sleep tonight and you thought about, hey, I I joined a data brick session. What did Carol and Valeria talk about multi-tendency uh and Terraform and DAP, how do I use it? If your memor is getting blurred, I hope you at least remember the three things
[24:38] when you're designing your next multi-tendency. And that's a key learnings we learned in our journey. So first if you are scaling your organization scale ownership. So you will have new teams comes in your instant might be hey let's just add another workspace but
[24:54] post a second deep breath create attendant instead define the ownership and let the ownership be the execution layer within the shared foundation. And if you are scaling your operation model, uh remember that platform should be in charge of building guard rails and
[25:10] enable self-service as your default choice. Everyone oneoff setup, every manual approval will become friction for your future scouting. So enable self-service wherever possible. And last, if you want no silos, AI and analytics collaboration,
[25:26] productize your data. You have infrastructure and teams comes and go. But what we've observed is the data usually outlive all of them. So your governance should not sit within your workspace or storage account. It should live with the data. So can it be safely
[25:42] discovered, reused and shared across team at scale. And thank you so much. Now I will open the floor for questions.
Learn more about the Databricks Data and AI platform.
The information provided herein is for general informational purposes only and may not reflect the most current product capabilities or configurations.