Phar Lakehouse: Automating Data Governance at Scale with Terraform
Summary
- Schwarz Digits Spain, part of the Schwarz Group whose best-known brand is Lidl, built the Phar Lakehouse platform on Databricks Unity Catalog, Delta Lake, and Terraform to manage data governance at a scale of 400 million events per day from its Lidl Plus loyalty and e-commerce ecosystems.
- Terraform automates the creation of production-ready medallion architectures as a reusable deployment pattern across three layers—catalog governance, cloud storage, and authentication—provisioning new catalogs in minutes rather than weeks without manually policing incoming data sources.
- In its first year of production, the team deployed over 50 catalogs and built more than 2,000 tables, with automated governance dramatically reducing support tickets and creating a scalable engine for AI workloads that require high-quality, consistently governed data.
Phar Lakehouse: Automating Data Governance at Scale with Terraform

As organizations scale data ingestion from hundreds of millions to billions of events daily, manual governance becomes impossible. this video shows how Schwarz Digits Spain standardized data platform deployment using Databricks Unity Catalog, Delta Lake, and Terraform, automating the creation of production-ready medallion architectures. Learn how infrastructure as code transformed fragmented storage, authentication, and governance layers into a replicable black box that deploys new catalogs in minutes rather than weeks.
The Phar Lakehouse platform separates concerns into three layers: catalog governance with Unity Catalog (defining schemas, permissions, and access controls), cloud storage with security and networking, and authentication connectors linking Databricks to cloud providers. By standardizing on this architecture and automating deployment with Terraform configuration files, the team deployed 50+ catalogs and built 2,000+ tables in the first year, achieving governance at scale with dramatically reduced support tickets.
🤝
Chapters
00:00Introduction: Data Quality and Governance Challenge02:06Lidl Loyalty Program: Event-Driven Architecture at 400M Events Daily03:32Scaling Challenge: Integrating E-commerce Complexity04:35Current Architecture: Phar Data Market with Single Catalog Medallion05:40Solution: Phar Lakehouse Data Mesh with Automated Governance07:00Automation with Terraform: The Black Box Deployment Pattern08:21Three-Layer Architecture: Catalog, Storage, and Authentication10:12Medallion Flow and Flexible Schemas: Bronze, Silver, Gold, PII12:08Results: One Year of Production Benefits and Metrics14:00Conclusion: Governance as Discipline for AI Success
FAQs
What is the Phar Lakehouse platform?
Phar Lakehouse is a data mesh platform built by Schwarz Digits Spain on Databricks Unity Catalog, Delta Lake, and Terraform to manage automated governance for the Lidl Plus loyalty program and the full Lidl e-commerce ecosystem. The platform standardizes data ingestion and governance as a replicable deployment pattern, creating new governed catalogs in minutes rather than weeks.
How does Terraform automate data governance in the Phar Lakehouse?
Terraform configuration files encode the entire deployment pattern as a self-contained black box that provisions all three architectural layers—Unity Catalog schemas and permissions, cloud storage security and networking, and authentication connectors linking Databricks to cloud providers—in a single automated step. This ensures every new catalog follows the same security and access control standards without manual governance work.
What is the three-layer architecture used in the Phar Lakehouse?
The Phar Lakehouse separates concerns into a catalog governance layer (Unity Catalog defining schemas, permissions, and access controls), a cloud storage layer (security and networking on AWS S3), and an authentication connector layer (linking Databricks to cloud providers). Separating these concerns allows each layer to be updated or replaced independently without affecting the others.
What results did Schwarz Digits Spain achieve with automated governance on Databricks?
In its first year of production, the Phar Lakehouse deployed more than 50 catalogs and built over 2,000 tables, with governance automatically enforced at every layer through the Terraform deployment pattern. Automated deployment dramatically reduced support tickets and enabled the team to scale data onboarding without manually reviewing and policing each incoming data source.
Full transcript
[00:08] Okay, I think we can now start. Great. Listen to me. Great. So, data is the food for AI. I'm sure that you wouldn't want to eat low-quality food. You shouldn't feed your low AI model with low-quality data.
[00:25] In many industry, the only way to get AI to work is to have high-quality data. This quote from Andrew Ng perfectly rephrases the classic garbage in, garbage out reality that we all face.
[00:42] But I love this analogy because it goes much deeper than its surface level. If you flip it around, it reveals a harsh truth about modern data platforms. Just like the daily food environments that we navigate,
[00:58] we often don't have a choice about what data is coming to our lighthouse. When you cannot choose every single raw ingredient that is coming to your ecosystem, your only real line of defense is how you govern, process, and
[01:14] filter it. And this is exactly where our architectural journey begins. I'm Leandro. I'm working as a senior data engineer at SFR Digital Spain. And today, I want to take you behind the scenes of this adventure. And like in any good adventure, it
[01:31] begins with a challenge. In this case, a massive organizational shift. One that brought a sudden and an overwhelming wave of new data sources coming in.
[01:49] I will share with you how we um decided to instead of going that manually police these incoming sources, why we decided to go with and turn our our entire focus to automated governance.
[02:06] So, this is how we turn a potential data swamp into a scalable engine for AI. But, first things first, let me tell you who we are. We are part of the Fast Group, one of
[02:21] the biggest retailing companies in Europe, and our most famous brand is Lidl. A few years ago, we decided to launch our loyalty program for the Lidl supermarkets. This loyalty program was called Lidl
[02:37] Plus, and it was based in an application mobile. In that moment, we decided to go with an event-driven architecture, which basically is based in our data contract API, and we basically could control what kind of data was coming to our lakehouse.
[02:56] It was a great decision because it could help us grow to the huge network that you see on the screen. We were receiving 400 million of events every day, what is about 640 GB coming in every day to our lakehouse, and in that moment, our lakehouse had a
[03:12] stored 5 TB of data. So, you can imagine that's a lot of data, right? But, remember, that was only for the loyalty program. So, imagine how we felt when we were told that the entire Lidl e-commerce ecosystem was coming to our platform, too.
[03:32] We needed to take a big step forward. The loyalty program was easy for us because we knew it very well, and even with the event-driven architecture, we could control the events that was coming to our lakehouse. But, the Lidl e-commerce was a different story.
[03:47] It was an already working machine. They had their own ways of moving data, and even they had their own business records and business requirements. And more importantly, the data that they were managing were unpredictable for us.
[04:03] So, in in this case, this was a huge challenge for us because we needed to integrate this all this data from different complex and cross-domain requirements. We know we knew it
[04:19] what what worked for the loyalty program would not work here. We needed a completely way of thinking. But before digging into the details about this transformation and how we solved this challenge, I want to show you the
[04:35] technical details of the architectures that we have right now and the one that we were using for the loyalty program. So, here is the far data market. The far data market is a centralized setup with one big catalog. So, all the
[04:50] events that we were receiving, we put it into a bronze layer. Then we had an team was activating it and moving it to a silver layer. And finally, we have several gold layers, okay? So, these gold layers were like playground for our consumers. Um basically, this were they
[05:08] created all the reports and calculated the KPIs and more importantly, where they were storing the tables created for the AI models. This architecture worked perfectly for a long time. But for the e-commerce,
[05:25] was completely different. We know it that it will not be able to handle the entire volume that we were facing. So, if you could not control the data that is coming in, what could control?
[05:40] Of course, I gave you a hint earlier, you need to put the focus on governance. And that's why we decided to go and create the Farley house. This architecture you can think it more or less like a data mesh.
[05:56] Okay, so instead of having one big catalog we deploy a catalog per each team's purpose. Basically, what they need to do. They get automatically built a metadata architecture ready to use. So, what they they will receive it's a
[06:13] bronze layer to put there to process the data, a silver layer to clean it, and a single gold layer in this case. This um feels very this this flow feels very
[06:29] natural for our users because it guides them to to do the things the right way. And another important huge improvement for our side is that in the Farley data market, we had the storage and unity catalog tools separated in two different
[06:44] repositories. For Farley house, we decided to put both together in the exact same automated pipeline. And it was of course a totally game changer for us. So, how do you actually deploy
[07:00] this standard setup for every single team? The answer is simple, automation. And if you want to build a standardized architecture at a scale, you cannot do it manually. It takes too much time and it creates too many errors that you need you need
[07:17] to rely on infrastructure as a code tool. In our case, we choose Terraform. And by using this tool, we can automatically build the exact same structure for every single team. It makes things fast, clean, and
[07:33] reliable. So, to explain you this automation, I will follow a top-down approach. So, here is the black box, okay? This black box is a reusable module that does all the heavy lifting for us. The best part is how is
[07:49] how easy is to use because when our teams needs a new catalog, the only thing that we need to do is to fill out the single file, the configuration file, which is a TF file, and we have one for each catalog that we have deployed.
[08:05] So, basically, we send this TF file to our Terraform project, it will read it, and it will deploy automatically the infrastructure comparing the current state and the desired state.
[08:21] This, as you can see, black box is divided into three different layers, and if we take a look inside this black box, we will see three the three different layers that I mentioned before. So, first, we have the catalog layer. This is where all the data bricks stuff is deployed. Basically, the catalogs, the schemas,
[08:38] and the users' permission, which is also called grant. With this user permission, we decide a design uh air back model based on user group. So, for each schema, bronze, silver, and
[08:53] gold, we have one reader group, one writer group, and an owner group. That means that if a user is part of member of a reader group, then it will have the select privileges, but not the
[09:09] modify one. This is very important for something that I will talk later. Secondly, we have the storage layer. This is where we exac- actually deploy the storage account that holds the data
[09:24] from the Delta tables. In this module, we are also deploying the configuration for security and networking, so we can keep the data safe. Finally, we We the authentication layer. This module works as a bridge between
[09:39] Databricks and the cloud. So this is where we are deploying the storage credentials and the data access connectors so Databricks can read and write the data in in the cloud. By separating these three layers our project keeps
[09:57] organized reliable reliable and of course secure. So in short this is exactly what gets built automatically every time that a a team needs a new catalog. So to recap the flow
[10:12] users put the raw data into bronze, they clean it and move it to silver and then create their business reports calculating the KPIs and deploying the AI models inside the gold schema.
[10:28] Another quick note here. Our users are also developing their AI models function and all these components are also managed by governed by Unity Catalog in this case. On top of the medallion architecture we the user can also request a version
[10:46] for PII. So that means that we can keep the sensitive and non-sensitive data isolated from each other because as you can see in the diagram every schema has their own storage account and this is very important to keep both use cases
[11:03] separated. We deploy more than just a medallion architecture. We can also deploy an utilities schema so this is a safe place for configuration files. And basically all the different data that is not related to the business
[11:20] directly so our users don't mix with the main delta tables. And of course we know that some users has unique need. They can also request the data mart.
[11:35] But no matter what they ask for, all the components follow the same rule. They They are being centralized managed by this data for project. So, we can stick to our main value.
[11:53] So far, I have a a talk about our big uh organizational shift and the challenges that appear with that. And also the hard decisions that we made by using automation. Now, like any good story, it's time for the resolution, no?
[12:08] And so now, let's look to the results of how this new model actually performed for us after 1 full year in production. We achieved three main benefits. The first one,
[12:24] we stand we created a standardized foundation. So, every team now gets the exact same structure when they ask for a a new catalog. This standardized foundation also is important because we can connect easily
[12:39] between workspaces. So, nowadays, our main business catalog is being accessed by 65 different workspaces. The second, um we could achieve an automated scaling.
[12:55] When a team needs a new catalog, it gets deployed in a few minutes. Currently, we could deploy more than 50 new catalogs and our users have built more than 2,000 tables only in a year.
[13:11] And finally, we could achieve security at scale. Because of this airbag model that we designed it, we could easily connect to our roles management system in inside the company. So, if a user needs to get access for a
[13:29] Google schema, this user can directly request this role through this a portal and the data owner will approve it. This means that our users, our data teams, get auto-autonomy instantly.
[13:45] And of course, for the platform team, means way fewer support tickets. So, to close my presentation, then be let's bring back again the food analogy.
[14:00] Having a data good governance it's like having a good diet strategy. It takes discipline. You need at least measure the calories that you are consuming. S- If data is the food for AI, as Andrew Ng
[14:18] said, then governan- is the discipline that takes that keeps it from eating junk food. When you have a massive amount of data sources coming in, things can easily go
[14:34] wild. This where the automated governance makes the difference between a data platform that helps AI models run successfully and one that fails because of the bad ingredients.
[14:54] If you keep your data governance smart and disciplined, your AI models will be smart and healthy, too. So, I'll leave you with a final question. Are you governing your AI's diet or just hoping for a healthy model?
[15:09] Thank you so much.
Learn more about the Databricks Data and AI platform.
The information provided herein is for general informational purposes only and may not reflect the most current product capabilities or configurations.