Stop Arguing About Numbers: Scaling Governance With Unity Catalog and Metrics
Summary
- Zalando built a three-layer architecture combining Unity Catalog for identity-based governance, metric views defined in YAML and managed through GitHub pull requests, and Genie for natural-language analytics to eliminate conflicting numbers across 50+ million customers' worth of data.
- The dual catalog pattern — private factory catalogs for team experimentation and shared showroom catalogs for governed production access — gives Zalando's 15,000 employees the freedom to iterate while maintaining a trusted enterprise data layer.
- Metric views are organized into a hierarchy of organizational, business-local, and private team definitions, with automated data classification and GDPR-compliant dynamic views enabling fast, safe data sharing across hundreds of teams.
Stop Arguing About Numbers: Scaling Governance With Unity Catalog and Metrics

At Zalando's scale, serving 50M European customers via microservices, data governance became a data swamp. With hundreds of teams owning separate data products, definitions diverged: marketing dashboards showed different net revenue than finance reports. Without a unified semantic layer, the organization could not build a single source of truth across dashboards, notebooks, and AI agents.
Learn how Zalando's three-layer architecture combines Unity Catalog for identity-based governance, metric views as a semantic layer built in YAML and GitHub, and Genie for natural language analytics. Discover the dual catalog pattern: private factory catalogs for iteration and experimentation, shared catalogs for governed data access. See how automated classification, dynamic views with GDPR compliance, and dimensional modeling, proven techniques from Ralph Kimball, enable teams to move fast while maintaining trust and consistency.
🤝
Chapters
00:00Introduction: Fabian and Mukram from Zalando01:10Who is Zalando: Scale, History, and Mission02:31The Problem: Matrix Divergence and Data Silos at Scale04:59The Solution: Three Layers of Governance, Semantics, and AI06:03Before Unity Catalog: Resource-Centric Permissions Chaos06:52After Unity Catalog: Identity-Based Governance Model07:38The Dual Catalog Pattern: Factory and Showroom09:51Dynamic Views: GDPR Compliance and Access Control10:56Automatic Classification: Democratizing Data Without Risk12:00Auditability and Lineage: Full Traceability14:45The Automated Workflow: Pull Request to Production15:33Configuration as Code: Sharing Data with Pull Requests18:29Enterprise Data Warehouse: Zalando Commons Metrics Model19:18Standardizing Metrics: From Legacy BI Tools to Lakehouse21:29Dimensional Modeling: Dimensions, Facts, and Metrics23:04Practical Example: Adidas Sneakers and Complex Queries25:00Metric Views: YAML Configuration and Databricks Integration26:22Central Metric Views in GitHub: Governance by Design27:50Hierarchy of Metric Views: Broad, Business-Local, Private31:17Custom Extensions: Owning Teams, Attributes, and Formats33:56GitHub Benefits: LLM Review Bot and Test Environments35:19Key Takeaways: Patterns, Automation, and Mandates
FAQs
Why was Zalando getting different numbers from different dashboards?
At Zalando's scale, with hundreds of teams each owning separate data products in a microservices architecture, metric definitions drifted — marketing dashboards showed different net revenue figures than finance reports. Without a unified semantic layer, there was no mechanism to enforce consistent calculations across notebooks, dashboards, and AI agents.
What is the dual catalog pattern in Unity Catalog?
The dual catalog pattern separates data into private factory catalogs, where teams iterate and experiment freely, and shared showroom catalogs, where governed and approved data products are exposed for broader access. Zalando uses this pattern to balance the speed of domain-team innovation with the consistency and trust required for enterprise-wide analytics.
How does Zalando manage metric definitions using GitHub?
Zalando stores metric views as YAML configuration files in GitHub, using pull requests as the governance workflow for approving and publishing new or updated metric definitions. This approach provides full auditability, an LLM review bot for automated quality checks, and the ability to test metrics in isolated environments before promoting them to production.
How does Zalando use Databricks Genie for natural language analytics?
Genie is connected to Zalando's metric view layer, enabling business users to ask natural-language questions that are answered using consistently-defined, governed metrics rather than raw tables. This ensures that even complex queries across multiple markets and brands return results grounded in the same definitions used in official financial reporting.
Full transcript
[00:07] Thanks everyone for coming. Thanks for not getting sunk in into the lunch coma. I hope you are all uh for our exciting talk. Uh my name is Fabian. I'm from Zalando from Germany and unfortunately not with me directly here, but virtually I have my colleague Mukaram.
[00:22] Hey Mukaram. Hi. Uh thanks everyone. Can you hear me? Yes, perfect. Okay, perfect. Thanks, Fabian. Um hi everyone. Thanks for having us. As Fabian said, my name is Mukaram. Together with Fabian
[00:38] as principal engineer at Zalando. So, what we'll be covering today is detailed in this slide basically. Um we'll start with a brief introduction to Zalando and the scale we operate at. And then we'll look at why teams ended up um with
[00:55] different numbers before moving to Unity Catalog as our governance foundation. And finally, we will cover our enterprise data warehousing model and show how shared metrics views create consistent uh and reusable matrices.
[01:10] That's that. Who we are. Uh well, Zalando was founded in Berlin in 2008. So, you could say it it is now in its teenage years. And uh since then it has grown into a leading European fashion and lifestyle platform with around 15,000 employees
[01:28] and more than 50 million customers across Europe. And a lot has happened since it started as an online shoe retailer. Um by 2010, it had expanded into fashion and new markets. In 2014, we went
[01:43] public. That was a huge day for our company. And by 2015, it had started evolving toward a platform business. And that will actually uh define the direction we are heading towards. And from 2019 onwards, um
[02:00] Zalando focused on continuously evolving that platform. Um in 2023, it launched new customer experiences including stories and the Zalando assistant program. Um by the way, if you haven't tried it, it's worth trying. Give it a try.
[02:16] And um in 2025, um again a big milestone, the company continued expanding it of course across different markets, but also acquired About You. Um and that actually um is in the direction what we'll be
[02:31] covering through today. And the challenges that comes when you have more data points coming from different sources. Um so, why are we still arguing about numbers? Well, it has a lot to do with the scale we operate at um and the challenges it
[02:49] bring with that. Um as I just said, um at Zalando, we orchestrate a massive digital ecosystem that consists of millions of that connects actually a millions of customers with the thousands of uh brands and partners across Europe.
[03:04] And every customer interaction generates uh data. And operating at this scale actually comes with a unique set of challenges. Uh first and foremost is that our data landscape is vast and complex that's spread by thousands of microservices
[03:19] that stream petabytes of data points into our central data lake. And this architecture actually allowed us to scale rapidly. Uh but at the same time, it made governance actually quite challenging and also blurred the distinction between the analytical world and the um
[03:36] transactional world. Um so, that's where the complication actually started happening. And for years, we strived actually for a distributed approach to solve this by decentralizing the ownership to domain teams uh who could actually manage their own data products,
[03:52] but that also complicated the governance because we you we you have hundreds of teams managing their own data products, how exactly they share, and how exactly everything works. And on top of that actually, we had a absence of a unified layer
[04:09] that just amplified the complexity and the consequences we get from such a scale. And what happened that we have a matrix divergence problem. Um what does that mean? Actually, we have the cases where we say that the
[04:25] marketing dashboard is showing a a different net revenue than a finance report, for example. And the the the main cause was that the the matrices were living in silos, and it was difficult to govern them and ensure that they are discoverable and trustworthy
[04:41] for a liability throughout their life cycle. And those are the challenges and the issues that were actually we were facing back then. So, how do we enable team autonomy while maintaining a trusted data foundation? Um our path to trusted AI-powered
[04:59] analytics has three layers. First, we have the Unity Catalog that provides the govern foundation. It gives us the centralized access control, discoverability, lineage, a consistent way to publish trusted data. And that's of course the foundation of,
[05:16] you know, what we going to talk today. Um exactly. And the second is the metrics views and that at the semantic layer actually for us. They allow us to define the business logic once and reuse it consistently across dashboards, notebooks,
[05:31] pipelines, self-service analytics, whatever else you have. And finally comes the Genie layer which brings actually a natural language analytics on top of those govern definitions. And this helps users explore data more easily while still relying on trusted metrics and business
[05:47] context for example. And together these three layers actually connect governance, semantics, and AI so that teams move faster without sacrificing consistency.
[06:03] Well, let's just start first looking into the governance foundation with Unity Catalog and later Fabin will cover the rest of the parts. Uh what we had actually before the Unity Catalog? We had our access model that was largely resource-centric. Um permissions were spread across
[06:20] thousands of S3 resources, I am roles, bucket policies, and there were so many exceptions. All lot of them. And reviews were manual and time-consuming and keeping everything aligned became increasingly difficult. And as the platform grew,
[06:36] this model became prone to configuration drift and was no longer scalable. And actually often resulted in in slowing down the the engineering teams. And that's what actually led us to rethink the whole approach all together. And what we get actually with the Unity
[06:52] Catalog now, we moved from resource-centric permissions to identity-based governance. What does it mean? We have now the access policies that are tied to users and groups, um making them reusable across data assets.
[07:07] And reviews remain um federated to the relevant teams, but automation makes the process more consistent. Um this model is easier to operate, audit, um and also to adopt as our organization and data landscape evolves.
[07:23] Um and it just not only um scales well, but also solves all of the problems I mentioned actually earlier. And this sets the very good foundation for the rest of the things we'll be talking about. So, how does exactly this model works?
[07:38] We coined actually um a term um and we have the dual catalog pattern, um, we call it. And what exactly we do with it is that we balance actually team autonomy with the governance sharing. So, what we have
[07:55] what I mentioned earlier, we had we we started with the extremely distributed approach where we said the the we decentralized everything. We said that the the domain teams owns the full production as well as the sharing responsibility. How does that happen?
[08:13] Sorry. Okay. Um, and the and the and the sharing was also with the with the domain teams, but that complicated the governance. But with the with the dual catalog pattern actually we have first the the private catalog, we call it the factory because it lets the domain teams do whatever they want to
[08:29] do. They it lets them iterate, it lets them actually do as much experimentation as they want, um, and, um, without being blocked by permissions or other issues or other dependencies on other teams or the central team. And this factory consists of private catalogs where,
[08:46] as I said, domain teams build and iterate independently through self-service, um, that is offered centrally. And once the data is actually ready for broader use, it moves into what we call the showroom actually. Um, it's a shared catalog, um, where access is centrally
[09:04] governed through dynamic views. So, on one end side, when teams, you know, have an ideation, they start working on a product, they iterate, they develop something, they have within their catalog their private catalogs, and, um, um, once they think that whatever they
[09:19] have produced is worth sharing with other teams to actually build on top of that more usable, uh, useful products, then they they think to share it. And that's where actually the shared catalog comes into picture. And the shared catalog actually, um, gives producers the
[09:34] flexibility while providing consumers with a trusted and consistent experience. And that's exactly what the the main foundation of our grand foundation with Unity Catalog is actually. But um the question is that why did we opt for dynamic views?
[09:51] Well, dynamic views are the main control point for the access control we have in the shared catalog. First, they allow us to enforce custom process process rules for GDPR and anti-trust requirements. Um we inject
[10:06] this logic directly into the view definition using functions you might be familiar as is the counter group member, for example. And this means that the access to sensitive data such as email addresses, for example, or first name or addresses in general is evaluated
[10:22] dynamically based on the user's identity and authorization. So, the moment you as a user actually make query to the data, the dynamic view actually have the logic in place to evaluate whether or not you are allowed to access that particular column or row. And that's actually becomes a single
[10:39] front for everybody in the company to have a unified way to consume something. And on the runtime, the access decisions are happening. Second, actually every column is classified automatically. So, no sensitive non-sensitive columns can be broadly made available by
[10:56] default, which reduces the manual process altogether and helps democratize data without compromising compliance. So, let's imagine actually what I had said earlier. Um a domain team actually produces a data. They experiment that
[11:11] and they think that this data is useful for other teams to build their products on. They start sharing it, and this data product does not have any personally linkable information that associates with an individual. That means that it can just be made available to anybody in the
[11:28] company to start using it, and they They have to wait for the compliance approval and for the access approvals. And that's exactly where this flexibility actually comes. The automatic classification process we have lets you actually automatically identify whether or not there is a sensitive information. And if
[11:45] there is not, the access is automatically granted to everybody who's allowed to do to access that. Thirdly, because all cross-team access is flow through the centrally managed views as I explained, we maintain a full auditability
[12:00] auditability. We can face which policies allowed a particular user to access a specific row or column. And this is super helpful in terms of many scenarios, but particularly think of it like duplication for example,
[12:17] a schema change you want to apply to a product. And with this lineage information, the usage information actually you always have the possibility to look into and actually effectively see who exactly going to be impacted by your changes um you wanted to make to your product.
[12:33] And finally, we product tries reliable insight reliable insights. We do not silently hide sensitive columns and return partial results. As I mentioned earlier, we have the dynamic views that have the logic in place that filter the rows and columns based on the
[12:49] the user identity. And this is actually transparent. It does not actually silently fails or gives you a partial data. It actually makes it clear to you that why you are not able to access certain parts of data rows or columns.
[13:05] And um this actually not only make it make sure actually nobody's using the wrong or partial data to drive business decisions, but at the same time actually ensures that everyone understands what exactly they are lacking.
[13:21] And this is actually what enables us to not only on one end side actually I trade fast within those domain teams, but at the same time use a fast process to share data across teams within Zalando. And in case there is no personally
[13:37] linkable information, it's actually by default kind of with automatic classification in place, and in some cases with the little assistance from the owning team. Um and in case of actually linkable personal information, there is a process that actually fully assisted with the automation to make that happen as well. What you see on on
[13:54] the right hand side is actually um our team A, for example, which produces data in their private catalog. They are able to access both versions of that, the one they have in their own private catalog, the one that is exposed in the shared catalog as a dynamic view. Uh but the team B is not able to access
[14:11] their private catalog. So, no other team can access another team's private catalog. The only way another team can access another team's data sets is via the shared catalog. That's very important point. And this shared catalog actually is the the central
[14:29] main point of our governance framework we have enabled. Um that's that. Um next I would explain how exactly this workflow um looks like. Um as I said that we have um enabled quite some automation to speed
[14:45] up the process from production to a to enable teams to iterate as fast as possible to effectively uh speed up the the sharing process without compromising uh quality. And this process is actually um as I explain
[15:00] here is uh based on different steps. So, first the the the team that produces the data, um they iterate, they finally decide to share the data once they have or they see the value in the data they want to share. And uh once a domain team
[15:18] develops the data in its private catalog, uh when the data set is ready they wanted to share it. The team submits the table and view configuration through a pull request. So, they this configuration looks like the following. It It points to their private catalog
[15:33] plus the name of the the view they want to have in the and shared catalog. If I go back to the slide, um so in the green box, there there resides their private catalog private object. In the dynamic view, the object they want to expose to everybody else. So, basically this config file
[15:49] actually basically one liner or two liner, which has a link between the dynamic view and the private table that they want to expose. And the once the team submits that table and view configuration through a pull request to a central
[16:05] repo, this configuration actually goes through automated checks to ensure actually there is no duplication happening. For example, we don't want that we have six different views dynamic views available on six different table or more different tables of sales because that
[16:22] actually then leads to the same problem as we wanted to try to address the matrix divergence problem. And this actually enables us and makes the discoverability easier as well as the identification of such duplication easier as well.
[16:37] And once that configuration goes through automated checks and peer reviews, where you know, ideally that any duplication has been identified or any other issues that could or should be addressed are properly addressed, um that pull request is approved. And the then the
[16:53] automatically once that it's approved and merged, the platform provisions the one dynamic views automatically. And within few seconds or minutes, actually we have the private catalog objects from the domain teams exposed into the shared catalog. That if does
[17:09] not have any personally linkable information, is automatically classified, will be available to the whole right away without anybody have to do anything. And then actually the consumers start
[17:25] accessing that in case it's there are some linkable information. For example, a table has X percentage of linkable information and Y percentage of non-linkable information, they would anyhow have access to actually non-linkable information straight right away. But for the linkable information,
[17:41] they have to follow the the request process. And finally consumers start accessing that via the shared catalog via the shared dynamic view we have in the shared catalog. And that's what actually we had on the
[17:57] Unity Catalog that sets the foundation for our governance and we have framework we have established. Now with that governance foundation in place, I will now hand over to Fabian who will take us through the enterprise data warehouse model and shared dynamic views.
[18:14] Great. Thank you, Mukaram. Uh quick change here on the screen. Uh No worries, Mukaram is still with us, but that you we have a little bit bigger uh picture here on the screen. Okay, so thanks, Mukaram.
[18:29] Exactly. Now we have all this data, we have these thousands of tables nicely organized, compliant access, no one needs to worry anymore about um is something unsecured, are we leaking any customer data, etc.
[18:45] But we still have the question, okay, now I want to report on revenue. Now I want report for different markets, but how do I do that? What are the aggregation rules? I have now a table, but the table just contains some numbers, and I don't know how to choose now the right business KPI, maybe the right um column from a table. Everything starts
[19:02] with GMV something. What should I now use to have an official reporting foundation? We need therefore to align on the KPIs and metric definitions. Um, and we did that with what we coined enterprise data warehouse model, which
[19:18] we call internally our main one Zalando Commons, because they are the common metrics for Zalando all aligned in there. Historically, we managed those within our legacy BI tools. And this one with the help of those we standardized thousands of metrics and dimensions to
[19:35] ensure a single source of truth. What exactly is that now? How can I imagine that? Well, the term fabric of data products floats sometimes around, right? So, that we have a holistic perspective on the business. And we have built this for over a decade. We are extending it still
[19:52] and updating it. And for self always with the focus of self-service analytics. So, analysts should go in there, just select their metrics and dimensions, and immediately have a source of truth selected and can use that to provide a clear predictable image of the business.
[20:09] And we're actually not reinventing the wheel here. A lot of the stuff we are doing is decades old insights and data modeling practices uh promoted back then by Ralph Kimball already. The core benefits of all of this is
[20:24] obviously consistent query results. No matter who is using that agent, human, etc. We have consistency across all the definitions. We have reusable reference implementations. So, these dimensions someone wants to integrate a new fact table, they just need to have the right
[20:40] join key relationships set up and immediately profit profit from all the previously defined conformed dimensions. We federate the ownership to the teams which are usually the metrics are very close to the data. So,
[20:56] because the data teams anyways are talking probably to the business about the definition of a certain metric. So, usually those teams also own the metrics and the dimensions. And last but not least, very important in the last years, of course, as I mentioned, this is really helpful for AI agents. I think it
[21:13] was mentioned in the keynote that there's a lot of things for AI agents, but the basics are still the semantics which need to be defined, which are um uh helping to increase the accuracy.
[21:29] Just a quick recap on what all of all of this means, and what does relational data modeling by Ralph Kimball means? So, we have dimensions, which are basically describing the groups of hierarchies and descriptors. Um for example, what actually is the single source of truth for what an article is, or what does a customer
[21:45] mean? Uh what is a customer uh definition? What are is the master data for that customer? Uh how are we splitting up um our sales channels, markets, etc., and even some primitives like dates and countries we actually have to uh align
[22:01] because it's very surprising but people can write a date in very different formats, right? Eventually. Or uh is now the international code which we should use for Great Britain, GB or UK? If you look at the ISO standard, it's actually GB counterintuitive.
[22:18] So, all of these things are aligned in the dimensions so that everyone speaks the same language. Then as mentioned, we also have the facts uh organized in fact tables, usually quantifiable measurements of a business process. So, how many items are in an order?
[22:33] Uh how many um for which sales price did this order um go to the customer? How many shipments? Um how many parcels or packages are in one shipment? How many shipments are going out, etc. So, usually numeric values which can be
[22:49] aggregated. And how those are aggregated are defined eventually by the metrics. So, these are then the standardized aggregation rules for consistent KPIs. What does this mean in practice? So, let's take a
[23:04] practical example. The question, how many Adidas sneakers did we sell to customers who placed their first order in Austria yesterday? What do we need for that? Well, we first need a metric, which is based on a fact most likely. So, this is a very simple example. There's somewhere a column
[23:21] called sold items before return in the fact table. And on top of that, you have in this case a very simple aggregation rule. It's just a sum on top. But, this obviously can spiral into way more complex definitions of a metric.
[23:37] And then we need some more information. We need the dimension information because we want to know, okay, we ask for the first order. So, let's say this is a new customer. And you would think, ah, wait, new customer, very simple. I just look if they didn't have an order before. But, this is also where maybe some uh it's a good example to
[23:53] understand how this needs to be aligned in the business because for Zalando, a new customer is someone who didn't order in the last 12 months. Because if they ordered before, they have been making a long enough break that we say we reactivated that customer. Other dimension examples for this example are now, because we want to
[24:09] filter only on sneakers, the category. And uh since we only want to focus on Austria, we need also the market. All of this is eventually linked in with surrogate keys. So, surrogate keys are artificial keys we add to all the data um
[24:24] to be independent of any business uh specification of the keys, like who knows if a if a SKU um is really helpful in the data warehouse context. And also this way we can use the best data type which the data warehouse needs to do the joins actually eventually efficiently.
[24:44] Good. Let's get uh let's move to the practical part. This was all nice in theory, but how do we use this eventually in Databricks and metric views? Quick recap. I guess everyone who's sitting here knows roughly what metric views are. But, just uh if you don't uh these metric views are
[25:00] centralized framework provided by Databricks for defining and managing business metrics and dimensions in order to standardize calculation formulas. So, exactly what I was just talking about but the offering now to implement this in the technology. They are sitting usually between the data sources in
[25:17] Unity Catalog. In fact, metric views are also Unity Catalog objects actually. And then you can use them in all those downstream consumers in dashboards, in Genie spaces, or even outside of Databricks. You can just query them via SQL. And um
[25:33] which makes all of that very flexible. Which is also one of the biggest changes of metric views compared to the legacy BI tools because usually those were very locked in those tools. They are defined using all of us favorite configuration language YAML
[25:49] which makes it easy to maintain also in Git and one interesting aspect from a governance perspective is or pure metric views so to say. I think it's always said yeah, you can align your business metrics but well, if
[26:05] 10 people define a GMV in 10 different metric views, what have you gained with that? We actually need to also define them once and encourage everyone to reuse them in multiple places so that we really have one definition which is everywhere used. But how do we do that? Well,
[26:22] we pull again our shared catalog which Moqadam explained earlier. We actually put the metric views in there in that metric in that shared catalog. No one has any right rights directly. So, everything needs to go through our central GitHub repository where we define those metric views. And
[26:38] um everything is then from there deployed with the service principal to the shared catalog. And with that with those metric definitions being in GitHub, we can make sure that our governance is strictly applied, which I
[26:54] will come to in a second what that exactly means. On the right side, we can see a little bit how that then looks like logically, how the data is accessed. So, the users can share data. The blue box is what Mukaram explained, that complex configuration of all the compliant data is hidden in that
[27:11] blue box. And we have allow shared metric views only on shared data. This way, we can also make sure that all these compliance rules, which we explained earlier, are actually enforced. And eventually, you use those metric views in SQL editor or in dashboards or in Genie or outside of Databricks.
[27:34] We had We quickly found out, and this also ties in into this like, "Yeah, but everyone can just create a metric view, and then you don't have any standardization at all." In order to make a central central metric view really work, you need to probably set it up very broad. So,
[27:50] you need to set up some ground foundation. Revenue is a very broad definition, right? What if some teams actually need some localized version of that? So, we had to come up with a little bit of a hierarchy concept. And the first level is what I just explained, these really broad
[28:06] wide metric views, which are applicable to the whole business. So, a revenue KPI or the fact table for order positions is very like everything is in there. Every vertical, every market uh uh since the beginning of time. These are usually the ground truth. They
[28:23] are usually on basically one-to-one table relations from the fact table to one metric view or the conform dimension to one metric view. And since those are later on reused everywhere, we have the most strict validation actually for this one. So, right now, we actually have one person
[28:39] overseeing all of the metric views which are created in Zalando, which we will adjust when it becomes a bottleneck, but we also don't expect this one to scale to infinity. Couple hundred metric views maybe is probably enough to define this very central source of truth. So, let's
[28:54] see how long we will uh need to go through one person who has very deep experience of and the mandate to actually control what should go in there. Obviously, with the help of a lot of uh of a council. One thing uh quickly also, the central
[29:10] relationships are also defined here. Um currently, we do this with uh joins, which we will see in an example in a second in the metric views. Um we might want to reshuffle this a little bit. Um Databricks is actually working on uh a relationships feature for metric views,
[29:26] which will slightly adjust uh how the relationships are defined, but they will eventually still be also in that first-level layer. Then, now let's say the sneaker vertical in uh Zalando comes along and says, "Yeah, but
[29:42] I don't want to always sit in every dashboard which would create always a sneaker filter this, and I actually we always just need last 12 months. I want to have some business local uh version of all of that." So, that's what we call currently second-level metric view, business local metric view. The name is
[29:58] not yet fully fixed. We're experimenting with that project and trying to see because that obviously is very special to the teams. What do they actually need? What do they want to combine? Et cetera. But eventually, these are the localized insights. These are the filtered and combined perspectives what
[30:14] the teams then actually need. Um they can also use directly the first-level metric views for very generic use cases or for genius basis, for example. Usually, genius is good in figuring out, "Ah, I can filter on market uh Germany or Austria, for
[30:29] example." But if you want to have this uh reused uh a lot of times, then probably a second-level metric view is good for you. To complete the picture, we mentioned already people have their private catalog as well, where they well they can just create things there, right? But
[30:46] um we don't care because as Mukaram earlier said, no one can actually access this private catalog except for the team themselves. So this is uh okay for us that they can do whatever they want in there. They can use this for quick iteration and prototyping seeing how does Metric you works um maybe building
[31:01] a test dashboard etc. Let's have a look at an example of a metric view uh in Zalando. Uh I'm not going to go through everything in here. I just want to explain those red boxes because this is what you don't find in the Databricks documentation because
[31:17] these are our custom extensions how we call them. Um because we have this in GitHub, our CI/CD pipeline basically takes this YAML, tears it apart actually, and modifies it before deploying eventually to Databricks. The first one is pretty easy um because we are deploying this with a service
[31:32] principle, it actually says for all these shared metric views in Databricks right now owner the service principle, which is not that nice actually. And we actually use specify here in GitHub um the owning team as a uh field in the YAML directly, which is
[31:48] very close then to the definition and was for us the best way to do this. This field basically the CI/CD pipeline takes out, converts this actually into a tag, and deploys this then to Databricks because if you tell Databricks deploy this YAML with owning team, it will actually fail.
[32:04] The second box is um very small but very complex actually. So if you have ever built a metric view and set up a join and that you have it like in this case uh combined a fact table uh joined a fact table to a fact table a conform dimension, you would actually
[32:19] need to redefine all these dimension attributes from that conform dimension again in the fact table. Which in our case doesn't make any sense. We have everything already once defined. Uh all these comments are there. What if someone changes the comment there and then the comment down here? So we
[32:34] enabled this little flag so to say, which actually this is the very simple version of it. You can also say I need a prefix. I want to prefix the comment. I only want to import this and that field. And this way the owner of the fact table metric view doesn't need to redefine
[32:51] again all the dates, all the other attributes from the conform dimensions. Last but not least, at the bottom we see the format, which also if you look at this, so there is a format field in Databricks in the specification for metric views, but it doesn't look like this. This we simplified this a lot, so
[33:08] you just say currency or number or string. We actually I think we have five values right now. Because um this is the central reporting model of Zalando. So um and we are a European company, so we don't have the need to
[33:23] have any other currency than euros in there. So why do every why does everyone now need to specify the euro in there? So that's we stripped out, basically hidden this behind this one field, which just says currency.
[33:39] I'm going to jump quickly over this. Um there are some benefits to doing all of this in GitHub. So I want to quickly to also have some times for Q&A um jump to the next slide. Um I will leave this here for 10 seconds. There a whole bunch of um benefits to having all of this in
[33:56] GitHub in my opinion or in our opinion, um which you can see here. One thing because we run this pipelines eventually is um uh and one thing which is really helpful for the metric view creators inside Zalando, we have an LLM review bot
[34:11] running in our CDP uh CI/CD pipeline, which is basically getting the diff from the pull request, all the existing metric views, and a couple of instructions {{}slash} skills uh on uh which naming conventions do we want to have, what do we mean with our central data model Zalando comments, and
[34:28] basically gives that to the LLM and says, "Hey, analyze this. Does this uh change make any sense? And give us a green or red flag to indicate to the reviewers how easy this can be probably merged or not. Another benefit is we're actually
[34:43] deploying for each pull request an individual, we call it environment, to Databricks directly on top of the production data because it's read only only. So, it's pretty safe to actually do that. And this way the people can actually test their changes already on top of
[34:59] production data. Maybe set up a test dashboard or Genie space or something and see if the results are exactly as they expect. All right. Quick summary. What should you actually take away from all of what we just talked about?
[35:19] We think or we really like our dual catalog pattern. It's really easy for people to understand, ah, private catalog, I can do whatever I want. It's restricted, I cannot share it. And then with the And this way we enable their teams' autonomy and they can fully experiment there. And
[35:35] with the shared catalog we still have the compliant access for everyone in the whole company and this worked really well for us. Automate as much as possible. You've seen we talked about automation pull requests here, CI/CD pipelines, et cetera. LLM reviews.
[35:51] These are really helping us to keep the feedback loop really short and quick. And we don't want any team to have single point of review bottlenecks. This is a tricky one. Um one thing some people are always
[36:07] surprised when they come to Zalando they are asking, wait, you have the mandate actually to say there is a central definition? Get Get that mandate somehow. It involves talking to people, and it involves maybe talking to leads and upper management, but this is really important
[36:24] because we then get the feedback, ah, I don't want to align." Then I need to talk to this team and we're like, "Yes, you need to talk to this team." Technically, creation creating all those metric views is very simple, actually. It's just a bunch of YAML Genie code writes it to you in 1 minute, basically. You just give it some instruction.
[36:40] The difficulty is in an aligning ownerships, who should own what, what is the actual definitions, etc. Leverage as much as possible which is already there from Databricks, just extend it where your company needs it.
[36:55] So, we needed some of these governance layers when constant talks actually um with Databricks also to even remove some of these customizations which we did. Um but in general, um there is already a lot of that there. So, uh don't try to
[37:11] reinvent the wheel on that one and also don't try to reinvent the wheel on concepts which have been there decades ago. I just saw on the opposite side yesterday a talk um from some senior staff engineers here at Databricks and they how to optimize your data warehouse and they were saying
[37:27] use star schemas, use dimensional modeling. It's still super valid. This new Lakehouse RT super engine which they announced now is working really good with it. So, um this is really um yeah, invest into proven solutions. There's no need to really
[37:43] um invent something new here. With that being said, we actually wrote a blog post on this content what we just presented and a little bit more also on our Genie uh voyage. So, feel free to check that out on the Databricks blog. Um here's the link to that.
[38:00] And with that being said, thank you for your attention and I think we have 2 minutes for questions still. All right, thank you so much and enjoy the rest of the conference.
Learn more about the Databricks Data and AI platform.
The information provided herein is for general informational purposes only and may not reflect the most current product capabilities or configurations.