Scale Trust, Not Headcount: AI-Powered Semantic Engineering at Southern Glazer’s Wine & Spirits
Summary
- Southern Glazer's Wine and Spirits, operating across 47 US states with thousands of brands and millions of transactions, inherited thousands of inconsistent metric definitions through organic growth and acquisitions, causing the same business question to return different answers in different tools.
- The company rationalized over 20,000 active reports across Tableau, Business Objects, and legacy platforms down to 600 certified metrics, implementing a unified semantic layer with data products and Genie Ontology as enterprise context on the Databricks Data and AI platform.
- An AI-native pipeline using analyst and engineer agents automates metric alignment and business rule updates, delivering new metrics at approximately $0.39 per report with high adoption and compounding value as the certified library grows.
Scale Trust, Not Headcount: AI-Powered Semantic Engineering at Southern Glazer’s Wine & Spirits

Trust in business decisions erode when reports, AI Agents, and operational systems provide different values for the same KPIs. The solution is a unified semantic fabric in which KPIs are managed as versioned, testable, and reusable assets—like software—implemented entirely within Databricks without added tooling overhead. Organizations no longer need hordes of data and semantic engineers, manually coding pipelines and maintaining definitions. With AI agents automating data transformations, semantic alignment, and business rule updates, organizations can reduce manual engineering effort, accelerate delivery, and lower operating costs while improving data trust. this video showcases how Southern Glazer's Wine & Spirits is implementing an AI-augmented approach to build their Metrics Bank and Semantic Layer in Databricks. You'll discover real-world results: reduced manual engineering effort, accelerated time-to-insight, and scaled decision-making capabilities across the enterprise.
Talk By: Anoop Kuriakose, Managing Director, AI & Data, Deloitte ; Sankar Chidambaram, Sr. Director Data and AI, Southern Glazers Wine and Spirits ;
Chapters
00:00Introduction: scale trust, not headcount01:02Southern Glazer's scale: 47 states, thousands of brands01:54The data Babel problem: thousands of dialects, no single truth03:44Three design principles: earn your place, trust, friction-free access04:14Rationalization: from 20,000 reports to 600 certified metrics06:08Process mapping: metrics tied to business value drivers08:22Semantic layer with data products on Databricks09:55AI-native analytics: Genie Ontology as enterprise context11:00Architecture: analyst agent and engineer agent pipeline12:44Enterprise context: the shared brain for all agents14:15Demo: submitting a business requirement in plain language18:49Confidence scoring and guardrails22:14Production scaling: RAG, guardrails, memory for repeated errors25:25Results: $0.39 per report, high adoption, compounding value
FAQs
Why did Southern Glazer's Wine and Spirits need a semantic layer on Databricks?
Southern Glazer's grew through mergers and acquisitions, with each acquired business bringing its own reporting systems and metric definitions, resulting in thousands of inconsistent dialects for the same KPIs like revenue and on-time delivery. The company needed a unified semantic layer so any tool returns the same consistent answer for the same question.
How did Southern Glazer's reduce 20,000 reports to 600 certified metrics?
The team rationalized their reporting landscape by mapping existing reports to business value drivers and identifying which metrics were genuinely distinct versus duplicates or variants of the same measure. This reduced over 20,000 active reports down to 600 certified, authoritative metrics that serve as the single source of truth.
What is Genie Ontology and how does Southern Glazer's use it?
Genie Ontology serves as the enterprise context layer — the shared brain — that all AI agents in the Southern Glazer's architecture reference to understand the business meaning of data. It is used in this video as the foundation for an AI-native analytics system where analyst and engineer agents generate and validate metrics from business requirements expressed in plain language.
What cost and adoption results did Southern Glazer's achieve with their AI semantic approach?
Southern Glazer's achieved a cost of approximately $0.39 per report with their AI-augmented semantic engineering approach, compared to the manual engineering effort previously required. The system has seen high adoption and delivers compounding value as the certified metrics library grows and more teams consume it through a consistent interface.
Full transcript
[00:08] Good afternoon everyone. Welcome. I am Anub Kurakos. I lead the data and AI business for retail and consumer goods sector within Deote. And I have u Shankar Chitaram here with whom um we
[00:21] partnered over the last last year um on building something great. We'll we'll share the story here. Chang >> thanks Aloo uh scale trust not headcount
[00:35] that's the whole talk everything that you are about to see has come down to one idea when the business asks for more insights the answer should be engineering excellence where the
[00:48] platform scales and the teams doesn't have to over next 25 minutes we are going to show you how we made that
[01:02] For those who don't know us, Southern Glaces Wine and Spirits is the world's preeeminent distributor of beverages, building brands for moments that matter.
[01:14] The multi-generational familyowned organization has operations in 47 US states and in Canada.
[01:27] and as well as brokerage services through southern glazes travel retail sales and exports division in the Caribbean, Central and South Americas.
[01:40] We operate at massive scale, thousands of brands, hundreds of c hundreds of thousands of customers and millions of transactions. Normally a company with this magnitude
[01:54] produces significant data volumes. Historically our challenge is not in the form of data acquisition but in establishing the consensus in its
[02:07] interpretation.
[02:21] We grew two ways organically as well as through merger and acquisitions. Each merger and acquisitions own bring its own system of records, own set of reporting structures and own set of metrics definitions.
[02:35] Each speaks its own truth. So you know the story of Babel. One language unified people and many languages cause chaos. We have not inherited many languages. We
[02:49] have inherited thousands of dialects. Each speaks its own. What is the customer? What is the revenue? What is on-time delivery? Like that. So the problem we are set to solve is
[03:04] one common standard unified language that any tool can access get the consistent answer. We had over 20,000 active reports
[03:18] running across Tableau, business objects and a handful of legacy platforms. Each BA systems governs itself. What
[03:30] does it mean? For the same question a user ask in two different tools. Most likely there are two different answers.
[03:44] Next one. And whilst before we kind of while we realized the problem recognized the problem before we wrote a single line of code we agreed on three guiding principles.
[03:59] Every decision after that the architecture the tooling the rollout everything traces back to these three principles. The first one insights are earned not inherited.
[04:14] So we asked every metric, every report, every data asset in the organization to demonstrate its value. The criteria is simple. The criteria is
[04:27] very straightforward. Does it power a core business process versus move a strategic goal? Anything that is not part of a core business process or it is not producing a
[04:40] measurable outcome that is not serving us is adding cost and complexity to to the system. So we made a disciplined choice to retire what no longer earns
[04:52] its place. So the second one is trust is non-negotiable. That's the foundation. Whether you ask through your questions through a
[05:05] dashboard or an application or an API or an AI agent, the user should get a consistent answer. One metric, one definition everywhere.
[05:17] Without this consistency, every analytics initiative fails and every A initiative even fails faster. The third one is frictionless access drives adoption.
[05:29] The best semantic world, the best semantic model that we can design means nothing if no one is using it. So friction creates workarounds and workarounds are breaking the governance.
[05:42] A broken governance means your data is not ready for AI. So we set a high bar. Analytics should feel as effortless as how we are using
[05:54] the apps on our phone. So for the implementation details I'm handing it over to Anu to walk us through the details. >> So the three three guiding principles that Shanka set up front came to life
[06:08] through three capabilities. The first one which grounded every metric, every insight in business processes came to life through a metric bank.
[06:20] The non-negotiable aspect of trust came to life through a semantic layer with data products and the seamless frictionless user experience came to
[06:32] life through AI native products and I will walk you through in each one of these three things in in detail over the next few minutes. So when we started the work, we did
[06:47] something. We did we went through the the business process and the business value drivers for Southern Glacers. So what you're seeing on the right side of the screen is a is a breakdown of the business process, the key process areas,
[07:00] level two, level three processes. What you see on the other side are the value drivers. what how does the the strategy of the organization break down to actions and and metrics that measure
[07:13] that? Then we overlaid the thousands of reports and the associated metrics insights and the data assets that that fed those metrics and insights on onto
[07:27] the the processes and the value drivers. And what we learned from that exercise was was very interesting. Several of the reports, metrics and the data assets
[07:39] were duplicated, weren't really earned to use uh use terms. They didn't earn their place in the business. So we rationalized what was duplicate and we distilled
[07:54] um a set of 600 u metrics that are tied to core business processes and value drivers and those are the metrics that that southern glacers needs to manage
[08:07] its business. So that is how we ensure that every metric every insight is uh earns its place. Now talking about trust. So now we know
[08:22] through the metrics bank we know what are the metrics to measure the definitions of the metrics. But how do you ensure that the data that feeds those metrics are trustable? For that we built data products
[08:35] organized by functional areas as you see on the on the left side of the screen like sales orders and invoices is an area, inventory, finance, supply chain and so on. By these areas, we organized
[08:49] data products for example uh sales performance as a product and we built it on the semantic layer using datab bricks metric views um so that it is easily accessible by any consuming
[09:03] applications. So none of the the logic the calculation sit in the consumption layer all of that sits in the semantic layer with the data products. Now we
[09:15] spoke about trust. How can you trust this? What drives trust? As you see on the right hand side, right hand panel of this page, the data products are owned. There are there are named owners. the
[09:28] SLAs's, there's a contract that establishes the freshness, the quality standards, the accuracy of the data so that it is a certified data product that end users use and not a a random table
[09:42] or an Excel spreadsheet that lives on somebody's desktop. So that is that ownership and accountability is paramount for trust.
[09:55] Why are we doing all of this? And we're doing all of this to empower users with insights. Do users care about a metric bank or a semantic layer of data products? No.
[10:07] What do they care about? They care about the questions that they have and the answers that they get and they want seamless user experience otherwise they won't use it. So that was made possible through this insights portal where end
[10:20] users can come in and ask their goals, ask their questions in natural language and there are agents behind this which Shunga will walk you through in great detail that build the insights
[10:34] that are founded on the the metrics bank and the data products who which sit on the semantic layer. So this is this is the portal through which all of this came to life.
[10:46] Sango, why don't you peel the curtain and walk them through the the details a little bit.
[11:00] Okay, thanks
[11:13] login to the system similar to how they log into any other system in the organization. It takes them to an UI interface. It's a very straightforward interface. It's not a SQL interface.
[11:25] There are no BA tools. There are no tickets. In this UI interface, they submit their requirements in a native langu in a natural language just plain English
[11:37] here. For example, show me the on-time delivery by region for last quarter. Something like that. They submit it. When everything gets done, they are
[11:49] getting the insights back in whatever form they wanted. It can be a PowerBI dashboard. It can be an uh it can be a kind of an API endpoint or an agent kind of an answer like the chatbot kind of an answer. In this demo, we are going to show you the PowerBI version of our
[12:02] application. So, what is happening in between that is our agentic pipeline. Consider this is a team of two specialist. One is the business analyst agent
[12:15] similar to a human analyst. Its job is to write what needs to be built. It understands the request and then enrich that with the content and then pass it
[12:30] to the engineering agent. The engineer agent gets the requirement from the analyst agent. and then do three major functions. Build, test, and deploy. These three
[12:44] major functions that it's doing. Here is the portion that adds Yeah, here is the portion that adds to the trust to the whole equation
[12:58] that is both agents are reading from a single shared brain. We call that as an enterprise context. It starts with a metric bank where the KPIs are all
[13:11] defined here. What does it mean? How it is being calculated? Why it matters? That kind of information is sitting there in our matrix bank. The second one is the data products. This is the
[13:24] production ready data organized by business functions and then ready to consume for users. So these two play a major role to add to the trust factor. The third one is a
[13:36] consumption style guide where it feeds the engineer agent to make sure that the product that the engineer agent is building is consistent and on brand.
[13:49] So nothing is a black box here. the whole stuff the logs the metrics the data lineage and guard rails even the the fin side of things everything is
[14:02] observed and powered by ML flow in this whole equation. So the next one is we are going to a small demo and how exactly we kind of put all these
[14:15] things together and I would like you to open this. Yeah. Okay, full screen please. Okay, so this is the interface normally a business
[14:28] user is getting and they come come to this interface and submit a business requirement number or an identity for a report and then it has list down to
[14:42] certain options are here on the reporting templates are coming on the right hand side. If you know your question, you are not sure about what template that you wanted to build, you can always choose auto. That means you are letting the LLM to decide everything
[14:55] based on the other architecture that I have explained. So the first one is the auto that LLM decides everything. The second one is the summary tab where it combines with the summarized data with the summarized
[15:07] charts and interfaces. Third one is pretty detailed. It has tiles as well as uh this has tiles as well as the graph charts as well as bar charts as well as the tables
[15:19] and most of our business users are techsavvy and they know how exactly Excel works and the exactly other systems powerba kind of systems works they said that build the report in a table format and we'll figure it out the
[15:32] visuals for those audience we are having a table interface here they can pick and choose the table interface it only the table data table, summarize table will be there on the reporting structure. The last one is a multi-tab. Each report has
[15:45] multiple tabs that is in progress right now. We are letting the users to figure it out how many tabs each reports, each dashboard contains and they can actually give that. This is the goals submission
[15:58] area where the users can submit their goals in a plain text actually.
[16:15] So everything has everything that has come down to once they submit their requirements it has come down to their what they can literally see what they are kind of submitting and how exactly it is being executed. It is a multi-step process.
[16:28] The first one is the business analyst agent and that kind of get the requirements persist the requirement in the BR form and then started enriching the YAML. The YAML is the structure
[16:41] throughout the process the YAML is being enriched by both business analyst agent and the engineer agent as it can as it accumulate more additional details to that. Once the data semantics are
[16:54] created in datab bricks in the YAML form a metric view has been queried based on that then the next step is building the PowerBI template the PBIT and then
[17:06] followed by the PowerBI executable. So we have once the job is done the user gets the status that the job is completed and they can click the report button to see the report the dashboard
[17:20] that they would like to see. There are couple of actions though throughout the process. We are showing the fullest transparency to the users.
[17:32] What is their initial request? What do they actually want to build and how exactly these agents are acting and enriching the content? So they can simply see that. So the whole YAML structure is appearing
[17:45] here like this. They can edit it. they can approve it and they can submit it so that there are no surprises at the end for the user. For a business user if it
[17:57] is too much for to read a YAML like this this long we also put markdowns on the on the top of the YAML and completely show the documentation like this in a plain text format.
[18:10] This contains the metadata and the data sources. What are all the columns? What are the measures the users care about and what are the relationship between the backend tables? How many filters are
[18:22] going to appear on the how many filters are going to appear on the dashboard and what are all the visuals the dashboard is going to contain and what are the visual goals based on the requirement that the user has submitted in the
[18:35] beginning. So the version of the report and all the metadata connection details are appearing here as well. The user can completely see this. The next action they can do that is the confidence score. How exactly we know
[18:49] that the LLMs are generating the right content that we would like to see. So this is the confidence score. This has come down to text SQL metrics and few other metrics behind the scenes like word and semantic search all those
[19:02] things based on that we put the value creation here and anything with a over over 80% confidence is the best report best outcome that the user can get it
[19:14] and they can also see that what are all the metrics that are above 80% what are all the issues they can see what is the confidence level that LLM is suggesting that whether the user can trust the output or they have to resubmit it. So
[19:27] this is the screen that provides the confidence level. The next screen, next action they can do that is the cost. So normally LLMs burn tokens very
[19:40] quietly when the when the admin shares the dashboard at the end of the month most of our we'll figure it out that most of our budget can has gone for burning the tokens. Right? So here in this case if
[19:53] you are looking into here we are providing super transparency on how exactly the tokens are getting burnt how many of input tokens how many of output tokens and what is the cost per request.
[20:06] Overall we have figured out that on an average to build a report in this methodology it cost around 39 cents to 50 cents for us.
[20:20] So with the four LM calls and this many latencies. Okay. When everything is done there is a feedback loop we would like to the user the next.
[20:34] So here is the screen. Normally the user sees that at the end of when once everything is done where they can rate the report if it is super accurate they can readily use it will be a fivestar or
[20:47] if it is anything is missing in terms of categories or measures or visuals or slices or layout they can simply highlight that what is missing here and then what are all the issues found
[21:01] and we are also accepting the suggestions from the users. So everything is getting persisted in the in the unity catalog and they can always resubmit it. It will cost them
[21:14] the same amount but they can always resubmit it to get a 87% confidence score to 90% confidence score something like that. So this is the feedback loop that we have given it to the users.
[21:33] with that we just m everything and uh powerba report something like that will come up so it's not a fancy report we just masked everything as it it contains our data so this is the whole demo that I would like
[21:45] to show you we can go to the next slide all of this built by the the two agents that you saw
[22:02] Okay, it's very straightforward to say that everything is worked fine. It's giving us 90% confidence. The initial prototype took couple of weeks for us to put together everything and show a proof of
[22:14] concept through our leadership. But the real challenge started like this. when we want to scale it for production that is where things started breaking.
[22:26] The first one is the context window overflow. Our base model has 100 plus tables. A model with 100 plus tables on a data model won't fit for a single prompt.
[22:38] So we have decided to load or retrain what is relevant for each request. We have achieved through rag and phase to
[22:51] make sure that only relevant tables are retrieved not the entire data model is getting retrieved for each request. The second one is the next big one is the hallucination. Initially when we want to scale it the
[23:04] lnm started hallucinating the column by itself joins by itself. It even create an entry by entity by itself and create visuals of its own. when we tried to
[23:16] export the report it started breaking. So all those hallucinations are really kind of taking efforts for us to fix it. We set a stage for six guard rails. Each
[23:28] time a guardrail will be established enabled based on the request that user has provided. How exactly the YAML has been created the AML file has been created and then how the visuals are
[23:40] created. The gates are ensuring that nothing is broken while we are shipping it to the user. Repeating mistakes. First you a report has been generated. There are some fancy errors are there. So you fix
[23:55] it on Monday. When you run it the same thing on Tuesday, same error appears again. So we added memory to the Unity catalog just to make sure that repetitive mistakes can be avoided. So that's the third big thing. And the next
[24:09] one is mostly on the model side. How can we trust the output? You're producing a report does not mean the report is correct. So it has to have the right numbers
[24:22] and all the prompts that the user is giving. They are expecting every single thing. For example, I want a trending chart. How exactly the sales is trending for past quarter? How is my on-time
[24:34] delivery for the past quarter by region? If they are expecting a totally a different chart, if the chart is broken, users will be less interested to use this tool. So to do that, we put the LLM as the judge. That's the confidence score that I was showing. The last one
[24:47] is the cost visibility. As I mentioned, the tokens are being burnt are quiet. We don't even realize how many tokens that we are submitting unless we see the invoice at the month end. So the phops
[24:59] tracker has been enabling us to make sure that there is a daily dashboard, there is a weekly dashboard, there's a monthly dashboard, even the finess tracker on the report that that is showing you how many tokens are being burnt for each request. What is the cost
[25:12] per report that we are normally burning? So these are all the challenges. The next one is the value delivered. Right? The first one is an engineering
[25:25] excellence. We have automated the heavy lifting. Mostly when there is a requirement is coming on our way. The first question will be where is my data? Is my data organized
[25:37] properly with right quality metrics, right guard rails, everything. So that heavy lifting is totally automated that helps us to do the time to market the
[25:49] speed. what it takes weeks to figure it out where is my data is the data has high quality and can I simply trust it to build the dashboard that I wanted to see is taking in minutes right now and
[26:03] the next one is the access when we are giving a business user something a effortless access where they don't have to see a friction that means the
[26:16] access is adoption is very high for us So they don't have to wait in line, create a ticket, wait in line for the ticket to be served, for them to access the contents. The last one is the
[26:28] compounding value. The more the system is being used, the more smart it gets. Next slide. That's all. We are ready for uh ready
[26:42] for questions. I think that's all. Yeah.
Learn more about the Databricks Data and AI platform.
The information provided herein is for general informational purposes only and may not reflect the most current product capabilities or configurations.