Associate 360: Building Workforce Intelligence with Databricks Genie and Unity Catalog
Summary
- Mars built Associate 360—a unified workforce intelligence platform on the Databricks Data and AI platform using medallion architecture, a common data model across hire-to-retire data, and semantic views—to give HR leaders natural language access to people analytics for 140,000+ associates.
- Semantic views with detailed business context and link tables were critical design choices that improve Genie query accuracy and prevent hallucinations, demonstrating that AI-ready data foundations require meticulous metadata work before enabling conversational interfaces.
- Unity Catalog provides governance and row-level security for sensitive employee data, enabling compliance with data privacy requirements while allowing the platform to scale across Mars acquisitions and business segments.
Associate 360: Building Workforce Intelligence with Databricks Genie and Unity Catalog

HR teams across enterprises struggle with fragmented data sources, siloed dashboards, and the need for data engineering support to answer even simple workforce questions. Mars, with over 140,000 associates across confectionery, pet care, and food businesses, faced complex challenges scaling workforce analytics across acquisitions. Building Associate 360, a unified workforce intelligence platform with Databricks Genie and Unity Catalog, Mars transformed how leaders make HR decisions using natural language queries instead of static dashboards.
this video demonstrates how to build an AI-ready data foundation using medallion architecture, a common data model across hire-to-retire data, and semantic views with business context to prevent hallucinations. You'll learn how Unity Catalog enables governance and row-level security for sensitive employee data, how link tables improve Genie query accuracy, and how to design for both human and agent interaction. The demo shows conversational analytics in action: from attrition trends to identifying high-performer retention risks.
🤝
Chapters
00:00Opening and Rachel's Introduction01:49Tiger Analytics Partnership02:55Mars Company Overview04:30People Analytics Transformation Journey06:42Associate 360 Vision and Three Key Asks08:54Governance Challenges and Requirements09:59Associate 360 Analogy and Vision11:02Three Problems to Solve12:55Solution Architecture: Medallion and Common Data Model14:49Semantic Views and Context Management16:56Link Table: Solving Performance and Hallucinations19:00The Meticulous Data Work Behind Genie21:10The Power of Context for AI24:12Before and After: Dashboards to Conversational AI26:09Acquisition and Scalability Impact27:32Genie Capabilities and Examples29:24Future Roadmap: Visualizations and Personalization30:45Predictive Insights and Demo Introduction32:08Demo: Natural Language HR Queries34:32Demo: Analytics and Insights Discovery36:40Demo Results and Next Level Capabilities
FAQs
What is Associate 360 at Mars?
Associate 360 is a unified workforce intelligence platform built by Mars on the Databricks Data and AI platform that consolidates people data across the hire-to-retire employee lifecycle for 140,000+ associates across confectionery, pet care, and food businesses. It replaces fragmented static dashboards with conversational analytics powered by Genie.
How does Mars prevent hallucinations in Genie-powered HR analytics?
Mars designed semantic views with rich business context and created link tables that improve Genie's ability to join data correctly without guessing at relationships. The semantic layer explains business terms, HR-specific definitions, and data lineage to the AI, grounding responses in accurate data rather than inferred patterns.
How does Unity Catalog enable governance for sensitive HR data at Mars?
Unity Catalog enforces row-level security on employee data so that managers and HR leaders can only access associate records within their authorized scope. The HR Data Officer at Mars manages data security, compliance with data privacy regulations, and data quality as part of the Associate 360 platform.
How does the Associate 360 platform scale across Mars acquisitions?
The platform's common data model across hire-to-retire data is designed to onboard new business units and acquisition data without requiring custom integrations for each new source. This scalability was a key design requirement given Mars's history of acquisitions, allowing new associate data to be governed and made queryable within the existing framework.
Full transcript
[00:09] So, let me just start off with um the disclaimer of the forward-looking statements. So, if there's anything in here um that is um not meant to be any guarantee or promise, right? As you all know. And so, a lot of what I talk about later on in the presentation and my partner here will be aspirational. So, I
[00:25] just want to make sure that you have this disclaimer. All right. So, welcome to building workforce intelligence with Unity Catalog and Databricks Genie. Behind every data point, there is a person like myself. I am Rachel Bellino. I am the HR Data Officer at Mars. I have been with
[00:43] Mars for 8 years now and was part of this uh wave of external talent talent that Mars brought into the organization to really build out digital transformation at Mars. And so, for the past 6 years, I've been working on the people analytics transformation, which
[01:00] at the core of it is very much how do you build the data foundations for people analytics at scale. So, um just a little bit more about my role. So, as HR Data Officer at Mars, I have four areas that I handle. One is
[01:15] data quality for our analytics use cases uh with the people data. The second is compliance with data privacy of our employee data. And at Mars, we call our employees associates. So, you'll hear me use that word a lot. Um and then the third area is um
[01:33] data enablement. And that's everything to do with the data engineering as well as the acquisition of data that we use to enable our analytics and AI use cases using our associate data. And then um the fourth um but not least
[01:49] is data security. So, all of the security of our people data resides with me from a functional um ownership as well as in partnership with IT. So, at Mars, we have a large ecosystem of partners that we use to accomplish our objectives, and one of them here is
[02:07] with me today, Tiger Analytics, and I'll turn it over to Sachin to introduce himself. Thank you, Rachel. Hello, everyone. Uh, name is Sachin, Sachin Prabhat. As Rachel mentioned, we are from Tiger Analytics. Uh, it's a company, uh, that is in existence for 12 years, and we have been
[02:23] growing very, very fast. Uh, just give you a data point. Six years ago, I started with Tiger, we were around 600 people, and, uh, fast forward to 5 years or 6 years later, we are around 600 plus people. Uh, we work with, uh, lot of service, you know, like, uh, inter-industries,
[02:38] CPG being the largest, and Mars is a customer for 6 years, and it's a strategic customer for us. Uh, and this is a solution that we are very proud of because it really kind of, uh, focuses on, you know, what does it take to address the real problems within the enterprise and bring value to the
[02:55] front using cutting-edge technology. So, it's the business first, not the technology first. But, we are so happy that we are being able to use some of the cutting-edge database technologies like the Genie, the Unity Catalog, uh, to kind of bring to the, you know, uh, fruition. Looking forward to the conversation, and happy to take questions.
[03:11] Okay. Okay, so, a little bit more about Mars. Um, we have, we're a privately owned company by the Mars family. We were founded in 1911, so, we are over 100 years old, a century in the making, and, um,
[03:26] we have three, uh, main areas of our business. One is confectionery and food, so, many of you may know Snickers, you may know M&M's. Um, I think my favorite, um, uh, chocolate are actually the Dove chocolates, and also our recent recent
[03:42] acquisition of Hotel Chocolat. So, if you haven't tried Hotel Chocolat chocolate, please do cuz it's it's amazing. Uh, we also have a food business, um, and the favorite brand that I have in that area that I love, um, I eat it every day, it's 90 seconds in the
[03:57] microwave, that's Ben's Original Rice. And so, I highly recommend the brown basmati rice, which pretty much goes with everything. We also have a veterinary health business, which is our vet hospital, so you'll know some of the brands there as VCA, as well as Banfield, as well as pet
[04:14] diagnostics, such as Antech. Um, and then very recently, you may have heard in the news that we have acquired Kellenova, so that really expanded our snacking portfolio. And in that area, we have great brands like Pringles, as well as Pop-Tarts.
[04:30] All right. So, uh, so what did we do here in terms of building workforce intelligence? So, um, if I I'm just going to rewind back in time because I was actually here at the summit back in 2023 presenting on the
[04:48] people analytics, um, transformation that we were on. And back then, what we were doing was we had implemented the medallion architecture that I'm sure many of you are familiar with, which is bronze, silver, gold. That was really a big part, um, of our journey to go from
[05:05] a data swamp to an actual, um, really well curated data structure, um, for all of our data. And so, once we had done that, one of the things I remember literally being here at the summit doing the presentation, I
[05:22] remember Mate presenting on Unity Catalog and talking about how Unity Catalog is the governance layer and how that's really, um, fundamental to pretty much all of the future functionality that Databricks was going to release. And at the time, I was really struggling
[05:38] with even though we had, um, built even though we had started ingesting data from multiple different HR data sources into, um, our Azure platform. I was struggling as an HR data officer on how to govern that. So, we actually through one of our
[05:54] consulting partners, we had migrate we migrated off from ADLS to implementing Unity Catalog to really govern all of our data. And so, one of the things that I like to tell people is that in our data platform for people
[06:11] analytics, if it isn't cataloged in Unity Catalog, then it's to me, I always like to say it doesn't exist. Right? So, I actually go through and I say to my engineers, "Please go find those files that people just tucked away in ADLS. And if there are any, please make sure
[06:26] that we catalog it so that we have that full visibility into all of the data on our data platform." So, so in in about 2024, one of the ask that came along in terms of building out this data foundation was
[06:42] we had spent a lot of time at that point in building what I called reusable data sets. So, a lot of our data that we were bringing into Azure and processing through Databricks, we were finding that it was the same data that we were actually processing.
[06:57] So, we implemented what we called reusable data sets. However, what I found was it wasn't really stitched together or unified in a common data model. So, in order for us to to move towards being AI ready and moving
[07:14] ourselves towards a conversational AI experience, I'd said to myself, "How can we actually make that shift?" So, all from with all of the data that we had, we wanted to make sure that we had along the employee life cycle or the associate life cycle, which existed in
[07:31] reusable data sets, how did how do we actually normalize that into a data structure from hire to retire so that it's all governed in one place, and then we can also build on top of that with conversational AI. So, that was part of the ask. The second was um
[07:48] on the subject of conversational AI, how do we enable natural language interaction for our end users? So, if you think about your um your Databricks footprint, which, you know, I I know that um Genie 1 is going to radically change all of this, but today in your organization,
[08:06] at least in our organization, the interaction with the data on Databricks is really with our data and analytics teams, not so much our end users or our stakeholders out in the business. But, we really wanted to um last year bring that change about, which is how do we
[08:24] offer that now it's a Genie 1 experience, but back then last year, it was how do we bring that chat GPT-like experience to our end users? And then the third part of this was scalability. So, working for a large
[08:39] enterprise such as Mars um and dealing with the data uh volume that we have, as well as all of the different acquisitions that the enterprise is doing at any point in time, it's how do you when you when you have an acquisition or when you acquire a new
[08:54] data set, how do you actually bring it into a governed data structure at scale and continuously and repeatedly and stitch that together with what you already have? So, that was also the challenge at the time. And then last but not least is obviously governance of the
[09:10] data. So, um one of the great things I love about Unity Catalog is the lineage. So, you can always basically say, "Okay, where did this data come from? How was it transformed along the way? And then where is it ultimately landing in terms
[09:25] of your use cases um that you've built on your data platform?" In addition, um we also built out a lot of our data quality um using Databricks notebooks um to actually look through and tell us what is the quality of our data because
[09:42] ultimately for people analytics, you know, your stakeholders are relying on that data and those insights that they're deriving. So, it's really important to be able to show them the quality of the data. So, that that was the ask. Um, but when you think about, um, what we were trying to achieve, last
[09:59] year I was talking about it a lot as the associate 360. And when I went around talking to people about what that meant, the I found the best analogy was to tell them about the customer 360 with marketing. So, how do
[10:14] you have that transversal view of your employee or your associates across their entire life cycle with all of the different components of their data. And in the HR space, when we talk about all of the different components, we're talking about all of the different
[10:30] domains in there such as recruiting, such as performance, such as comp compensation, as well as the core HR data, as well as any development or any learning. So, how do we bring that all together into a people-first data foundation that, um, we can really use
[10:47] for AI. Perfect. Easy peasy, right? Yeah. But yeah, I know I think that was a big aha moment when Rachel and we and the team were discussing that
[11:02] how to bring this to life, right? And I think that analog that we hit upon as Rachel said was hey, you know, I think we do customer 360 all the time. How about we kind of flip that, look inward, and kind of see that how I can kind of do the same thing for
[11:18] our own associates, right? And that really was the genesis of the solution. And the three problems that we were looking to solve. The first one was how I can stitch the whole journey from hire to retire for the associate. And I should be able to kind of get a view of
[11:34] any of those slices at any point in time. You know, so if I want to know today that 2 years ago what was Rachel's job profile, what kind of, you know, like certifications they did, you know, how she was feeling at that time, all of
[11:49] that I should be able to, you know, like extract and get a very easy answer for that. So that is the first one that we are trying to solve. The second problem that we are trying to solve is how do I design a system or a model, if you will, that's not only used by people, human
[12:07] like us, but also by agents. So how do you design the data model for the agents? Because, you know, we want it it to be a self-serving model where Genie can kind of engage and then people can engage with the Genie and, you know, like kind of get those kind of answers at their fingertips. So that is the second thing that we are trying to solve.
[12:23] The third thing that we are trying to solve was how do I ensure that it is super safe and super super secure and super private to you. You would think what is HR data we are talking, right? It has lot of sensitive information, people's compensation, their, you know,
[12:39] like ratings, all kind of things. You definitely don't want the situation where like say something gets exposed in a wrong way, right? Somebody sees something that they are not supposed to see. There are a lot of those security and the, you know, privacy implications. So how do you kind of bake that into the solution? So that
[12:55] is the third problem that we are trying to solve. And those three kind of the areas that we kind of focused on and that's what I'll talk a little bit about is that how we kind of solve that using Genie and the Unity Catalog as part of the solution. So as Rachel mentioned, we definitely followed the medallion
[13:11] architecture and some of that was already there right in the ecosystem, right? However, what was missing was hey, you know, if I have to kind of have this solution that stitches all together, I need to have a holistic view, a 360 view of the associate and I should be able to kind of have the agents very easily engage with that data
[13:27] set and being able to kind of find the answers and quickly kind of like without hallucinating get to the core of that answer. So, we kind of built some additional consumption views on top of that straight regional data sets, right? You okay? Need water?
[13:44] Uh Sure. There's water on the right side if you step out, you know. Uh So, so you know that that that was the consumption layer that we built, right? It was set of the common data model that we built, you know. Now, why did we build that common data model? That's another thing I want to highlight is because we wanted to also
[14:00] have a level of abstraction so that it is getting data from different sources, right? It's getting data from from your learning system. It's getting data from your you know like human resource management system. It's getting data from even some of the unstructured data, right? Uh surveys and
[14:16] things like that, right? So, if I have to stitch it all together, I want to kind of build a solution so that if tomorrow some of those changes, right? Let's say for example, I move from say SuccessFactors to some other solution, right? Uh or for survey I want to move from right say you know one salt to another salt. I should still be able to
[14:32] kind of right like interact with this layer of abstraction and being able to get the answers and trust those answers, right? So, that was the kind of genesis for this common data model that we built. And then on top of that we built this semantic views. So, semantic views is where it was very important to give the context to all these data assets that
[14:49] we're building, right? Um and that was our salt to ensure that there is no hallucination, right? by the you know agents at their kind of right like engaging with the data. And there's like a two levels. One is obviously the context for the data itself, uh there are lot lot of internal
[15:05] technology terminology terms. Each company has it, right? So, for example, in Mars we have a segment called MGS, you know. Uh now, if I tell the agent agent comes across MGS, by itself it won't know what it is, right? Unless you can give that context, right? So, there's a context to the data
[15:20] and there's a context to the business itself, you know. Uh so, those kind of things we had kind of bake that into the solution. And that's where we kind of used a couple of, you know, like solvers. And I can talk a little bit later. One of the solvers giving the, you know, like instruction
[15:35] to Genie, so we use the Genie instructions to kind of make sure that it's being kind of get some context there. We also built some context-centric tables, where we kind of captured a lot of business context. And you know, for example, right? I think when you're
[15:50] getting all these data sources, like employee ID, location ID, etc. What that really means, you know, and having those simple easy to understand right views is what we created. And then we had the those agent spaces that we created on top. And those agent spaces were tied to the
[16:07] personas. So, for example, if I'm a learning manager, I could have a different kind of questions that I want to interact with the data for. If I'm a recruiter, I'll have a different kind of needs and different kind of questions and different kind of data that I want to interact with, right? So, we created these dedicated
[16:23] Genie spaces for those personas. And we kind of created those guardrails. Now, these guardrails are very important so that learning manager does not stumble upon certain things or access certain things that they are not supposed to access, right? So, that is where the whole privacy and security is
[16:39] very very important. So, the next thing that we did was on the Unity Catalog side, we leveraged that kind of build a lot of governance constructs and a lot of security constructs. So, that right only the right people can access the right data. And that's all kind of right fueled by the Unity Catalog here.
[16:56] Now, all of that when we kind of brought together, it actually created an ecosystem where as I engage with the data, I will have the My Genie space. Through that I'll interact. We also kind of build a web layer you know, on top of that, right? So, that is a good user
[17:11] experience, right? But then eventually it kind of interact with a particular Genie space, ask the question. And towards the end of the presentation we show you a quick demo, right? Like to kind of give you a sense of how that works out. Uh and at that time it will kind of understand the business context of the question. It will then kind of figure
[17:26] out, okay, which consumption views I have to kind of access and kind of get back the answer, and you can kind of go through that thread, right? Like you know, it retains the context, it maintains the context, and then it kind of it gets you to the point from the data to the insights. Can we go to the next one?
[17:42] So, one thing that we struggled with, if you will, and I want to kind of just highlight that, that as you think of associate 360, where I can have like this whole stitching that we talked about, how do we I make it easy to kind of
[17:57] understand that whole associate journey, right? Uh so, because you could have the same associate at different points in time, they could have different levels, they could have different feedback, they could have different training, certification, etc. So, how do we kind of tie it all together? So, one of the solve that we
[18:12] had was the construct of a link table. And the idea of a link table was really that, you know, I think we can have like a time-phased view, right? Of the associate journey, and we can kind of capture all the surrogate keys there, right? Slice by the time, and it is kind of goes as low as the day level, right?
[18:27] So, that I can kind of have the lowest granularity the day level to know what was happening, and it can kind of obviously from there can roll up to the week, the periods, the months, years, things like that, right? And that is kind of crucial to the solve that we have, because now as the query starts,
[18:44] Genie first comes to the link table, right? Based on the nature of the question, and then kind of from there it gets guided to the, right? Like set of other tables, right? The surrogate tables, if you will, and kind of then gets surface the answer. So, that was crucial in kind of bringing down A, the hallucination, and making sure that it
[19:00] can kind of it get to the right data really quickly, because earlier, right? We were struggling with some of the speed, the performance. I'm pretty sure you would have seen that similarly when you're engaging Genie that, you know, sometimes performance could be a little bit of a uh roadblock. This was very, very helpful in kind of, you know, improving the performance as well as the accuracy.
[19:17] I I just want to add a build here cuz um I'm I'm hoping this is resonating with the audience. You know, this type of data work is pretty detailed and pretty meticulous because when you think about just bringing in ingesting the data into your data platform and bringing that through your
[19:34] bronze, silver, gold medallion architecture, what happens over time is um do you really have the right data model that is coming to life with all of that data that you are ingesting?
[19:49] Are you Do you have a normalized data model where that allows you to not actually replicate the data over and over and over again so that you're going to get actually different interpretations of the data? So, the the large chunk of work that we
[20:05] did here was we actually had to break down, you know, what we had built from a reasonable data set standpoint. We had to architect what that data model looked like, and then we had to basically take the data and transform it or re-ingest it, and then put it into this construct.
[20:22] What I found along the way was, even though we had done that and it was looking super clean, we actually had to figure out what would be important for that conversational AI layer as we went along. So, as we were doing this, we
[20:37] were using Genie along the way to interrogate it to see if we were getting the results that ultimately we were looking for. And that meant that we actually had to bring the link table in because when we first actually put Genie on top of it, then
[20:53] this was last summer, it was we weren't getting the results we were looking for. So, that's that's what I would say one of the lessons that we learned is, if you are building out a common data a like this so that you can have clean quality data for your use cases, you
[21:10] actually want to put the AI on top of it as you would go go along to make sure that you're getting the results that you're looking for. And so it it all boils down to context. You've heard about it, you know, for 2 days now as part of the keynotes, how important context is. And this is This is context,
[21:28] you know, that we're looking at. And it's about getting it right so that you can get the the insights that you're looking for. So I can't emphasize how important, you know, this work is. So a lot of um I I see at the conferences is the end result, you know, and a lot of
[21:44] companies talking about how they're leveraging um Genie like ourselves, but um look to your data teams. You know, look to your data teams that are doing this fundamental work to actually make the data AI ready so that you can get to
[22:00] what is being showcased, you know, what as as right, as the wonderful output. Yeah. Yeah. Go to the next one. You want to cover that? Or So So yeah, so this is um
[22:15] you know, this is uh Genie and we had called it Associate 360 because the backbone of that is what we've built. And this is just illustrative. You've You've seen it and we're actually going to show you a demo. But basically how it works is you can interact with it like a chatbot, but I
[22:32] know it's more than just a chatbot and it's evolving very quickly. But last year, you know, that we were trying to I would say mimic or replicate that experience that people have in their personal lives when they use a chat GPT um or a Claude. And so it allows um the
[22:49] end user to ask in natural language, you know, questions of the data. Um it's querying This is the important part and I know I just talked about that in terms of the data the work that the data teams are doing, but it's querying governed views. So, um last year we were actually
[23:06] spending a lot of time looking at the SQL that was being generated by Genie to make sure that it was referencing the right tables. Um we also had to create some views off of the tables to make sure because the SQL wasn't able to handle it. So, we said the more
[23:21] efficient way was to create views against that. Um and then in addition last year we actually had to um bundle it with a visualization agent because at the time the visualizations that um were being produced um were uh not quite what
[23:38] we would have liked from a user experience. So, we bundled it with a separate visualization agent at the time. And there again context comes into play. You have You have to give instructions. You have to give instructions to make sure that when you're um asking the question and the
[23:55] and the user's asking, "Can you produce a chart for this?" that you're giving it the right context and the right instruction to give something that's actually meaningful for the user. And then the last part is summarization. So, um when it comes to insights, you know, that that that narrative that you
[24:12] give to the user is extremely important. So, again last year when we were prototyping this, we actually used a summarization agent that gave that provided a text narrative to the end user to bring it more to life.
[24:32] Okay. So, what is the the from and the to? What is the before and the after here with what we did? So, um I neglected to mention, you know, that how many dashboards we had or still have. I know um um I'm hopefully speaking um to many people in this world that will
[24:47] agree with me, but many users still want dashboards and will still ask you for dashboards. However, part of the vision here was shifting away from that. You heard me say conversational AI many, many times. So, um we still have dashboards today and that's fine. You
[25:03] know, people will always want that and so being able to have a tool that that UK will answer the question, but maybe the user's going to say, "Well, okay, can you create me like a little quick dashboard so that I can always reference it?" That's all possible today. But,
[25:18] part of the before is all of our products, which are dashboards, had all been siloed, looking at very specific um data sets for a specific domain. So, with Associate 360, because it cuts across and unifies all of the different
[25:36] subdomains of data, that it actually obviates now this fragmentation that you have in the dashboards. Now, users our users can actually ask questions that would involve not just, for example, your recruiting or your talent acquisition data, but it would also
[25:52] involve like learning data or performance data. Um in addition, um this is you know, when you when you ask questions about metrics, you always want, hopefully, the same answer returned uh at all times. So, that again is part of the
[26:09] foundational aspects of developing a common data model. It becomes then your definitive data source. And so, any metrics and calculations you put on top of that will all produce, hopefully, the same results. Um but then, you know, that
[26:24] will also get me into the metrics governance, which is a whole separate uh conversation. Uh also, uh being in an environment like Mars, where we are um engaged in acquisitions, and I'm sure we'll be doing more acquisitions in the future, goes back again to how do you onboard
[26:41] the data, acquire data in a way, do we have a structure in place such that any net new data can come in and have a place, a structured place to go into? So, this really allows us to accelerate that work once we actually get to it.
[26:56] And then the last part is analyst tickets to to answer questions. So, in our people analytics team, we have the role of business translators. And so, they were relying previously on a lot of data analysts on our teams to actually
[27:12] get answers out of the data, but now they're able to just actually ask Genie those same questions and it gets them the results like super fast and we don't have to have an analyst spending, you know, days looking into it because now you can actually just ask Genie.
[27:32] All right. So, um you know, this this is all about again getting to workforce planning at the speed of a question. I won't spend a lot of time on this because we actually want to demo to you what Genie looks like, but these are just illustrative examples where now with putting Genie or
[27:48] conversational AI on top of what we've already built, these are type of the These are the types of questions that Genie can actually handle. So, one of the things that I'll talk about later on is the visualizations. So, the bar charts, the graphs. And so, again, that's all
[28:05] running off of our underlying common common data model view. Um and then also there's um the aspect of an org hierarchy. So, I'm sure in every one of your organizations, you've got what's called an organization hierarchy. So, that's basically, you
[28:21] know, like divisions, units, departments, etc. So, building that out in the common data model, now you can actually ask questions of different levels of the organization and you can get to that um with that underlying data. And then
[28:37] the last one is um using unstructured data. So, um you can put unstructured data, you know, in Unity Catalog as you all know in in your data lakes and your data platform in Databricks. Um And so this is aspirational for us as we're in
[28:52] the middle of an acquisition is to be able to think about what does engagement look like for newly acquired entities. But all of that is now possible. And also I know that here we talk about the summarization agent again which I talked about, but I I presume that that will
[29:08] all nicely be baked into Genie 1. Yeah. Okay, so plans for the future. This is the last slide before we actually get into a demo. So you can see that we've built a foundation, the CDM.
[29:24] We also have Unity Catalog. We have Genie as well as previously last year summarization agents, but you know, now now it's all baked in hopefully. But where I'd like to talk about in terms of the road map is going back to
[29:39] the visualization. So one of the things that I want to make sure that we get right for our end users are those visualizations. And so I really do think that boils down to the testing that we'll do with Genie to see what happens with the
[29:57] visualizations as they're produced. Because you want that experience to be good for your end users. You want them to get the right visualizations that actually are meaningful for them for their insights. So that's one thing that I think we'll be spending a lot of
[30:12] time on to make sure are there instructions that we can give it that are The next one is personalized so persona based. I think it's really important as you're thinking through the experience for your end users using Genie is to think about the personas. Think about the roles that are actually
[30:28] using it. So I'll give you an example. So we have HR business partners that we expect will be using this a lot as well as finance business partners. And so how do you fine tune that experience with Genie to be able to get them what's relevant and what's meaningful for them.
[30:45] So, I'm looking forward to really digging in deep and figuring out how we can actually personalize that um that interaction that they'll have with Genie right off the bat so that um they get a positive experience from the get-go cuz we all know that um
[31:02] whenever we release any AI tool, it is very much about adoption and change management um and not just about the technology. And then last but not least is um going to predictive as well as prescriptive um um insights. Now, here I will just offer
[31:19] the caution, you know, as HR data officer, right? Always work with your legal department in terms of um predictive insights or prescriptive insights, especially in the area of HR because any um decisions that are made um with, you know, any insights that are
[31:35] produced um really should be um done with caution in partnership with legal teams. So, that's but it is something that we're looking forward to because um we always talk about actionable insights, right? And this this is just not just about the insights but what actions
[31:51] could you take, possibly take um with with the analysis that you're getting. Okay? Time for demo? Yeah. All right. Okay. So, we'll do a quick demo. Just uh one thing I wanted to kind of call out is that this is based on the synthetic data
[32:08] uh for of course all these reasons, right? It's you know, like we want to kind of be very very careful with these data assets and all. But, it will give you a flavor of how the solution, right? Like uh kind of thinks through. Let me start from the beginning and I'll kind of be pausing in between just kind of give you a sense of what it we are
[32:24] trying to do here. So, So, the first question I'm going to ask it. It moved a little bit fast. But, it is a very generic question. You're saying, "Tell me about attrition." You know? So, as you can understand, that's not a
[32:40] very detailed question, right? It could mean different things. But, what I want to highlight is that there is enough context that's baked into the solution, that's being able to kind of siphon through that question and kind of understand based on the personalized space, right? That what it could mean.
[32:56] So, let's go through it. So, it's first thing understanding the question, finding the right data. And then it starts to do the analysis. And it says that, "Hey, you know, I think I believe that you're asking about the attrition rate." And it is talking about, you know, how it has ranged from, you know, X% to Y%
[33:13] between 2018 and 2024. Now, some of you might be thinking, "Okay, why 2018 to 2024?" That is the duration for which we have loaded the data, you know, just so that you know, right? Uh so, because it it was not being specified what time frame and all, so it kind of took that whole time frame as a starting point,
[33:28] right? Uh now, it also kind of goes beyond that. It says, "Hey, the highest attrition rate was 9.6% in 2022, while the lowest was 7.0% in 2024." So, it is based on the prior context and the learnings, it knows that those are kind of the things that generally I should be
[33:44] associating with attrition, because otherwise it is not really going to be a lot more meaningful, right? And then it's also going beyond that. It's saying, "Hey, you know, I'm also seeing that the attrition rate is generally decreasing after 2022 with a notable drop in 2023 and 2024." So, within that summary, it has gone through a lot of
[33:59] those, you know, gyrations and being able to kind of give you a lot of information. Of course, then follows that with the tabular forms, so you can kind of have those details and all. And you can then kind of start, "Okay, I have a decent understanding of what's going on. But, now I want to kind of go deeper, right?" And I'm saying, "Okay,
[34:15] calculate the average attrition from that window, 2018-2024, and give me the countries where it is higher than average, you know, in 2024." So, now I'm trying to get a little bit more specific. Uh so, obviously it interprets that question. And then it says that, "Hey, you know, I think these are the two countries, United States and
[34:32] China, where I see, you know, like a higher attrition rate than the average in 2024. So, US is 21%, China is 9.6%, right? Um and then you say, "Okay, this is good. I can see that, okay, these have higher, but is
[34:47] that a Like can I know more, right?" So, now I'm going to ask the next question that, "Hey, you know, tell me what percentage of that is really high performers?" Because I think, obviously, if I'm losing more of high performers, then that is not as good a thing. Uh so, that question kind
[35:03] of start to get me towards some insights. So, it start to say, "Okay, you know, I think in the percentage of attrition made up of high performers is 5% China and 5.3% US." This means that the small proportion of employees leaving both countries are considered high performers, right? And it goes beyond that. It say, "Just
[35:18] wait for that question." It say, "Right, would you prefer Just go back a little bit. Would you like to see the attrition percentage for high performers using different performance So, why that question, right? Because sometimes it can see that, "Okay, are you looking for
[35:35] a certain way of calculating it, right? Are you looking for average across those window or do you want to kind of do it number by year and then kind of like look at the data?" So, that is the kind of the intelligence that it has. So, it's kind of like looking for that user, you know, uh response before kind of proceeding further.
[35:51] And now I'm going to ask the next question, which is From 2018 to 2024, what percentage of company attritions are high performers, right? Uh so, why I asked that question? You know, I wanted you to compare, you know, for US and China, is it higher or lower than the company level? And it kind of
[36:07] goes into that and it says that, you know, I think let me just get the answer here. So, what it's saying is that, you know, I think the overall uh the percentage is, you know, like better in US and China compared to your, you know, overall company level percentage. So,
[36:24] while it is higher attrition in those two markets, but you know your most of those attrition is not really that of the high performers. Now, that's your insight. What you want to do with it, that's the action that you can take. But, we just want to give you a sense of what are the
[36:40] different levels of uh you know questioning that you can do, right? You we start with a very abstract question. It figured out the context. It start kind of going deep and it can kind of work with you. Where we are, as Rachel mentioned, is that still it is kind of hovering around the what questions. It's
[36:55] not getting to the why question, like why attrition is like say high in this market and all, right? And that is the next level where we want to get to and we are so excited with all the announcement that's coming from Databricks this year, you know, with the whole thing around, you know, like ontology and you know, like the Genie one and the Genie agents. It
[37:12] actually create that space where I can kind of build that layer, right, of sophistication that I can now get to the why answers, sorry, why questions and get those answers and that will be the real true unlock. But, hopefully this gives you a sense of how the interaction and what is all the context. One of the
[37:27] thing that we have been saying is that you know, if you refer to that famous Google paper that where they talked about attention is all you need, we are saying that context is all you need, you know. So. All right, thank you so much. That was our presentation. We'll happy to take questions. Thank you so much.
Learn more about the Databricks Data and AI platform.
The information provided herein is for general informational purposes only and may not reflect the most current product capabilities or configurations.