Skip to main content

Accelerating analytics with Databricks AI/BI and metadata

Summary

  • After migrating from Snowflake to Databricks, Addepar's 12-person analytics team discovered that rich Unity Catalog metadata is the critical foundation for AI-powered analytics, with metadata strategy improvements lifting AI agent accuracy from 21% to 95–99%.
  • Addepar adopted a three-phase approach—planning, execution, and retrospective—using Databricks AI/BI, Genie, and metric views as a semantic layer to translate complex semantic models and deliver consistent answers across dashboards and agents.
  • Three metadata strategies—markdown links, column contracts, and narrative context—are highlighted as practical techniques that amplify agent accuracy and compress analytics work from weeks to minutes.

Accelerating analytics with Databricks AI/BI and metadata

Watch: Accelerating analytics with Databricks AI/BI and metadata
After migrating from Snowflake to Databricks, Addepar discovered that AI accuracy depends entirely on data foundations. this video reveals the practitioner's roadmap for accelerating speed-to-value by building analytics directly in Databricks using AI/BI, Genie, and metric views as a semantic layer.
Learn why rich Unity Catalog metadata is critical for AI-powered analytics, how metric views ensure consistent answers across dashboards and agents, and proven techniques for translating complex semantic models with AI Dev Kit and Genie Code. Addepar's three-phase approach, planning, execution, and retrospective recommendations, shows how metadata strategies like markdown links, column contracts, and narrative context amplify AI agent accuracy from 21% to 95-99%, transforming analytics from weeks of manual work to minutes of intelligent queries.
🤝

Chapters

FAQs

How did Addepar improve AI agent accuracy from 21% to 95–99%?

Addepar achieved this accuracy improvement by investing heavily in Unity Catalog metadata strategies including markdown links, column contracts, and narrative context. This video shows how enriching the data foundation with meaningful metadata dramatically amplifies AI agent accuracy for analytics queries on the Databricks Data and AI platform.

What are metric views and why are they important for AI-powered analytics?

Metric views serve as a semantic layer in Databricks that ensure consistent answers across dashboards, agents, and Genie queries by encoding business definitions and calculation logic in one place. Addepar adopted metric views early in their migration and found them critical for maintaining consistency as they expanded AI-powered analytics across over 10 business domains.

What is the Databricks AI Dev Kit and how did Addepar use it for semantic translation?

The Databricks AI Dev Kit provided a breakthrough for translating complex semantic models into Databricks-native constructs, reducing what would otherwise have been lengthy manual migration work. Addepar used it as part of their three-phase migration approach to accelerate the planning and execution of analytics built on the Databricks Data and AI platform.

What is Addepar and what scale of data does it manage?

Addepar is a global data and AI platform serving investment professionals across more than 60 markets worldwide, with over 1,400 client investment firms including private banks and single-family offices. The platform hosts over $9 trillion USD in assets and is supported by a team of over 3,200 full-time employees across nine global offices.

Full transcript

[00:08] All right, good morning everyone. Uh, just testing. Can you all hear me? Sounds good. All right, perfect. Uh, my name's Chris Krosinski. I'm a analytics manager at Adapar. Um, and I'm going to be talking about accelerating the speed to value of analytics with data bricks AI, BI, and agents.
[00:30] Uh just a reminder uh of course you probably heard it about a hundred times by now but you'll get a survey. Please comp uh complete it if you have a minute or two. Um so quickly just a little bit about Adapar if you're unfamiliar. Um Adapar is a global data and AI platform. Uh we
[00:48] empower investment professionals to turn complex financial information into actionable intelligence. Um we host over $9 trillion USD worth of assets um on the Adapar platform. We have over 1,400
[01:06] um client uh investment firms, private banks, single family offices, um firms of that nature. Uh we cover over 60 markets worldwide and we have over,200 full-time employees across nine offices
[01:22] uh globally. Um just a little bit about the team that I work on, analytics at Adapar. Uh we are a growing uh 12p person analytics team. Uh we're more of a centralized
[01:38] team um at Adapar. It's been an evolution probably how most analytics uh teams um evolve um especially within the last 5 10 years at most companies probably. Um it's a combination of analysts and engineers. Um again
[01:55] distributed globally. Um and a little bit about the scope of what we build. Um so we handle basically endtoend analytics. Um job ingestion jobs, data ingestion jobs, pipelines, um data
[02:11] transformations. Um, we build out the semantic layer that then supports dashboards. Um, and more so now agents, Genie, um, and of course, uh, good old spreadsheets. I don't think those are
[02:26] ever going to die. There's always a spreadsheet out there or two. Um, we support over 10 business domains. Um, some examples include product, go to market, finance, uh, multiple operations teams. Um so
[02:42] teams and departments of that nature. Um and some of the use cases uh that we support. Um it's pretty broad uh product uh usage, customer 360 analytics, revenue billing um analytics, customer support cases and conversational
[02:59] analytics uh more so recently uh with AI enabling that. So I'm going to quickly talk about some of the uh factors that um we considered
[03:14] on the analytics team at Adapar for why we chose to start building out analytics in data bricks AI uh BI. So this is we were just on the heels of completing our um migration ingestion of
[03:32] analytical data sets into data bricks. Um we wanted to try to quickly build out analytics where the data actually lives um which was on data bricks. It's not only us um putting data on data bricks but teams across the company. And that's
[03:47] one of the main benefits getting as much of the company's data as you can on a single platform. So you don't have to worry about uh multiple you know cloud warehouses to a degree um and deal with um issues that may come come with that.
[04:03] Um some of the other nice things um that we considered was the AIBI um products features. They inherit Unity catalog governance automatically. Um so you don't have to worry about kind of
[04:19] managing permissions and you know access to data and dashboards across the warehouse itself and then within a BI application or multiple tools afterwards. It's all kind of handled right there within data bricks. Um what was also very appealing um at the time
[04:36] that we're weighing these decisions was um Genie and the ability for agents to unlock self-service analytics um especially for non-technical users. I think you know getting the non-technical um persona involved in the process has
[04:52] kind of been a challenge through the years but we kind of saw some openings here the non-technicals finally had a chance to participate. Um also like that metric views at the time was a new feature coming out. So the existing analytics platform that
[05:08] we're using um it had a codified semantic layer. Um really like that because you know it's it avoids or reduces you know having different logic among different dashboards. Um the classic case is you know two people
[05:23] bring two different versions of a dashboard to the same meeting and instead of talking about you know the insights it's argument over whose dashboard is right um so semantic glare helps to reduce that um and data bricks had metric views coming out so we like
[05:39] that um but you know if if a team or you know down the road uh we would decide you know maybe we do need um an external tool uh we saw that the metric views are based in the YMO. So that kind of preserves some flexibility. You can use
[05:57] still use for you know some teams the um AIBI products but some teams might need some like uh niche features in another tool they uh uh sorry uh data bricks maintains uh external partner BI tools.
[06:12] Um so you could you know a team could subscribe to one of those tools and uh satisfy one of their niche uh use cases. So I'm going to talk about sort of like the three phases of setting up our
[06:31] analytics on uh data bricks. um it involved planning um actually executing and then I'll go over a little bit of a retro um for you know if we were to do this again um how would we go about it or recommend to a friend or you guys you know what to do
[06:47] next um so first thing we have to do is evaluate uh our BI or dashboard and semantic uh model landscape the lay of the land out there who's using what what semantic models and dashboards are not being used. Um, luckily we're tap we're
[07:05] able to tap into some system activity, identify, you know, the content and the models that are most being used. Um, but you can't just rely on that alone. You know, there might be a very critical dashboard that's used by one person, right? So, it's probably going to have
[07:20] some low low usage, but it's critical to the business. So, you do have to or we recommend definitely consulting with teams and users to ensure you know nothing's falling through the cracks there. Um so communicating with your your stakeholders very critical here um to ensure like a uh seamless planning
[07:37] process and that's going to influence your execution. Um we researched various uh semantic model translation options um again because in data bricks it's YAML based um and the platform we were using had
[07:53] its own proprietary uh semantic model language. So there was going to be a translation process. Um one option was to manually do this which was not appealing. Um but luckily you know what a time it' be alive with AI solutions
[08:10] being out there that were quickly researching you know could we do this with AI? Could it um help supercharge accelerate a translation process? Um and we did enough research where we felt confident um that could be so. Um and
[08:27] finally we organized and planned the migration of the semantic models and dashboards um into initially into a couple different phases or tanches about four to five groups of dashboards and tranches based on priority and how often
[08:42] they were used. Uh next the execution phase. Um, so we're I think we were pretty early adopters of metric views. Um, we started we started researching them as soon as they came out last fall.
[08:59] Um, the biggest challenge really was the most complex part was um the translation process. Um, so we were use we're trying out different agents. Um, at the time was data bricks assistant. It was okay.
[09:14] I didn't really have much knowledge about um our existing BI platform. It was very general, but um there's a lot of nuances and niche use cases where it just it wasn't doing a good job at that time. Um then we started to go, you
[09:30] know, try do some trial and error with Claude and we found ourselves going back and forth. Um data bricks assistant renamed itself with Genie Code. Um so definitely a lot of trial and error. We did work a lot with the data bricks um account team and the support they were
[09:46] providing to kind of help figure it out. Um because of that we decided to at least migrate an initial batch batch of dashboards um using just SQL behind the scenes for uh underlying logic of the widgets um just to at least start
[10:03] testing the dashboards out themselves while while we figured out the AI part. Um what did help was the data bricks people they turned us on to the um AI devkit. Um so that was really a winning solution. Um
[10:20] you know it had all the agent skills and knowledge right there in the dev kit and data bricks seems to maintain it with the latest features. So, if you haven't checked out the dev kit, heavily recommend doing so, especially if you plan or want to just um play around with
[10:36] possibly doing a migration yourself. Um we found that the agents were doing a better job um with or yeah doing a better job translating
[10:51] semantic models um for models that you were using data that we had metadata um in data bricks for um and our theory was you know it was able to translate and understand what that data was and just
[11:07] make a better um semantic model initially from the first go at it. Um for the models that were using data that's undefined, uh it was more of a back and forth conversation um before it was able to like uh get the
[11:22] semantics into a shape that we were satisfied with and producing the same results against um when comparing against our initial or current analytics platform. Um so that was a benchmark uh for us. We wanted to make sure that the results were the same. Um, if they were
[11:39] different, that meant there's probably a logic difference. Um, and that that was a problem. Um, so our analysis, sorry, our analysis and experience pointed to Genie Code um actually getting better, especially over the last
[11:56] two to three uh months um versus uh Claude. Um, and it's probably going to be, you know, your mileage may vary on which one's uh, better, but um, as you're going to keep seeing throughout the presentation, I'm going to be
[12:12] pointing to back to metadata and how like foundational it is in this work. Um, whatever learnings we captured for the agents being able to handle little nuances and those back and forth conversations, we're trying to capture
[12:27] that understanding and capture into instructions which um were turned into like agent skills for both Claude and Genie code. Um, and that ensured more like deterministic uh translation results and repeatable uh results.
[12:49] Um, so we're still going through the migration, but we've learned a lot. Um, so as a as of a current retro, um, here are some recommendations if you're thinking about doing this yourself. Um, first I would say focus on good data foundations. Um, that's always been true. I don't think AI is really a
[13:06] shortcut around that. Um, it's going to influence your results with AI 100%. Um so metadata times three absolutely here. Um I would say it's metadata is no longer a nice to have. I think you know
[13:21] maybe 10 years ago uh it was something that you do your data ingestion and metadata. We'll come back to it you know when we have time. Uh no longer. So I would say um having it there now is going to make your life easier on the
[13:36] next steps with um creating that semantic model, dashboards, agents. Um you're going to have a better easier time with it. You can use AI to help you generate metadata. Um you know that that is
[13:53] certainly an option. Um I would advise ex exploring you know um going into a notebook having a conversation with Genie code um asking it to understand the data and suggest some metadata. Um but of course um in
[14:09] capital there you must uh have a human review this. Um just from our results you know we we always had to edit something. uh we've never seen just 100% great metadata generated by I uh AI
[14:27] there's every company is going to have its own nuances with the data doesn't matter if it's like a well-known system like Jira or or Salesforce how you use Jira is different how than how the person sitting next to you might use Jira right so those nuances are going to
[14:43] impact uh things and those nuances are the metadata it's not just a definition of the field but like what's that field mean at your company? You know, um that's the that's the difference and could be the key maker.
[15:02] Um for metric views, um it was a new capability evolving quickly. Um again like what we liked there is and what we've always liked I mean metric views are a semantic layer. You define your metrics once and use it downstream across various content dashboards, genies, um AI agents, uh whether it's
[15:20] Genie code or even claude. Um AI will build again AI will build a better semantic model. Um if you have a solid data foundation,
[15:37] uh so this is just a quote from a very popular blog post that again you probably heard quoted. um 100 times um by now. Um it's the anthropic article and how they describe themselves being able to uh enable self-service analytics. So it's Claude or I would say any AI
[15:54] agent uh is exceptionally useful for closing the gap, right? For drafting column descriptions, proposing metrics, but again um there's a recommendation for curation and ownership um being managed by humans.
[16:12] So these three metadata strategies to consider they I put these here before the summit. I think ontologies may deprecate or make some of this obsolete but you might not go or jump into ontologies like right away. So
[16:27] still consider it. Um markdown links. Um if anyone's familiar with Obsidian, think about like backlinking between, you know, the various um things like tables, schemas, fields. Um you could use markdown in the comments section um
[16:44] in data bricks. Um and you could form those links or um include code or going into like column contracts. Um indicate what columns are a foreign key. Um what which columns to join on and
[17:00] what's the cardality going to be on that join. Um these are things that I don't think have I've typically seen as being part of metadata, but heavily consider it because um it's going to influence and make your results with AI. it's going to tell it um how to use that data
[17:18] um better. It's going to make your metadata richer. Um and third, narrative context. So, it's not just about what the field is, but also who typically uses that data, that table, that field, um
[17:34] what process it supports, what business department, um what questions it answers. um typically I haven't seen that in metadata especially in like docs of uh like a Salesforce or a Jira um but it's going to also be unique at your own
[17:49] company so that it's an opportunity to enrich with that extra um context. Uh just a quick poll um for a little intermission here. Who would say that at
[18:05] least 50% of their warehouse, what have you, is defined with metadata? I see a hand going up. I'm impressed. Very impressed. Another one. All right. How about 98%
[18:23] hands down? My hand is down. Um, how about 25%. Any more hands go up? uh a little bit more but a big majority is still down. So I mean I think that's
[18:40] an opportunity view that as glass half full. That is a big opportunity to add metadata make it richer and that's going to increase um AI accuracy which you've been hearing about through the rest of the summit especially in the keynotes. Right.
[19:02] Quickly going to now go over just uh skills for genie and AI. I'm going to quickly go through some quotes. Um definitely recommend reading the article for anyone who hasn't. Um so these are some stats that stood out to me. Without uh agent skills,
[19:17] um agents were only about 21% accurate. Adding skills and here I'm going to say skills within data bricks world that would include metadata. I would say that includes semantics and actual skills. um
[19:34] then instructions and other nuances that wouldn't neatly fit inside metadata or semantics itself. Uh sorry I went too fast. So went from accuracy went from 21% to 95 to 99%.
[19:50] um through metadata, skill, semantics. Um but it's not just set it and forget it necessarily either, right? Um they argue for maintenance. If you just leave it alone, in their experience, it went from 95 back down to 65 in a month. So
[20:10] this might be a new job title um coming up too. just maintaining skills, metadata, semantics, context. Uh so kind of going through skills quickly for sake of time. What and why?
[20:26] Um what's a skill? A set sets of uh procedural knowledge um for various data sources um to consult in what order, how to navigate ambiguous data, and what a finished analysis might even look like.
[20:43] Um, it it's kind of compliant with dry, don't repeat yourself. Um, skills should represent the best versions of your prompts and instructions. Um, you should share them with your team. Level up everyone's AI capabilities and it's a good way to promote standardization. You
[21:00] know, the same instructions said the same way, hopefully getting similar results. So some examples of you know where we've been using agent skills um certainly for
[21:16] codifying the nuances we've been discovering with translating semantic uh layer um root cause analysis you know that can go down deep you may have some procedures where you want a consistent experience
[21:31] um consistent dashboard design so green always means good red always means had um and data even like data set discovery. Next two slides quickly is just examples of a data set discovery brief um skill that we've been using. Um
[21:48] you kind of see like how we're saying, you know, these are the things the data or this is what we're trying to find out. What can the data answer? Um what can it answer? What data would benefit from being, you know, joined to this data to enable more answers?
[22:09] Um just a little bit further down. There's more to the skill. Um I'll try posting what I can to LinkedIn just so you know other others can copy um if they want. Um just some quick tips. You can ask you know Genie Code or Claude or any agent
[22:24] to create skills for you. I think even in like Atlassian I've seen that they they uh support skills with Robo. Have Robo create a skill within Atlassian. uh within Jira or Confluence um after you do a great or build a good thing if
[22:39] you want to make you know an instruction or like save those instructions but in a generic way um that's what we do we'll you know create some analysis want to create some instructions so we can do it again um so it just ask Claude or um
[22:54] Genie to kind of create generic form of those instructions um and again help standardize like different configurations of tools Um, at the end, assuming this deck will be able to be shared, just have some
[23:10] resources that could be helpful for you guys to look into further. Um, and that's it. Just final reminder to remember to uh complete the survey. Thank you.

Learn more about the Databricks Data and AI platform.

The information provided herein is for general informational purposes only and may not reflect the most current product capabilities or configurations.