Skip to main content

Genie at Scale: Atlassian's 0-to-95 Analytics Playbook

Summary

  • Atlassian scaled Databricks Genie from a pilot to a production capability serving tens of thousands of queries monthly by focusing on metadata quality as the primary driver of accuracy, ahead of model selection.
  • A hub-and-spoke scaling model routes natural language queries to domain-specific Genie spaces within Atlassian's existing user workflows, including Rovo and Confluence, reducing the need for users to learn new tools.
  • Atlassian's five-point playbook — anchored in pilot design, success criteria, context enrichment, few-shot prompting, and organizational scaling — provides a practical framework other teams can apply to their own Genie deployments.

Genie at Scale: Atlassian's 0-to-95 Analytics Playbook

Watch: Genie at Scale: Atlassian's 0-to-95 Analytics Playbook
Scaling self-serve analytics to hundreds of users without overwhelming data teams requires more than deploying an LLM. Atlassian shares its practical 0-to-95 playbook that took Databricks Genie from pilot to trusted production capability serving tens of thousands of queries monthly.
Walk through the complete journey from Data Lake 2.0 transformation to hub-and-spoke scaling models. Learn why metadata quality drives accuracy more than model selection, how to design Genie spaces as products with success criteria, how to route queries to the right domain-specific spaces within user workflows like Rovo and Confluence, and what challenges surfaced on the path to production. Discover the five-point playbook for building high-accuracy self-serve analytics that drives adoption at enterprise scale.
🤝

Chapters

FAQs

How did Atlassian scale Databricks Genie to production?

Atlassian scaled Genie from a scrappy pilot to a production capability used daily by hundreds of Atlassians by prioritizing metadata quality, designing domain-specific Genie spaces with clear success criteria, and integrating query routing directly into existing user workflows like Rovo and Confluence. This video shares the full five-point playbook they developed along the way.

Why does metadata quality matter more than model selection for Genie accuracy?

Atlassian found that the accuracy of Databricks Genie's text-to-SQL responses is more heavily influenced by the quality of table and column metadata — descriptions, business definitions, and examples — than by the choice of underlying language model. Investing in rich metadata creates the context Genie needs to generate correct queries without requiring users to know SQL.

What is Atlassian's hub-and-spoke model for Genie deployment?

The hub-and-spoke model organizes Genie spaces around specific data domains, with a central routing layer directing user queries to the most relevant domain space. This allows Atlassian to maintain accuracy by keeping each space focused while scaling across many business areas, and enables integration with products like Rovo so users can access Genie within their existing workflows.

What were the biggest challenges Atlassian faced taking Genie to production?

This video focuses on the hard realities of Genie adoption, including meeting an MVP accuracy bar, earning user trust, handling data complexity from mergers and acquisitions, and driving adoption at scale beyond the initial pilot cohort. Atlassian found that building trust through consistent accuracy and embedding Genie into existing workflows were as critical as the technical implementation.

Full transcript

[00:07] Good morning everyone. How are you all feeling today? Are you all enjoying the summit? Awesome. Awesome. All right. So, welcome. My name is Manav Trivedi. I am a senior product manager at Atlassian driving the Robo chat experiences. And with me today is Prakash Reddy who's the
[00:24] head of our data engineering platform and is driving a bunch of the initiatives that we're going to talk to you all about today. So, we've got some fans out here. Uh all right. So, um a few things to to keep in mind about this conversation. So, a lot
[00:41] of the AI talks will tell you all about things that are promising, things that are in the future, things that get you excited. We're going to look at the other way. We're going to tell you all things that were hard, things that we actually faced and how we got through it. We're going to tell you all about how we
[00:56] took Data Bricks Genie from a scrappy pilot into something that was at a production grade, things that hundreds of Atlassians use on a daily basis. And at the end of it, we're going to stick through to the Atlassian values of playing as a team and we're going to
[01:12] we're going to give you all the playbook of how you all can leverage our playbook and try to implement it at your companies uh the following week after the summit. So, uh let's get into it.
[01:27] So, here's a map of how we're going to spend some time together. We're going to tell you all a bit about Atlassian and its internal data, uh the products we have, the customers we serve, some of the challenges that we also face with our mergers and acquisitions. Then we're going to tell you all about
[01:42] why we really need self-serve analytics. Why do we see that the data assets that all of us have today don't seem to meet the needs that where where our users are. And then we're going to talk about some of the solutions about it. What are the things that we tried to do? Where did we
[01:58] fail at? What did we learn? And how Genie sort of came in and helped us with it. And then we're going to talk about the challenge after that. Cuz meeting an MVP bar is great, but driving adoption at scale and getting trust of these internal users is another thing. And so we're going to spend some time on that
[02:13] as well. And lastly, we're going to give you the entire playbook, so you can take it away from there. And I'm going to pass it off to Prakash, who's going to tell us about Atlassian and its data. Hey everyone. Nice to meet you all. Uh before we talk about AI-powered analytics and our journey uh with
[02:30] adopting Genie, I wanted to talk about the landscape and the environment that any AI agent has to operate in. Because how easy or difficult it is going to be to implement a self-serve analytics totally depends on that. So, let's talk about Atlassian's data.
[02:46] Atlassian is a 20-year-old company. We serve more than 300,000 paid customers ranging from startups, small medium businesses to large enterprises. Across a broad portfolio of products spanning software delivery, service management, collaboration, strategy, and AI. These products run across cloud,
[03:02] data center, and historically server deployments connected through a large ecosystem of first-party, second-party, and third-party integrations through marketplace. So, now here is what 20 years of growth actually looks like. In terms of data terms, a single customer journey can touch multiple
[03:18] products, multiple deployment models, billing systems, support systems, marketplace apps, and partner integrations. Each of these systems has its own identifiers, its own schemas, its own ownership boundaries, and semantic definitions of what usage, customer, account um and success means.
[03:36] If you ask five teams, "How many active users do you have?" and ask them to slice it by the products and by the deployment models, you'll probably get five different SQL queries with different nuances baked in there. So, the problem was not that we lacked the data, but the problem was we had
[03:52] enormous amounts of valuable data spread across multiple systems, multiple tools, and with many valid interpretations. And that's the complexity that any self-serve at Atlassian had to unfortunately navigate. All right. If you look at this image very closely, many of you can see
[04:09] uh the pain and challenge that I'm going to talk about right now. I'm guessing many of you have faced this in the past, and some of you face it today as well. What you're looking at is our legacy data lake, our world before rich governance, before curated data products, and before we could realistically think or even talk about
[04:25] self-serve analytics at scale. On the left, you see all the sources of data flowing into our data lake, first-party data from Atlassian suite of products, events, telemetry, behavioral data, etc. Second-party data such as firmographics, technographics, and which would typically tell us how customers
[04:40] use our products and what kind of deployments they have, if they were using data center or server behind the firewall. Uh and third-party enterprise SaaS vendors, there are many of them that we have. This is not the final list. Uh that enables our business functions to run behind the scenes within corporate
[04:56] Atlassian. Now, mergers and acquisitions particularly uh pose a very different set of unique challenges when we have to integrate a lot of 3P applications, their own data warehouses into our own lake. And um here is one thing, right? This is not unique to Atlassian. And if you've been
[05:13] at a company for a long time, I'm sure you all can relate to this, right? Every data lake foundation starts really clean, then an acquisition happens, then a reorg, then a new billing system that you migrate into, or a new platform migration, or a modernization project,
[05:28] right? And then teams who build temporary pipelines, temporary data products for point-in-time analysis, continues to stay for years to go, right? So, what happens is it makes the lake over a period of time um instead of being it a clean uh library of the data, it becomes a large
[05:44] warehouse with full of useful things, useful data to find, but hard to understand, hard to trust. So, what happened in our case? We had We saw a lot of duplicate data. We saw a lot of conflicting metrics, a lot of tribal knowledge, unclear ownership, and
[06:00] multiple versions of truth. With this state of data, we didn't feel comfortable putting any natural language on top of that. AI would amplify bad data, and we didn't want to put bad data in a shiny chart interface, right? So, before Genie could
[06:16] become a production grade analytics experience for us, we wanted to do the unglamorous but essential work. Transform this lake into a governed and owned and trusted data lake. Fast forward, what you see here is is our transformed and governed data lake.
[06:32] We call it as data lake 2.0, built on Databricks, governed through Unity Catalog, and designed from ground up to be AI ready. So, what changed? Standardized ingestion, standardized declarative pipelines, modularized ETL logic using DBT, and
[06:47] from there on, as you see, data flows into the bronze, silver, and gold medallion data architectures. So, the gold layer is our analytics-ready layer, where curated aggregate data layer that downstream teams can trust and and depend on. So, this transformation, the boring unglamorous multi-quarters work
[07:04] that we had spent our efforts in, built us good governance maturity. We built certified gold tables on top of that, and this is what set the stage for us to kind of ask the next set of questions. Now that we have a governed and trusted data lake, how can we put this in the
[07:20] hands of all the Atlassian employees at scale? And that is our story for our self-serve. Over to you, Manav. All right. So, Prakash just told us about the scale, thousands of tables, a governed lake house, years of investment. But let me tell you all a
[07:37] really interesting thing. What happens when an executive gets stuck because their dashboard doesn't answer their question. They send a Slack message to someone from the data team. And I'm guessing that this problem is not unique to Atlassian and a lot of you all face this today as well.
[07:53] It's not got to do with the data quality of what you've invested in. It's also not got to do with the quantity of data that you have. It's actually got to do with the access of data. What's happening here is that the platform that we've all invested in can
[08:10] only be as valuable as the number of people that it directly serves. And for a lot of us here today, that number is extremely tiny. And so when we look at it from uh a persona-based perspective, it basically breaks down to three key personas.
[08:26] It's our business user who, let's say they have an urgent question that needs to get uh answered. Um and so they would end up filing a ticket with the data team. Now, the data team is busy working on other things. They can't just drop things, scramble, and try to figure things out. And so it takes time.
[08:43] And this time leads to delays in decisions to be made. Sometimes, what happens is that decisions cannot wait and decisions then end up being made on instinct. Then it's also our data teams. It's not got to do with a skill problem. It's got
[08:59] to do with the productivity trap that a lot of our data teams are part of. Half of their time is spent in serving ad hoc requests that they get from multiple teams. And so the time that they could have spent working on strategic initiatives gets spent on working on
[09:15] these ad hoc tickets that they get. And lastly, it's our executives. We have hundreds of dashboards. Sometimes, we even build curated dashboards for a lot of initiatives that our executives care about. But dashboards quickly turn into answering yesterday's versions of
[09:31] today's questions. The ceiling the the the window that these dashboards are represented to be turned into a ceiling for generating any of these insights. And all of this starts to talk about cost. What is the status quo that we're trying to drive
[09:47] here and why do we all get impacted? So when we look at it from the business user as the executive as well, who basically has to wait for these insights that come in, we're basically wasting time. It's costing the company money. For the data teams, it's also costing a
[10:03] lot of money cuz we're scrambling them. And then they aren't able to spend time on things that are more important to the company as well. And this is why we feel that the LLMs were the answer for us. Self-serve analyst analytics was the reason why we went after it.
[10:19] But it also led to a problem in the landscape. When we look at it a couple of years ago, about 50 to 60% was an acceptable accuracy rate. That meant four out of the 10 questions that someone asks an LLM, it just got wrong. But you can't
[10:35] drive adoption, reliability, and scale at that rate. Since then, we see that our LLMs have gotten dramatically better. Databricks came up with Genie as well, which provided the context infrastructure that we were looking for. And so when we started looking into Databricks Genie, it wasn't something
[10:52] that we thought of it skeptically. It we looked at it hopefully. We knew some of these questions that we wanted to answer and so we looked at it in a way that can we actually adopt it in a way and scale it from there. And so before we just dive into Genie, I'm going to hand it off to Prakash who's going to tell us a bit about the
[11:08] the industry out here today. So the North Star was very simple. Anyone in the company, regardless of the SQL skills, regardless of the technical background, should be able to ask a data question in plain simple English and get a reliable, trustworthy answer without
[11:24] filing a ticket, without knowing what table to query, and without waiting. We already had the platform. We had done the modernization. We had all the data sitting in data bricks. Uh the gold tables were there. Uh governance is also there. What was missing was the last mile, at least
[11:39] that's what we thought. The conversational layer that made all of this data accessible to the 90% of the company that can write a SQL query. Now, we didn't just wake up one morning and said, "Let's just go use Genie." While we are busy transforming our lake to 2.0, we spent a better part of that
[11:56] year studying the industry where it was heading. We talked to other enterprises who were experimenting and we we read every piece of paper that we could find on ground or text to SQL uh at that point in time. And quite honestly, it was very overwhelming. The space was moving fast. The benchmarks were impressive on paper, and everyone
[12:12] had a demo that worked magically on the five tables that they had on just like a hello world program. But it was hard to differentiate between the hype versus the reality. The reference case that caught our attention was Query GPT that you can see here. Uh Uber published Query GPT in September of 2024, their internal uh
[12:27] text to SQL translation layer that operated on a scale of about 20,000 plus tables. That number mattered to us, but more importantly, the blog was very honest. They didn't They did didn't just tell us the wins that they had, but they openly talked about the challenges that they faced. For example, disambiguation
[12:44] when multiple tables could answer the same questions. Accuracy drops when you have complex joints, hallucinated column names, uh context window limitations, and and so forth. Um This honesty told us two things. One, this was genuinely a hard problem and not something that we
[12:59] could solve as a hack day project or or a just a prompt our way out of this. And second, the challenges that they uh described were exactly the challenges that we were facing as well on our own data lake. We also came across a couple of patterns that were emerging across the industry during the same period.
[13:15] Um pattern number one, intent-based uh text to SQL. Where LLMs parse the user's question, detects the intent what the user's asking for such as a metric, dimension, a time range, maps it to a pre-built SQL template uh based on intent mapping, and submits
[13:31] gives you the result, right? It seemed pretty fast, seemed deterministic, but broke the moment when someone asked a question that was not anticipated and was not mapped for. Pattern number two, rag-based SQL generation. This was the most popular one during that time. LLM parses the user question, embeds it,
[13:47] does a semantic search to find any and all relevant metadata that's stored in the rag or vector store, stuffed it all into a prompt, sent it to the LLM to generate a query. But, the quality of the answer was entirely determined by the quality of the rag that was maintained. If your metadata was thin or
[14:03] ambiguous, the same problem. You'll get a wrong SQL and wrong answer. Pattern number three is agentic SQL, the model that doesn't just generate one query, it plans, it generates, it validates, and retries step-by-step. If the SQL fails, it loops back. It definitely was more
[14:19] robust compared to the pattern one and pattern two, but slower and harder to debug when it went wrong. There are some open source that we were exploring as well during that time, like Ren AI, Vanna AI, and a couple of other open source solutions out there. Each had their own strengths and their own
[14:34] weaknesses, but they all hit the same ceiling. We also came across the Bird Bench benchmarks and found out that without heavy domain-specific and context set, accuracy plateaus at about 70%. Um and
[14:49] that's like a three out of 10 failure rate if you think about it, right? Across open source and emerging patterns, we quickly learned that model alone would not give us the results that we really wanted. We needed to invest even further on hardening our data and metadata around
[15:05] it. And treat context as a first-class citizen. And before we found the tool or we built the tool, um we we did what every engineers would do. We first thought of building our own solution. Uh and we built it ourselves. And we And
[15:20] what we learned from this attempt is exactly what made us ready for Genie when we finally saw it. So, with that, I'll hand it over to Manav, who was the product manager behind the product we built internally. Manav, over to you. Thanks, Prakash. All right. So, the vision was straightforward. Any
[15:35] Atlassian employee, regardless of their technical proficiency, regardless of how much they had information about SQL, should be able to ask a question in plain English and get an answer from the governed lakehouse. We called it Data 360. Um
[15:51] it's basically a 360-degree view of Atlassian's data that was open to any employee in the company. And we looked at four domains to start it off with. We looked at finance, growth, our internal people information, and ecosystem. And there were three specific criteria for
[16:06] why we even selected these domains. The first one was that these were some of the loudest domains as well. They had the highest number of ad-hoc tickets that our data teams were receiving, and we knew that it is an important thing for our business to solve. The second thing was that these were some of our most motivated stakeholders
[16:23] as well. They knew the problem that they were facing on their end, and they were passionate enough to explore something that was challenging. They knew that there were some risks involved, and it is a learning opportunity for all of us, but they were in it. And that is also a recommendation that we have to all of you all that if you're looking for
[16:38] internal stakeholders, find folks that are passionate and who would be partnering with you all along this journey out here. The third thing, as uh Prakash mentioned, it's got to do with the quality of the metadata that was available as well. These were some of the domains that had the richest set of metadatas that we had available, things
[16:55] that we knew we could build self-serve analytics on top of, and we didn't have to spend a lot of time trying to clean up or migrating any of the datasets. So, the way we built it was that um we leveraged the patterns one and two that Prakash just spoke to you about. We had a hybrid architecture of an intent-based
[17:13] approach mapped with some of the the the rag-based approaches as well. We had a vector DB that was mapped with a bunch of the the internal internal metadata. We had a few business definitions as well that we defined in few shot prompting. And so, let me give you an example of what happens when someone
[17:29] asks a question. Let's say someone asks a question, "What was the MAU for Jira in the last quarter? Can you compare it with the last two quarters as well?" What happens here is that as the query comes in, we would do a semantic match internally. And then, we would pass off
[17:44] this information to the LLM that would then generate the SQL query, call our data bricks, and then summarize the information that it got back. That was the way it worked and we built it on Atlassian's internal AI platform called Rovo. There were few things that we learned as
[18:00] part of it. Things that we learned that a lot of vendor demos couldn't explain it to us. Things that a lot of theoretical papers couldn't explain to us as well. The first thing was that metadata was extremely critical. It was more important than the underlying LLM underneath it as well and it it it was
[18:16] the primary factor that drove accuracy. And so, the richness of the column definitions that you have has got everything to do with how well the results you're going to get from Data Bricks Genie and any other LLM that you have as part of it. The second thing was that scaling became a real problem. The
[18:32] business definitions that our finance team had were very different from what our growth team had. And so, what we thought would turn into a scalable pattern actually turned into a repeatable problem that we had to keep copying and pasting and trying to figure things out and it was problems that we hadn't thought of. The third thing was
[18:48] that maintenance became brutal. What that meant is any change in our column metadata upstream broke the agent directly downstream. Results didn't come in or results were incorrect. And maintenance turned into someone's full-time job that we hadn't really accounted for. And that's where Genie
[19:04] came in. Genie provided us with the context infrastructure that we were looking for. And that's what Prakash is going to tell you all about uh how you adopted it. Okay. So, last year we were sitting right here where you are, quite literally. We were in the audience of
[19:20] the Databricks Data and AI Summit 2025, and when we saw Genie demonstrated on stage, something clicked. Not because it was flashy, but because we recognized our requirements in the product. So, we walked out of last year's summit and went all in. We committed engineering resources the following week
[19:37] because we weren't guessing anymore. We knew what we needed and we had just seen it. So, what made Genie different from what we had built ourselves, the data physicist you saw before? Three things. First, Genie spaces. This is a big one. Genie has a concept of curated governed
[19:52] spaces where you as a user can define the scope, the tables, the metric definitions, sample questions, instructions, few-shot prompting, you name it. All scoped to a specific use case. This helped us to keep the scope tight and really bounded. Second, uh
[20:08] native text to SQL. It's not a bolted-on integration. It's core to how Genie works. Because it sits directly on Databricks, it respects the Unity Catalog governance right out of the box. So, things like row-level security, column masking, um
[20:23] access policies, etc., all of that stuff is something you get out of free, right? So, users only see what they can they have entitlement for. Third, no separate infrastructure to maintain on top of of Genie, right? Genie focus pretty much on your engineering energy on just the the thing that gives
[20:38] you the most accuracy, right? Building the data context, building the business context, rather than the plumbing underneath that we were doing before. Now, we didn't just turn it on and hope that it would do magic for us, right? We designed the pilot like a product launch. We picked four domains: finance, growth,
[20:55] people, and ecosystem. And by the And the way we picked these domains are very deliberate. Highest amount of ad hoc data requests that we were getting, most engaged business stakeholders, and uh highest or best documented data, right? These four domains check the boxes and scored the highest in our minds. So, for
[21:11] each domain, we built dedicated Genie spaces on top of gold certified tables only, just the ones that had the richest metadata, and the ones that could answer the most relevant business questions. Well-scoped Genie spaces of less than 10 tables or under, and we kept it really
[21:27] focused. And critically, we defined success criteria before we launched. Um things like what accuracy rate did we need to hit? How much dau and mau would be a meaningful signal? How many data request tickets would need to disappear for us to kind of it's
[21:42] worth scaling for us? We wrote all these questions down and the answers down for those. We agreed on them with our stakeholders, and we treated each space as a product and a success bar that we held ourselves accountable for. Once we launched the pilot, we instrumented telemetry and measured
[21:57] everything. We obsessively gathered feedback, and for the first 60 days, we treated every user interaction as an input. And we learned uh a quite a bit about it. So, our learnings are in these three buckets, right? First, Genie gave us roughly 60 to 70% accuracy right out
[22:14] of the box with zero tuning, zero custom instructions, no sample queries, no incremental enrichment of metadata. Just pointed it to the tables and gave it a go. Then we layered on top of that our context work. We improved the descriptions of the tables better. We
[22:29] added more clearer uh metric definitions. We added few-shot prompting, expressions. Whatever Genie could offer us to put context in, we we improved upon that, right? That uh brought our accuracy to about 80-90% depending on the use cases and the kind of questions you were asking on them.
[22:46] And the second thing that surprised us was the non-technical users. Users who were using Genie as part of the pilot, they actually liked seeing the Genie's chain of thoughts and the SQL that was emitting at the end of the question. We assumed that our business users would want to kind of uh
[23:02] want it hidden or so, right? But, they didn't showing the generated query built trust because even if they couldn't read every line of the generated SQL query, they knew that this was a real system that was doing real work and that it wasn't a black box. And third, scope matters more
[23:19] than anything. Spaces with fewer than 10 tables focused on a single domain, single use case, atomic use cases consistently outperformed more ambiguous ones. So, this became our golden rule. Small focused beats large and comprehensive spaces.
[23:35] We did have some issues and struggles during our way, right? So, complex multi-step questions, the kind of questions that require several joins, complex analytical queries on top of that, or analytical functions definitely showed us poor
[23:50] results, poor accuracy scores. Model Some models some questions would get it right, sometimes would get it wrong, right? Session memory was a bit unreliable. You teach Genie something on a conversation like a preference or a correction, and in the next session all of that would vanish if you had not
[24:06] stored it as an instruction. This was frustrating for users because they expected Genie to learn. We also hit context limits. There's only so much instructions, so many table descriptions, so many sample queries you can add in a given Genie space before you fight with token window.
[24:22] That forced us to be ruthlessly selective about what context we wanted to add and what we wanted to remove. And lastly, our partnership with Databricks. This wasn't a typical partnership that you'd know, and I want to be direct about this, right? Because this part matters if you are evaluating
[24:37] Genie today. This is not a vendor relationship, it was a partnership and we had direct line to Databricks product teams. We filed bugs, they got prioritized. Our pilot was directly influencing the Genie product strategy, and the collaboration with our marketplace that
[24:53] we're going to talk about in the next slides, thanks to Databricks team, It made It just made us possible. So, the pilot worked and we had several domains live. Accuracy was in that 80-90% range. Users were interested and the trust was building. And then exactly
[25:08] what you'd predict happened. Every other team in the company wanted in. They wanted a Genie space, too. Scaling suddenly became bottleneck because we were a small team who were trying to scale this for the entire company. So, our success created demand that we couldn't scale with a lean and central
[25:24] team that we had. So, we chose to change the model that we were operating under. We started with a hub and spoke approach. The hub, sitting in the center, was our program enablement team, a small team. They owned five things: standards, tooling, governance, playbook, and benchmark quality.
[25:39] They didn't own the actual work, which the actual work was done by the spokes. They didn't own the Genie spaces for other teams. They owned the leverage that let other domain teams succeed on their own. Spokes, spokes were the domain teams themselves, just like growth, finance, the examples that you see here. Spokes owned the outcome and
[25:56] hub gave them the tools to succeed without becoming a bottleneck. Now, underneath all of these things, they built something that turned out to be just as important as this organizational change that we did. We identified champions within each domains who weren't not not necessarily power
[26:12] users or data users, but they were power users, right? They were the folks who were uh who became our internal evangelists. We ran show and tells every 2 weeks where teams shared what they were building, what they were working, what it what wasn't working, and every reusable pattern across the different domains, such as a good set of
[26:28] instruction, a clever sample query, uh a meta template that improved accuracy. Every of these things flowed back into our shared playbook that the hub maintained. So, the system got smarter for every domain that we onboarded into Genie.
[26:43] Now, the model solved for the organizational scaling problem, but we started to observe a new set of problem, a user experience problem. Because now we had about 10s and 15s and 50s of Genie spaces spread across the company, but users didn't know which Genie space to go and ask the question.
[27:00] And that friction was starting to kill the adoption. And what we did next, Manav is going to talk about that. Awesome. Thank you, folks. Um so, you know, this is an interesting problem. Imagine this use case. You build Genie, you try to drive adoption out there, and
[27:17] you have a bunch of these users who are like, "Hey, you've built five Genie spaces for my business domain, but which one do I talk to?" And so, the the the friction of selecting which Genie space to talk to was actually killing the magic of giving you the act- the the the data insights
[27:33] out there itself. For solving this, we looked at things internally. Atlassian has its AI-powered platform called Rovo, along with the Confluence infrastructure for Teamwork Graph. We have information for every Jira ticket that you use, the Confluence documents
[27:49] you're on, who you are within the company, what domain you're on, what Slack messages you're on, what Google Drive information you have, all of it. And so, we decided to work on this together. We mapped our Confluence infrastructure with Teamwork Graph, along with the agent infrastructure on Rovo, along and the conversation APIs of
[28:07] Databricks Genie. And what that meant was that users didn't even have to leave their work experience. They could just open up a Confluence document along with the Rovo agent and just ask the question right there itself. And that was the magic that we added.
[28:23] So, what would happen here is when a user would then ask a question to the Rovo agent, it would have the context about who they are, what documents they're on, what domains they're working on, and then it would effectively route that question to the appropriate Genie space. We had
[28:40] information about what are the descriptors for each of the spaces and the column data within it. And so, when a finance user would ask a question about revenue versus a growth product manager that asked a question about revenue, those two meant very different things. But it was on the agent to
[28:55] effectively route them to the appropriate genie space to answer that question. And then what we did was once we passed it the the query to the right space, it was all on genie through its conversation APIs that it actually fetched the information and provided it back to us. And that was the magic that
[29:10] we're talking about. That the user didn't even have to leave their space. Most AI tools expect the users to come to where the data is. We flipped it. We brought the data to where the users were working and made it easier for them to access these insights. And we didn't
[29:25] restrict this internally. The Databricks query runner agent is actually available on the Atlassian marketplace for you to leverage. And so if you folks have are in a company that is an Atlassian and a Databricks customer, you can get advantage of this
[29:42] out of the box. It's got the same architecture that I just spoke to you about. It's got Rovo, Teamwork Graph, and the conversation APIs of genie that you can get out of the box. You don't have to spend time building everything that we just spoke to you about out of the box. You can Sorry, you don't have
[29:57] to spend time building it yourself, but you can take advantage of it out of the box. And with that, I just want to focus on one specific domain, finance. And it's critical because the numbers that a finance analyst would fetch
[30:13] cannot the accuracy for that has to be critical. It cannot be different because it affects our bottom line numbers directly. And so let's look at what their journey was before genie. Let's say they had to do a bunch of analysis. They They would open up multiple SQL tables. They would open up
[30:29] a bunch of Excel sheets, do complex joins, and then if they had any question, they would have to file a ticket with someone from the data team, wait about it, collaborate with them, and then figure out what that answer is. With genie and with Rovo, this took a matter of seconds. They would open up
[30:45] the Confluence doc where they were writing the report on. They would just ask the question and get these insights immediately. But, there were three things that we again had to look at to make this successful. The first thing was that we had to be very deliberate about the type of spaces that we created. Finance as a domain is
[31:02] extremely large as Prakash has mentioned to you about, we had to be very tight vertically. So, there were multiple spaces that we had for like procurement and the revenue space and all of that and we had to be deliberate about it. The second thing was also that we had to ensure that the underlying metadata was
[31:17] was appropriate. It was acceptable and it was passing the standards that we had uh for it. The third thing was that we really needed some internal champions. We needed subject matter experts that were bought in and would that that would stress test test the the the the agent
[31:32] and the underlying infrastructure that we built. And so, what we did was that for any space that we were establishing, we had these champions on board it first. We had them take it to the reins. We had them break it and then we fixed those things and only then we drove adoption after that and and that's where
[31:48] we got the trust from. And so, with that, I want to leave you all with this question. If you're thinking about establishing self-serve analytics within your company, think about who these stakeholders are that you're trying to serve. What are the top questions that they keep asking your data teams on a Monday
[32:04] morning? And probably that's the set of things that you want to go after and try to solve in the first place. And with that, I'm going to hand it off to Prakash who's going to share the full playbook of what we're going to we spoke about today and take us home. All right, we're coming to the close to our session. So, all right, let me bring
[32:21] this all together. We've covered a lot of ground today. The complexity of data estate, governed foundation we built, the industry research, the home-grown attempt, the Genie pilot, the scaling model, the query runner, the marketplace launch. But, I don't want you to leave this room
[32:37] thinking this was a great Atlassian story. I want you to leave thinking I know exactly what to do next Monday morning. So, here's the playbook. Number one, treat your genie spaces as products and not experiments. We define success criteria, a named owner, a launch readiness bar, a feedback loop. If you
[32:53] wouldn't ship a product in your company without an acceptance criteria, don't ship a genie space without them, either. Number two, invest in metadata quality before you invest in anything else. This is a single's heavy highest leverage thing you can do to improve your accuracy more than model selection
[33:10] itself. More than prompt engineering, more than clever instructions, rich, precise, governed metadata, data like table descriptions, column descriptions, metric glossaries, business context, etc. This is the work that moves your accuracy from 70% to 95% depending on
[33:25] first-order or second-order or third-order questions that you ask to it. Three, build a benchmark evaluation loop from day one. Create a golden set questions with known answers validated by our data owners, business owners. Run it Run your space against it before launch and every other
[33:44] week after launch. Track your accuracy over the period of time. This is important. You cannot improve what you cannot measure and you cannot build trust without proof. Number four, scale your genie program with using a hub and spoke model if you
[33:59] don't already have one. The central team enables and governs the domain teams own and operate. A shared playbook that compounds. Do not let the central team to become the bottleneck. And the last one, embed your AI where the work happens and people work. This
[34:14] is really important, right? The best interface is no new interface. If you want adoption, don't ask people in your company to learn a new tool, open a new tab, or remember a new URL. Bring the capability where the people already do their work and meet the user where they
[34:30] are. Now, these five points will get you to production. But, I want you to leave with where we think this is going next. The pattern that we're seeing is a convergence on context layers as critical infrastructure for AI analytics.
[34:45] What do I mean by that? Uh today when you set up Genie space, you manually write instructions, you add the table descriptions, you create sample queries. It works for a scale of five to 50 Genie spaces across a distributed model. But, what happens when you have 100 or 200?
[35:01] Context is not static. It drifts. It lives in one tool, but it may not live in the other. And when a business rule changes, let's say a definition of active customers or active users change, you will have to manually find every Genie space and update it or keep it in sync, right? That is not sustainable and that's not
[35:18] scalable. And the industry knows it. So, you're seeing it in the emergence of semantic layers, ontologies, business glossaries. You saw today morning's keynote as well about it. Uh that are machine readable and open standards like NCIP that can transport and allow this context to go around different tools.
[35:34] At Atlassian, we are investing in something we call it as a context layer, um a context compiler of sorts, where a system where data owners, domain stewards can author, govern, certify, and activate the metadata and context across several surfaces, Genie being one of them. Our source of truth for
[35:50] context, which is governed the same way we govern data and tables today. And critically, this includes a feedback loop. When a user flags an answer as wrong, when uh the benchmark score drops, we bring this feedback loop back into the context layer and keep it up to date all the time, like a flywheel effect.
[36:06] Not static and continuously improving. That said, thank you. Thank you everyone for being here. A big thanks to Databricks team who has been amazing partners for us to get here where we are.

Learn more about the Databricks Data and AI platform.

The information provided herein is for general informational purposes only and may not reflect the most current product capabilities or configurations.