Skip to main content

Workday Sales Companion: Building 500k Tasks per Hour with Databricks Apps and Agents

Summary

  • Workday's Sales Companion scaled to 4,500+ active users and attributed over $400 million in incremental annual contract value in Q1 alone, built on the Databricks Data and AI platform integrating 59 data sources and 1.5 billion data points.
  • The platform processes 500,000+ tasks per hour using Databricks Apps, MLflow, and horizontal scaling, freeing data scientists from infrastructure management and enabling conversational AI experiences for sales, customer success, and finance teams.
  • Workday's agent architecture evolved from standalone ML models through LangGraph multi-agent orchestration to GraphRAG for long-horizon planning, with Unity Catalog providing governance and identity management throughout all generations.

Workday Sales Companion: Building 500k Tasks per Hour with Databricks Apps and Agents

Watch: Workday Sales Companion: Building 500k Tasks per Hour with Databricks Apps and Agents
Workday built a conversational AI platform using Databricks to deliver instant business insights directly to sales, customer success, and finance teams. By shifting from year-long custom builds to a standardized Databricks approach, Workday delivered three production AI applications in a single year while scaling Sales Companion to 4,500+ active users and attributed over $400 million in incremental annual contract value in Q1 alone.
Learn how Workday integrated 59 data sources through a metadata-driven architecture supporting both structured analytics agents and unstructured retrieval agents using GraphRAG for long-horizon planning. Databricks Apps, MLflow, Horizontal Scaling, and governance through Unity Catalog enabled conversational experiences at 500,000+ tasks per hour, freeing data scientists from infrastructure and allowing the business to extract 30 years of sales expertise without hiring consultants.
🤝

Chapters

FAQs

How does Workday's Sales Companion use Databricks for conversational AI?

Sales Companion is built on the Databricks Data and AI platform and integrates 59 data sources to give sales, customer success, and finance teams instant access to business insights through a conversational interface. The system processes over 500,000 tasks per hour using Databricks Apps, MLflow, and horizontal scaling, enabling real-time responses without data scientists managing infrastructure overhead.

What is GraphRAG and how does Workday use it for long-horizon planning?

GraphRAG is a retrieval-augmented generation technique that uses graph structures to connect entities and relationships across large document and data collections, enabling reasoning over complex interdependencies. Workday uses it in Sales Companion to support long-horizon planning queries that require connecting information across multiple data sources and reasoning over extended time horizons.

What revenue impact has Workday's Sales Companion AI delivered?

Workday attributed over $400 million in incremental annual contract value in a single quarter to Sales Companion, demonstrating measurable business ROI from the conversational AI platform. The application reached more than 4,500 active users, delivering sales insights and account intelligence at a scale that was not previously possible with custom-built tooling.

How did Workday shift from year-long custom AI builds to faster delivery?

By standardizing on the Databricks Data and AI platform, Workday moved from year-long custom build cycles to delivering three production AI applications in a single year. The shared infrastructure — including Databricks Apps, MLflow, horizontal scaling, and Unity Catalog governance — allowed teams to reuse components and focus on business logic rather than rebuilding platform engineering from scratch.

Full transcript

[00:08] All right, perfect timing as everyone trickles in here. Um I was asked to throw this in here, forward-looking statement. Obviously, uh if there's any lawyers in the room, I think we have to leave it up there for 10 seconds or so. Um so, feel free to read it. But, now that that's out of the way, um
[00:24] I'm thrilled to be at my first Databricks AI Summit um and talk about how we're leveraging this amazing platform at Workday. My name is Taylor Swett. I lead our our global customer operations, um technology, and architecture team at Workday, which is a really long way and
[00:40] fancy way of saying that my team manages all the tech and technology that touches either our field or our ops teams. Um we also oversee all of our go-to-market AI use cases, as well as our seller experience. Plus, probably 100 other things that my team does behind the scenes as well. Um
[00:56] but, I'm super lucky and grateful to work with some brilliant workmates uh like Tamel here on some amazing products. So, Tamel, you want to introduce yourself real quick? Hey, I'm Tamel. I'm the senior principal ML engineer data science and innovation team in Workday. So, we help Taylor make more sales, right? It's simple as that. Thank you.
[01:13] Awesome. So, quick question. Uh how many of you have either heard of Workday or hopefully will be putting uh whatever large expense reports through the platform this week. Uh I would imagine most of you maybe a lost receipt or two this weekend. Um but, for everyone that
[01:29] hasn't heard of Workday or doesn't use our platform, um it all comes back to one idea for us, which is leading our customers forever forward. What that really means is we're the AI platform that manages people, money, and agents for the world's leading organizations. Some of the most
[01:45] important things that that any company deals with. We cover over 65% of the Fortune 500. We have over 75 million users on our platform. And we process trillions of transactions a year, both agentic and just regular business process.
[02:01] Um and we're having agents help incredible customers like Databricks um reach their full potential and grow quickly. Um other customers like Salesforce able to build on our platform with Extend. And you see like a customer like Chipotle in the middle there gets back
[02:17] to kind of our bread and butter of HR and cutting their time to hire by 75% with some of our agents. But our most important mission as part of our chapter four, now that Aneel's back as our CEO, and it's really rethinking how the world works. And it kind of challenged us internally
[02:32] cuz as a company if we were going to do that, we kind of had to go back to the drawing board of how we work internally as well. So what does that look like in reality? Well, for a company our size turns out that's a multi-billion dollar challenge. Um and our go-to-market approach was
[02:48] rapidly changing over the past year couple of years. Um we've had 40 plus sales tools after cutting a bunch, adding a bunch of new ones. Uh there's 200 plus dashboards that exist in some way, shape or form. Uh 30,000 enablement contents out there
[03:04] after after deprecating a bunch of old stuff. 10,000 plus contents out there that are our customer facing. Uh and the most challenging thing for us was the acquisitions. Um eight companies were acquired over the last few years. And as we grew, that
[03:20] required new updates and changes to processes every week. Um every one of those things came with great intentions. They all kind of helped in their own way, kind of in their own kind of silo. But together, all that capability created a new kind of challenge for us,
[03:35] which was synthesizing it. And we had a person in the middle, you know, the rep, the CSM, whoever, trying to turn it all into the next sale or renewal themselves. And that turned out to be the billion-dollar opportunity for us and where we kind of have have gone on this journey.
[03:52] So, our solution was simple yet ambitious. How do we throw AI in the mix here and and kind of solve this at scale? Um and our product the hypothesis there was really what if we swapped out the human in the middle for an AI companion?
[04:08] Obviously, we're an HR company having humans and and people in and kind of human in the loop is really obviously what Workday has been known for and and that was a really important thing to us of how do we maximize the value we add to every workmate that we have? This was a long journey for us. It
[04:24] started with a a conversation two two and a half years ago with one of our SVPs saying, you know, as ChatGPT was becoming more popular and all that stuff like, hey, why don't we just make Workday GPT? And that was kind of the internal code name that that we started
[04:39] on this uh this mission. Obviously, branding didn't love that and we came up with with companion there, which I think summarizes it really well. We started small, to be honest. We started actually with a 200-person beta that we called sales companion was the first iteration of this. And it was
[04:56] only trained on our acquisitions. So, it was really good at answering kind of products on our our latest products that our reps were expected to sell the next day. Um and what happened was that worked really well. So, then I was like, all right, how do we throw more data at this thing
[05:11] and see if that works? We wound up dumping our entire enablement repository in there. That worked really well. Obviously, new issues that Cam will talk about in a minute, but that started to scale really well and effectively. The next logical step was let's throw
[05:26] more data at this thing and see what happens. So, we started bringing Salesforce and other platforms. I think I at right now I think there's 59 different data sources that we're bringing to companion. I think the really important thing we've learned on this journey is that it's not just throw a bunch of data in there and
[05:42] kind of hope for the best. We've curated this by hand. So, you can see, you know, obviously we're touching all the objects, all the Salesforce objects, everything that matters, but one that sticks out in particular is is external news. It's super important for the field. We tell everyone we know a lot about our
[05:58] business internally. We don't know as much about our customers, obviously, and that was always on the human to go figure out. But, we've curated all of this all of these data points, all 1.5 billion data points, where you know, my example is basically we don't bring in, you know, Target's got a Black Friday sale this week as relevant
[06:15] news for them. So, all of that's filtered out and it's all news related to people, capital, financial decisions there. Um the harder challenge for us is the business end of this equation was really how do we extract the business context
[06:30] and kind of persona engineering of the best of the best. You know, my team and I would go interview our top sellers. It's like, "What makes you really good at your job?" And And the responses we got were kind of I don't know. I'm just good at it, you know. It's like it was really hard to do that. Or if we think of our
[06:46] executives, our senior executives with, you know, decades of experience, how do we extract 30 years of gut experience hitting your number and how to have AI kind of interpret it and think like that. I think for us on the business side, again, it got back to
[07:02] nobody knows our business better than us and that's why we went on this journey of building it ourselves. As our CEO put it at the time, our his goal was I never want to hear let me get back to you again. That's the killer in any deal. It's the killer internally where somebody's got to go look something up. And that was our
[07:17] mission and and again, we've been on that mission and and I think doing a pretty good job the last couple years. Um unfortunately, I can't actually show you what our product looks like internally. Obviously, there's a lot of confidential work data in there. Um but, I did, obviously being an AI convention, burn as many Gemini tokens
[07:34] as I had access to making these videos. Uh and I want to give you a quick sense of what it looked like. I couldn't figure out how to get sound in it, so there is no sound. Um and you can probably tell which part Gemini did and which part I did. Um but it is one of kind of the most important use cases, I think, where it's
[07:50] like you know, reps, CSMs, whoever are kind of prepping for a call, prepping for a meeting, and I'll talk about exactly the use cases that they're using it for. But it gives them instant access, whether on site, whether it's prepping beforehand, whether it's strategizing and prioritizing their territory. And
[08:07] it's really built the dream rep, the dream CSM, the dream persona in a box, and it's and it's kind of taken the this field by storm. You know, with any tool, I'm sure you imagine there's a lot of adoption issues in a company our size. People like to do it their way.
[08:23] They have their own thoughts about things. This is the one tool that I think if we took away from people, we'd they'd be coming at us with pitchforks. And it's and it's been a natural adoption. Um what that actually looks like in reality, and obviously I had to hide some of the confidential parts of these
[08:39] screenshots, but what you saw in that video was our I think our, you know, most advanced agent, which is our account strategy agent. It's been years in the making of, okay, how do we go from we do a bunch of research to you to we synthesize it to how do we actually start to project forward on an account?
[08:57] And it was built for the entire account team. We have components of measuring their AI readiness. Um we have sentiment analysis on every contact that would matter on the deal. Um but it goes way beyond that. Um and it's not just for our sellers. Our CSMs
[09:12] have the ability to have companion kind of flag risk on their accounts and churn risk and stuff like that, like 18 months before a human would ever see it. And it's kind of that overarching tower on their accounts when they have too many things to do. Um the really interesting thing for us
[09:28] was how quickly it expanded to teams outside of even our control of of GCO there. Our marketing team's in there. Our product teams in there getting insights on on things that they never had the ability to do. And it's kind of it almost cannibalizing other
[09:44] spend for those teams. I can't tell you how many conversations I have where somebody was going to go spend a or go have a consulting company, you know, analyze customer sentiment and all that kind of stuff. And now our product team, including our our chief product officer, just goes in there and and asks about
[10:01] every call recorded or every email ever sent or or what are customers saying about XYZ and they're using that data to pivot narratives in real time. And it's again, it's it's kind of gone well beyond our wildest dreams in a great way and become like a core internal product for
[10:17] us. Um Which brings me to the most exciting part cuz, you know, that's where we get challenged. I'm sure you've all seen it, maybe experienced it personally, definitely read about it on LinkedIn. But the most challenging part with a lot of these AI initiatives is is they look great, they pass the eyeball test. A
[10:34] week later, nobody really asks questions and everyone kind of goes on to to do their thing. And for us, the biggest challenge was how do we justify the spend on this and and kind of really prove it and get the support we needed. Um We knew the companion story was so much bigger than the productivity gains and
[10:51] all the kind of standard stuff we have, which again, great part of the story. We're giving back almost 10 hours a week to the to the field here. But we knew the story was bigger and we knew it was driving revenue. As my team was kind of interviewing sellers, we always kind of heard those anecdotal comments of you know, CIO, uh I remember
[11:08] was like, you know, I've never had a vendor show up to a meeting more prepared. And it was like, you almost knew the questions we were going to ask before before we asked them. And obviously our reps, you know, prepared for that with Companion. But last quarter, I think is kind of where it pivoted for us and where Tamil
[11:24] and the team has done an amazing job of really tracing this. And if you look at the last quarter in Q1, it it's attributed to over $400 million in Ingram incremental ACV for us. We're seeing it across the entire deal cycle. Everything from flagging new areas of opportunities
[11:40] by over 10% moving deals faster by 14% and even post sale increasing our renewal up sales by almost 8%. The really big challenge for this is when you know, we throw big numbers like that in front of our executives. Obviously, it looks great. But then next question is like, you know,
[11:57] how real are those numbers, right? And we spent a lot of time iterating through this to make sure they were really actual attribute attributions where we filtered out the noise. So, if you know, you asked a question about Sonar, our AI product, but then sold them hired score. Like we don't count
[12:14] that deal. Like we we found something semantically in the system where Companion recommended something. We saw what happened in Salesforce in the next 48 hours and then kind of watch that deal progress. And for us, that's been a game-changer. We're excited to see where this product continues to go.
[12:31] Um and with that, I'd love to pass it to Tamil to to talk about how they built it. Thanks, Taylor. Awesome. Taylor going to walk through how the field uses and what's the value it brings to the
[12:46] ecosystem as a company. Now, we are going to walk through how did we arrive at this? For most of them, the agent talk started, okay, build a AI wrapper and let's take it from there. But for us, it really started like in 2022
[13:03] when we started building basic ML models. Like a propensity to sell models. What product to recommend to what customer. The churn models, right? And our marketing qualified lead models, right? What even to qualify for a marketing lead. That's where it started with us.
[13:19] And then it slowly progressed to your our first drag. And then slowly to a agentic workflows, the one on Taylor showed like account strategy workflow. And then now to your full-blown agentic companion.
[13:36] Today, that one of the agents uh maybe we'll see later in the slide. They can ask questions. They can actually run for hours, right? Even they started doing a causal analysis that's only a data scientist can do. They can actually do it here and project their sales. And all of this
[13:52] happened because the agent learned to use the existing data that we curated, existing futures and models that we curated. And just agent calls them as tools. All of them are built on the same ACL fabric, same observability, the ML flow
[14:07] uh layer. All along the whole way, right? Both for our ML models and for our agents. Nothing was thrown away. Today, our core framework built on LangGraph, which internally called Magneto, more than 700,000 people are using it.
[14:25] And we are trying to add like 10,000 more to it. And five different versions of the companions built on top of this framework. With over 5,000 ground of features and countless number of data sources. We started with interesting third third
[14:41] party sources like news. And we are in the process of adding like social media and others to this. And we are going to look at how we arrived at it. You know, the current state of it and then we will look at how did we get there. Today, if you look at the actual problem
[14:57] statement that Taylor told, we have countless countless acquisitions. Each one of them bringing their own CRM database. Some of them we even we have even never heard of. And their own content data.
[15:13] And their own product adoption data and and their own product telemetry data. Everything is flowing in uh to our data lake. That's where where we started collecting. If you look at the shape of the data, right? We were already doing very good in the structured data analysis as a data science team.
[15:29] Then suddenly there's this large influx of data that we already gathered, the unstructured data of data, we we have a way to kind of enable them to the users, right? So this is the architecture we arrived at today. That's how we look at it. Your structured data that's actually
[15:46] doing your data science work and the unstructured agent that's actually doing all of your unstructured uh data like searching your vector store or looking up a news, anything of it. On the unstructured side of things, it's a little bit inspired by a graph rag uh
[16:03] from Microsoft paper. It looks at all of your call transcripts, video transcripts, or the emails you sent, any of your enablement content that's out there in your Confluence, Sales make, whatever you you store your data to,
[16:19] your Google Drive of the nature, everything comes through the unstructured uh agent. It uses a inbuilt graph and a a modified version of a graph rag and a vector sets and a MongoDB lookups to actually achieve this.
[16:38] And on the structured side, uh we are a data science team. So what we built is a data science agent, a maybe a junior data science agent, I would call it, that actually does your all of your customer analytics or any of your advanced analytics that's needed for your structured of things, right? And there is a supervisor that combines these two and brings
[16:56] the whole insight to the question, right? It's not just the data, it's just not your next best action, it's a whole insight, everything combined together. Your predictions and the explanation of the predictions and how well the prediction is actually going to work in the field for them, curated very
[17:13] personally to each individual seller, right? So that's our architecture.
[17:33] Taylor said it was like a 2 and 1/2 year journey. And we didn't arrive at that architecture by accident, not overnight. It took us almost a year. When we delivered our first rack, a simple rack, it was a huge hit. Then we set out to actually implement a full-blown supervisor in early 2025
[17:50] using a framework called DSPy. It was a very rigid supervisor critic loop. It never learned how to break out of the loop. So, we had to quickly pivot. Then we adapted actually we built like four or five different
[18:09] rack use cases or rack agents routed through a central DSPy router or a classification router based on the product or the field of query, it kind of routes to these multiple different agents. That was our early go-live in April 2025. Then each quarter we kind of figured out why
[18:26] our actual first supervisor failed. If you look at it, just a spoiler, the first supervisor didn't know when to stop, right? It didn't have a concept of reflection. It can't It was constantly critiquing on its own throwing out all of its work.
[18:45] Then by mid of 2025 we introduced a coding agent, a data science agent, that can answer all of your data science equations or any of your advanced analytics questions. What's the trajectory for my sales for the next quarter? Or is there a causal causal relationship between giving a discount of this percentage is going to improve
[19:01] my lift uh this percentage. You don't have to go to a human. There was agent right sitting right there too answering all of these questions. Then we introduced uh graph rack inspired by Microsoft Paper, our own version of GraphRock uh built on um
[19:17] uh NetworkX and later Potato CU Graph, that actually brought in the depth that's needed to the agent. We didn't realize this was actually one of the biggest blockers for your long-horizon agents, right? This became our knowledge source where a long-running or a planner can
[19:32] actually go to and plan your multi-step task, right? It can go to a depth, it can go to a page rank, it can actually go to a community summary to figure out what is that around a particular user ask and actually plan this out very well. Then we actually refactored the code uh late 2025.
[19:49] Our first agent actually took like 3-4 months for us to uh deliver. The second agent we actually delivered in like 4 weeks. The third agent we actually delivered in 3 weeks. So we were able to arrive at a modular structure that actually made us go faster and faster. By the end of 2025,
[20:06] we had like six different GraphRock agents and one coding agent behind the scenes, and we had a problem routing between all of these agents. The classification was uh hit and miss with a large classification uh labels. So we have to go back and fix our
[20:22] supervisor. By the time, industry was already adopting a concept called a reflection. That was the one final missing piece. When we added that reflection to the actual agent along with your GraphRock, it actually unblocked the long-horizon planning. It was able to kind of run for hours. I think
[20:38] Taylor showed one of the screenshot that one ran for like 3 hours 29 minutes. It's there in the screenshot that actually produced the results uh the user was asking for. Uh and a fun fact, Taylor was one of the one who was actually always stress testing my test system. He ran for like
[20:53] 20 hours straight. It actually came back. I don't know whether it's good or not. We were able to actually make this long-horizon planning work well and kind of extend without losing its context, right? That was the hard part for us. If you look at the
[21:09] parallel industry parallel how it evolved, right? We started with basic LangChain. We didn't even start with LangGraph. Actually, we didn't even start with LangChain. We started with the DSPy. It's a very simplified version of like a private private top style framework that we use to actually build agents back then.
[21:26] Today, we still use that for our most of our eval frameworks. All of our eval frameworks, we use DSPy. And I think Databricks also adopted that, and that's the eval framework Databricks exposes today in Databricks today. And then there was a hyper graph tool AI and other tools.
[21:43] Uh we were actually exposed exploring those to kind of do our coding agent. I think we settled on uh Hugging Face's small agent. It's actually very really really really good coder. They had the right uh abstractions to actually handle our coding logics.
[21:58] We implemented a remote Spark executor. That's what we use today to connect to our Databricks clusters today to execute any of your either your data data workload or ML workload. It can actually connect to a Databricks cluster and execute it for you. And then the graph rags.
[22:14] Uh with respect to graph rag, we started with a simple CPU bound graph rag because this graph is a knowledge graph. It doesn't change a lot because your product documentation is not going to be updated every single day. But since because of the way we actually
[22:31] orchestrated, the graph itself was taking like 12 seconds on a on a whole. So, you have to actually do a graph walk to figure out what depth of uh the nodes you have to actually pull in. So, we have to kind of load it into a
[22:46] CU graph library from Nvidia and throw it in. It actually went from 12 seconds to like 100 milliseconds. On an average today, agent uses 12 to 13 uh graph lookups or graph locks or a page rank, whatever the uh graph related terminologies
[23:03] to actually do the planning and execution that we were able to kind of save off more than 10 12 seconds that really resulted in the experience user experience. The one other thing we realized in graph is there are a lot of hot nodes. Let's say if someone is looking for Workday, it's connected to
[23:18] everything that you know of. It actually brought down our data bricks vector store. Uh we had to kind of work around that and we kind of evolved with what the industry was going. Then the refactor and then the finally the deep research agents we started enabling at end of 2025 and finally the
[23:41] reflection that got us to where we are today. Now, we look at what we did and how how we did it and what is the role that Databricks really played in enabling it. As a matter of observed, we started in Databricks
[23:57] curating all of this data to actually build our classical ML models. So, we already created a huge set of data, not just the snapshots. We were actually bringing the CDC data from all of your metrics what changed when all of your CRM systems. We have the whole history of
[24:13] this data just to do your feature engineering time travel. So, we have all of the data living in data lake. And all of my models that we already built was lived in data lake and out of all of the failed models specifically, I cannot stress this enough, right? This is most
[24:29] of the things where you business. It's just not the successful things that you got you get to implement. You need to know what not to tell your agent, right? What are bad data sets that you cannot take it? What are the features that didn't work out that you cannot actually take it to your agent.
[24:44] It's going to be going to simply spin in circles and not get anywhere to you, right? So, that failures the years of years worth of our email registry actually guided us to actually curate the data that actually we can actually give it to a agent to be successful to answer any questions. And finally,
[25:00] this is actually covered in the next slide. As a data science team, we are not a UI developers or authentication or your infrastructure developers, right? Databricks actually brought in the whole user identity the enterprise has. What are the groups the sellers belong to?
[25:16] What's the group they were C-suite belong to? What's the actual access they have all of these financial data sets that are actually mapped directly in your Databricks? And we were able to kind of enable it just like that just by actually mapping a row filter on the tables that you want to execute.
[25:32] And with the OBO from Databricks, it was much easier for me to enable this. So, these three principles are core, I would say, why we actually started leaning towards Databricks more.
[25:49] And as for the second one that that I explained. Before the era of LLMs, we tried to we actually built a ML model, and we are trying to explain this ML model to a particular user. We were telling, "Okay, this particular account is going to churn, but why is going to churn?" We were dealing with LIME and SHAP values
[26:06] to explain them why it's going to churn. Just to do that one single widget, right? We spent almost a year, right? We need to kind of learn React, bring in external help to actually bring in the front end, and we need to start building
[26:22] graph APIs on top of replicating the data from Databricks to uh some no SQL store to actually expose these APIs or specifically our predictions and our SHAP and LIME functions to just build a one single
[26:38] UI, right? Even though apart from the actual UI that we the UI skills we learned, everything else got erased by apps, right? With apps, I don't have to build any one of this, right? I don't have to build identity. I actually I didn't even write a single authentication and authorization code in
[26:54] my current companion story, right? All of them is taken care by Databricks. So, these two enabled us actually to deploy our first agent in Databricks. Since we come from the ML flow
[27:09] standardized ecosystem, our first agent naturally ended up as a ML flow artifact. We actually served out of Databricks ML flow serving. It gets a start because we were having like one or two rag agents and we later added graph rag.
[27:25] It was okay. So, the response time for within 1 or 2 minutes, we were able to actually sustain. We were actually able to scale because you can actually auto scale the ML flow systems. Then we started adding coding agents. It was actually pushing the overall response times to upward of 3 plus
[27:41] minutes. So, that's where where the ML flow actually has It has a hard time out of like 3 minutes and the overall ecosystem we were trying to kind of set every single time out possible. But, at the max we are pushing to like maybe 5 minutes at the max and that's
[27:57] that's a hit and miss always. So, the ML flow from Databricks and the early version of apps where we actually hosted our UI enabled our first agent.
[28:15] Then, the second iteration of apps with horizontal scaling truly enabled a long horizon agents for us, right? Now, the agent can actually run for hours because there's no time out in the apps.
[28:31] And with horizontal scaling enabled, you get to choose how much of the pods you want to schedule for your workload, right? How much of data actually you can actually bring in from your data systems to in memory for you to context engineer and send it to LLMs. You can actually choose the memory.
[28:47] Right? And you don't have to limited by the constraint of the ML flow serving infrastructure. We kind of implemented our own streaming to actually make this work off of ML flow. None of them has to be done anymore. So, every single side of that is wiped
[29:04] up by Databricks apps. And with the recent Databricks horizontal scaling, it became our actual platform. The whole Databricks became our platform.
[29:19] Just walk with me on my left, right? I don't have to build the SSO or the Okta identity, not the not even the row-level ACLs. Actually, from the get-go, we never implemented our own tracing. It was already there for us. We were using the tracing for our classical ML models. That just got extended to my LLM models.
[29:35] And we already had the ML flow registry to versioning. And the vector search came out of Databricks. And it had a very good secrets management. And I can just this enough, the spot execution. You'll You'll see later in the slide why I'm talking about it. It's just that a
[29:51] huge amount of data is locked behind your a single dashboard that user actually cannot ask anything about it, right? This truly enabled what user can actually go do by themselves, right? And all this with
[30:06] enterprise grade audit and logging compliance and AI gateway with your FinOps and our runtime security enforcements, everything governed in one data lake. So, Databricks actually technically become my platform team. All of my data scientists and ML engineers can go to I don't have to go
[30:22] to my platform team altogether. What we rather built is we spent iterating over our business logics, our prompts. We built our custom prompt engineering library around DSPy and our orchestration agent orchestration. I went through the whole year of learning, right? We We started from V1 and ended up on
[30:38] V5. But with between V1 and V5, we went through hundreds of iterations to just make this work. So, we spend get to spend time on improving the agent that actually matters for the matters for the business than actually building out the platform.
[30:59] So, that's the hypothesis I just showed you, right? This is the actual picture of how this hypothesis looks like, right? The one thing I want to highlight on the right side is the 500k task per hour, right? So, this is a thing that we never knew the users actually needed to execute. This must have actually get their
[31:15] questions answered. None of them has to go through a data analyst or a data scientist to get their actual data brought to life for them, right? This is a 500k task per hour, not even a day. The amount of volume, the scalability, the Databricks enabled on the massive
[31:31] data that we curated, it's actually the true enablement for the business. So, below the waterline, if I look at all of them in the black, every single one of them is a Databricks product, right?
[31:47] In my mind, just the app is the tip of what the users see. The whole ecosystem is below the iceberg, like below the waterline. So, in essence, we didn't build a stack, we actually composed the platform out of
[32:04] Databricks that kind of evolved, and Databricks actually put up with us to kind of enable this. We'll talk about that in the next time. So, the things that didn't work well in Databricks, where
[32:19] Databricks actually showed up to us and enabled us to move more, right? wrong. Let's talk about what didn't work in Databricks.
[32:35] If you look at the classical cluster Databricks for a data science workloads where one or two data scientists connect to a cluster and run their email model. But when we actually started enabling more users, they were asking so many questions, we actually hit up concurrent limit on the
[32:50] Databricks cluster, right? It actually hard set to 150. We had to actually work with the Databricks product team to enable it to move it to 1,000 and 2,000. And the same thing for MLflow. We worked with the app team, so they actually enabled us one of the first private preview to kind of
[33:05] have this horizontal scaling tested out because we were hitting a wall on MLflow. And the same thing for the trace UI. If the trace is like 40, 50 runs, the trace is good, but with the however long runs, you can easily accumulate like 4,000 4,000 run trace. That's like 5 GB of
[33:21] data just the trace alone. That's still not loading, right? I don't think anyone actually solved it. We are actually working close with Databricks team to kind of enable this.
[33:38] The other part of the ecosystem that users didn't get to see, right? I got a nice UI, yes, there's agent answering all my questions, what's needed, but how do we get to that? What's behind that that's actually enabling it? Every single one of it is powered by Databricks. It's just not powered by Databricks. It's just the core ML
[33:55] principles embedded in it in all the way, right? Just to take the news, right? Your 6 months worth of news accounts for like 3 billion tokens. I cannot just take that and throw it out to cloud. You're going to burn like hundreds of thousands of
[34:11] dollars, right? We ended up actually fine-tuning a smaller LLM and extracted our pre-filter the data and we had to get to your vector store. Actually, you wouldn't believe, we actually built a a custom offline vector store on Spark to actually just process
[34:27] the news. Today, there is a account strategy agent that actually depends on a separate news gathering agent that actually runs in a spark streaming mode to actually get this done. It's not all land graph, but the whole back and power along with the spark everything is what it made it possible
[34:44] for us, right? The rest of it is simple. It's just your extension of your email practice, right? Every single thing we didn't just take it for granted. We have to really prove it what sinking strategy to use. And this sinking strategy actually changes from source to source.
[35:01] For most of the well documented sources in enterprise where interquartile actually works well. I just a tip, you don't have to go for page by page sinking, go for interquartile. It actually cuts your token cost by like 30% but keeping the same quality. And I want to actually bring up a graph
[35:17] and deploy it on a GPU that's actually possible in data bricks. And I want to actually parse my PDF. It's already available in data bricks AI parse, right? And we were using Tecton before now got acquired by data bricks. It's running on spark that's the backbone of what we all do.
[35:40] So this this is the just latest of what's running in production. If you look at it, most of the questions get answered like two to five minutes, but if you look at the long tail like the one goes over 30 minutes. They are if you look at these are very deep questions. It needs to run for hours like
[35:56] especially the executive narrative, right? This is not possible before. Maybe the persona is simple, but you need to actually hire a full-blown data analyst to just do that. Now it's just a one-hour work on
[36:12] our companion and you're good to go. Just to finish it off, none of this was possible for four years ago because we are building a simple email model. And none of this possible four months ago
[36:27] without the horizontal scaling in the apps. When I started I started with showing up how long it takes to run the longest running query on my companions. The app was the easy part, right? The
[36:44] ecosystem was the story. And thank you Cody and Databricks team and Taylor and business team for supporting us. We're open for the questions.
[37:03] And I know I got only 3 minutes, that's why I cut it short and I had like another 30 slides to go through. If you have I didn't What I didn't cover is I didn't cover how did we test it, right? Everyone so tells you how to build an agent, right? How do you test and evolve? The evolution comes from your actual test
[37:18] harness that you're going to build. I talked about DSP boy. Go and take a look at it. That's going to be the bread and butter of what you want after you start building your agents, right? We actually started building our harness first than our agents. That was the first use case came to us. I have this agent. I want to make a go
[37:35] go no-go decision to production. How would I do it? That's what we built first. Uh if you get access to it, take a look at it, but I was I'm able to kind of answer your questions and how did you actually go through it? What are the decisions that we made to actually make this happen. It's all in the backup slides. Thanks.
[37:52] Woo!

Learn more about the Databricks Data and AI platform.

The information provided herein is for general informational purposes only and may not reflect the most current product capabilities or configurations.