Skip to main content

Scaling Genie for Production: GSK's Enterprise AI/BI Factory

Summary

  • GSK Global Supply Chain scaled Databricks Genie from a proof-of-concept to a production enterprise AI/BI factory serving 37 manufacturing sites, implementing semantic caching, tiered query execution, and full request tracing to meet enterprise performance requirements.
  • The team reduced Genie response latency from 34 seconds to 22 seconds and cut Genie throughput load by 64% through a tiered system that routes simple queries to SQL tools and a semantic cache before escalating to Genie for complex questions.
  • Automated CI/CD for Genie spaces, gold-standard SQL benchmarking, and cross-functional co-creation with business teams were the practices that transformed skeptical stakeholders into enthusiastic adopters across the supply chain organization.

Scaling Genie for Production: GSK's Enterprise AI/BI Factory

Watch: Scaling Genie for Production: GSK's Enterprise AI/BI Factory
Most organizations struggle to move Databricks Genie from successful pilot to production at scale, facing data contextualization challenges, unacceptable latency, and user adoption barriers. At GSK Global Supply Chain, with 37 manufacturing sites and thousands of decision-makers across regulated environments, a proof-of-concept wasn't enough. The challenge was moving from one-off conversational AI to an enterprise AI/BI factory serving diverse personas asking different questions across complex supply chain domains.
GSK built a production-grade Genie platform powered by Databricks' managed semantic layer, achieving 10x scale through semantic caching, tiered query execution, full request tracing, and automated CI/CD for Genie spaces. They instrumented Genie's multihop pipeline to understand latency sources, implemented a dashboard-based observability system, and reduced latency from 34 seconds to 22 seconds while cutting Genie throughput by 64%. Codified best practices, benchmarking against gold-standard SQL sets, and cross-functional co-creation with business teams transformed skepticism into enthusiastic adoption across the organization.
🤝

Chapters

FAQs

How did GSK reduce Databricks Genie response latency from 34 to 22 seconds?

GSK implemented a tiered query system that routes simple or repeated questions to SQL tools and a semantic cache before reaching Genie, reserving Genie for complex multi-hop questions. This reduced the load on Genie by 64% and brought average latency down from 34 seconds to 22 seconds.

What is semantic caching and how does it help Genie scale at GSK?

Semantic caching stores the results of previous Genie queries and matches incoming questions to cached responses when they are semantically equivalent. By serving cached answers for common questions, it reduces Genie invocations and improves response time for frequently asked supply chain questions.

What challenges did GSK face when moving Genie from pilot to production?

GSK faced data contextualization challenges, unacceptable latency at scale, and user adoption barriers when moving from a proof-of-concept to serving thousands of decision-makers across 37 manufacturing sites in regulated environments. A one-off conversational AI demo was not sufficient for the diverse personas asking different questions across complex supply chain domains.

How does GSK use CI/CD to manage Genie spaces at enterprise scale?

GSK built an automated CI/CD pipeline for Genie spaces that manages the full lifecycle from development through deployment. Combined with benchmarking against gold-standard SQL sets to validate answer quality, this allows the team to update and improve Genie spaces reliably without manual deployment steps.

Full transcript

[00:10] Um welcome everybody. Um hopefully maybe after the after party we give you a bit of energy with our talk. Um so today we'll be talking about um what we have done with GSK global supply chain how we've been using Genie and how we've
[00:27] been scaling it. uh that's a very important problem. So how do we go from a prototype into an enterprise ready solution? My name is Virjini. I'm Jenni director at GSK uh working for global supply
[00:45] chain uh for the past two and a half years and two and a half year we were not talking about agents but now I'm moving more and more into agents. Um I have a very technical background. I was data scientist, machine learning engineer, AI engineer, now Gen AI
[01:02] engineer. Um, and a bit of a fun fact about me, I was um at the Spark Summit in 2016 and uh at the Spark Summit in 2016 in Brussel because I'm I'm from Belgium. They took a picture of me and they put
[01:19] it on their banner for the 2017 Spark Summit. So I counted I contacted um Zabriggs. Well, Spark at the time and I told them, "You owe me a free ticket for the for the Spark Summit 2017 because I'm on your banner." So, I went into the
[01:36] Spark Summit uh 2017 for free uh thanks to u to this. Um and with me to tell you this story, I have also shitish from from data bricks uh with me. Hi everyone, I'm Shat. I work as a
[01:54] senior AI engineer at datab bricks part of the AI FD team. Uh a little about me, I'm based out of Goa. It's a beautiful coastal place out of India. Uh I have a background in mathematics and computing
[02:10] and I have over 10 years of experience building AI ML solution. It's really nice to be here today. Yeah. Um so before we go into the juicy technical um uh stuff for today, I want
[02:27] to give you a little bit of a view of what's the products, what are our challenges, what is GSK because maybe you also are new to uh the farmer world. After that we will be very quickly giving you an overview of the architecture uh really briefly and then
[02:45] we will go into some of the challenges that we had and how we've solved it uh together with data bricks and a little conclusion at the end of the presentation. So GSK GSK is a pharmaceutical company
[03:02] international. We have 37 uh manufacturing sites across the globe. Uh we produce uh vaccines uh specialty medicines and also general medicines. Um so we are really a huge
[03:20] company and uh we have lots of challenges because all of these different products they have their own life cycle they have their own uh definitions and each plant is um operating a little bit um on their own.
[03:37] So our challenges are really global. They are about scaling and making sure that whatever we built is uh fit for purpose. So how do you make a product generic like for all your plants but do you make sure that the it is actually
[03:53] useful for your uhmemes for your factories. So um the supply chain um at GSK is quite complex. You have uh well topics such as demand and planning. So how are
[04:09] we what what do we need to produce? What's the demand? what should we actually start uh producing? Uh things like um manufacturing and production. So um uh questions like um am I efficient?
[04:25] Um am I uh am I not making any scraps? Do I plan do I produce the right products? Uh things like that are the questions that these uh people are asking. Uh the logistic, how do I move those different batches across my supply
[04:41] chain? uh things like do I have enough product in my warehouse. So uh in the supply chain of GSK and I think it's a general problem in the supply chain we have different personas and they will ask different question along the journey
[04:58] right like the uh manufacturing people will not ask the same question as the quality people because quality is responsible uh for making all our product u compliant right first time no impact on patient. So all of these
[05:14] questions is what we wanted to answer uh with our agents with uh our products and um across all our sites. So you can imagine there is the complexity of the supply chain but there is also the complexity of all the sites that we
[05:31] have. Um, and when we built this product, we were not just like thinking about we're going to build one agent and it's going to be um deployed and and and that's it. We were really thinking about how can we
[05:49] rethink um how can we change the way users are actually interacting with the supply chain? How can they now consume information differently? How can it be much faster and much more insightful
[06:05] than before? So, um you may recognize those levels. They are from OpenAI maybe two years ago. that is what um Samman um released and I've adapted them for the supply chain because I think they are
[06:20] really interesting and um so the level one is really just about asking question um about the supply chain and all of these uh questions that I've just um talked about. The level two is about going one step further and it's about
[06:37] doing some root causing automatically. It's about generating those insights uh out of the box, not having always the human trigger uh kind of those those questions. Um and then level three is really about acting,
[06:53] right? How do you act? How do you make sure that um the the next best action, the things that were recommended can be actually applied? Uh I think we are almost there. um within GSK global supply chain. We certainly have done level one and two
[07:10] and we're working actively on level three and well level four and five I think they are more and more within reach with changing the supply chain improving it innovating it like auto autonomously right right um
[07:27] so how does that translate concretely um well if you are at the site uh if you are a manufacturing person you have questions like and and they are very concrete, very practical like um there was an issue uh that was uh detected
[07:45] like a change change control or something or deviation and the question is what should we do? What's the impact downstream? Um do we need to basically take an action or am I going to be out
[08:00] of stock out of this? So thanks to our product and all our agents, we are able to answer those questions um much much faster than before and take an action to avoid um some further impact.
[08:16] So our product looks like this. It's a very simple and sleek interface where um everybody can just uh connect and interact with the supply chain. They see this first interface and
[08:32] they have first of all these next best action reports which they can um they can consult. It's about um all our different performance um uh indicators and it's also personalized. So if you
[08:49] are a different uh persona, if you're a different site, you'll have your own report. And um one fun fact a bit about this is that before we did the reporting we were just releasing conversational AI
[09:06] and people were like oh um this is cool but uh well I'm always asking the same question so I'd like you to automate my questions. So we did a reporting and now um actually they are telling me this is too much information vini you should
[09:23] give me a conversational AI. So we we just have now we just have both we have we have the report but we also have our um conversational agents that are there to um either deep dive or um uh help
[09:40] them figure it out because like um those agents they generate so much information and this is very valuable but sometimes you're a little bit lost in how you should um interpret or interact with this um this this information.
[09:57] So this is our product is really like one single pane of glass where we deploy all our agents and where the supply chain can really interact with those uh different agents. Um but very quickly how did we implement
[10:15] that? What's our uh infrastructure behind? I think this is very classical and you will recognize it. uh because it it's nothing nothing special but we have our code orange this is our data platform they provision all the
[10:32] infrastructure for us and then on top of it we have uh a common data model with really a lots of tables in there and on top of it we have implemented genie lots of genie spaces well I heard that they are called genie agents now so this is
[10:49] an outdated slide um so but genie agents and then next to it. We have all our tools that the agents can use to do anomaly detection um to do report generation all of these type of tools and we have all the connectors we
[11:06] connect for example with viva because this is an important um tool within uh the supply chain and then on top uh we have uh started working on ontology and then all of these agent deic layer.
[11:23] And of course this um this we always kind of like don't see that part of the iceberg but u we spend a lot of time on our data. We spent a lot of time on u making sure that the data that we would
[11:39] feed to the agents was really good and for that we've used the KPI library from datab bricks as well. Um so lots of work on on on on the data and um but well
[11:55] this is not that easy. It's it's really a cool product. It's now there. It's live. It's in production. We have lots of user but we had challenges. And today we want to talk about the different challenges that we had. And the first thing is how do you do this right? How
[12:11] do you create agents that are actually giving you uh insights based on your data when your data is like 37,000 plus tables? It's kind of like where do you start? How do you do that? Like um how
[12:26] will you be able to answer all these questions from your user? So um I'll let you now talk a bit more about these different uh challenges. Thank you. Thanks Vjini. Uh so like Virginia said
[12:43] our solution to the contextualization challenge was using Genie data bricks managed semantic layer for structured data which allowed us to move into production with our supply chain AI assistant within 3 months. However,
[12:58] while working with Genie, we realized that a sloppy Genie space gives us sloppy, wrong and confident responses. So the first thing that we did was instead of treating a space setup as a one-off, we codified this into a
[13:14] checklist and best practices that every space within the global supply chain had to follow. Now let's take a look at some of those key takeaways and antiatterns we learned. There are four things that really
[13:29] matter. Scope tight. Push business logic within the data layer. teach in SQL and always benchmark and let me show you each of these along how it usually goes wrong within organizations.
[13:46] So the first learning what build one genie space uh per business domain uh with clear target personas and centered around the uh questions the users actually ask and compared this to the
[14:02] antipattern where we have one space trying to cover manufacturing forecasting orders customers all at once. The second learning if your business logic lives in process the model needs to rederive it every single time. So
[14:18] instead we try to model it within the data layer as much as possible. So think uh materialized views joins gold tables etc. uh the antiattern on the right that you see is uh we have pros inside genie
[14:35] instructions and if we have uh sequence of such instruction it leads to context rot and quality drift. The third learning teach through SQL. So one good parameterized SQL query is
[14:51] better than having three long paragraphs trying to explain the same thing in plain text English. And as you can see, contrast this on the right with a bad instruction which is trying to explain cold chain stockout logic in plain English. And honestly, no one reading
[15:08] this can understand it reliably and neither can the language models. And finally, always use benchmarks. So we recommend uh having 2250 uh gold standard set uh SQL set within your genie spaces and then aligning uh
[15:26] any change to the space based on the benchmarks and not uh solely based on the creators testing. Now compared to a custom texttosql solution the reason Genie was so easy to adopt is that the UI is extremely
[15:43] streamlined and intuitive. So what that means is it almost nudges the creators to follow the best practices in contextualizing the data in the right way and especially it's quite easy to work with uh especially for
[16:00] business users andmemes who are usually less technical and here's what the whole process looks like in a single loop. So we start with business definition, the data models, business logic, the jargon. We prepare a
[16:17] semantic model on well doumented UC tables, metadata, contextualize the space with the instructions, SQL examples, guiding questions and joins. Finally, we benchmark this against our
[16:33] gold standard SQL set. And this benchmarking loops back into the contextualization until until the business is happy with the response. And only then does the genie gets consumed within our agents via APIs.
[16:50] Now we replicated this across multiple domains, wired it to real agents, and we finally shipped. And the moment the real users and real agents hit it at scale, we ran head first into our next challenge that uh virt.
[17:16] So um once we had our GE started, the business was like so excited and we were like, "Oh my god, this is what we need and we need to scale it." You've seen the supply chain, it's really big. Um and we started scaling it but uh we quickly ran into some latency problem and um so that means the question would
[17:33] be answered in more than one minute and people were used to chat GPT that gives you an answer in like couple of seconds. So uh the users were really like complaining about uh the latency. So again here we implemented a a bunch of
[17:50] um solutions together to uh to resolve this problem. Yeah. So again when we started shipping uh agents with Genie what we saw as engineers was longtail latency. So our P50 our P90s they were shooting up. We
[18:08] were observing gateway timeouts on complex queries. We also observe frequent workspace rate limits the 429s. Additionally, promoting a genie space from dev environment to higher environment was completely manual. There
[18:23] was no versioning or roll back controls built in. On the other hand, what our users saw was was much simpler and worse. They would see long wait times, sessions that just died and similar routine
[18:38] reporting question taking 20 plus second asked each and every single time. And the reason why this was so painful to fix was when a session timed out, we couldn't attribute the load. Which genie space, which user, which time frame,
[18:56] timeout and dead session could have been warehouse queuing, genie's internal retries, workspace rate limits, or even general service degradation. So we moved to an approach that could instrument the full path and trace every
[19:14] interaction within Genie at a request level. Now Genie is a multihop system which means that once a request reaches Genie, it goes uh it it performs a sequence of steps or hops uh before the
[19:29] response is returned. And we wanted to instrument these hops to better understand the source of latency and then be able to optimize it. So the first hop that we instrumented was the SQL generation time. The time it takes Genie to synthesize a valid SQL
[19:47] statement based on the contextualized Genie space. The second hop the warehouse execution time the time it takes GD to send the synthesized SQL query to a warehouse and return the response. So if the warehouse execution time is high, it could mean
[20:04] that either the queries generated by uh Genie are not optimized or maybe the warehouse capacity has become the bottleneck. And then the third hop the retry counts. Uh the way Genie works internally is if
[20:19] the SQL statement errors out against a warehouse, it internally attempts to retry it to fix it and then retry. So a higher retry count could mean the space is not set up right leading Genie into making mistakes while generating the SQL
[20:36] query. And while these and while this trace gave us uh a view of request level uh distribution of latency each agent within global supply chain was different which meant the structure of these
[20:52] traces was also different. However, we needed a global and aggregated view of latency and traffic distribution across all our genie spaces and workspaces. So, we built these dedicated dashboard
[21:07] using system audit tables. One for QPM, queries per minute to help us understand if we were hitting any workspace limits and then attribute it to a particular space, user or a time frame.
[21:23] Second dashboard we built was for latency so that we could identify the spikes and correlate those with the request level trace and finally debug it. Now let's have a uh look at how we build
[21:38] these dashboards. So datab bricks audit tables had all the information we really needed to build a quick and fast dashboard. So for QPM we filter the relevant events and actions corresponding to uh the create and start
[21:55] conversation. While for latency we use window functions to grab the first and last event timestamp difference them to give to get an approximate per message latency. And the reason this works so well is because Genie APIs are
[22:10] essentially polling. So when you when you submit a request to Genie, you need to pull it until the response is ready. And what that means is each poll action creates a corresponding event in the audit tables. And differencing the first
[22:26] and last event basically gives us the time we spend waiting for the response. Now once we identified these outliers from the dashboard, we used the message ID and conversation ID and mapped them back to the uh request level trace that
[22:42] we just uh saw uh and debugged it further. But now why was the longtail latency so bad? So a typical agent in global supply chain orchestrates over multiple genie spaces. Each of these genie space is
[22:59] attached as a tool. So, Genie being an expensive uh and uh you know very involved step uh each query even if it's a simple query it's a query that has been asked multiple times
[23:17] does a full natural language uh to SQL to warehouse roundtrip and as we scaled uh to multiple agents multiple genie spaces it meant that genie became our bottleneck due to how often we were
[23:33] relying on it. So to prevent this, we came up with this tiered system that we put in front of Genie that only allows Genie to do the work that it really needs to do. So
[23:48] there are there are three main mechanisms in this tiered system in order of cost. The first one being the SQL tools. So for our highest frequency well-known questions, we expose them as plain uh plain text template uh SQL
[24:06] statements where an LLM just needs to populate the parameters to the template SQL statement and and it gets the SQL and it just sends it to the warehouse to get the fresh copy of the relevant data.
[24:23] The second tier that we are calling retrieval augmented few short text to SQL. So this is the semantic cache. We keep a vector DB of user query to SQL pairs uh which which can uh either be populated online or offline. On a new
[24:42] question we do a semantic search and we route it to different execution paths based on the similarity score. So a similarity score of.99 or above we treat the query as essentially the same that we have answered before in the past which means
[24:59] we pull the SQL statement from the cache and we send it directly to the warehouse for execution between roughly.8 to.99 sec uh uh score uh we say the queries are close but they
[25:15] are not exactly identical. So we pull these top k similar neighbors and we feed it to a generalpurpose LLM and let it generate the SQL statement. But remember we still ground this generalpurpose LLM in the data semantics
[25:33] that we created for Genie which means providing it access to the table names, columns, descriptions, join information etc. And finally uh if the similarity score is less than8
[25:48] we we don't rely on the cache at all and we route those queries to genie itself and when a query reaches Genie it is further used to populate the vector DB
[26:09] and let's uh have a look at how this system performs and how we monitor it as the cache evolves. So because we learned our lesson on observability early, this hybrid tiered system is fully instrumented too. So we monitor every trace and importantly
[26:26] which execution path and the execution time it took using custom scorers. So when we monitor execution path and execution time that helps us monitor how the latency uh of the system is affected
[26:43] uh and how the cache rate of the system is as the cache evolves with time. We also combine this uh with scorers like correctness to benchmark the uh system, tune the thresholds and ensure that our
[26:59] uh the quality of our responses stays consistent. So the the threshold values that I just uh mentioned in this tier, they are typically tuned for each use case based on those benchmarks.
[27:20] And here are our actual benchmarks. So we benchmark four configuration starting with the genie baseline uh adding in the SQL tools uh and then the exact match semantic cache and finally the retrieval augmented text to SQL. So we saw
[27:36] latencies drop from 34 second to 22 second and approximate 33% improvement. And we also saw the number of queries reaching Genie reduce by 64% from baseline to our full tiered system.
[27:53] However, we also saw a drop of 3 to 4% in our correctness metric. uh and and we think that's a fair trade given a one/ird of our latency and 2/3 uh savings in our genie throughput
[28:09] capacity which actually translates to lot of cost savings for our applications. An interesting thing to note here is between the configuration two and three when we add the exact match or the semantic cache we initially see a jump
[28:26] in our median latencies and the reason for this is once we introduce the semantic cache it means that for each request we pay an overhead of doing the semantic lookup and therefore it's very critical to
[28:42] monitor the cache rate of the system And ultimately the relative improvement in latency that we see with this system depends on the cache hit rate. So TLDDR we didn't make Genie faster. We
[28:57] just stopped asking it questions that it didn't really need to answer. Now our next challenge was to be able to manage life cycle of a genie space across different environments and we
[29:13] turned our genie space life cycle into proper git operations. So in dev we would have our engineers and our thememes build these genie spaces and we would use the genie export APIs to
[29:28] serialize these spaces and uh along with any ACL configuration or per environment transformation. So these per environment transformations are typically the warehouse ids um catalog names, schema
[29:44] names that might change across environments. Once this is pushed to the main branch, we have a dispatch workflow that applies these per environment transformation and imports the genie space into the higher
[29:59] environment. So, Genie Space is now versioned, reviewed, tested and promoted just like any other piece of software. No, no more uh going through the UI in production environment and replicating the setup.
[30:16] One thing to note is that although this illustration just shows our CI/CD uh for genie operations, but we we also have our agents that use these genie spaces and we publish those agents along with the agent artifacts to a common
[30:33] registry and which is how we share it within uh GSK organization. And that brings us to our third challenge. Yes, thank you. Because um there are a lot of technical challenges, but they
[30:50] are actually more and more easily solved and you you must see all of these tools are popping up every day. So, one challenge that we sometimes forget is the people challenge. So, not everybody is aware of um of of all of those new AI
[31:08] tools and not everybody is trusting them. And especially at the beginning of our journey which is like one and a half year ago when we started adopting Genie well a lot of people were not really trusting these uh agents and trusting
[31:24] what they would uh give to um to to to those agents. So it's very important to have a strategy around adoption and GSK has done a a very good job since the beginning because since um at the
[31:41] inception of Chad GPT we created our own Chad GPT which is called G it's the personal assistant and it's been rolled out to all our uh employees as really a way a sponsorship towards AI right and
[31:57] everybody now has also So in their objective to use more and more AI. So that's at the global level what we've been doing. But uh when we created this product these genie rooms this is really a co-creation because um um the AI team
[32:14] the data team we know very well about those technologies but we don't know as much as um the business ourmemes um which question should be answered. So uh since the beginning we've put these guard rails thanks to the CI/CD pipeline
[32:30] we have a very good govern um way of working so we can now really share those genie rooms and let our theme train them themselves so they become actor they become really um co-creator with us of
[32:47] those uh agents um we have also lots of feedback loops um not sure if you saw in our UI but we use the feedback button. So, we're really capturing user feedback and we act upon it and um we want to
[33:02] make it an an autonomous loop but we're not yet there. Um and what we've done is really a lot of um training as well of all ourmemes uh explaining how they should use it, what they should be doing with it. Um so all of these actions led us to a lot
[33:20] more trust in the agents that we are building and it's a journey right like when we started people were uh a little bit less um happy or less trustworthy of those tools but now they just like want us to deploy more and to cover the whole
[33:37] supply chain with those type of agents. Um just to conclude this is just the beginning and um what you've seen is the quick wins for GSK because we are after all a manufacturing company and um it's
[33:55] nice to talk to your data but we want to really act we want to manage our supply chain our factories with AI and uh make sure we actually transform the way we are working not just automating the PowerPoint And those things that people
[34:12] are doing now, we are really thinking and building the foundation for our future self um self-managed self autonomous uh supply chain. Um and just to finish this talk, I'm
[34:27] very proud to be part of this journey and to have really an impact on our uh our patient because we've been talking a lot about the technical stuff and what we've been doing and the performance. But at the end of the day, we should not
[34:43] forget like we are going to bring um our product faster to patient and that's what matters, right? Um and I want to thank all my colleagues that are not here but that really um certainly are
[34:58] part of the journey. Um just to I'm not going to name them all but it's certainly there was the data team involved the uh AI team the platform team really very important uh to support us in in our journey and finally all the
[35:15] business um uh the business user have been very good allies in this journey and data bricks that helped us really um improve and um accelerate on our journey here with uh with Genie and with product.
[35:31] Um, this is the end of our talk and we have five minutes. So, if you have any question, there are two mics and you can come to the mic and ask your question. Don't be shy.

Learn more about the Databricks Data and AI platform.

The information provided herein is for general informational purposes only and may not reflect the most current product capabilities or configurations.