Skip to main content

Building Production Voice Agents with Databricks and AI Gateway

Summary

  • Italgas, Europe's largest gas distributor, replaced a rigid IVR system dropping 20,000 calls per month with a conversational AI voice agent built with Cluster Reply on Databricks, cutting dropped calls by 50% and increasing automated resolution by 40%.
  • The layered architecture combines real-time LLM interactions for natural language understanding with specialized vertical agents for quotations, CRM operations, and FAQs, using Databricks Unity Catalog for knowledge governance, MLflow for performance tracking, and MCP for secure tool integrations.
  • The voice agent routes requests intelligently to specialized agents and integrates with live CRM data, freeing human operators to handle only the complex cases that genuinely require human judgment.

Building Production Voice Agents with Databricks and AI Gateway

Watch: Building Production Voice Agents with Databricks and AI Gateway
Europe's largest gas distributor faced a critical challenge: traditional IVR systems were dropping 60% of 120,000 monthly calls and could only automatically resolve 80%. To replace this rigid, user-hostile system, Italgas and Cluster Reply built a conversational AI voice agent on Databricks that understands natural language, routes requests intelligently, and integrates with live CRM data.
Learn how the solution uses Databricks Unity Catalog for governed knowledge management, MLflow for performance tracking, and MCP for secure tool integrations. See the layered architecture combining real-time LLM interactions with specialized vertical agents for quotations, CRM operations, and FAQs. Discover how this approach cut dropped calls by 50% and increased automated resolution by 40%, all while freeing human operators to handle only complex requests.
🤝

Chapters

FAQs

Why did Italgas replace its IVR system with a conversational AI voice agent?

Italgas's traditional IVR system was dropping 20,000 calls per month because callers found the menu-driven experience frustrating and abandoned calls before reaching resolution. The rigid structure could not handle natural language requests, forcing callers into predefined categories that often did not match their actual problem.

How does the Italgas voice agent architecture work?

The architecture layers real-time LLM interactions for natural language understanding on top of specialized vertical agents for specific tasks such as quotations, CRM operations, and FAQs. Databricks Unity Catalog governs the knowledge base the agents draw on, MLflow tracks performance over time, and MCP provides secure integrations to live backend systems including the CRM.

What results did Italgas achieve after deploying the AI voice agent?

After deploying the Databricks-powered voice agent, Italgas cut dropped calls by 50% and increased automated call resolution by 40%. Human operators were freed from handling routine inquiries and could focus exclusively on complex cases that genuinely require human judgment, improving both customer experience and operator efficiency.

How does Databricks support production voice agent deployments?

Databricks provides the end-to-end infrastructure for the Italgas voice agent: Unity Catalog governs the knowledge base used for retrieval, MLflow tracks model performance and enables monitoring over time, and MCP secures tool integrations to external systems like CRM. This combination gives the team a governed, observable platform for a production system handling thousands of daily customer calls.

Full transcript

[00:10] So, you know those moment where you have issue with a company or you're not sure how to change your appointment for some services or you know, like the bill
[00:25] doesn't feel correct and you need to verify something and so you you look for a solution on a website or on the internet and you can't find it. So, what you do? Well, maybe not my generation anymore because we don't believe in calls,
[00:42] just texting, but still a lot of people call a contact center. And so you you enter the virtual assistant mode and what happen?
[00:58] Well, you start talking with the again, a virtual assistant. So, something not human, it's kind of automated and they start asking you question. So, you have to do to select a choice, they present you to
[01:15] different choice and then sometimes you can do it by selecting the numbers, sometimes by voice, but it gets long and tiring and then you go on and it feel ask you questions and you
[01:33] still have to make choices and you go on and on and on. And then two things usually happen. You either get super frustrated and you drop the call and
[01:48] and you go you go on with your life or you find a way to talk to a human, to an operator. But then what happened is that you have to start all over again and explain to the to the operator what it was going on,
[02:04] what was your problem, and hopefully it will help you. And all of these resulted in 20,000 drop calls every month.
[02:20] You can see that it's not a feasible number. So I'm Serena Adelini. I'm the head of data and AI architecture at Italgas, and this is Nicola Giorcelli, data and AI lead at
[02:35] Reply Cluster. So what we are going to to show you today is how we solved this problem in hopefully solved this problem in our company, and we will walk you through what were the issue and how we solve it,
[02:54] give you hopefully some nice insight on the technical side, and some takeaways, some lessons learned, and where we are heading in the future. Before entering into details, just a brief introduction of our companies.
[03:11] Italgas uh is the first gas distributor in Europe. Uh Italgas is a group. Its core business is gas distribution, as I said. Uh but we also have other company in the
[03:27] group. For example, we have water services company. We have an energy services company that does optimization for um industrial, mostly industrial, um environment. And then we have Blue
[03:42] Digit, which is the company that I work for. Blue Digit is the tech company of the group. It was born around 5 years ago by detaching the IT department of Italgas
[03:57] and becoming its own standalone company. We still kind of are the IT department of the whole group, but we also provide with advisory services and we also
[04:13] create, realize proprietary solution that we sell to third-party customer. And Class Reply is the company I I work for. It's the system integrator that is the partner of Blue Reply. So, we've
[04:28] been working together since 2018 when we started with the Italgas data platform. So, Class Reply is one of the 400 company between the Reply group. Reply group is a worldwide SI right now, which is was
[04:44] born in Italy and most of the activities are in Italy. Class Reply itself it's an Italian-based company and it's a tech company in the Reply group, which focuses on Microsoft technology. So,
[05:00] we are in Azure and on-premise stack and we were the first company to get all the six solution partner designation. I was of course part of the data and AI solution designations. And we brought the Databricks solution
[05:16] for to achieve designation. Um the other major partnership we have of course it's the Databricks one. Um we started our partnership right around when we started consulting with Italgas. So, uh in the days of the Delta Lake
[05:34] uh general availability. And we grew our partnership together. Right now, I'm leading a a center of of excellence uh composed of roughly 30 35 people uh in Italy. Everyone is a certified
[05:52] Databricks and at some point we were the first company in Italy by number of champions. We don't have any sort of industry focus. We are cross-industry. We are just a technology-focused company. And right now, just this year we deployed
[06:07] more than 40 projects and solutions in the Databricks ecosystem. So, before starting with a solution, the very annoying virtual assistant that was
[06:23] talking before, um it's called I V R, which stand for interactive voice response. So, for those that are not familiar of what an I V R system is, um let's see a bit into the details.
[06:40] Basically, you have a first node that starts when you answer when the uh virtual assistant answer the call. It's the entry point and so when it starts, you get presented with a number
[06:57] of option. As I said it as I mentioned before, you can either select with a with a number on your phone
[07:12] or uh by saying um some words. And it can recognize the the word that you say. Once you enter a branch, what happen is that you are presented with other and other and other
[07:28] options. So, uh you can't cross from you can't um go through all the branches if you realize that you choose the wrong path, you have to start over. If want to go
[07:43] ahead and jump some step ahead, you can't do it. It's very rigid and very fixed. And it can get very complex. So, to give you an example, this is the some of the branches that we
[08:01] had in production at some point. So, as you can see, it gets very, very complicated. And if you get uh into one of those branches that spread for so long, you can imagine how frustrating it can
[08:17] be. Um I mean, I'm still talking and it didn't finish to run all through the branches, so and through the leaves. Uh so, what were the main weaknesses? I mean, it's easy from what I present to
[08:34] understand what were what they were. First of all, it's really rigid, as I mentioned before. You can't jump from a branch to another. You have to stay on the path. It takes It might take a very long time
[08:49] to solve your issue, so it gets really annoying. It can handle a limited number of uh complex requests. Not before Not because it's not possible to do it, but mostly because
[09:05] it gets really complicated, and so people just gave up before. And uh the other thing that you can imagine how complex it is to maintain and do um modification of all the system when you
[09:22] have to handle something that enormous like the branch that I showed you before. All these uh All these problem affected uh the number of calls that were answered.
[09:38] And out of 100 120,000 calls per month that we have, 60% were dropped. And 80% only 80% were solved automatically. So, you can see these
[09:55] numbers are not good. And so, we had to do something. And the solution that we came up with was the voice solution, a voice bot solution. And really the idea behind the voice bot solution was to
[10:10] basically erase what the VR was doing. So, uh, remove all the complexity and the rigidity of the solution that the VR was. And these were the main key capabilities that we were looking for. So, the voice interaction, it should be as natural as
[10:27] possible. So, you can have a back and forth with the the chatbot. You can interrupt it and you can uh, recontextualize your question if you see that the chatbot doesn't respond as you would expect. And of course, we have a usual
[10:44] knowledge. So, we have a an FAQ knowledge and a couple of specified knowledges for all the main problems that the contact center must resolve. And of course, the last one is to be real time. So, we get all the data at
[11:01] the real time. We ingest them in data base and we have MC APIs that wraps all the APIs that we need to, for example, create a new appointment or to make a quarter in the CRM of the Italgas company.
[11:21] To really simplify our solution, we, uh, started with a a layer solution. So, a layer architecture. The first one is the experience layer. So, the one that is direct and directly interacting with the user. Right now, we as we mentioned, there is a telephone number and the user can call
[11:38] and the uh voice bot will answer. Underneath that, there is the LLM layer. So, we have a GPT real time that is provisioned by Microsoft in in Azure Foundry. And we use a web app that basically plugs the real
[11:56] time regardless and all the layers underneath that. The main core of the solution is of course the cognitive layer. So, we can see that the cognitive layer is basically what enables the chatbot to think and
[12:14] retrieve all the information that you need. So, we have a memory and context management using Lakehouse, which is as you might expect something that you have for basically all the chatbots. And then we have all the MCPs wrapped in Databricks. So, we have
[12:30] native MCPs for the retrieval and the uh data APIs. And then we have uh an external MCPs for for example the CRM. Underneath all of that, we have the observability stack. So, we use MFlow
[12:47] for all the tracing, so all the calls. And we have a a custom solution to track the uh interaction uh within the application itself and gather all the feedbacks that we need to improve our system.
[13:07] Let's dive a little bit more deeply in our solution and let's see why Databricks was cho- chosen to be our backbone. So, this is a bit more in depth. Um again, the idea is to have everything
[13:22] modular or as modular as possible. And we use Databricks to basically as a the AI factory of the voice bot itself. So again, starting from the left, we have the experience layer. Uh we are
[13:38] planning on adding new experiences, new way to interact with the contact center uh that are more familiar to, for example, me and Serena. So using WhatsApp is the obvious one for European. And then we also are trying to create the a section in the public uh web app
[13:56] or website of Italgas, so you can interact with the contact center within a web application. The LLM layer uh we split it in two. So basically, we have the fast real-time interaction with the
[14:11] ChatGPT, with GPT-5 real-time, and the web app that I I was mentioning before. But underneath all of that, there are specialized vertical agents that does like the quotation part, the app part, the payment part, and other stuff that we are planning in the
[14:28] future. All the MCPs, as I said, are managed in within Databricks, and then uh all the data and the knowledge is also managed by Databricks, and we are storing everything within Unity Catalog.
[14:47] So how these two uh kind of solution coexist. Um Again, we wanted to be able to have a let's say pretty natural conversation in the voice bot. So the idea is to have the basic part in there. So the authentication, some FAQs,
[15:04] and some really common problems are directly plugged to the GPT real-time. And then we have an intent classificator. So every time the intent of the call doesn't align with the common problems,
[15:19] the orchestrator hand it over to the other agents. So, the other agents again are the specialized one and we build them to be called by the device automation, but also for other use cases. Right now, we have again the quotation
[15:36] one and the CRM one, which is the one we are using for updating or creating new appointments for the client. So, why would did we choose Databricks? The main focus is that we already have a
[15:52] pretty mature platform in there. Over the years, we built a comprehensive data platform where we have 95% of Italgas data already ingested. Most of them is batch, but we have also real-time and near real-time flows. And we also created the AI platform in
[16:09] there. So, all the classic machine learning is in Databricks and also all the uh early chatbots were ready in Databricks. And of course, Mosaic and MFlow a pretty good and comprehensive stuff for the AI solutions.
[16:24] Everything right now is stuck to MFlow, where you're registering every AI solution in there and we are using the gateway and the serving endpoint provided by Mosaic AI. And la- last but not least, Unity Catalog for us it's probably one
[16:40] of the most important and most used product within the Databricks ecosystem. It allows us to centralize and manage all the our access or all the lineage. And we are also using many metadata to enforce our solution and make it
[16:57] trackable for all the platform. So, let's go back to the problem statement that Serena was describing uh at the start of our
[17:13] presentation. So, let's see our client that calls again our new voice bot solution. Again, it's not real time and it's there is no video because it was in Italian so it didn't make sense. But I I walk you through the basic idea
[17:29] of the solution. So, our client calls again and right now we have a new voice bot that presents itself and starts speaking what the option are for the client. And the client can already speak to it
[17:45] and see what kind of response uh it gets. If the response doesn't checks out for the client, he she can already uh speak again and the voice bot will stop talking.
[18:05] When the client finishes speaking and the chatbot understands what the in new intent is or the new question is, it started again again to speak and hopefully the conversation progresses.
[18:20] So, how does it do that? Again, I think it's pretty clear right now but the idea is is to have a couple of different actions. We have the common FAQs and the common knowledge questions. So, let's say that the client ask what a PDR is. A PDR is basically the meter in Italian.
[18:36] Um the chatbot goes through the knowledge and it basically does a a retrieval and speaks back to to the client. Instead, if it does want a for example a new appointment or a change of appointment, we go to the MCP tools and
[18:53] we get the CRM to speak with the the voice bot and we pay the appointment with more authentication and more parameters. So, let's see how uh this new solution
[19:11] fit um compared to what we had before. Cuz I mean it's cool to to experiment and do uh really cool stuff with technology, but at the end of the day we have to run a business, so it needs to be effective.
[19:28] And the results are pretty good. Um we are in a kind of beta testing moment where we uh we are in production, but just uh with uh
[19:43] a number of call. It will be fully in production, so handling all the 100 um 20,000 calls uh in September, so almost there. But the results are really good.
[19:59] We There was a cut of 50% of dropped call. And uh the automated resolution called grew by 40%. So we are really um looking forward to see what September
[20:15] numbers will look like. Uh but again the the beta testing users are doing pretty great. To close this uh this talk, I would like to give you some takeaways, some
[20:31] some point some pointers remember. The first of all is, you know, IVR was a great technology, was really advanced uh at some point, but now um has a lot of limitation. So
[20:49] we can do better, I would say. The Agentic solution uh allow you to have a better solution a better uh user experience for your customers. So uh you know
[21:05] there is a a limited number of things that we can do in a day. So, uh beside the drop call, we had a problem that our war operator were um they had to handle too many calls. So, they were
[21:22] this wasn't good enough because uh they were handling very trivial problems and uh they struggled to answer to all the calls. So, with the Agenty version of IVR, we are now able to
[21:39] um let our operator only do uh complex really complex task that just a human can do at the moment. And uh the other thing is uh I think nowadays you can still understand that it's a
[21:55] synthetic voice, but I'm quite sure quite positive that we will reach a point where for the user will feel exactly like talking to a human.
[22:11] Uh the other things we had in meta in really good feedbacks. Um User were handling calls in a um handling calls in a much better way. Um they stay on call. They are not
[22:27] getting frustrated. So, the sent the sentiment analysis is really good on this side. Um I mean, we know that feedback is really important. Uh So, we are we are trying to uh grow uh on that
[22:44] part. So, we are trying to understand what uh is missing. And so, the next steps are of course answering to the clusters of questions that we are not handling at the moment. But at the same times we are, as Nicola said, trying to
[23:00] look to other solution where we can uh reuse what we already built. So, all the agent handling appointment and uh frequent asked questions or anything that we are doing at the moment. Uh and what we will do
[23:19] uh in a different way. So, integrating with WhatsApp or uh our website or something else that we will see maybe a a virtual assistant. know. Uh but what happens
[23:34] what will happen is that uh thanks to the architecture that we have, the data will be stay in the same place. It will be handled exactly in the same way. Uh so, with the same grants and permission
[23:50] and um security guardrails. And I will just do the integration part with the front end. So, this is pretty cool.
[24:06] I know it was a very unfair uh time slot and it's the last day. So, uh thank you very much for coming and listen to us. I hope it was a at least enjoyable and uh interesting. And if you have any question, we are more than welcome to answer.

Learn more about the Databricks Data and AI platform.

The information provided herein is for general informational purposes only and may not reflect the most current product capabilities or configurations.