Freshworks' Enterprise AI Platform: Building Freddy AI with Databricks
Summary
- Freshworks standardized their data infrastructure on the Databricks Data and AI platform with Delta Lake for unified ingestion from 40-plus siloed sources, Unity Catalog for governance and lineage, and MLflow for end-to-end ML lifecycle management to power Freddy AI across 75,000 customers in 120 countries.
- Their four-pillar strategy — Govern, Ingest, Democratize, Agentify — delivered 4 to 5x faster model training, automatic lineage tracking, and zero-ETL architectures while scaling from 1,000-plus internal users to fully autonomous agents.
- The video covers Freshworks' progression from data chaos with 40 disconnected sources to a governed platform supporting autonomous agents that complete end-to-end workflows without human intervention, using Databricks Agent Bricks and Genie.
Freshworks' Enterprise AI Platform: Building Freddy AI with Databricks

Enterprise AI requires unified data infrastructure, not just agents. Freshworks' challenge: serve 75,000 customers across 120 countries with Freddy AI while managing 40+ siloed data sources, non-production ML models, and governance gaps. Their solution: standardize on Databricks with Delta Lake for unified ingestion, Unity Catalog for governance and lineage, and MLflow for end-to-end ML lifecycle management.
Learn how Freshworks scaled from data chaos to a governed platform supporting 1,000+ internal users and autonomous agents. Discover their four-pillar strategy: Govern, Ingest, Democratize, Agentify. This delivers 4 to 5x faster model training, automatic lineage tracking, and zero-ETL architectures. Sreedhar Gade and Prem Kumar Patturaj explain the governance, model routing, and architectural decisions that make enterprise AI safe and scalable.
🤝
Chapters
00:00Freshworks' AI Native Platform at Scale01:16From Assistants to Autonomous Agents03:10Architecture: Medallion, Unity Catalog, MLflow08:58Past, Present, Future: Four Pillars Strategy11:41Building Scalable Data Pipelines13:18Data Democratization with Genie14:38Autonomous Agents with Agent Bricks17:37Lakebase, MLflow 3.2, AI Gateway Innovation21:34Governance and Safety Guardrails
FAQs
What is Freddy AI and how does Freshworks power it with Databricks?
Freddy AI is Freshworks' AI product that powers autonomous workflows across their 75,000 customers in 120 countries in employee and customer experience applications. Freshworks built the data foundation for Freddy AI on the Databricks Data and AI platform, using Delta Lake for ingestion, Unity Catalog for governance and lineage, and MLflow for model lifecycle management.
What were Freshworks' main data challenges before standardizing on Databricks?
Freshworks was managing data from 40-plus siloed sources in disparate formats, making it difficult for internal teams and customers to consume data reliably. Their ML models were not production-ready because the pipelines feeding them were too complex for legal and security teams to audit for lineage and traceability.
What is Freshworks' four-pillar data and AI strategy?
Freshworks organized their transformation around four pillars: Govern (unified governance with Unity Catalog), Ingest (scalable pipelines from all sources with Delta Lake), Democratize (self-service data access for 1,000-plus internal users via Genie), and Agentify (deploying autonomous agents with Agent Bricks that complete full end-to-end tasks without human intervention).
How much faster did model training become after Freshworks adopted the Databricks Data and AI platform?
Freshworks achieved 4 to 5x faster model training after standardizing on the Databricks Data and AI platform with MLflow managing the full model lifecycle. The unified platform also enabled automatic lineage tracking and zero-ETL architectures that reduced the data preparation overhead previously required before models could be trained.
Full transcript
[00:08] Good afternoon everyone. This is Shri Shriare. I lead cloud and data engineering for Freshworks and I also have Wasan joining me and uh two of us are going to kind of open up our playbook sort of on how we democratize uh data and AI at Freshworks. So but before I jump into the crux of the
[00:24] matter u I'll also talk a little bit about uh Freshworks. So Freshworks started in India 2010 uh with very humble beginnings as a single product company and uh focusing on Asian markets but then very quickly we have uh
[00:39] you know um grown so well that you know today we are like a multi-tenant and uh you know we serve our products in close to 120 countries in 40 different languages and uh serve our products in employee and customer experience fields. So that's the uh you know the u the
[00:55] customer base is is is growing as we speak but uh what is more notable on the right side is uh we are the a native company right now. So every interactions every workflows uh that are in our products right now getting transformed using AI and with AI.
[01:16] So the whole premise of this presentation is all about customers are no longer uh looking for an assistant AI solutions. Um this is not about just being a co-pilot but this is about being a completely autonomous agent which automates their workflows also completely takes not just the retrieval
[01:32] part of it but also the actions and then completely end to end tasks to achieve the outcome that customers are looking for and that's exactly the direction that we're actually building our products with and uh so but before I um you know jump into the actual architecture and how we actually went ahead and solve solving our problems so
[01:49] here are the three uh you know problem statements that we looked at. So when we started our journey of our delta lake the data was not ready 40 different sources multiple different tools um disparate uh formats and that made it so complex for internal teams as well as
[02:06] our customers to actually consume the data right and uh also if you look at it the uh uh the uh ML right the the whole the models that we are building internally also are consuming from outside it it is really if you look at it they are basically is not production
[02:22] ready. So uh uh when I say that the whole pipelines that we were talking about they were actually like uh too complex for us. So it is not very easy for our legal or security departments to uh go back in time and look at the lineage or some kind of a traceability for us to actually uh really justify
[02:39] whether are we actually building a safe or not right and and uh due to these two reasons we have low productivity across the teams and which is not very obvious at that time but now when we actually look back in hindsight uh we could have actually done lot more things than what
[02:54] we did in the past. So I will uh invite Wasan to quickly uh come on and you know talk about how did we leverage databicks partnership to build our data lake and uh delta lake which actually helped us
[03:10] you know address most of these uh you know challenges that we talking about. Yeah, thank you Shri. Um so I'll just uh walk you through at a high level how uh freshwork solved for the three problems that uh Shri talked about right
[03:26] um this is a high level architecture of u their uh entire data platform um the foundation of this architecture sits in the open lake house um they have plethora of sources 40 plus sources that get um the the problem that you saw
[03:44] previously ly shri was highlighting right they have 40 plus sources u it's been a siloed experience uh if they have to build AI they have to reach out multiple sources so fresh figured this out like way back in 2122122
[04:00] that they needed to unify this they needed a foundational uh uh data layer so this medallion architecture was set up uh with open lakehouse and data bricks jobs um were uh built to bring in data from multiple sources. It supported
[04:16] both real time and uh near realtime requirements along with the batch requirements. Um but the cherry on the cake was Unity catalog. Um to give you a sense of how uh the scale of Freshix uh is uh they operate across five plus
[04:32] global regions. Um they needed this kind of architecture setup across all these five regions and not just that they needed to integrate data between multiple regions. um get an aggregated view and all this complying to the local uh regulatory requirements. So how we u
[04:51] freshworks manages about uh if you're already familiar with data bricks workspaces freshworks manages about 100 plus workspaces that is across their business units and u region. So what unity catalog gave them was a easier way to share data across regions. Previously
[05:08] they used to duplicate data from one region to another region. They had to copy it, create a duplication process but Unity catalog gave them an easy way to share data without duplication. It gave u trust to the data. So now what
[05:24] happens is when you see the data flowing through um and getting uh written into multiple layers the bronze, silver and gold um there is automatic lineage that is built. So anyone that's using the data they can actually get the birds of view of where this data came from, how
[05:40] it's being transformed, how it's being used. So it it helped in discovering data uh enabling trust in the users and so we all we saw that ML was not production ready that was one another issue that we talked about. Um so
[05:57] Freshworks leveraged ML flow to uh manage their end toend life cycle. uh they um starting from data analysis, explorative data analysis, uh feature engineering, training of the models and um inferring them they it was a
[06:13] discontinued process in the legacy system. With MLflow, they streamlined the entire process. Now um Fresh Freshworks actually builds tens and thousands of machine learning models for uh customized for every customer. So this MLflow uh framework gave them an
[06:31] easy way to build u automated systems right from data to their uh end inferencing and not just that their internal and external team started leveraging the data seamlessly. So u uh
[06:46] the external teams who actually use freshworks get analytics on top of it. They get to know um for example if uh anyone who is using one of their products fresh service would get to see the analytics on top of it how many tickets were raised how many tickets were resolved what were the high priority tickets and so on and so forth
[07:02] right those analytics started coming in uh easier uh the cleaner the data were um the machine learning and a models were able to give the insights u more faster and uh reliably and all their uh internal teams like product marketing
[07:19] and finance started leveraging this data and um enabled uh democratizing democratizing data across the uh internal use cases.
[07:36] So Freshworks bet was simple. They needed one platform, one governance layer and that would be leveraged by every single workload in their organization. And this bet actually paid off. Um they started the journey of uh setting up this uh unified lakehouse about 2022. They were able to unify 40
[07:53] plus sources uh build a medallion architecture. They had the source of truth of data. They had cleansed data redacted data um PI enforced PI policies enforced on it. Um and uh readily available for uh all the users.
[08:10] they they are now catering to more than thousand plus users internally who are actually actively using this uh platform and of course unity catalog is one of the um seamless way to give the metadata management and governance layer across
[08:25] this organization and the results were um the Freddy AI which is their u uh AI system that's called Freddy AI that uh became four to 5x faster in turning models for their end customers.
[08:42] Um and the entire ML flow which was standardized uh for uh the ML uh machine learning life cycle completely was automated. So onboarding new getting new machine learning models became uh easier and of course they were ready for the
[08:58] agentic revolution that came on came on. So now a and agent can leverage uh the data that is used. Now I'll give it back to Shri to uh talk about um their 2026 data strategy and so forth. Awesome. Fantastic. So, so when you
[09:14] looked at the architecture diagram right from the RDBMS the source of truth multiple different data sources to the pipelines um and the event system in the middle right Kafka and then the the consumers on the data side multiple more than 100 different uh workspaces across
[09:31] internal IT plus also the customerf facing uh you know both AI and non-Aspaces everything is completely revolutionized reduced duplication the performance and throughput is uh next level So we spoke about past, we also spoke about present. So let's talk about the future. Now where are we going from
[09:47] here? So we are looking at these four pillars you know govern, ingest, democratize and agentify uh as a mantra for our go forward strategy for our data and AI at freshworks. So let's go double click into each one. Right? So when you
[10:03] look at u governance essentially uh here is our fundamental philosophy is uh the trust is um non-negotiable and uh the agents which are not governed will not get shipped as simple as that right so in the unity catalog like you can see um
[10:19] you look at the arbback right so and also aback which is basically you know attribute based access control and we have the data classification where automatically you have you know pi redaction and then also data data tagging and we also have end to end
[10:35] lineage. Uh this is where uh it really helps your legal and security team to exactly how the model was trained, how that particular um answer was given. So these are the areas where you know we invested and then you know uh amazing amount of governance and then you know we continue to build on top of it and
[10:52] why did we do it because agents um you know was actually giving a keynote uh you know a few minutes ago in the executive launch where um when it comes to humans we are very very careful about you know the permissions and the back but when it comes to agents you know uh
[11:08] we don't think as much but they're super powerful they can actually go you know what I have access to this this this let me actually go and clean up there a lot of horror stories we are hearing across the board. So this is exactly what we are trying to preempt and kind of a proactively protect ourselves from uh this situation where ungoverned AI
[11:25] doesn't get shipped automatically and then you can see a nice scale you know with which we are actually uh uh operating at this point we are talking about thousand plus users um of data bricks or the solutions built on top of it uh with a goal of you know multiplying that over the next few years
[11:41] but also multiple you know uh more than five geographical regions across the world and 100 place workspaces across those regions. Now let's talk about uh data like pipelines. So the pipelines is all about uh you know the one is you have producers like which is source of
[11:58] truth all the different types of systems like you may have RDBMS systems you may have uh salesforce or any other source of truth you know where actually system of records that are creating data and then you have pipeline and then you have consumers on the like data lake or data
[12:14] warehouses. So the entire uh flow of pipelines is is actually made so easy with the lakeflow connect where uh there are like by default like 40 plus connectors are available so you don't have to hand code anything so plug and play and then suddenly you see the data getting um you know flowing into your
[12:30] lake data lake right and then there's also lakeflow designer where which kind of helps you make this whole journey much easier right they just drag and drop and then depending upon how your environment is um so you know in no time you'll start seeing data in your delta Right. And then you have a spark, right?
[12:46] This is where we do all the transformations at at a pabyte scale and everything actually works like really super smooth. So with that you can see some timelines there. Any new uh workspace onboarding used to take weeks for us especially the uh since it comes from a different system onboarding of that system just to make sure that the
[13:02] data is is massaged and like you know it's kind of transformed the way we need it and uh it's also onboarding to database and make sure that your consumers are able to consume. Now it uh you know is under like 48 hours right now and then we are reducing even even further. So agility is much much higher
[13:18] due to lake for pipelines and then democratization. So the genie uh the philosophy here is ensuring that every employee at freshworks becomes a data analyst. They can conversationally speak to genie. For example, if I'm a finance person, I'm FPNA. I I can basically talk
[13:34] to my data and then get some representations in terms of why uh these numbers are financial numbers are looking like this. you know if the the sales is up or down, why gross margins are the way they are, all those insights you can actually get just um you know simp through simple commands. So Genie
[13:50] kind of you know enables us that so we have a goal of you know deploying this across the company as well. So you also have this AIB static models like they continually learn about your organization and they enrich the output as we speak. Um we also have uh you know
[14:05] MCP right so this is how we um you know your genie also can integrate with your GitHub so essentially all the transformations that you're doing within datab bricks you can actually go and update the GitHub repository and then also ask your through genie uh you know
[14:22] u uh query that uh repository you know make some changes and then raise a pull request and execute it as well. So there is a a a small scale you know engineering SDLC pipeline built within data bricks in that space you know to automate uh
[14:38] various tasks within your workspaces as well. Um last but not least is agent bricks right. So like I said in the beginning it's not about assistive a technology but it's all about autonomous. So u
[14:53] agent bricks has been I think the good partner for us. Um you know Freshworks internally uh you know has been uh experimenting a lot in in agent space where like we have been building our own agentic uh platform and uh you know the MCP gateway skill library making sure
[15:08] that these agents in both our CX and EX space continue to uh deflect customer requests uh with completely bypassing human agents and and that in that journey I think we have been able to achieve together a lot. So especially you know if you look at it you know autogenerated system benchmarks right
[15:25] and then uh supervisor agent like especially watching other agents perform and then you know continuously make sure that you know they perform at the level are they drifting uh are they continue to you know performing at the level so what are the OKRs for those agents and are the outcomes are the ones that they
[15:41] exactly that they need to be right so um we learned a lot from the agent bricks and in the internal IT so we are kind of building vertical agents for finance finance um HR talent acquisition. So where there's a specific training and a
[15:57] specific data set that is being used for example recruitment and u through this agent bricks we are looking to you know accelerate some of those internal processes but at the same time uh it's kind of AI democratization very clearly across the organization it's not just engineering it's about um you know any
[16:15] teams who have never touched um never coded in their life never built uh you know any code in their life they're they're actually coming in and then looking at uh the workflows that they have and see how we can actually use AI to automate some of these workflows and then see the outcomes very very very uh
[16:30] quickly and and uh like I said in the the pipeline building so especially connecting like a workday or any other systems is so easy so the data is flowing very fast uh very quickly and seamlessly and you are able to actually uh train your agents to do exactly what
[16:46] you want them to kind of do so these are the four pillars is our uh core focus and priority uh you know this year and and uh like know early next year as we uh transform experience of our customers like I said we have close to 75,000 customers in 120 countries and probably
[17:03] like the millions and millions of human agents that were deployed across all these uh products now uh these agents that are coming in uh will help if the customer support a customer of my customer uh who's actually reaching out for help um you know irrespective of language irrespective of time zone
[17:20] irrespective of the training levels they're able to give same level of uh experience and they are able to learn they are able to you know get uh better over time as well. So um so then the next step right you must have heard about lakebase uh this is u another you
[17:37] know major innovation and uh uh the lot of announcements happened like yesterday and today uh on lakebase. So essentially this is an OLTP uh workloads that we have but then you know you have entire creation uh management governance and then scaling becomes very very smooth
[17:54] and uh seamless. So we have some experimental workloads on lakebase at freshworks where um I think it's a millisecond uh you know submillisecond level scale up scale down agility we are able to see which means if no one is using you basically scale it down almost
[18:09] to zero and if uh you you see sudden spikes you don't have this lag that kubernetes has basically scale up and scale down so uh generally you you keep it to a maximum or like 80th or 90th percentile capacity lot of wasted
[18:24] capacities is sitting there on your compute those things will uh eventually die down as we speak right so um so in in Freddy AI uh as well I think there is a lot of u you know persistent context is one area that we actually looking to
[18:40] build this is all these is able to token optimization and then you know how quickly we can empower our agents to you know uh move forward and no database tax I think that is um that has been a traditional challenge for all of us um especially source of truth. Uh how do
[18:56] you actually make it uh work like a zero ATL? So there's no lag between um all the transformations and entire pipeline and then also uh error rates especially if you're transforming so many times um then the source of truth and then your end lake generally don't match with each
[19:12] other. There are there's error rate that is actually introduced which will show up in your reports sometimes uh the responses and so on and so forth. But if the the lake and the you know uh database is the same essentially you have the one is you are saving on the storage you're also saving on the entire
[19:28] transformations u lag as well as it automatically becomes your uh zero ETL so u other couple of innovations that we have been also you know um historically working on is mlflow 3.2
[19:43] So all the self-hosted models um we kind of um train them on predacted customer data and uh these pipelines are are built using ML flow. So what basically happens is you have a data source training data and then you also have
[19:58] models and then they have certain cadence that they go through uh through this model and over time uh you also continue to look at how they are performing and is there like you know when when it is time for them to actually get trained. So uh MLflow has been I think uh very very helpful in that space right um it's also kind of a
[20:16] u like uh they're saying like you know database is actually multi cloud it doesn't matter it's a GCP or azure or AWS so you'll be able to actually um talk to all these uh models like you know likewise in every place uh one place and a gateway especially if you're
[20:31] an enterprise company using multiple models at play right and there are disparate models there could be frontier models as SLMs um some of them are self-hosted, self-built, some of them are actually managed models. So, all of them how do you actually manage them in a single place? So, the AI gateway is
[20:46] helping in that space. So, and then uh every model that is coming from outside or inside gets terminated and then your application is only talking to one endpoint but then that endpoint is actually intelligent enough to look at uh scale, cost, uh performance in some
[21:02] cases functionality as well. So you are able to intelligently switch between uh models in in the real time on the fly depending upon what type of workloads are coming in what type of u you know you can actually set some triggers in the cost um or you know some kind of a
[21:17] temperature control and those are the settings that you can actually have. So it's a kind of a smart model load balancer sort of. So this is also like something that uh uh we have been uh you know experimenting with
[21:34] right. So at Freshworks we believe that uh governance and uh trust is imperative right everything um when we initially started our a journey it was an afterthought but I think last few years we have been I think building it inside out so um that is something is uh it is in the catalog but it is actually it is
[21:49] not governed so we want to make sure that every single asset that is there in both the data as well as the a side is actually goes through our guard rails so we partnered with the uh AWS Microsoft prompt safety filters, AWS guardrails and also homegrown uh safety filters
[22:05] that actually track every token goes in and out of um our product and and we look for multiple different areas you know safety, privacy and or back any kind of traceability issues uh and then they are logged or blocked depending upon you know severity of that particular issue right so
[22:23] uh to sum it all up right um so the we have realized immense value till date like like you can see you know um four to 5x in the custom model training multiple data sources like more than 40 50 got uh unified you know with uh everything is coming to our delta lake
[22:39] and also uh we have a uh you know we have a path forward in terms of how we want to actually manage uh essentially otherwise initially for all of these solutions we are talking about agent framework ML MLOps framework we are
[22:54] talking about guardrails um ML gateway for each of them you have to go to a different company or a different solution or you have to build hand like you know hand build something but in this case it is it's all available in single place and then uh remember it's your data is where uh AI makes more
[23:11] sense and then data plus AI is in a single place so the data movement is actually very minimal and then that's what gives you maximum value and in 2026 we are talking about multiple investments in the space so Genie goes across most departments so that's how we are democratizing data uh AI and data
[23:27] within the company so all those PowerBI reports, all those Tableau reports, um you know, uh all those PPTs. So eventually I think they will disappear and these are like a live realtime conversational analytics is what going to take over and uh and then last but
[23:43] not least, so this is a 360 partnership. So we are customer for data bricks and databicks is customer for us. So their entire fleet of ITSM uh and employee service management and IT service management runs on Freshworks. Uh and likewise all of our internal and
[23:59] external data warehouse and data lake and delta lake actually runs on uh data pix.
Learn more about the Databricks Data and AI platform.
The information provided herein is for general informational purposes only and may not reflect the most current product capabilities or configurations.