Building AI-Powered Financial Products: Hybrid ML and GenAI at Caixa
Summary
- Caixa Econômica Federal, Latin America's largest public bank with over 150 million customers, combined machine learning propensity scoring with LLM-based digital twin validation to increase offer conversion rates from 3–5% to 21%.
- The hybrid AI pipeline runs on the Databricks Data and AI platform using medallion architecture for data preparation, MLflow for model versioning, and Unity Catalog for secure model registry and governance.
- Auditable decision logs and fairness monitoring ensure regulatory compliance, demonstrating how ethical AI and rigorous data governance drive both customer trust and business results in financial services.
Building AI-Powered Financial Products: Hybrid ML and GenAI at Caixa

Personalized financial offers require balancing customer interest with suitability and responsibility. Caixa Econômica Federal, Latin America's largest public bank, built a hybrid AI pipeline combining machine learning for propensity scoring with generative AI for ethical validation. This talk presents real-world results from their My Wallet app, where behavioral segmentation, Random Forest modeling, and LLM-based digital twins increased conversion rates from 3-5% to 21% while filtering offers for genuine suitability.
You'll learn how Caixa implemented this at scale using Databricks: medallion architecture for data preparation, MLflow for model versioning and governance, Unity Catalog for secure model registry, and the AI Playground for comparing large language models. The system includes auditable decision logs for regulatory compliance and fairness monitoring. This case study demonstrates how ethical AI and rigorous data governance drive both customer trust and business results in financial services.
🤝
Chapters
00:00Introduction and Caixa Overview05:45The Opportunity: Open Finance and Regulations07:23The Technical Approach: Digital Twins08:12The Challenge: Precision Over Spray-and-Pray09:04Building Digital Twins10:10Machine Learning Pipeline: Segmentation and Propensity13:23LLM-Based Digital Twin Validation15:52Results: Filtering and Conversion Success21:34Technical Architecture with Databricks22:40ML Engine, GenAI, and Governance Layer24:05MLflow and Model Registry27:21Key Takeaways and Future Roadmap
FAQs
What is a digital twin in the context of AI for financial services?
In this video, a digital twin refers to an LLM-based model that simulates a customer's likely responses to financial offers, acting as a validation layer after ML-based propensity scoring. It checks whether a proposed offer is genuinely suitable for a specific customer before the offer is extended.
How did Caixa increase conversion rates using AI on Databricks?
Caixa combined behavioral segmentation, Random Forest propensity modeling, and LLM-based digital twin validation to filter offers for genuine suitability, lifting conversion rates from 3–5% to 21%. The system was deployed through their My Wallet app and built on the Databricks Data and AI platform.
How does Caixa handle regulatory compliance for AI-driven financial decisions?
The system includes auditable decision logs that create a traceable record of why each offer was made or filtered, supporting regulatory compliance and fairness monitoring. MLflow model versioning and Unity Catalog provide governance over the models used in the pipeline.
What is the role of the Databricks AI Playground in Caixa's project?
The AI Playground was used to compare large language models during the design of the digital twin validation layer, helping the team select the most appropriate LLM for their use case. This evaluation step was part of the broader architecture built on the Databricks Data and AI platform.
Full transcript
[00:09] So, good morning everyone. It's great to be here today and share this story with all of you. But before I start, I would like to ask you a quick question. So, raise your hands if you've ever heard about Caixa Econômica Federal, please. So, few hands. I see some Brazilians in
[00:27] the room, right? So, what makes this interesting is that we are talking about one of the largest financial institutions in Latin America. So, uh Caixa is responsible for uh looking for many governments and
[00:43] social programs in Brazil. And uh we are here to tell you a real story of what happens when you combine micro offers with data and AI at scale. So, I'm Jessica Oliveira and I'm solutions
[00:58] architect at Databricks and responsible for Caixa. And I'm here today with Andre to help me tell this story. He's a data scientist at Caixa and one of the key people responsible for this project. So, and together we'll walk through the journey, the results, and the impact we
[01:16] achieved with this project. But before we dive into the solution, I would like to to introduce you Caixa and some of their operations, right? So, first of all, Caixa it's a public bank. And you will understand why they need to
[01:32] be. So, they are not just a traditional bank. They're actually the backbone of social and public policies in Brazil. So, they're responsible for managing and distributing social and public governments uh benefits for millions of
[01:49] customers, right? And let's take a look at the numbers behind this operation. So, Caixa was founded in 1861 and more than a century and a half later, it remains deeply connected to the daily lives of Brazilians. And today
[02:06] they have more than 150 million customers. So, that's more than 70% of the entire population of Brazil. And uh they manage trillions of reais in assets, okay? And
[02:21] what is really impressive about Caixa is that their magnitude is not measured only by the customers they have, but also visible in their physical presence across the country. So, they have more than 2,500
[02:37] uh physical branches in the country and more than 26,000 correspondent locations providing services such as bill payments, deposits, uh balance inquiries. But in a country as vast and diverse as Brazil, sometimes
[02:56] uh this network extends traditional branches, okay? So, it can include even trucks and boats. So, you might be wondering, trucks? Yeah. Caixa operates three mobile
[03:13] uh branches mounted on trucks that can be delivered wherever they are needed most. So, uh once on site, this kind of branch offers many of the services available in a traditional branch, but it also serves
[03:30] as a marketing platform for bringing some programs like debt renegotiation and uh and support for micro uh interpreters and also, for example, uh rural credit
[03:46] programs for farmers in locations far away from the centers of the city. And if the trucks weren't surprising enough, there is more. They also have two boat branches that navigate through the Amazon rainforest rivers. So, these
[04:03] two floating branches navigate these rivers and connected the most remote areas from Brazil to the financial system. So, picture uh picture that now. Uh for some customers of Caixa, the nearest uh branch it's not in the next
[04:21] uh in the next street or in next city or close. It's a boat traveling through the rivers in the Amazon. So, this is really impressive. And during the pandemic the the COVID pandemic, right? They
[04:39] had they was defined by the government as the the bank responsible for distributing and managing social benefits to the people uh affected by the lockdown. And at this moment, they opened
[04:55] uh more than 100 million accounts, digital accounts, in few months. And the most impressive number here, it's not the number of accounts that they opened. It's the people behind them. So, at this moment, more than 67 million Brazilians
[05:13] had have access to the financial system for the very first time. So, uh moving from a cash-based reality to a digital system and financial system for the very first time. So, by the end of the program, they have
[05:28] distributed more than almost 300 billion reais to to the people, right? And why now? Let's talk about a little bit more about the project and why we are the why the project is is dealing today.
[05:45] So, uh over the last years the environment around Caixa has changed a lot. So, they're among the top three institutions receiving the highest number of open finance consents in Brazil. And
[06:02] open finance there has achieved have achieved a new phase. So, customer is now allowed to share a much more broader information such as investments,
[06:17] insurance and for example foreign exchange data. And with this amount of data Caixa had a really challenge to deal with all this data, but they there also another
[06:34] challenge in in in in the in the project, right? The other one is related to the new regulations data privacy. So, with this new regulations Caixa had to deal with responsibility with more responsibility with this all
[06:51] this data, right? So, this opened a new opportunity for them. So, the opportunity to use a more deeply information of the customers to build a real personalized experience for the customers. And
[07:07] for this method to run this at the scale you need a modern and a data and AI platform to be able to deal with all these challenges and with all these data regulations and privacy regulations. And that's why I'm calling to the stage
[07:23] Andrea to explain to he's deeply connected to this project. So, he will explain what they did to to this program to this project. Thank you, Jessica. Hello, everybody. First, I want to give a special thanks to my boss Rafael Monareto. He designed
[07:39] in this project and managed it starts to end. I just this English speaker and the the bricklayer, later bricklayer. Also, he it's the master degree thesis. It's a good person.
[07:55] Uh in Caixa, like uh other banks, we have a mortgage campaigns. And for example, we have um million contacts in a campaign and that we we needed to to do the things in the
[08:12] And in a regular campaign, we have three to five percent conversion rate. And we we thought how we we can achieve more precision in our campaigns cuz brute force is great friction. Flooding a customer base with generic offers destroy trust and saturate digital
[08:29] channels. Because uh propensity it's different than eligibility. We have models that show that customers have propensity to accept an offer.
[08:45] But uh this customers wants this offer, this offers is suitability for him. We need a cognitive filter that understand the human contest.
[09:04] And we did this with we thought to create a digital twin of our customers. What is a digital twin? In many industries like um Tesla industry, have a a sewing machine and and the production line.
[09:19] And we have a copy of this in a computer. And so in the in this industry, for example, uh have sensors in this machine that read the machine in real time and and see you when about to break and
[09:36] have a about to have a maintenance. In the airplane industry also test the the wings in and just out to wing and then the rocket to we test the wings.
[09:52] And why we don't do this to the final sun dream? Cuz we we can't and how we achieve this? Cuz we we can't pick our customers and plug sensors and see what here are thinking.
[10:10] So we did it in three steps. Step one, behavior segmentation. We used K-means to clusterize the customers and 11 groups. We looking for liquid available balance,
[10:27] monthly income, cash flow, consolidated balances, financial maturity, investment history, and so variables to understand our customer and classified.
[10:51] so we picked all these variables and run a a machine learning in random forest. And we have a accuracy of 83 percent. When I do my models, I like to the accuracy standard 8% to
[11:07] 90% cuz if you run a model with 100% of accuracy, something is wrong. It's over fitting. And
[11:25] And why we use the random forest not XGBoost or logistic regression? We try a a of models. And we pick it random forest cuz random forest uh have how to plot the decision tree. And
[11:41] random forest is a model that uh create many decisions trees of the customers. And it's and our customer go through a path or the path. For example, um and we have a
[11:56] a customer that for example have a balance of 5,000. They go through a path and after I don't know, uh 30 years old they go to other path. And we can see this in random forest. It's easy to see. It's easy to show to the
[12:12] stakeholders. And the bank we have a strong vulnerability in the isolated environment. So, it's good. And I have a AUC of 0.5%. KS of 0.19.
[12:35] AUC is is a metric that measures the ranking. And KS measures the counter. And this is the the chair of the cake. The this is how
[12:50] train. We run all this uh in the data bricks. Uh I remember when I became to go to this this project. I tried to download uh
[13:06] a model of hugging face and it didn't work. And I showed the project to to Jessica. And no, uh see the the playground section of the data bricks. You can use AI directly in data bricks. And I want to thank you cuz this project
[13:23] it's because you you have helped us a It's really important. So, in it is a step three, we use a digital twin with a LLM.
[13:39] We have inputs. We pick the clusters that start once. The means of the dispersions. Max and min. And create a profile of the customers.
[13:54] Cuz um we don't run the prompt to each customer individually. It's expensive. We run in a batch of 300 customers for a prompt. To save tokens cuz tokens they use
[14:11] expensive, isn't it? And you you take the financial snapshot of the customers in a CSV file and send to this AI.
[14:28] With uh And after just go to the the prompt. We don't need to teach uh AI what's a digital twin. Each already know what's this. We don't need to teach the the priest
[14:45] how to preach. So, uh we asked she, "Does this offer make sense for this specific user right now?" So, we put all this data
[15:02] and send it to respond in a CSV file also with uh the ID of the customer, yes, no, and uh justification. Uh why this offer is relevant or
[15:18] why is not. And uh he started to to say yes and no and justify it's amazing like no don't don't proceed for this customer because he don't have balance or no it's
[15:34] so young it's or it's not her profile or yes yes you can proceed this customers already have investments with us.
[15:52] Based on my reality today does this offer make sense for me? They I make this question for the customers and available in real time. This is a human filter bringing suitability to the process because
[16:09] propensity is different than suitability. So the numbers. We start with 400,000 customers
[16:24] that are from our app my wallet an app that use open finance to show the customers our data from Caixa and other banks. And we
[16:40] we pass it through the ML model. And the ML model say that 200,000 customers are propensed.
[16:58] And after we put a filter before the AI only for the people have who have available balance to do the investment. We we don't want to offer investment for people don't have balance it don't make sense make sense. And the
[17:13] AI in the last step the are qualified 6 and 1/2 thousand of customers. But you you can think why you are struggling our base.
[17:32] That's uh we can do 200,000 offers. Because uh we live in a world uh full of ads. We install an app on in our mobile phones and we go through it and
[17:49] and disable the notifications. We don't want to know our customers disable the notifications. We want we can reach them. We don't want to be marked as spam. So, we are qualifying our targets cuz in
[18:06] it is rather uh wisdom is to intuitive the customers' wish. We must be on time to understand our customers and give relevant offers.
[18:22] We can't We cannot more spray and pray. We must be precision and the ML engine and they I give this precision and this accuracy.
[18:43] And after this goes through our app not in a section called my recommendations and the customer can see the recommendations app and go personifications, SMS message, and can be other digital channels also.
[19:06] And the results are amazing. 21% uh of the people accepted the offers against 3 to 5% in a traditional benchmark. We have a great asset service jump. We give them more accuracy offers to our customers.
[19:23] We end up with spray in the behave uh governance of this. And what we know what's next? Oh, we have
[19:38] uh current limitations. Synthetic AB based on prediction probabilities, like house validations, metrics uh limitation selection rates, no formal fairness, magic blend of the ads.
[19:54] Twin scope capture of financial behavior but misses emotions or physical matter app. We don't have yet a way to understand the auto or customer are feeling.
[20:10] Selection bias model of the trainers are restricted to active apples on lines for the customers that opt for open finance, but we want to expand to our customers. And what's next? While I am talking here, my tasks are growing in Brazil.
[20:31] Uh expansion extend the pipeline to credit answers and other products. I will have modeling implement models for incremental campaign impacts. Federated learning enable cross
[20:47] institutional learning safely. Causal AB real random tests with conversions and return of MS KPIs.
[21:03] And it's just going to know we will speak how we build this inside Databricks. Is it fair, Jessica? Thank you, Andre. It's always inspiring to see real projects with real results.
[21:18] And what is funny is that I'm SA responsible for Caixa, but I'm also their customers. So they're they're customers. So while we are here, probably there is a model from Andre analyzing my data and think about my next offer. So this is really good. And
[21:34] now we are going to understand how did they make all this possible using Databricks and our platform, right? So this is our technical blueprint. So the architecture is here. And in the ingestion part, we combine
[21:50] many different data. So the open finance data with more than 57 million API calls with internal systems data. For example, CRM, credit cards, uh customer complaints. And then all this data went to uh
[22:06] medallion architecture. So the raw data is in bronze. In silver, we had the an anonymized and uh treated, cleaned, filtered data in silver. And finally, uh the customer 360 view in the gold
[22:23] layer here. Next, all this data feeds into our two engines. The first one, the machine learning engine, it was responsible for the segmentation of the customers. And the GenAI, it's where uh is which they
[22:40] they have the smart agent that decide, for example, if they may or may not or may not offer some product to to the customer. And as you can see here, this is a really important part. They have a unified governance layer to all of this
[22:57] project. So since the ingestion from the data treatment to the engines from machine learning and gen AI using the same governance interface, right? And this is a really important part of the project, right?
[23:12] Uh how can they deal with the data? So I always say uh great model with bad data will rarely succeed. But a simple model with excellent data can bring powerful uh powerful success and results to to
[23:29] our customers. So Databricks SQL was the hero of this part. The team, the analyst teams, and the data processing teams was able to create all the the the treatment of the data and reliable pipelines using the the Databricks interface. And uh
[23:48] now we can move to the model part, right? So with the data ready, the data scientist could build the models using MLflow. So remember when Andrea mentioned it, uh comparing different algorithms such random forest and XGBoost, MLflow was
[24:05] there to do, for example, the comparison side by side of all the models. So they could see all the the the results, the accuracy, and other metrics to decide what is the best model. And once the model was defined, it can be
[24:23] uh it will register the Unity Catalog. So Unity Catalog is a central repository for the approved and the champion models to to all the projects, right? And another part, very important, that they used it MLflow, it was to track
[24:40] everything. So legal team and auditors, they need to know why AI is how the AI is making the decisions. So MLflow was very important here because it tracks every interactions. So sees the
[24:57] questions, the answers, and all the experiments, all all of this is tracked and is presented in the history log, a complete history log. So, if the the auditor asks you, for example,
[25:12] why an offer was blocked for a customer, they can look to the history log, get this information, and send to the auditors. So, this is really important to productionize uh projects like this one. Uh Andre
[25:28] talked about the AI playground, but this is was very important, also. So, in this part, and nowadays with many AI models available, it's really important that you can choose the best one and with
[25:44] lower cost to solve a specific problem. So, they use the AI playground to put different AI models side by side, also, and ask the same question with the same prompts or different ones to choose which model
[25:59] will answer the best the their questions and with the lower cost lower cost, right? And what is interesting here is where you define your model, when you choose the the right model, you can export all the code to an agent, for
[26:15] example, and after that, register it in in Unity catalog to use and to share with your your partners. So, to wrap up the architecture here, uh everything converges into three main pillars,
[26:31] right? The first one, uh using MLflow to versioning, but not only the code, but also the questions, the answers, and everything every interaction with this agent. The second one, security, so having interface to
[26:47] secure everything and to audit everything is very important. And the third one, the monitoring. So, what is really important here also is for example monitoring data drift. So, using the MLflow you can be alerted
[27:02] when any any important changes happening with your data. So, this is really important also. And to finish here, we have three takeaways that you can take home today. So, the first one is
[27:21] hybrid AI works. So, combining machine learning traditional machine learning for numbers and generative AI to do an ethical future is really great and works for a scale as you can see in the
[27:36] Caixa's reality. The second one is scale demands a unique platform. So, imagine if you had to create an access policy to data, another one to data processing, another one to ML, another one to
[27:52] to gen AI or AI. So, this you you will be you will you will be more time doing these policies than actually creating solutions and taking results in the end. And the third one that I consider very important is that
[28:10] protecting your customer and being responsible could increase your number because as you can as we proved here, we they protecting their customers, they would they were able to increase the conversion from 3% to 21%. So, this is
[28:29] really important here. And now we have a new challenge here since the World Cup is happening here in the US. This one is really difficult, I know, but cross bring it cross bring us to to get our sixth star, okay? Now, we are
[28:46] open to questions. If you had any questions, we are here. And thank you for your presence. And thank you Andrea and Kaisha and everyone who is here.
Learn more about the Databricks Data and AI platform.
The information provided herein is for general informational purposes only and may not reflect the most current product capabilities or configurations.