Santander's Digital Transformation: Migrating 16 Petabytes to Cloud for Data-Driven Banking
Summary
- Santander Brazil completed an 18-month migration of 16 petabytes from on-premise Cloudera and ScyllaDB to Databricks on Azure, scaling from 8 to 35 data domains while reducing data access lead time by 40 to 50 percent.
- The transformation succeeded by securing business ownership before technology decisions — creating new roles of data owners and stewards — and embedding federated governance into the organization so that data quality became a business responsibility, not only an IT concern.
- A medallion architecture with Delta Lake, security by design, and four years of preparation achieved zero business disruptions during migration and grew the active user base from 1,000 to nearly 5,000 while positioning the bank for AI-driven innovation.
Santander's Digital Transformation: Migrating 16 Petabytes to Cloud for Data-Driven Banking

Santander Brazil's Move to Cloud initiative demonstrates how a 140-year-old bank transformed its relationship with data through federated governance, business ownership, and domain-driven architecture. The program migrated 16 petabytes from on-premise Cloudera and ScyllaDB to Databricks on Azure, scaling from 8 initial data domains to 35 while reducing lead time for data access by 40 to 50 percent and growing from 1,000 to nearly 5,000 active users. Built on principles of business engagement before technology decisions, creating new roles like data owners and stewards, and using medallion architecture with Delta Lake, the platform established trusted, reusable data assets across customers, channels, products, and services. Two years of preparation, 100-plus team coordination, and emphasis on security embedded by design ensured zero business disruptions during the 18-month migration while positioning the bank for AI-driven innovation.
🤝
Chapters
00:00Data Transformation Vision and Strategy02:17Establishing Business Buy-In and Ownership04:20Data Governance and Federated Approach07:54Self-Service and Security by Design10:03Driving Value Through Strategic Use Cases13:51Data Domains and Ontology19:01Cloud Migration and 16 Petabytes21:11Medallion Architecture and Data Products24:17Technical Architecture: Databricks and Azure28:55Migration Work Streams and Execution32:35Scale and Results: 3,000 Users Trained34:03Lessons Learned: FinOps, Governance, Education
FAQs
What did Santander Brazil's Move to Cloud initiative involve?
Santander Brazil's Move to Cloud initiative migrated 16 petabytes of banking data from on-premise Cloudera and ScyllaDB to Databricks on Azure over 18 months, scaling from 8 to 35 data domains and growing from 1,000 to nearly 5,000 active users. The program required two years of preparation, coordination across more than 100 team members, and zero business disruptions throughout the migration.
How did Santander Brazil secure business buy-in for their data transformation?
Santander Brazil made business engagement a prerequisite to technology decisions, requiring that business stakeholders see clear value in governing and treating data before investments were made. This involved creating new organizational roles including data owners and stewards, embedding data strategy into business unit planning, and demonstrating ROI through strategic use cases before scaling.
What governance model did Santander Brazil implement?
Santander Brazil adopted a federated governance model that distributes data ownership to domain teams covering customers, channels, products, and services while maintaining central standards for quality, security, and lineage. Security was embedded by design from the outset, and the Databricks Data and AI platform provides the technical foundation across all 35 data domains.
What results did Santander Brazil achieve from its data transformation?
Santander Brazil reduced data access lead time by 40 to 50 percent, scaled from 8 to 35 data domains, and grew the active user base from 1,000 to nearly 5,000. The migration completed with zero business disruptions, trained over 3,000 users, and established a trusted, reusable data asset foundation that positions the bank for AI-driven innovation.
Full transcript
[00:09] So, good afternoon everyone. We are very pleased here to be here to share with you our experience in turning Banco Santander into a data-driven organization. So, today Rodrigo here and I will share with you what has been the strategy deployed in Brazil and the past
[00:25] four years so that we can today say that we have completed our migration but more than this, we have actually established a data-driven organization throughout the business. So, before we get started, we wanted to know
[00:40] a little bit about the public here, the audience. So, please raise your hands if you are part of the technology teams in your companies. Yeah, you are an advantage here. So, please let me know if you are from
[00:58] the data or CDO organization. Thank you. And do we have any business representatives here?
[01:13] Great, great to have you here. So, this is about the stats that we want to change. We need more business like him involved in the process of taking advantage, uh, extracting value from the data that you have inside your companies. And that's what we are going to share with
[01:29] you how we have done it in Brazil and hopefully this is insightful for some of you and we'll be happy to share more afterwards. So,
[01:45] four years ago we started our program in Brazil and we had a very clear decision where we would like to enable the capability in order to have the business to take the best advantage of their data producing uh, uh creating journeys,
[02:00] creating experiences that would enhance the relationship they had with Banco Santander, and creating more value to our stakeholders. So, while we the the the vision is pretty simple. Every decision
[02:17] should be based on trusted data. And as you know, we have more than 100 years as a bank, so we have we have a lot of legacy systems. And all the business should be empowered by AI. In order to do that, we would go to
[02:34] the basics and work with the organization of the data that we had across different silos inside Santander and make it worth and make it work so that we could extract data value from analytics and AI solutions.
[02:53] So, in order to do that, we needed to have the business on board. Why we had to have the business on board? Because to do the work that we have to do, we need investment. And they the business has to see value from this from treating data, governing data, not
[03:12] only seeing the regulatory aspects that we usually have when we are talking about data governance. So, we needed them to have embedded data in their strategy and decision making. Okay. And in order to do this, we needed to
[03:29] scale to be able to scale. So, we didn't want to have CDO centralizing all the services to make data available to the business. We wanted them to lead this process, and so we opted for a business lab federation in order to create all
[03:45] the data infrastructure foundation inside the bank. And mostly, we have to do it very responsibly because we are highly regulated even more nowadays, and we wanted to make sure that we had security embedded by design and
[04:03] regulatory compliance as we needed. So, how did we do it? We wanted to build the foundations to have this transformation going through. Uh and we had to drive value in order to do this. So, the first
[04:20] step that we took was we needed to sit down with the business, identify how we could create value to them through data. So, before we even got started with the technology aspects, the data governance aspect, we sat down with the leader uh
[04:37] senior leadership of the business and identified with them, discussed with them what was uh the struggles, what were the opportunities that were not being uh taken to the next level so that we could identify what kind of solutions we
[04:52] were looking for and what kind of data they needed to do this. And only when we did we completed this step, we started to work on data governance and technology because we needed to have the business on board as the first step to
[05:09] this transformation journey. And they would have to be accountable for the whole journey along with us. Usually, you see business saying, "Okay, I'm going to ask the CDO and we they are going to take all the job. They and if they are not successful, I'm just going
[05:26] to say, "Okay, you didn't deliver the model that I needed. You failed." And the thing that we knew for sure is you can create a great model, but if you don't have the knowledge about the business, what are the surroundings, how it operates, usually, you don't get to
[05:43] be as successful as you you be if you had them along with you throughout the journey. So, by that, what we've done, uh transition then from only business consumers to owners of this data strategy, they would be the ones to
[06:00] prioritize all the investment needed in order to make this work. Uh usually this is the first struggle that we have when we are trying to do what we've done. And they would also um put the effort from their own people and
[06:16] put them to work together with us. So, it's a little bit of changing the ways of working. Uh and this was a key point for us to move ahead. Once we had the strategy uh
[06:31] design, we moved to the data governance. So, as I said before, we didn't want to do it centralized because we didn't believe this could work at the scale we needed. So, what we did was we created new roles. I'm going to go through this in
[06:48] detail in the next in the next slides. And what we accomplished was have the business led data ownership. We created a strong and robust framework in order to guide every business to do the
[07:06] data domain construction by themselves. And once we've done that, we would have the trusted and reusable data that we needed. Moving forward and only then we would talk be more focused on the
[07:22] architecture. We knew that the architecture that we had wouldn't be able to scale. So, for sure we needed to move to the cloud, and Ohid is going to go through that in details for you guys of the tech organizations.
[07:38] And in order to scale really scale and have the federated governance model, we would have to have self-service self-service capabilities so that every team could run their work by themselves,
[07:54] and that's what we achieved, and you're going to see how it changed the landscape at the end of the day. And this is also what accelerated the course of our journey because about 2
[08:09] years ago we had this federation working very um uh smooth and very well handled. Uh when And when we started to work on the cloud, these teams were very new to the
[08:26] new technology, and even though they didn't have previous training, once we finished the program, they were able to work properly on cloud. And security embedded.
[08:41] Uh we are very highly regulated, not only locally in Brazil, but we also have to comply with all the European Central Bank regulations as well, as we are part of a global uh banking. And for that, we would have to have not
[08:59] only a very fine access control in place because what we envisioned was you would consume data the same way you buy soap at a marketplace, and we needed to secure that data was
[09:14] going to be granted to only the purpose that uh was actually asked it for. So, we would have data sharing and consumption done properly. And of course, none of this is possible
[09:30] unless you're working uh with people. We run a change management program so that all the business that were being embarked in this process would have the training, uh will learn about data, how to work
[09:47] with uh data governance, how to work with cloud, uh how to work with Databricks, and also we revised all the process that we had uh back then. So,
[10:03] let's see here how we did it. The first thing that we did was driving value and culture. Usually, what you have is the is the business asking for you to build your models. I need to control
[10:19] churn. I need to uh have pricing, uh better pricing strategy. And they usually don't get along in the right to develop those models. So, what we did is pretty simple, but you have to have the discipline to
[10:35] make it through. So, once I said before, we had established all the opportunities that they had. They were from very uh feasible to less feasible, but high impact.
[10:51] And we needed to know what were the big big plays for them. Not always you start with the big prize. Sometimes you start with the small prize, but that it's very feasible so that you can make them see the value coming through in order to
[11:07] move to the next one that it's a little bit more difficult, but they are on board already and you have uh the sponsorship that you need. So, as you see here, I cannot with you actual details of what we have
[11:22] implemented in Santander in terms of use cases, but it doesn't go very different from this. So, we will have the senior leadership sitting down having the ownership that if we do this, if we do the model, if we do the
[11:39] the work, they would be accountable, not the CDO, to deliver the business outcome that they were promising. And this makes a whole difference because they are the ones that have to explain
[11:54] what has been the failure or has been the success to achieve the goals that we had. And mostly because, as I said before, the model doesn't do the work by itself. We need to change the surroundings in order to do this.
[12:11] So, the first thing, identify, then we needed to prioritize those use cases so that we know what goes what comes first. Then we have to have the assignment of the business ownership, identify
[12:27] what are the KPIs of success. And then we have to monitor this. Usually people drop it in the first year so that you don't see all the value coming through. And this is a mistake. And what we did is for the entire
[12:43] three waves that we have run before in Brazil, we followed and monitored all the results and as we were seeing things not going the way we needed, we made the adjustments necessary in order to really present the value that we
[13:00] had promised. And then, once you have the teams going, the business team feeling more comfortable to do it by themselves, you don't have to be monitoring it anymore. And you start to see a change
[13:15] in the landscape of the business teams. They start to bring more people that knows how to work with data, that they can actually run the programs by themselves. And when we look back today, uh right when we began, they would all ask
[13:33] me to work with them and put groups together so that they could develop. And nowadays, we are not necessary in this process anymore. They do it by themselves. And before we we started data governance
[13:51] uh program, we had the the use cases defined. And the way we've done it was not to develop all data domains at once. What we did was we created an ontology where we had
[14:08] the bank translated in five big groups of context. Here you see customers, channels, products, and services, uh transversal functionings, and offering in relationship with customers.
[14:24] But what we have behind this, it's 35 data domains that are in place nowadays. And we didn't start them all together. We started only with eight. Those were the eight that would deliver all the value that we needed for the first wave
[14:41] of case use cases. And in this sense, it was easier to convince the business to make the investment we needed. Before we got started with the rest of the uh the data domains, we would be delivering all the value we have
[14:57] promised before. So, for the second wave, we reached 18 data domains. We were already with uh momentum. We had to work with more investments. We also changed the
[15:12] approach from product to customer view. We created a customer 360 view also to enable the use cases we needed. And to do this, we created new roles in the company. We created a data owner.
[15:29] This role usually goes to the head of the business. They own the prioritizations of all investments. They know everything that it's being uh strategically prioritized in their in their in their pipelines.
[15:45] Uh they know what are the struggles. They have the power to put the right people to work along with us. And they would be the ones to provide a vision and the strategy for the data domain. Along with him, he has the right arm
[16:01] that is the data steward that is the one responsible to execute, ensure that we have all the framework of data governance uh applied to the data domains. And along with them, they would have the data custodian that it's usually people
[16:17] from the IT team that would be responsible to enable and sustain the data domains in a data data-day basis, okay? As an outcome, we come up with strategic alignment because we wouldn't be creating data that wouldn't be used
[16:33] after all. It would be something that would be creating value. They would be reusable. The second wave of cases was much um faster than the first round um because we reused a lot of the data that we produced before.
[16:49] And mostly, they were trusted information where we usually have the uh duplication, but never the same information at the end of the day. This was produced with the rules of the business. So, it was a reliable
[17:06] uh qualified data with all the data controls that we needed. And we had risk and compliance addressed as well because we have all management for access to the day to the data and regulatory compliance as well.
[17:23] So, the outcome here is more than just put into a 35 data domains. We produced data that was actually working for the business, creating value every day. And but when once we reach at the 35 data
[17:39] domains, they didn't need us to conduct the work anymore. They were doing by themselves. They created the business and they we one person earlier before earlier today asked at me, "How did you do it?" Yes, we had third parties
[17:56] at the first at the very beginning, but then the teams uh they were transformed and part of the third parties joined Centene there to be part of our teams and the other ones were hired with
[18:11] the correct skills and some people were trained to take place in this position. And what was the challenge? As a 140-year-old bank, we have a lot of legacy systems.
[18:28] Uh everybody was pretty comfortable where they were. They didn't want to go to the cloud environment. At first, they didn't want they were okay with the silos that they have. So, it took us a very
[18:45] a lot of resiliency in order to make the change and convince them that bringing out to the cloud under our strong governance would be the best option for them to prepare themselves to take advantage of
[19:01] AI at its most. So, uh here so that you know uh he's going to go in details uh later. We moved 16 petabytes of data to the lake.
[19:17] We had a lot more. 16 was the only piece that we saw real value to take to the cloud. A lot was purged before we did the work. And to do this and not compromise the business,
[19:33] not to make anything fail after the migration, we had the job done by the data domains. They were evaluating what data had value, what was not important, what had to be moved to archive to a
[19:49] cold layer so that we wouldn't be spending. So, they had all this education in order to make this possible before we even done it. It took us 2 years of preparation, training these teams, moving the pieces
[20:07] together so that we hit at the big bank for the migration. We didn't do it before because the teams wouldn't be ready and we would have impacts for the business. Here all the closing of the bank was completely
[20:22] safe during the whole migration. We didn't stop any business any day during this migration. It's very complicated because you don't have all the mapping before. There was a lot
[20:38] of validation, there was a lot of negotiation in order to do this. It sounds pretty easy looking at the slide, but it's not an easy job to execute to be quite honest. But we did it in 18 months
[20:55] doing all the work along with more than 100 teams. They were in a day-to-day working with us to put together the data domains and doing the migration with us. So, the approach that
[21:11] we did we used the medallion um approach. So, everything would arise at the bronze layer, but everything was already being cataloged with dictionary, lineage, all the taggings that we needed
[21:27] to control costs, all the tags that we needed to know uh to reply the security information. Everything we needed was done. The service that we had for self-service was no longer done by the CDO, but the
[21:45] data domains will do their own ingestion, and O'Hita's going to go through this in a minute. And uh once we reached this point, we had the data products the data products that we needed. Uh all the
[22:01] It's good to share with you a little bit. So, all closing of the bank is done using data domains right now. All the CRM campaigns are using data domains. The customer view that we have they use the data domains. So, everything is done
[22:18] right now at Banco Santander using this data. We still have some silos with legacy systems that once we have stabilized and we have learned how to manage costs uh the proper in in a very proper way,
[22:35] we are now moving to get the rest of the legacy data to the cloud so that we end up only using the cloud as the the source of information, the single source of truth. And this makes it
[22:50] uh possible to really generate value. Uh it takes a long time to get the business to really say they own it. Nowadays, we see in the very beginning, nobody wanted to be a data owner, and now they are very interested in running
[23:07] this role because they know that this will be uh important in order to develop the AI strategy that they are eager in for. So, moving forward, that's the the change that we did. We
[23:24] had used to have 1,000 users at the very beginning. We are almost reaching 5,000 users on the Databricks right now. Uh we reduced dramatically the time the
[23:39] lead time for any uh since uh data access also in ML Ops. Uh we saw the time to to market reduced by 40%, 50%, and all data assets are being cataloged
[23:57] nowadays automatically with self-service service. And this is what I had to share, and I'm going to hand to Hider, who's going to go over more the tech aspects.
[24:17] Good afternoon, everyone. My name is Rodrigo Oreira. Today, I'm going to share our journey of migrating from our on-premise data lake to Databricks. And I will do that from a tech perspective. I divided this presentation into four
[24:34] main parts. First, I will explain why we decide to move to the cloud and the benefits that we achieved. Then, I will briefly walk you through our target architecture.
[24:50] After that, I will share how we executed the migration. And if I will close it with the main lessons we learned during this journey.
[25:13] Our previous data lake platform was basically on a on-premise Hadoop environment. And over time, this platform started to create important limitations for us. The first challenge was the high cost
[25:28] costs of operating, maintaining, and renewing this platform. The second challenge was the fragmented experience for our users. Data engineers, data analysts, and data scientists had to work across different
[25:46] tools and different experiences. The third challenge was performance. We are facing bottlenecks, especially as the volume of data and the number of your workloads continued to grow.
[26:02] And if finally, scalability was limited by our on-premise infrastructure capacity. In other words, growing this platform required required required new hardware acquisition, which was expensive and slow.
[26:22] So, moving to cloud was a strategic decision to improve scalability, efficiency, and the user experience. After the migra- the migration, we achieved several important benefits.
[26:41] The first one was cost efficiency. By moving to a managed cloud-based platform, we are able to reduce our operational costs and the hardware acquisition. The second benefit was better better cost transparency.
[26:58] Now, we can segregate costs by business unity, which brings more visibility and accountability to cloud usage.
[27:15] The third benefit was reducing the time required to make data available for consumption. And this is very important because it directly impacts the speed of analytics and decision making.
[27:32] We also improve resiliency, availability, and disaster recovery capabilities. And if finally, we improved the overall experience for data engineers, data analysts, and data science teams.
[27:49] Now, let me briefly walk you through our target architecture. The core of our modern data lake platform is basically on Azure and Databricks. We use Databricks to support batch and
[28:05] stream ingestion, data processing, and consumption. Our architecture follows a medallion approach using Delta as the foundation for storage and processing. We also adopt Unity Catalog as a key
[28:23] component for governance and security. Around the platform, we enabled observability, audit, FinOps, and billing capabilities. This
[28:38] This was important in not only to run the platform, but also to manage it properly at scale. Our goal was to create a modern data and AI platform. A platform that is scalable, governed,
[28:55] resilient, and easier for users to consume. Let's talk about the migration. To execute the migration, we organized the strategy into five work streams.
[29:16] The first work stream was data. Here, the objective was to move data and enable enabling ingestion in both environments, our legacy on-premise platform and the new cloud platform. This parallel ingestion was important to
[29:31] support the transition period. The second work stream was application. In these work streams, we mapped the dependencies between apps, defined the
[29:48] migration waves, and migrate the applications to Databricks. The third work stream was users. Here, we enabled sandbox environments, onboarded the users, and supported their
[30:05] journey into the new Databricks environment. The fourth work stream was governance. Here, we mapped the existing access permissions, defined the new access groups, established the naming standards, and enabled Unity Catalog.
[30:24] And the fifth work stream was education. Education was the foundation of the migration. It was responsible for training engineers, users, and teams, so they could adopt the new platform
[30:39] with confidence. Of course, during this journey, we faced several challenges. In the data work stream, the main challenge was the large volume of data.
[30:56] We had to migrate years of historical data. And because of network limitations, we could not rely only on data transfer through the internet. So, we also used a physical data transfer
[31:13] with data box. In the application work stream, one of the main challenge was prioritization. We had many apps, many dependencies, and many business teams involved. So, defining the migration waves was a
[31:31] critical part of the process. In the users work stream, the challenge was adoption. We need to engage the users
[31:47] and encourage them to start using the new environment. This required a lot of communication, support, and the hands-on enablement. In governance, one important challenge was the creation
[32:02] and acceptance of new domain roles. For example, data owner and data steward, like I like I said. These roles were essential to make governance sustainable in the new platform. And finally, in the education, the
[32:18] challenge was scale. We had to train and enable a community of more than 3,000 people. To give you a sense of the scale of this migration, I would like to share some numbers.
[32:35] In the data work streams, we used more than 20 Azure data box and we migrate more than 16 petabytes of data. In the apps work stream, we migrate more than 70 applications.
[32:52] This required dependency mapping, migration waves, validation, and coordination with several teams. In users work stream, we migrate more than 3 thousand users to the new platform.
[33:16] In governance, we created new access groups and granted more than 6,000 permissions to applications and users. And in education, we certified more than 250 employees. We also delivered a lot of training
[33:32] sessions, workshops, and the hands-on labs to support the adoption and scale. These numbers show that this was not not only a tech migration. It was also a large transform
[33:47] organization transformation. To conclude, I would like to share the main lessons we learned. The first lesson is that FinOps is essential.
[34:03] It It is not just a discipline. It's a key capability to operate a data lake successfully in the cloud. The second lesson is that data and process sanitization before the migration can
[34:20] significantly reduce costs. Migrating everything as is is not always the best approach. The third lesson is that coach tagging is critical. It increases visibility, improves
[34:37] accountability, encourages responsible cloud usage. The fourth lesson is that data governance must must be established before the migration, not after it.
[34:52] Governance needs to be part of the foundation. And the fifth lesson is about executive sponsorship. In a migration of this scale, a strong engagement from senior leadership is
[35:07] essential. It helps to create alignment across business areas, remove blockers, and ensure that teams prioritize the migration alongside their day-to-day responsibilities. And finally, education was the key
[35:25] success factor. Training the teams, supporting the users, and building confidence in the new platform made a huge difference. So, in summary, this migration was not only about moving data and applications
[35:41] to cloud. It was about to create a modern, scalable, and governed data and AI platform. Thank you, everyone.
Learn more about the Databricks Data and AI platform.
The information provided herein is for general informational purposes only and may not reflect the most current product capabilities or configurations.