Skip to main content

Modernizing Mission-Critical Data at Barclays: Migration Without Illusions

Summary

  • Barclays migrated its equity data warehouse from Netezza to a cloud-native Databricks architecture across four regions, handling 1.8 petabytes of data, 18,000 daily jobs, and up to 2.5 billion orders without disrupting a mission-critical, highly regulated financial platform.
  • Rather than a naive lift-and-shift or a full rewrite, Barclays chose a deliberate middle path that preserved core business logic while modernizing the technology stack, using Delta Lake for ingestion, Unity Catalog for governance, and Photon optimization to reduce costs by 30-40%.
  • Spatial join performance improved by up to 95%, cross-region data sharing was enabled, and Power BI replaced the legacy consumption layer — with key lessons around cluster configuration, CICD compatibility, and the hidden complexity inside legacy systems.

Modernizing Mission-Critical Data at Barclays: Migration Without Illusions

Watch: Modernizing Mission-Critical Data at Barclays: Migration Without Illusions
Transitioning a mission-critical data platform is rarely straightforward, especially when legacy systems already work and regulatory requirements are strict. Barclays moved its equity data warehouse from Netezza to a cloud-native Databricks architecture across four regions, handling 1.8 petabytes of data, 18,000 daily jobs, and up to 2.5 billion orders. The modernization balanced aggressive timelines with avoiding naive lift-and-shift approaches, creating a deliberate middle path that preserved core business logic while modernizing technology infrastructure. Using Delta Lake for reliable ingestion, Unity Catalog for governance and lineage, and Power BI for consumption modernization, Barclays improved spatial join performance by up to 95%, reduced cost by 30-40% through Photon optimization and cluster configuration, and maintained complete CICD compatibility across regions while enabling cross-region data sharing.
🤝

Chapters

FAQs

Why did Barclays choose Databricks for its equity data warehouse migration?

Barclays selected Databricks as the replacement for their Netezza equity data warehouse after the legacy platform reached end of life, drawn by the cloud-native architecture's ability to handle their scale of 1.8 petabytes, 18,000 daily jobs, and 2.5 billion orders. The Databricks Data and AI platform provided the combination of performance, governance through Unity Catalog, and cost efficiency needed for a mission-critical financial workload.

What was Barclays's 'middle path' migration strategy?

Rather than a pure lift-and-shift that would carry forward technical debt or a full rewrite that would risk business disruption, Barclays chose a middle path that preserved core business logic while modernizing the surrounding technology infrastructure. This approach allowed the team to deliver on aggressive timelines while avoiding naive assumptions that often lead to migration failures.

What performance and cost improvements did Barclays achieve?

Barclays improved spatial join performance by up to 95% and reduced platform costs by 30-40% through Photon engine optimization and careful cluster configuration tuning on the Databricks Data and AI platform. These gains were achieved while maintaining complete CICD compatibility across four regions and enabling new cross-region data sharing capabilities.

What were the key lessons Barclays learned during the Netezza to Databricks migration?

Barclays found that legacy systems contain far more hidden complexity than initially apparent, and that cluster configuration requires careful tuning to achieve target performance rather than relying on defaults. The team also emphasized the importance of CICD compatibility and the value of treating migration as a modernization opportunity rather than a pure technology replacement.

Full transcript

[00:08] My name is Harsha. I run data platforms and data management at Barclays for the markets division. Barclays is a universal global bank. Um we serve consumers like people like you and me, like
[00:23] ordinary people, right? Uh we're pretty big in the UK when it comes to the retail bank. Uh payments we are pretty big, right? Uh but here in the US you probably know us in the markets division from the Lehman takeover
[00:38] during the financial crisis. Or you probably know our cards business, right? So we have had some famous cards, right? Like the A American Airlines cards and things like that. Or if nothing else, if you're from the East Coast you'll know us from Barclays Center. It's funny enough, like I went to this
[00:54] dinner and somebody in front of me sitting was asking me what do I do at Barclays Center? I'm like, "I don't do anything at Barclays Center." Cuz on the badge it was printed Barclays Center instead of Barclays Bank.
[01:10] Well, anyway, Barclays Center is something that we sponsor. We have our name on it. Like you will probably see some famous games there with uh with the football. I think there's quite a few uh games at the Barclays Center as well, right? So we we we are definitely like one of the bulge
[01:25] bulge bracket banks, right? When it comes to markets, we are top five in every asset class that we deal with. Um Barclays has a legacy. It's a 300-year-old bank, right? Um but we've always been very very innovative
[01:41] when it comes to like, you know, doing things that are like, you know, uh at the cutting edge technology. Uh otherwise we wouldn't have survived for 300 years, right? 350? Anyone go for 400?
[01:59] Um where I was going with this was a question actually to all of you. I've already given the clue. Who do you think put put the first ATM uh out there which dispenses cash? Yeah, bingo. They're very smart crowd here.
[02:17] Otherwise, I won't be talking about it, isn't it? Um it's actually not true, right? Like uh there were quite a few players who were in who were like tinkering with the cash machines at the same time as we were. Right? Uh next question. Does anyone
[02:33] know when that was, how long that was? Like uh I know the ATM was once upon a time like cutting-edge technology for banking. We are no more. I don't think I have been to an ATM for the last 5 years. But anyway, like any guesses as to how long the ATMs have existed, when the first ATM was put on the
[02:51] Sorry? No, wrong. Try. No. 71? Try harder, guys. Come on. No, 71 lost, so you can't go to 75 then. It was actually 1967.
[03:07] Right? Maybe my English colleagues can tell me like, you know, where in England was it? Enfield, actually. Yeah. Yeah. Yeah. Um
[03:23] Uh what I was trying to say is that like during the '60s, there were quite a few people like trying to put out ATMs, cash dispensing machines. But the point is it's execution. It's not just about innovation. You need to execute on that innovation. So, we were the first ones. We won the race in putting it out dispensing cash in 1967.
[03:40] That's why largely, if you go Google, it'll say it's Barclays who put the ATM machine first. Right? So, what does my team do? Like my team like with for for markets division, uh we do foundational data, right? Data foundations which powers trade trading,
[03:57] right? As I was telling you, we are top five in every asset class. What that means is like billions of orders flow through our systems, right? When it comes to cash equities, we do some of the most complex deals, uh that means like technology and data are at the center of everything that the
[04:13] bank does. And yeah, we are innovating, we are modernizing. As you probably have heard, like everything is about AI nowadays, but we are going to talk about more about modernizing the data platform itself. Cuz garbage in is garbage out, right? So, your AI will do at stealth
[04:30] what is rubbish if you don't get your data platforms data foundations right. So, we are doing multiple things, but today we are here to talk about a journey that we undertook with Databricks in partnership, right? To modernize one of
[04:48] our big data platforms. That's equity data warehouse. Thank you. So, basically, what is equity data
[05:03] warehouse? So, it holds, if you think about the equities business, it's cash equities, options, futures, equity derivatives deals that we do on equities. So, everything that runs on exchanges, SpaceX, everyone saw that, right? Like the big one, it's all in the
[05:19] news. Uh what that in turn means is that like we get hundreds of millions of people putting in orders that comes through our systems. We are one of the big prime brokers. That means people put their flows through us either for us to execute or for us to like, you know, uh
[05:35] help them get the best rates, whatever, whatnot. We also have what we call as dark pools or alternative trading systems, right? So, all these flows go through this particular equity derivatives warehouse. Uh it holds trade data, orders data, executions data, all the market data
[05:51] that you need, all the instrument static data, right? So, large amounts of data. Think about it. Giving you some appreciation for scale, we are talking about 1.8 petabytes of data uh that we hold. Cuz we being highly regulated means that like 10 years worth
[06:07] of data regulator can come in, ask for this data any point in time, right? Um 18,000 jobs. To give you some scale, right? We probably like process about 1 terabytes of data every day. That means
[06:23] anywhere between 1 to 2 and 1/2 billion orders flow through the system, right? So, these are like individual orders, right? Like for multiple different instruments. So, let me just stick to the previous slide. Sure. Thank you.
[06:38] So, what kind of like data what kind of like analytics come out of this? Like, you know, traders use this for making trading decisions, sales people use it for pitching, you know, businesses to their clients. Um and also like all your regulatory reports come out of this. We are, as I
[06:54] said, heavily regulated. Multiple regulators around the world, right? Have different requirements. The as I was telling you, like we we process about a billion billion and a half records orders every every day. Half of that will get reported to FINRA, the the the
[07:09] regulator here in the US, right? Every every day that is, and it has to be done by a certain time. Um other than that, compliance, surveillance to show that like we're not spoofing, we're not like manipulating the market and things like that, right? So, it's a critical cog in our
[07:26] infrastructure, right? So, this particular platform was running on what we call as an appliance, as many of you will be familiar. It was from our Netezza. If I see some people who've been here in the industry quite a few years. So, you
[07:43] would know NetApp so it is over the last 20 years ever since Lehman acquisition prior to Lehman acquisition Lehman was using it. All right, it has served us well, but I it came to a stage uh wherein we had to modernize. We needed to get to
[07:59] the new new world. The reasons were very simple, right? Like we were constantly at um at risk of failure. See, the appliances had come to an end of life. Uh support was falling off this year. Would that would mean that we would have to buy new
[08:15] appliances, but also what it means is these appliances are not plug-and-play, right? They are vertically scaling systems majority of the time. We were running at 92 to 95% capacity right? At any given point in time. So,
[08:31] with this kind of volumes throughput required and the throughput and latency requirements, you can imagine how much hand-holding the team should have done, right? When um when liberation day was announced for example, right? So, uh uh
[08:47] So, the amount of stress on the system and the people around it was super high, right? So, uh we we as it was end of life hardware would fail, natural, right? Comes to 5 years end of life hardware fails. The problem was this hardware
[09:04] cannot even be bought from the primary market. You would have to go to secondary market to buy it. Right? So, uh scale was one thing and the pressures from the business the business is growing, right? Like equities business likes in APAC was saying that they want to go 2x prime business wanted
[09:20] to go 3x one day 10x on the other the other day. So, predictability in the business is not there. It's and you need to have that kind of agility if you want to be in the top five if you want to be able to like compete. So, they so to to be able to scale is
[09:36] very very important. And then there was this whole innovation ceiling, right? You couldn't run in easily run AI ML. You couldn't couldn't run like you know, different use cases that uh different kinds of analytics that you want to run on these platforms cuz you
[09:52] had to prioritize data loading so that like all the data that that that huge amount of data gets loaded, processed in time so that your regulatory requirements are being met. So, all your other analytical use cases, your decision-making use cases take a secondary backseat because they are
[10:08] revenue generation. So, compliance is key for all of us, right? Like if you don't comply, like basically regulatory will come and say stop this business. You're not allowed to do this business anymore. So, all these things led us to like you know, we need to modernize, right? Um
[10:27] So, what as as as I touched upon some of the things that we wanted in a modern platform was elasticity, right? More than elastic What I mean by elasticity elasticity is ability to scale horizontally and through easy click of buttons rather than like you know, somebody from IBM having to come
[10:44] in the middle of the night and having to put some hardware onto the onto the rack, So, we wanted it to be cloud cloud-based. That was non-negotiable, right? Uh real-time requirements like you know, gone are the days when people were happy with like T+1 reporting. You people want
[11:01] to have like uh the analytics at their fingertips so that they're making the right decisions whether these are sales guys or trading guys, right? And based on like what is the wallet size, who is the customer that they're talking to, uh how important are they for us from like every customer is important, no
[11:16] doubt about it, but some customers are more important than others depending on like their wallet size in the market, their like relationship with us and all of those things, right? So, you need to have that information at your fingertips. Uh open architecture like you know, kind of commitment to
[11:33] uh open source is key. We did not want to get locked in once again, we're just coming out of a locked-in relationship, so so to say with the IBM for the last 20 years, and no point walking into another locked-in relationship. And also like open source what it does is like it allows people to innovate
[11:49] outside of that vendor relationship, and that helps you like build better platforms, better uh solutions for the business. Uh governance lineage, uh you know, it's taken for granted, but like we don't do that very well uh across the board,
[12:04] right? That's why data is what it is today, you know, all the uh different big organizations. Uh finally, I think like as I talked about innovation and business value, right? So as I said at the beginning of my talk
[12:20] it is inevitable, AI is here. If you're not doing it, you'll get left out. If you're doing it on rubbish data, bad data, or your platforms can't support it, you'll get left behind. So we needed a platform which was going to get give us the edge to compete with the best.
[12:43] So next like, you know, we were looking at various different platforms. Uh this is about how did we choose Databricks over the other platforms. We did like about 3 4 months of um uh of proof of proof of concept, POCing, MVPing, whatever you call it, right? So 5 years
[13:00] ago we had done that with like Snowflake. Uh we did that again with Snowflake, Databricks, and and a couple of other solutions, right? Um we were looking at a few things, right? One was I think 5 years ago Lakehouse architecture wasn't wasn't there, right?
[13:17] It wasn't mature. But when we started looking at it 2 years ago, the maturity was there. Um if you Google even 3 years ago, 1 year ago, uh to say, you know, typical data warehouse solution or you go ask your CTO team or the architecture
[13:34] architecture team, the answer would be like, you know, out of the box data warehouse, go with X. You want AI/ML, go for Databricks, right? with Aslan.
[13:49] I was telling him after the POC and we've done a few things and we chose to go with Databricks that I'm taking a contrarian view, right? Probably cost me my job. And he he, if you recall what he said, he basically said that like, "Harsha, we
[14:05] appreciate the thought process. I we appreciate where you're standing today, but it's my job to ensure that you'll not lose your job." Here I am today, a year and a half later, we have gone live in every region. So,
[14:22] we started with APAC, then we went to EMEA, then we did derivatives, and just last week we went live with Americas, our biggest business. Um So, we found with Databricks like all these things that I talked about, like
[14:37] they were taking the box in the right place, and they were willing to roll up their sleeves and work with us day in day out. That 3-months of POC that that that I talked about was not just like, "Let's try out a couple of things, right?" We pretty much went and tried everything. The team worked like
[14:53] probably every weekend, just like just like they did when we went live with What are weekends? Exactly. So, and Databricks was there with us, their engineers there were there with us trying out the the solutions, try trying to make it happen to show
[15:09] that like, you know, it was not like, "Trust me, bro, it's on a PPT. We'll take care of it," right? And hopefully Databricks will continue to be agile like that, and maybe you guys won't change like I mean, mean won't change after the IPO, right? So,
[15:25] that is definitely key. Like for us like to be able to prove something cuz it's so important that technology that we didn't want to risk it, right? Uh and lot of the with Unity Catalog like the enterprise governance comes out of the box.
[15:41] And then like you know lot and Databricks was known for their AI/ML. They've been known for that since many many years. We didn't not want an arbitrage of people saying if you want like you know traditional reporting where and dashboarding we go
[15:56] to this particular solution or this particular data warehouse and if I want AI/ML we go to another one. We wanted a unified platform. So, all of these things led to like you know our deciding that we'll go with Databricks, right? And also like I forgot to mention that we have a huge
[16:13] uh Hadoop estate which runs lot of Spark workloads. We wanted like a glide path which will give us like a you know a a a unified uh data platform for not just equities but like all of the markets business. So, with that let me hand over to
[16:29] Shashank from Databricks uh and my colleague Shivram from Barclays. Shivram actually owns equity data warehouse, the platform that we are talking about here today uh to take us through this journey of how it has been and how we managed to achieve this in
[16:46] 14 months. All right. All right. Uh thanks, Harsha. Um so, I think Harsha spoke about the decision to move to Databricks. I think which begs the question choosing the right transformation path.
[17:02] I think that's an important next step of the journey. So, a couple of things that stood out for me as Harsha was speaking about it. Before that like a quick intro about myself, I lead professional services for Databricks and I partnered very closely with Harsha, Shivram and rest of the
[17:17] team throughout this migration journey. Um a couple of things that popped up in my head as Harsha was describing in terms of business drivers for uh this transformation. First, uh the team was operating on a very real timeline constraint.
[17:32] You know, the legacy appliance was approaching end of life. The operational constraints due to growing data volumes were causing operational challenges every single day. Um and a very common response to this kind of pressure is to choose a purely lift and shift style kind of a migration
[17:48] approach. On the other hand, I also heard a few other business drivers for migration, such as a desire for future state AI capabilities. Uh you know, the need for real-time reporting, etc., which is a slightly different kind of a business driver. And it if you choose to optimize for that
[18:04] path, that's essentially a little more complexity, right? So, um which brings me to a core tension that exists in any large migration or modernization program, which is you know, in any migration, there has to be a balance between the
[18:19] amount of change risk one is willing to sign up for and the amount of uh future state value realization that one expects by the end of the migration journey. Notice that I didn't say after the migration journey, by the end of the migration journey.
[18:34] And uh what we have observed working with customers globally here at professional services is many teams tend to treat these two um sort of balancing factors as binary choices. And they tend to gravitate towards either one of the two extremes.
[18:51] Now, what stood out for us right from the outset was the amount of clarity Barclays had in terms of what they wanted to optimize for in this regard. So, Barclays were very clear that there's no way we can pull off a full-blown modernization, right? With the timeline constraints that we had. But they were equally intentional about
[19:07] the fact that what they did not want to do was to take the legacy implementation as is and replicate it on a modern platform like Databricks which operates fundamentally differently and expect it to work well magically at scale. Right? Um and that conscious decision or that
[19:25] uh deliberate middle path if you will, that became the North Star for many of the uh design choices, the implementation decisions that were taken down the road. And we'll see more of that, but one thing that I like to call out is it was also not a singular decision that was
[19:40] taken at the outset, right? Um if I tell you that we took this decision at the beginning and then everything downstream became super smooth, you're not going to trust me because that's not how it played out. Um the reality was this required maintaining a constant balance throughout the migration journey and
[19:56] that balance required making trade-offs, uh fine-tuning our approach iteratively, evolving the architecture and we're going to see all of that in an upcoming section, but uh before I go there, let's quickly look at what the middle path looked like in practice.
[20:12] Um what was preserved, what was modernized? So on a very high level, what was preserved was the legacy business logic layer and what was modernized was the technology infrastructure around that. So things such as the business uh schema, the data model,
[20:29] the core data processing workload logic, uh and the downstream reporting outputs, all of these were largely kept intact. Uh but what was uh what was modernized was a lot of components, tooling and technology components that you see on the right side of this slide.
[20:45] Um but again, this was So a lot of these components were uh modernized and replaced with Databricks native capabilities and Power BI. But this was not a technology change just for the sake of changing technology, right? Uh the goal here was not to arm
[21:01] people with tens of new tools and declare victory. Um apart from the operational and scale challenges that Harsha mentioned about earlier, one of the other goals of this uh technology transformation was to enable the new stack with a set of engineering
[21:18] best practices that were largely missing or very difficult to attain in the old stack. So, things such as modern CICD, DevOps best practices, uh governance and lineage, um as well as improved observability, et cetera, right? All of that.
[21:34] So, So, yeah, that that was one of the uh two of the major goals of the technology modernization exercise. Um coming back to the middle path approach, um I mentioned earlier that the team on the ground was committed to not propagating anti-patterns to the new
[21:50] stack. Um to that end, key business processes were reviewed from that lens and wherever obvious anti-patterns were identified or wherever designs were identified that were likely to yield sub-optimal outcomes on Databricks, those were
[22:06] identified as candidates for like a light modernization, if I can use that term, uh which was essentially adapting and aligning to Databricks best practices. So, again, as I mentioned, we're going to see some examples, but um before we go there, now that we have
[22:22] context around what was preserved, what was modernized, and what were the guiding principles behind that, I will hand it over to Shivram to cover what was our uh how did we set up the delivery model and how did the execution realities play out?
[22:40] Thanks, Shashank. Yeah, hi everyone. So, as uh Harsha and Shashank alluded to, right? So, we had the the strategy, we knew, I mean, how to execute the strategy. Now, the next stage was to plan the strategy execution, right? So, it's it's a global bank at the end of the day where, you know, we have our business in APAC, EMEA, and Americas.
[22:57] With the APAC business, I mean, they wanted to expand, you You their investment banking offering by 2x, 3x where the volumes will grow. I mean, threefold I mean for for the APAC business and with the current Netazza appliance that we had, I mean, you know, the business was constrained because we couldn't support that volume growth. So,
[23:14] that was one of the primary reasons where we chose APAC as the first driver for our Databricks migration. And what we were able to achieve was, I mean, within 4 months, I mean, we were live with Databricks. And with decommissioning Informatica, with decommissioning Business Objects, and you know, running all our workflows using PySpark Python. And uh SQL
[23:31] serverless SQL warehouse, I mean, on Databricks supporting those high volumes, I mean, for APAC business. And that was that basically gave us a launchpad for the rest of the regions to follow. Right? And uh the next region was EMEA where APAC, though it was high volume, the reporting was not that enormous.
[23:46] Whereas, when it came to EMEA, the volumes were mediocre, but the consume layer of the reporting was quite uh large, I mean, in EMEA, right? So, APAC gave us a launchpad to basically migrate our high volume workflows Databricks. And EMEA basically was where the consumption pattern or the consume layer
[24:02] was primarily tested. And that basically gave business an impetus to, you know, make informed decisions, right? With Business Objects, it was more like static reporting. Whereas, with Power BI, with the power of Databricks, you know, business had the leverage to, you know, run advanced analytics using Power
[24:17] BI and Copilot where, you know, they could ask questions and they could get responses to their to their business, you know, queries, I mean, in a timely manner rather than relying on technology, I mean, to get some of these uh results, right? Then, the next part was AMER to to Harsha's point. I mean, the volumes for AMER was fourfold, I mean,
[24:34] compared to APAC and uh EMEA where we primarily get close to a billion and a half messages on a daily basis, I mean, that we process. And for a billion and a half messages that we receive, we report 3x of that volume to various regulators, right? FINRA, SEC, and and other uh other, you know, compliance and other
[24:50] regulatory bodies, right? So, AMER was basically scale and uh the the consumption, you know, twofold, threefold, right? So, the last stage of our migration was EMR. So, with APAC we tested scale to a certain extent. With EMR we tested the consume layer. With EMR we basically migrated, you know, the
[25:05] entire workflow with 2x scale, I mean, with with Databricks and Power BI. So, how did we execute it? I mean, the the project was executed uh amongst three parties. So, one was obviously Databricks where they helped us, you know, define the architecture in the first
[25:20] place. And then we had, you know, Impetus as our uh vendor partner who basically helped with code acceleration. And uh and basically defined and implementing the architecture that Databricks came up with and where Barclays itself was, I mean, providing the business logic and the and the
[25:36] workflows, I mean, that were needed to migrate, I mean, to Databricks. So, it was a combination of three parties that helped basically, you know, execute the the the project as a whole, right? So, now we basically spoke about the execution plan, right? And Shashank and both Saharsha mentioned that we didn't
[25:52] want a lift and shift on what we do, right? So, there were some challenges that we saw, right? I mean, we came up with an initial architecture which was a lift and shift with what we had with Netezza for one of our near real-time workloads where the volumes are generally high. When we implemented that, I mean, we figured out that, you know, though
[26:08] Databricks was able to solve the problem for us, but the performance was Netezza was much better. So, then we had to change track on, you know, how the architecture pattern evolves, right? So, we had a small file problem, right? Where you basically have too many files, you basically use auto loader, you spin up a cluster, and then
[26:25] you execute the the workflow. The time that it took, you know, for an end-to-end execution of the workflow was quite large compared to what you would see in Netezza because Netezza everything was happening within that appliance. Right? So, that's where, I mean, we implemented uh grouping of the files, especially for the NRT workloads where
[26:41] we removed this cluster bootstrap problem where, you know, the the amount of time that it takes, you know, for the cluster to come up for the workflows to execute, that was primarily resolved with the the problem. And the second bit was we also uh introduced them in CDC CDFs, change data catalog change data feed feature,
[26:58] where it basically our data bricks was smart enough to basically understand what is a new data set that I need to load. I mean, you know, in the medallion architecture from stage to bronze, right? And to to gold from bronze to gold to bronze to silver to gold. And uh that basically helped, you know,
[27:14] increase our performance threefold, right? So, the workflows that were taking uh you know, 2.5 hours to complete them and was completing in 87 minutes and we have basically improved that even further now. I mean, with the AMR AMR workflows that we had, right? So, the these are some of the patterns,
[27:30] I mean, that we primarily implemented, you know, for for our NRT. And as as Shashank alluded to, right? I mean, all these things basically was improved over time as opposed to you know, just taking one architecture and just uh you know, executing it uh Apart from this, right? I mean, there are a few other uh aspects I mean, that
[27:46] we covered. So, the NRT workflow that worked in A pattern had anti-join patterns where the volumes were, you know, mediocre, but that didn't work for you. So, that's where as I mentioned, I mean, we use CDC. Some of the workflows like uh
[28:04] your Python logic that I mean, which was sequential I mean, we introduced PySpark, which basically helps with parallel execution. And then we had uh DevOps sort of things, right? I mean, where it was manual initially I mean, we basically executed uh we uh deployed the code using CICD I mean, which you know, improved the the the
[28:19] deployment you know, pattern that we had within the within the bank, right? And with all these enhancements I mean, whatever we did I mean, some of our workloads I mean, we got a 95% improvement in the execution time, right? Where, you know, the workflows that were running for 2 or 3 hours basically completed in 5 minutes,
[28:35] right? And that's precisely a combination of the compute and the architecture change I mean, that we did within the workflows. Now, all of these things worked fine, right? So, initially I mean, when we started executing the architecture, I mean, we didn't worry about the cost, right? But then, as we basically went live with APAC
[28:52] anemia, we saw that, I mean, cost is a factor I mean that we need to consider it very seriously because, you know, you know, you can quite easily, I mean, run over your cost target that you have. And we were seeing that with APAC anemia. That's when we started looking at, you know, our cluster configuration, I mean, you know,
[29:07] we didn't have serverless enabled within the bank for the right reasons because, obviously, Databricks is a SaaS platform. And some of the clients that we currently have are sensitive clients where they don't want their data to be, you know, exposed to the SaaS platform. So, bulk of the compute that we currently use was primarily AWS, you
[29:22] know, provided compute which is your job cluster, pool cluster, and all-purpose cluster, right? So, we primarily have used that, but we have used SQL Warehouse, you know, for some of the large SQL warehousing workloads. And the next journey, you know, for our challenge was, I mean, cost efficiency,
[29:38] right? So, what we did was we primarily look at looked at our cluster config. We looked at our CPU utilization on, you know, where where the where the current bottleneck is. And we basically re-jigged, you know, some of our CPU configuration, I mean, setting up the auto-scaling, setting up the termination time, you know, from 30 30 seconds to
[29:56] say 15 seconds, and then vice versa. And we basically were able to bring down our cost by 30 to 40% in APAC. And we can see that, you know, coming down even further with the reduction of Photon for some of the workflows where it's not needed. I mean, Photon is by default enabled for all workloads. We primarily disable Photon, which will which
[30:12] basically enabled us to bring down the cost further. So, one of the things that I just want to suggest is, right? I mean, though people consider Databricks as a cloud-based platform, you know, the the synergy that we saw with Netezza migration, right? Where bulk of the workloads that we had, the stored procedures, the queries, I mean, works
[30:27] 95% out of the box in Databricks. So, it's not a big big learning curve that is needed, I mean, if someone wants to, you know, migrate their workflows to Databricks. One of the challenges that we had was, you know, the the usage of Informatica and other ETL tools, that which was quite legacy. So, that's where
[30:42] we leverage, you know, the vendor partner and, you know, Claude and GitLab Duo to basically convert or migrate some of our workloads from Informatica to, you know, a Databricks compliant workload. And, you know, we we saw that, you know, the execution was was pretty fast over there.
[31:04] and the cost optimization and the cost factor is still evolving where, you know, now we are live with all three regions, but we still need to basically ensure that I mean we are running good in our run rate. So, one of the advantages probably that we had was I mean we knew what our run rate with Netezza was and we had to beat that run rate, right, by at least 15 to 20% where
[31:20] we had to demonstrate that, you know, we have chosen a platform that is cheaper to execute than what we see with a with an on-prem you know warehouse appliance, right? Because with an on-prem warehouse appliance, there are a lot of costs that are hidden where your network cost, your data center cost, I mean we don't get to see those costs, I
[31:36] mean, first hand. Whereas when it comes to cloud, I mean everything is, you know, transparent, right? Where, you know, you basically run a query, I mean you know that you know your query costs X amount of dollars. There was a person I mean who just ran a select star from a table, $70 for that query, right? So, every query is charged, right? So, we
[31:53] So, it's it's a mindset change that, you know, people need to understand, right? But at the end of the day, the the the the the architecture that we currently have, I mean, works. I mean we were able to demonstrate that, I mean, with the volumes that we see and with the performance that we were getting with Databricks and the the the platform
[32:09] works, right? So, that's precisely I mean where where we are today and uh from an outcome standpoint, I mean I just wanted to just elaborate on one of the key uh drivers, right? So, prior to prior to Databricks migration that was for APAC, I mean, pre uh
[32:27] September 2026, there was this liberation day event that happened in uh in US that basically resulted in a volume spike in in Asia. So, you basically saw 590 million that was the highest peak you know, transactions that we saw in APAC. Though Netty's are held up, I
[32:43] mean, loading those 590 million volumes, the team had to work 24/7 hand-holding the jobs, hand-holding the reports. And, you know, though the data loading happened without any problems, I mean, the reports were delayed by 3 to 4 hours, right? I mean, with manual hand-holding. Come September, we had this Middle East
[33:00] event that happened. That basically, you know, you see that, you know, the volume again spiked to 500 million, but you see the reporting line, you know, the reporting SLAs were still met. It's still static, right? And, there was no one, I mean, who was monitoring the jobs. Everything was, you know, seamless from our perspective. On an average, I
[33:16] mean, we see 200 to 250 million volume in APAC, but over the past few days for for past few weeks, I mean, we are seeing the volume spike up at 400, but it's just seamlessly working, right? Without any manual touch points that we have, right? So, that's the benefit or the beauty of, you know, the platform where, you know, we were able to scale
[33:32] up the APAC business, right? To support 2X 3X volumes. And, you know, enable them to, you know, expand their volume growth and business growth. Now, what the new platform made uh possible, right? So, I have both Harsha and Shashank, I mean, alluded on a few
[33:48] things. So, one was, I mean, 100% CICD compatibility, wherein, you know, with Informatica, I mean, we always had challenges with uh you know, deploying the platform in an automated manner. So, the deployment was manual. With Databricks, I mean, we have integrated GitLab with Databricks, wherein all your
[34:04] deployment is automated. Cross-region data sharing, I mean, this was one of the biggest factors for us. With Netty's, I mean, if I had to generate any global report, I had to copy data from all two regions into a third region, which was Americas, and generate all my reports out of America. So, there was a lot of data copy that
[34:19] was happening. With data sharing enabled in Databricks, we have basically eliminated the the data copy copies that we currently do. And, you know, the report that is generated out of US has a seamless access to the data that you have in uh uh APAC and EMEA and we have implemented
[34:34] that first time within Barclays, you know, with with uh Databricks Delta Sharing, right? The next is the consumption layer modernization where, you know, we were able to integrate Power BI with Databricks and we were able to reduce our reporting footprint by 60% with the, you know,
[34:49] combination of Power BI and Copilot. There were a lot of static reports that we generated. I mean, we were we were able to eliminate that. And you know, we were the business was able to self-serve a lot of these reporting capabilities using Power BI and Copilot. And faster decision-making, right? So, as I mentioned earlier, for every ask that we
[35:06] had before, you know, it was always a to and fro uh to and fro email from uh the business. I mean, that technology used to get so we have eliminated that uh business path and that's what the new platform has made us made possible. And Shashank primarily alluded to, you
[35:22] know, the the modernization without uh illusions. So, I just don't want to go through that again. Probably Harshavardhan, you can uh Sure. Uh I think it's worth like talking about what were the gotchas, right? Um We we did this whole thing in 14 months, as I said. So, incredible
[35:38] timeline pressures, right? Uh we had to shut down Netezza by end of this month, actually. Yeah. Uh and so we we as I was saying, the final phase of this went live last week and we are going to shut down Netezza this weekend.
[35:54] Uh the gotchas were basically FinOps, which uh Shivaram alluded to, right? At one point, we were running at what? 15,000 pounds a day, right? Um roughly like four times over the budget that we had.
[36:10] Um so, this this this happened because of three different reasons, right? We had an attribution problem. I e. like new players came on board and they they had those costs were getting attributed to this particular project. So, you need to be careful about that. Like, you know, federated model is always the best. So, every uh different business
[36:28] unit that comes in, every different use case that comes in, if you can, federate out the account so that they own the cost in their accounts. So, you don't have these uh you know, uh cost like leaking from one to the other, right? Uh the second thing was um
[36:43] optimization. Whenever you do these kind of pro- projects like optimization of cost is an iterative process. You need to like focus on it, iterate through continuously, and like keep squeezing the cost. The final thing was um you know, around the actual migration
[37:00] bubble, right? Because you are doing like four regions in in 14 months, that means that you're running UAT, you're running prod parallel, you're running prod, and and and a dev environment. You're running four environments. So, scaling that down as what as you go through these processes is super important.
[37:16] The second gotcha is like, you know, if you think that the vendors will do everything for you, however great they are, right? Uh Databricks has been a great partner, Impetus has been a great partner, but like, don't be under any illusion at the end of the day, you need to have the SMEs, you need to own the delivery cuz
[37:32] uh when once the project is done, like you need to uh hold a handhold it in production, right? So, you need to support it going forward. Uh apart from that, like you know, um architecture will continue to evolve. There is not going to be like, you know,
[37:47] when when you sit down and think about an architecture, as you are depending on the complexity of your use case, as you go through the cycles of like uh tackling different problems, your architecture will evolve. So, you need to be flexible and nimble enough to like change course,
[38:04] right? Anything else that I missed, guys? Pooja, you covered it. And perseverance and hard work pays off. Well, with that, we'll conclude. That was our journey, guys, and hopefully that has inspired some of you to take up
[38:19] the journey. It is possible to do it. And I wish all of you luck with your journey. Yeah.

Learn more about the Databricks Data and AI platform.

The information provided herein is for general informational purposes only and may not reflect the most current product capabilities or configurations.