Skip to main content

Serverless at Bank Scale: NAB's Complete Workload Migration

Summary

  • National Australia Bank migrated from a fragmented data stack spanning Teradata, Redshift, Snowflake, and Informatica to Databricks serverless across notebooks, jobs, SQL Warehouses, declarative pipelines, and model serving, with 1,500 users now on serverless SQL.
  • The migration uses Unity Catalog for fine-grained governance and lineage, Private Link for secure connectivity that keeps data off the public internet in a regulated financial environment, and a phased workload prioritization approach with Spark workloads uplifted to Spark Declarative Pipelines for incremental processing.
  • NAB is integrating Genie Code for AI-assisted analysis and the Databricks AI Gateway for AI workloads, with Lakewatch, Lakebase, and disaster recovery on the roadmap as the next wave of serverless adoption for the bank.

Serverless at Bank Scale: NAB's Complete Workload Migration

Watch: Serverless at Bank Scale: NAB's Complete Workload Migration
Serverless adoption at enterprise scale requires more than technology choice: it demands a strategic approach to governance, security, and incremental migration. National Australia Bank (NAB) transitioned from a fragmented data stack (Teradata, Redshift, Snowflake, Informatica) to Databricks Serverless across notebooks, jobs, SQL Warehouses, declarative pipelines, and model serving. This case study reveals how regulated financial institutions unlock serverless innovation while maintaining fine-grained governance and secure connectivity through Private Link.
Learn NAB's phased migration strategy, workload prioritization approach, and infrastructure design using Unity Catalog for governance, Private Link for connectivity, and Databricks AI Gateway for AI workloads. Discover hard-won lessons from uplifting Spark to declarative pipelines, scaling serverless notebooks to 1,500 users, and integrating Genie Code for AI-assisted analysis. Take home actionable patterns for cost reduction, faster MLOps cycles, improved SQL performance, and sustained query scale in regulated financial environments.
🤝

Chapters

FAQs

Why did National Australia Bank migrate to Databricks Serverless?

NAB's previous data stack was fragmented across Teradata, Redshift, Snowflake, and Informatica, creating governance gaps and slowing down innovation. Databricks Serverless provided a unified platform where compute scales automatically, enabling NAB to focus on platform-level innovation rather than infrastructure management.

How did NAB approach the migration from their legacy data stack to Databricks Serverless?

NAB took a phased, workload-prioritization approach, assessing each workload for serverless readiness and securing organizational approvals before transitioning. Spark workloads were uplifted to Spark Declarative Pipelines for incremental processing, and SQL workloads moved to serverless SQL Warehouses, scaling to 1,500 users.

How does NAB maintain security and governance on Databricks Serverless in a regulated environment?

NAB uses Unity Catalog for fine-grained governance, lineage, and access control across all data and AI assets. Private Link provides secure connectivity and egress control within NAB's regulated financial environment, ensuring that data does not traverse the public internet during processing.

What AI capabilities is NAB adding to its Databricks platform?

NAB is integrating Genie Code to provide AI-assisted development for analysts and the Databricks AI Gateway for managing AI workloads including model serving. The roadmap includes Lakewatch for security analytics, Lakebase for transactional applications, and disaster recovery support as the next phase of adoption.

Full transcript

[00:07] Good morning everybody. Uh thanks for making the time to be with us this morning. Uh my name is Tom McMe. I'm a solutions architect at Datab Bricks based in Melbourne. And today I'm going to be joined on stage with Alan Cuthbertson who's a senior manager and distinguished engineer at National Australia Bank. And today we're going to be talking a little bit about what it
[00:22] takes to adopt serless at scale. Um we're going to be talking about some new concepts uh from a data bricks perspective. And if you think about from a serverless standpoint, we haven't we've been talking about serless for a number of years now. Um but and it's but it's where our core key product
[00:37] innovation is continuing to take place. And so you've seen over the last couple of days some of some great announcements within the keynotes. And it's and this is actually being powered by serverless from a compute standpoint. And so today we're going to be sharing how one of our leading financial services customers,
[00:53] National Australia Bank, have approached their serverless adoption. How they've looked at taking advantage of this innovation that databris is continuing to uh release to to our customers. Now I think their journey provides some great proof points, some great lessons learned and some really uh good insights for
[01:09] hopefully you can take away today and apply to your organization tomorrow. I'm going to be spending a few minutes up front just to talk a little bit about what we're seeing from an industry perspective and then I'm going to hand it over to Alan who's going to take you through their journey and and share a little bit more about where they're looking to head to next.
[01:26] But first, I thought it'd be worthwhile just uh drawing your attention a little bit to a sound bite that's been um released or published as part of our recently published financial services outlook for 2026. And here we talk call out a major shift uh and in that being
[01:43] roughly 94% of customers are now piloting uh generative AI oric AI workflows across core business functions like cyber security risk pricing and uh and personalization. So when I see this number, I see uh this
[01:59] as no longer AI being a key differentiator or pilot piloting AI no longer being a key differentiator, but in fact where the competitive advantage is really starting to be shown with our in customers is the ability to deliver and execute. And so you heard from uh in
[02:16] the keynote today in fact where Casey drew to um to calling out how teams are stuck building infrastructure to run their agentic workloads. Now the key divide can be driven by both from a technical and a nontechnical aspect. Um but at some point uh within
[02:33] this journey for customers comes the key question around can the platform support their data and AI at scale. And this is where serverless plays a key role. Uh and so when you think about the traditional data stacks that have been adopted by customers over the years you
[02:48] know they they've been built for historic reporting uh and uh and retrospective batch analytics. not this continuous AIdriven uh workload patterns that we are seeing uh being attributed on the platforms today and we've seen at data bricks internally
[03:05] and also with our customers this change of the shape of how customers are using or seeing their platforms being utilized over the last six months where agents are now becoming one of the number one users that are driving queries within your platforms and what this means from
[03:20] a platform perspective is the requirements have shifted No longer you need to build a platform just for the personas around data engineers, analysts and scient
[03:37] with your data platform as well. And this is being driven from a continuous fashion also. So while this platform shift is probably not a surprise to you, um what one of the challenges that we see customers challenged with is how do they preserve or support this shape of
[03:52] queries and shape of workload patterns without fragmenting their their data estate and AI estate. You know, being able to apply consistent controls, auditability and lineage across their data that is being reasoned from across with with their agents.
[04:08] So serless can help simplify this by providing a execution fabric across any type of workload pattern. Think of it as that single execution layer. Um, additionally from a serverless perspective, it enables customers to take advantage of the pace of innovation that databris is continuously providing
[04:24] to customers. As releases are made available within the platform from a serverless perspective being evergreen and versionless, you can take advantage of these new features without having to think about the infrastructure upgrades that you need to take place from a classic perspective you might have to do
[04:39] today. So when customers first start their journey and thinking about how can they adopt serless within their uh with within their environment, we often spend a lot of time with customers up front to build the trust and confidence of of how
[04:56] um the platform will scale within their organization. So serless brings a shift in the shared responsibility model where we uh the boundary leans more towards data bricks as we're taking on the ownership and the run of the underlying infrastructure. So
[05:11] think of things like virtual machines uh the the DBR management uh network infrastructure and also scaling to support your workloads. Now at the center of this is uni catalog. So think about uni catalog as that connected tissue across governance
[05:26] for your data and AI. And so as you embark on your serverless journey, you know, it allows your teams to focus more on the things that really matter that is underpinned by by unity catalog. So these are things by like identity and access policies, audit and lineage, but
[05:43] also most importantly workload design as well. So, and as the agentic demand grows and the shape of the queries are continuing to I guess exponentially grow from an agentic reasoning perspective across your governed data um so too the
[05:59] guardrails need to scale with you on this and this is where serless can help. So let's now bring this to life a little bit and um in my mind I think you know the real proof in the pudding is actually hearing from a customer who's gone through this in the real world. uh
[06:15] and so what does it might take to adopt serless inside a bank across some uh critical workloads also but respecting the security controls and governance requirements that uh that a regulator organization need to adhere to. So I'd like to now invite Alan to the stage who's going to share through the journey
[06:31] about how they've approached serless adoption um working closely with their security and risk teams and scaling across uh their organization. Ellen over to you. Thanks Tom. Yeah, it's it's my pleasure to take you through what we're doing at National Australia Bank. Uh how we're using data
[06:47] bricks to build out our strategic data platform in a highly reg regulated industry. Uh so first to set the scene a little bit uh NAB is the second largest bank in Australia and the largest business bank. Uh for 170 years NAB's
[07:04] been helping our customers with their money. Uh we work with small medium enterprises and big businesses to help them start, run and grow. Uh we serve a customer base of 10 million at around 700 locations in Australia, New Zealand,
[07:21] and around the world. Uh I'm proud to be part of a large workforce as one of 38,000 colleagues. Um and ours is a large business. Uh we're truly datadriven. with a lot of data to work with and it's a critical resource in
[07:37] every part of the bank. So a few years ago we had a vision for a a brand new data platform uh and there are four key reasons why we decided to embark on this this journey. Um firstly we needed to simplify uh our platform at
[07:54] the time had dozens of technologies uh terod data red shift EMR snowflake informatica airflow the list goes on. Some of these were were modern at the time uh but the overall solution was
[08:09] fragile. It was complex and had scalability issues. So we needed to we needed to standardize uh patterns on the the system at the time were use case driven. This led to lots of different paths through the system and a choose your own adventure kind of style.
[08:27] We needed stability and reliability. Uh data transformation logic wasn't visible. We needed uh transparent lineage end to end and there was limited monitoring and alerting capabilities and uh finally scalability.
[08:44] We had to.
[09:00] Oh, good. Uh, yeah. Sorry. Scalability. We had to build a a system that would last for the next 5 to 10 years. Uh, so to build this data platform with kind of three building blocks. Firstly, a really clear data strategy that the business could understand. Uh, NAB's
[09:17] data strategy is to deliver data at speed, at scale, in a secure way. Uh, data is fundamental to how we serve our customers and how we operate. Data like electricity, what do we mean by this? It needs to be reliable and
[09:33] continuously available for everyone. Secondly, we needed leading edge technology that's futurep proofed. Uh data bricks is the cornerstone of how we're going to power NAB's data strategy. The powerful thing about using data bricks is that everything's in the
[09:48] same place compared to existing approaches where we had to rely on copying extraction integrations and managing all the the spaghetti in between. Finally, people. So Nab's got a lot of
[10:04] experience running cloud applications and platforms as well as traditional database and data warehousing skills but data bricks in itself is a skill. Uh so we've partner with data bricks not just on the platform but also on on training. We've made extensive use of the data
[10:21] bricks academy as well as uh classroom focus sessions to ensure that our analysts, machine learning engineers, platform engineers, data engineers all get a data brick certification before they get access to the platform. And we
[10:37] find that that combination of self-paced learning and classroom based gives the the right mix of kind of technical and NAB specific training and again uh building a data bricks platform needs uh multidisciplinary
[10:52] teams that are uh curious and keen to learn new skills. We've partnered really closely with our account teams as well as the the product teams in in San Fran and Amsterdam to ensure that our requirements feed into the the data bricks roadmap.
[11:14] So this is how we engineered ADA which is our lakehouse and how we get to to data like electricity. So if you look at this diagram from from left to right the the platform is only going to be as good as the data that we put in it. It's going to power our bankers and our customers and our
[11:30] strategy to be the most customer centric company in Australia and New Zealand. So we go to systems of record and extract everything of business value. We've got a hyper standardized small number of ingestion patterns to ingest from the
[11:45] top 200 sources in the bank. uh we use fiverr SAS uh for software as a service sources like Salesforce and workday for relational databases we use uh HVR um
[12:00] uh for non-invasive access to the database logs and that allows intraday change data capture uh this is a a game changer for us and for a lot of companies that are used to a kind of end of day picture in a traditional data warehouse and then for filebased sources
[12:17] We use autoloadader and datab bricks delta live tables or spark declarative pipelines as they're as they're known now and the lakehouse itself is built in a medallion architecture with bronze is raw like source all the conformance and
[12:35] transformations happen within within data bricks and we've we've gone for a conformed third normal form uh IBM reference model in silver and then our gold model is a semantic layer that's modeled for for business usage.
[12:55] At the heart of it all is is Unity Catalog. Um, Unity Catalog solves a number of problems for us for data governance, user access, data lineage, and making sure that people have access to the data they need when they need it, and they know what exactly what they're working with. Uh every attribute we load
[13:12] into our lakehouse uh gets metadata captured and tagged and we sync that metadata between unity catalog and calibra which is our enterprise data governance tool. Uh this also allows us to deploy our own fine grain access control.
[13:28] And on the right hand side from the consumption side we've got over 1,700 users of a greater than one pabyte lakehouse with pipelines running continuously. Uh we've got data scientists that are uh familiar with with Python and comfortable using the data bricks
[13:45] notebooks that's got ML flow baked right in there. Um and then we've got data analysts who are more used to querying relational database systems. They use uh serverless SQL warehouses that spin up and spin down as as peak rises and
[14:02] falls. We visualize all of our performance metrics using AIBI dashboards and PowerBI and now our users are are quering the data using natural language in Genie spaces.
[14:19] So as I said at the heart of the data our data bricks implementation is a unity catalog a single place to govern everything in data bricks including all of our regulated critical data elements. So the next logical step in our journey was serverless compute and allowing uh
[14:37] serverless data bricks compute to access data in our data plane and our other resources in the in the bank. Uh and as Tom mentioned adopting serverless opens up an array of innovation from data bricks. Uh in our case we took a methodical approach to migrating across
[14:53] to to serverless. And in the next few slides I'll I'll run through how we've gone about it. So before we could migrate anything um to serverless, we had to do a number of assessments as part of the approval process that involved a third party risk
[15:09] assessment, customer risk assessment, security architecture review, and we also had to work with one of our our regulators. Uh we had to prove a number of key controls were in place before we could get the green light for a box number two on the on the diagram here.
[15:26] things like internet onoff serverless egress control policies and budget policies. So private link lets traffic between your users and data bricks as well as between data bricks control plane and
[15:42] your data plane stay on the crow the cloud provider's private backbone with no public internet at any point. So that that's essential requirement for our regulated workloads and it lets us realize the full potential of serverless and streamline all of the platform
[15:59] capabilities that are going to use it. It's necessary for us to connect to other services in the bank like uh GitHub for code, artifactory for packages and hashior core vault for secrets. Uh so we had to work closely with data
[16:15] bricks and our security network and cloud architects to design and deploy it. Uh so once all the infrastructure was in place um we began a phased cut over of
[16:30] all the workloads that we identified as benefiting most from serverless. The first use case that we migrated was serverless SQL warehouses. And why would you choose serverless SQL over classic or pro? From the infrastructure management point of view in serverless
[16:47] data bricks just manages the compute fully whereas in pro you still got to uh manage minmax cluster sizes fiddle with knobs configure it the way you want it to be. From a a startup and latency point of view, uh serverless is
[17:04] typically faster from a cold start and has got better interactive responsiveness for the the users whereas Pro can be slower to start up uh depending how you've configured it. On autoscaling, um again serverless is
[17:19] more automatic and fine grained whereas Pro has got good autoscaling but you need to do more manual tuning. And from a cost model point of view, serless is great for like kind of bursty BI workloads. It results in less idle waste time. Pro
[17:36] can be cheaper for some steady predictable workloads, but if you tune it well, but it's easier to waste money if you oversize clusters and and leave them running. So what we did is we've migrated 1500 users across
[17:51] to to several SQL warehouses across 20 domains. Um, we segregate our users by domain. So each domain has one or more SQL warehouses. You set the t-shirt size for the new SQL warehouses. You do some simple config. You work with the users
[18:08] to move things across, monitor performance, then switch off the old stuff. It was a really seamless transition. For serverless notebooks, um serverless compute helps most with the developer
[18:23] speed and platform simplicity. You get instantish start with less waiting time to attach and start clusters. Um there's no cluster admin work. So there's no sizing, no patching, no DBR upgrades um for for most use
[18:41] cases. Uh it autoscales by demand. So it's better for spiky interactive workloads. It's also got better utilization. Uh so that avoids any idle uh personal clusters that are just sitting
[18:57] running out there. It's faster to onboard new analysts and engineers. They could you can just run notebooks without having to know cluster config and cluster setup and you get far few this thing works on my cluster type complaints. So to summarize um it's
[19:14] great for exploratory analysis, lightweight feature engineering, SQL, Python notebooks with intermittent activity and it's really good for large analyst communities. Couple of things to watch out for. Uh it can be less flexible um for kind of
[19:32] highly customized libraries uh system configs that kind of stuff. It can be costlier for long running always on a heavy compute versus dedicated clusters. You know, obviously uh you do have to check if you've got any specialized networking or security
[19:48] constraints or data that you want to access from the the serverless plane. But I'd say as a as a practical rule, uh default default all your notebooks to to serverless and only carve out uh dedicated clusters for really specialized heavy jobs.
[20:06] And again, we've we've migrated hundreds of analysts, data engineers, and data scientists across to use these. Now, for people that haven't used Delta live tables or or Spark declarative pipelines as they're called now, um this
[20:23] is this is kind of what they look like. You you declare them through code and then you see a graph like this in the data bricks UI. Uh so we began our journey by standardizing all of our bronze pipelines to use SDP. We
[20:38] onboarded 120 data sources in the first year and today 100% of our bronze pipelines run in in STP and about 50% have been migrated in in silver. Our long-term goal for uh declarative pipelines is end to end stream
[20:54] processing from bronze all the way through to all the way through to gold. Um, declarative pipelines enable incremental processing, built-in change data capture instead of reprocessing entire data sets. Pipelines pro only process new and change records. So that
[21:11] reduces cost, improves performance and reliability. Uh, LakeFlow's built-in CDC capabilities, including auto CDC, again eliminates the need to to hand code any slowly changing dimension logic that you've got in your in your Spark and
[21:27] SQL. So it just simplifies things and we've seen uh some real tangible outputs from this this process. Latency's dropped dramatically. We've got some uh some use cases now running 15 minutes end to end all the way through the the lakehouse. Uh success rates have
[21:42] improved. Uh data quality has improved. Onboarding engineers is faster because we've got one framework to to build on. And we've significantly reduced the cost of our kind of bronze and silver pipelines.
[22:01] So a little bit about Genie code. After we migrated all the the users to serverless SQL and serverless notebooks, uh we enabled Genie code for analysts, data engineers and and AI engineers.
[22:16] That was probably about six months ago. And now I looked at this the stats today and half half the people that access our platform use the Genie code this week. So everyone's almost everyone's using it. Uh the big advantage Genie Code's
[22:33] got over uh other coding assistants that we've got in in the bank is it's got access to everything that the users got access to in in Uni Unity catalog. So that gives a huge advantage over Claude for example. H we've had really good
[22:48] feedback from the the user community and they just they just want more. So it's landed really well. Uh serverless also powers our AI stack. So the data bricks AI gateway which is
[23:04] part of the mosaic AI model serving as a server serverless uh governed entry point for all your LLM and ML model traffic. So instead of letting every app call OpenAI or Anthropic or any internal models
[23:19] directly, all of the calls flow through a single data bricks managed gateway endpoint. It runs fully serverless, no clusters, no infrastructure to manage, tightly integrated with Unity catalog uh ML flow and all the data bricks
[23:36] governance. So this is a a kind of high level architecture of the AI capability that we've built. Uh there's two main goals that we set out to achieve with this. One's to build, test, and deploy a given use case all in the same place.
[23:51] Two is to be as self-service as possible um while still providing all the different capabilities needed to develop uh AI use cases in a safe and secure way. As you can see, we've used quite a few data bricks capabilities to provide a a
[24:07] kind of seamless integration between the different components that enhances the developer experience. Uh it means you can quickly and easily test and compare different LLMs against each other. And we're also utilizing data bricks apps to bring the application to the data.
[24:28] Yes. Yeah. So, we've said a few times today that that serverless gives us access to the the newest innovations that data bricks have on offer and we're we're constantly working with our account team to see what else we can use. Uh so, a few recent examples, first being uh Lakewatch, the the new kind of
[24:45] security incident and event management platform. All the jobs in Lakewatch are fully serverless. Uh so we've recently announced a design partnership with with data bricks on on this product and we've been working with them for the last the last few months and there's a there's a lot to come in this space. I think it's
[25:01] a massive opportunity for for kind of both both companies. So there kind of watch this space. Uh lake base we've heard a lot about this week. We're we're using it for um online feature stores and we're looking to migrate some of our ODS use cases across to to Lakebase
[25:21] and then from a multi-reion and multicloud disaster recovery. It clearly becomes a lot easier the more serverless processes that that you run. If you're not managing your own computing, you you give data bricks permission and access to more, then you
[25:38] can work closely with data bricks to manage your your DR solution. So that's something that we've got on the road map as well. Thanks very much. And back to Tom. Thanks. Good.
[25:57] Nice one. So just a few comments from myself before we close up the session and we'll take some questions. Um, so the next wave or the current current wave and the next wave of innovation is continuing to take place in serverless. Um, and so if you haven't started your serverless journey today, then I highly encourage you to reach out to your account teams and we're here to support
[26:13] you in addressing any questions you might have about it and building the trust and confidence around your serless adoption. Exactly what we've done with National Australia Bank to help them achieve some of their great success goals that Alan has spoken about. Um, and the platform is no longer just built for the traditional personas. Um so
[26:30] there is a new persona that you also need to think about on how you can scale your platform to support this and this is the agentic workflows powered by Genie agents or directly through MCP integrations. Um we're seeing a huge growth and shift in the shape of how uh
[26:45] data uh platforms and data bricks is being utilized to support customer use cases and also the real unlock in enabling seamless uh uh compute execution layer uh is uni catalog as well. And so this enables your teams to have confidence around the different
[27:01] types of compute that is being utilized to support the various type of workloads. Having consistent controls across each of those domains to enable you to have that full uh single place to have that control and and full visibility and auditability.
[27:17] So just wanted to also leave you with just a few reference points for potentially might be of interest for you to double click further today. Um there's a great blog article that our platform engineering team have released around a little bit about under the hood look at how they built serverless and this notion of versionless and also
[27:33] decoupling the the um the client to the to the backend compute. I'll encourage you to to double click on that if you're interested. And then also from a a higher level perspective, um looking at from an industry standpoint within the financial services industry, um the
[27:49] financial services outlook and key trends that might um resonate with yourself and and help you define some of the key um challenges or opportunities that you could potentially partner closely more closely with from a data brick standpoint.
[28:05] And lastly, we're very lucky to have National Australia Bank here presenting a number of sessions across the summit today. So this is one of three sessions that the team are spending their time and and sharing their their knowledge and learnings to everybody here. Um we've also got we had a session yesterday so that was recorded so that
[28:21] will be available to you as part of the virtual summit. And then shortly after this session here uh the National Australia Bank team are sharing their Genie adoption story in the Moscow West building. um if there's interest in um following us across there. Um but certainly that will be recorded and made
[28:37] available after the summit as well.

Learn more about the Databricks Data and AI platform.

The information provided herein is for general informational purposes only and may not reflect the most current product capabilities or configurations.