SAS to Databricks: How Great Southern Bank Democratized Data with Genie and Self-Service Analytics
Summary
- Great Southern Bank, a 420,000-customer Australian bank with 1,000 employees, rebuilt its entire data environment from scratch on Databricks after consolidating 85% of enterprise data onto a single platform with Unity Catalog governance.
- The modernization delivered measurable outcomes: regulatory return processing time dropped from days to automated hours, fraud investigation effort fell by 15%, and 2.6% additional lending capacity — representing hundreds of millions — was unlocked.
- Non-technical business users now assemble datasets in minutes through Genie Spaces instead of waiting days, and the bank achieved 100% program ROI in 2026 while deploying its first production AI agents.
SAS to Databricks: How Great Southern Bank Democratized Data with Genie and Self-Service Analytics

Great Southern Bank (420,000 customers, 1,000 employees, Australian market) faced a critical choice: maintain three fragmented legacy data warehouses or rebuild from scratch. Matt Cammack's team chose ground-up modernization on Databricks, consolidating 85% of enterprise data. By standardizing on a single platform with Unity Catalog governance, automated pipelines, and fine-grained access controls, GSB reduced time to deliver regulatory returns from days to automated hours, freed fraud investigation effort by 15%, and allocated 2.6% additional lending capacity (hundreds of millions).
Now, non-technical business users ask questions in plain language through Genie Spaces, assembling datasets in minutes instead of days. GSB deployed the first production agents, achieved 100% program ROI in 2026, and proved that smaller institutions can compete by executing faster on data insights than larger competitors.
Chapters
00:00Great Southern Bank SAS to Databricks Migration01:32GSB Overview: Customer-Owned Bank and Competitive Challenge02:36Legacy Problem: Fragmented Data Across Multiple Warehouses04:15Strategic Decision: Rebuild From Ground Up with Single Platform05:40Automation, Governance, and Fine-Grained Access Controls07:03Business Impact 2023 to 2026: ROI Through Data Foundation08:43AI Acceleration: Home Loan Retention Model and New Agents10:03Shift in How People Work: Genie for Non-Technical Users11:09Democratization: Competing Without Scale Through Data Speed12:55Key Learnings: Governance, Standardization, and People16:01James McNiff: Prioritization Framework for AI Use Cases17:08Databricks Banking Outcome Map: Growth, Risk, Efficiency18:28Value and Readiness Filter: High-Impact Use Cases21:11Home Loan Retention Race: Real-Time Decision Example23:05Genie Space Demo: Governance-Aware Query Answering24:45Complex Queries: Ranking Customers by Retention Risk26:07ML Explainability for Churn Prediction27:00Summary: Data Has Answers, Discipline Is Key to Shipping AI
FAQs
Why did Great Southern Bank choose to rebuild its data platform from the ground up?
The bank's data environment had grown organically into multiple independent warehouses with fragmented ownership and heavy reliance on manual processes, producing conflicting and inconsistent data. Facing new regulatory requirements as a significant financial institution with assets over $20 billion, GSB made a deliberate decision to start from scratch with a single strategic platform on Databricks rather than incrementally evolve its legacy systems.
What measurable business outcomes did Great Southern Bank achieve?
GSB reduced time to deliver regulatory returns from days to automated hours, freed 15% of fraud investigation effort, and unlocked 2.6% additional lending capacity representing hundreds of millions in value. The bank achieved 100% program ROI in 2026 after consolidating 85% of enterprise data onto Databricks.
How does Genie help non-technical employees at Great Southern Bank?
Business users can ask questions in plain language through Genie Spaces and assemble the datasets they need in minutes, rather than waiting days for the data team to fulfill a request. This shift has fundamentally changed how people work at the bank, reducing time spent preparing data and increasing time available for action and insight.
How does a smaller bank like Great Southern Bank use data to compete with larger institutions?
GSB's strategy is to execute faster on data insights than larger competitors, using speed as a competitive differentiator rather than trying to match their scale. With Databricks and AI tools, a 1,000-person bank has built production AI agents, automated compliance processes, and delivered self-service analytics — capabilities the bank believes let it punch well above its weight in the Australian market.
Full transcript
[00:08] Well, thank you very much for making it this far. Um, delighted to be with you. I'm just going to kick off with just a very short video.
[00:55] Hey, hey, hey.
[01:32] Uh, I'm Matt Kamak and I'm head of customer technology and data and AI at Great Southern Bank. Um, not to be confused with the Great Southern Bank here in the US, but we are one of Australia's largest customer-owned banks. And for over 75 years, we've been empowering our customers on their
[01:48] financial journey. To this day, we remain committed to helping all Australians own their own home. We serve over 420,000 Australians and but we operate with a fraction a tiny fraction of the size of the banks with which we
[02:06] compete. But we are still committed. We don't think our size limits our ability to compete and be provide competitive products and we've earned the reputation as being one of Australia's best banks for customer service.
[02:21] But we also have to be very realistic with only only a thousand employees. We don't have the scale with with the organizations that we compete with in Australia. So back in 2021, we saw an opportunity.
[02:36] We are a relatively simple bank and we asked ourselves a question. What if our data environment was simple too? Could data and AI give us a different way to compete? And that's the start of our journey.
[02:53] But back in 2021, we were candid about a problem. Our data environment was anything but simple. We faced a structural data challenge that will be familiar to many people in this room. Over time, our data environment had
[03:09] evolved organically. We had multiple data warehouses, each acting independently from each other. ownership was fragmented and there was a heavy reliance on manual processes even for our most critical activities.
[03:26] The result was exactly what you would expect. Our people had conflicting and inconsistent data to do their jobs. We had built an industry of factchecking as a result. And put very simply, we were
[03:42] spending far too much preparing time preparing data and not enough time on action or insight. In addition, we were becoming what the Australian regulator calls a significant financial institution as our assets went
[03:59] over 20 billion. The higher regulatory requirements of that milestone meant that it was essential that we had clean, clear, and trusted data. So we made a very deliberate decision
[04:15] not to incrementally evolve our legacy platforms but start again from the ground up literally scorch earth and we wanted to build a strategic asset and capability that really positioned the bank for the future.
[04:32] Now for a bank bank operating at our tiny little size that is not a trivial decision. But without it, we couldn't see how we can sustainably scale, strengthen governance, and more importantly bring the intelligence to our information and
[04:49] better unlock new value for our customers. So we committed to standardize and go all in on data bricks as our single govern data and AI platform. And that partnership with data bricks has been crucial for us because the innovation they've brought through
[05:05] the platform over time has helped us address many different challenges ensured that we successfully complete and as we've seen I think here this week um they continue to innovate on that platform. So over three years we executed on that
[05:22] decision. We've consolidated three platforms onto a one unifying enterprise data and we are now decommissioning those legacy environments. That's not only saved cost but has actually reduced complexity for us in a very real and
[05:40] tangible way. We chose to automate everything as we went on the platform. That obviously means we automated our pipelines, testing, monitoring, all those things. But for the business, manual spreadsheetbased work has been
[05:55] simplified and streamlined and that has created significant operational efficiencies. We built security and governance in from the ground up. We use Unity catalog for lineage and traceability and we have fine grained access controls that has
[06:13] actually helped us resolve a number of historical control issues that we had and we've even now be able to reset enterprise data risk for the organization here today. I'm pleased to say that every aspect of the bank, all functions
[06:29] from product, customer, finance, operations, regulatory, credit, risk, capital, you name it, every single aspect of our bank now feeds and flows through our data bricks environment. And this commitment to a single platform has been critical for us. By committing
[06:47] fully, every new use case and every investment we made contributes to our foundation and allows us to maximize reuse and has helped ensure that the benefits have compounded as we've gone.
[07:03] And as a consequence, we have seen real impact throughout the journey. By 2023, we delivered our first regulatory return. And that took a process that used to take days and multiple people and now happens completely automated in
[07:18] a number of hours. We found that but then as we migrated more every return and report we moved onto the platform we those benefits were repeated and that's actually created a lot of of efficiencies across the entire
[07:35] group. By 2024, we had embedded data governance and quality monitoring in fraud alone. This reduced overall effort of fraud investigation by over 15% as they spent less time colleating information and more time investigating.
[07:53] By 2025, that improved data quality meant that we were able to allocate capital with more confidence. We found an opportunity to release a significant amount of additional lending capacity which equates to about 2.6%
[08:10] of our entire book. And when you've got billions, that actually adds up to quite a large number. And I'm pleased to say that in 2026, we have actually achieved all the objectives of our business case and we've got a 100% return on our program
[08:28] investment. Now, the benefits don't end there. Um, we've with the foundations and governance that we've had in place, we're now able to pursue AI at a much
[08:43] faster pace. One example is as a bank, home lending is critical and core to who we are. But we had a problem. We had a home loan retention model that was built and managed by a third party. It was
[08:59] costly. It was opaque. And we couldn't adapt it and have the agility we needed as a bank. So a single analyst in the team rebuilt that model in-house on data bricks really within a couple of weeks and not just match but improved on its
[09:16] performance. This meant we now retain the IP and we can evolve the features and performance of that model as we move forward. Now we're accelerating our AI adoption and pursuing new value, looking specifically in how we can do better
[09:32] scenario planning, forecasting, and stress testing. And we are now about to deploy our first agents where we are automating quality assurance processes.
[09:47] But the real shift for us has been how our people now work with data. This is a quote from our senior manager of performance, product performance. Now, he's not a coder. He doesn't know how to write SQL or Python, but he does have a very intimate knowledge of our
[10:03] business. Previously, getting answers him for him meant downloading data, wrangling it in spreadsheets, often with data that was often days or even weeks old. Now using data bricks and AI he can
[10:20] actually work with the data directly. He uses Genie to ask questions in plain language. With Genie code he assembles data sets and actually visualizes the results. This has been transformative for us.
[10:36] Work that used to take literally days can now be achieved in minutes. And that not just frees up really important capacity but more importantly is unlocking creativity. So an example of this is that we used to
[10:53] when we go to decision forums now we can bring the data directly into decision forums and committees ask it questions in real time rather than relying on offline reports. So as I said earlier, we don't have the same scale as the institutions with
[11:09] which we compete. But with data bricks and AI, we're finding that capability is now being democratized. Historically, building this type of capability required significant scale. You needed big teams, you needed significant investment, and you need a
[11:25] lot of deep expertise. What's changed is that platform like data bricks has lowered that barrier and that's allowing an organization of our size to now compete more effectively and that's been at the heart of what going all in on data bricks has meant for us.
[11:48] So what have we learned along the way? So first of all there is an old adage that you know data is should sometimes should be considered a liability until it's proven to be an asset and I think there is an awful lot of truth in that. Building solid data governance and architectural foundations
[12:04] is an asset and is now allowing to us to move at pace with AI. Secondly committing to standardizing early. We partnered with each business team across the organization r rationalizing and simplifying their data
[12:21] operations as we transition them. This is probably I think the most important thing on a very long journey because without that discipline getting their buy in to come onto the platform um it really is essential to ensure that
[12:39] any fragmentation or trade-off you make where you don't build on your core platform actually compounds faster than you think. And thirdly, most of the value is not in the platform, it's in your people and how they work.
[12:55] Investing in the right talent and having consistent ways of working has been also key to how we've got to where we are. And I think the organizations that get this right don't just modernize their technology, they actually change how your bank works and how decisions get
[13:11] made. So hopefully what our story helps show is that a really small bank because even by
[13:26] Australian standards we're small um but how a really small bank can operate with the capability of a much much larger organization without often the say their same scale cost or complexity because I think as we move forward the advantage
[13:43] is going to be less about who has the most data. It is going to be about who can use their data fastest and with greatest data uh greatest confidence. And that for us is why we're proud to say we've gone all in on data bricks.
[13:58] Now to help bring this to life for you and bring a little bit of color in terms of how we've done this, I'd like to introduce James Mcniff from Data Bricks who's then going to give you a bit of a talk. Thank you.
[14:19] Awesome. Thanks a lot, Matt. It's been so good to see the journey that GSB have taken over the last two years and see it culminate being here in San Francisco on the stage is is really incredible. So, yeah, hats off to Matt and the team for such a a fantastic job. Let's see.
[14:38] So, I'd like to uh jump into one of the points that Matt just closed with there. advantage comes from how fast you can turn data into decisions. And what I would like to do for the next 15 minutes is spend a little bit time talking about how to do that in practice. And so Matt's team already did the hard part. They went all in. They made the
[14:53] deliberate choice to go all in. They build their uh government foundations. And I know that's exactly where a lot of you in the room here are heading today if you're not there already. Um, but one of the questions I get when I have a lot of customer conversations is, okay, we've got the platform, we have the data, but how do we know what to
[15:10] actually build first? Where should we actually start? And so, just um, for a little bit of fun, quick show of hands, how many of you leaving this conference have more AI ideas than you could possibly execute in one year? Just raise your hands.
[15:25] Yeah, it's a good amount. And that's sort of the problem. It's it's not it's not about capability. Um, and it definitely isn't about the platform. In my opinion, it's about actually knowing where to start and what to do first. So, yeah, I'm James. I'm a solutions architect at Data Bricks. And there's two things I'd love for you to take away
[15:41] from the next 15 minutes. One is a really dead simple way to know how to prioritize and what to choose. And two, I'll show you how to do that in practice uh with a real banking question uh sort of playing on an analyst that that doesn't have any SQL experience.
[16:01] So here's a bit of a number that frames the discussion. Uh we recently released our data bricks state of agents report. It was just this year. You could find it online as well if you're interested. There was a statistic in there from MIT technology review. What it said is that 67% of um of organizations they're experimenting with AI but they haven't
[16:17] actually deployed it in production. Only 19% have actually um productionized agents. And so if you just think about that statistic for a second, what that's essentially saying is that twothirds of people, you know, experimenting, but less than one in five have actually shipped anything to production in terms of agents and, you know, in terms of AI.
[16:35] So why does that actually matter? Well, it's it's not a capability gap. As I say, the models are more than good enough now. The platforms are there. In my opinion, it's actually prioritization that's the challenge. And so there are a lot of teams that are experimenting. They might have 10 unfinished pilots,
[16:52] but the problem is they're spreading themselves quite thin and not focusing on the one use case or the one project that would actually drive that bank forward. So, this slide is going to look like quite a lot and that is actually deliberate. Um, feel free to take a
[17:08] picture and uh the slides are available afterwards as well, I believe. But what this is, it's our data bricks um banking and payments outcome map. And so if there's one thing I'd like you to take away from this slide is that it's a business value map. It's not a technology map. It's not there's any
[17:23] architecture or specific tools. Um and the idea is that the map is created from all the work we're doing with um financial services organizations across the world with data bricks. And what it's basically saying is is that a box uh sorry a bank is either um you know
[17:41] they're either driving growth for example churn modeling uh hyperpersonalization um they're reducing risk through protecting the firm for example KYC anti-moneyaundering um you know fraud fraud analytics or uh they're operating more efficiently so
[17:58] you know uh faster and unfend close for example but why this framing matters is because it takes a sort of vague question, you know, what should we do with AI? And it turns it into a business question. Are we trying to grow? Are we trying to run leaner? Or are we trying to reduce risk? And that gives you a
[18:13] common language um with the people who actually hold the budget in your organization. So you can talk uh talk in their language before ever getting into specific products.
[18:28] But there is also a bit of a catch. So if everything's a high priority, then of course nothing is. So what you actually want to do is take that big wall of use cases that I just showed you and then you want to pass it through a simple filter. So just ask yourself two questions. Number one, how much value does this actually create for us? And
[18:44] two, how ready are we actually to ship it today with the people and skills that we have in house? Like can we actually do this right now? And what that will do is give you four boxes. So on the top left you've got um the high the high value and the high readiness. That's where the gold is. That's where you
[19:00] should start because you've got the overlap of we can do it and it's high value and there's going to be a big payoff if you can if you can nail that. Um on the top right you've got high value but for whatever reason you're not quite ready yet to execute it. So there's huge value there but perhaps
[19:15] there's some more data plumbing needed or some more model work. You can think of those as the area where you should invest to unlock more value um and come back to them but not necessarily in week one. In the bottom left you've got those quick wins. the lower value um but
[19:31] they're not likely to drive big headlines for the bank. You know, get some momentum, but unlikely to sort of, you know, generate huge news headlines, for example. Um and on the bottom right, you've got the backlog items. If you've been honest, these are the ones you should completely park and and not spend any time on. Now, obviously, I'll say
[19:48] that this is just an example. When you run this exercise yourself, the use cases will sort of fall in different boxes depending on, you know, your your organization. Um, I'll also say that this isn't just our framework. There's a lot of research out there as well to sort of back up this thinking. Uh, McKenzie, for example, um, they released
[20:04] a number of reports that essentially says the value doesn't come from redesigning everything all at once, but um, deliberately rewiring a few high impact workflows first. And so like focusing should be the the strategy.
[20:23] What what is also quite interesting is that of course it will look different for your organization, but what we have seen is that there actually is a lot of overlap with the items that fall into the top left, especially with self-service analytics. Um, and it's kind of a really easy first call. It's that the data is already on the platform. You don't actually have to
[20:39] build anything. And it's it's one of the very few things you could actually deliver that everybody in the organization can benefit from. It's not just a small team of analysts or a specific team. Everybody needs access to the data in a business and to be able to ask questions of it. So yeah, quite
[20:55] often it will end up in the number one spot of that high value, high readiness. And so yeah, what does this actually look like in practice? Um, this will be very familiar to many of you, especially if you work in in
[21:11] retail banking. Um, but essentially this is a race that plays out in in most retail banks. So, you know, several of your home loan customers are about to have their fixed rate mortgage expire. Um, and you know, as soon as that happens, a competitor could come and and snap them up and offer them a, you know,
[21:27] better offer, for example, and then you've, you know, you've lost that customer. So, the retention team in the bank has quite a narrow window to actually retain the customer. And what they'll do is is something very similar to the question here. Um, I'll just read out for you, but yeah, show me the high-risisk customers with fixed rates
[21:42] expiring in the next 30 days. How many of them are there? and what's the total value at risk. So if you think about like the old way around how you would actually answer this question and we'll take our example analyst, she hasn't got a huge amount of SQL experience. She's quite new to the team and the data is
[21:59] actually quite sensitive in its raw form. You're going to have lots of PII in there. Um you know and if you think about the data sources that the question needs to access, you know, it needs to hit the the core banking system to pull out the loans and the you know the expiry dates. um it'll have it'll need
[22:14] to hit a different system to pull in the retention risk scores for example maybe an ML store and so what actually happens is our analyst has to raise a ticket to go to a data engineering team and that team has to pull in all this data have to join them together and probably a third system which would be your you
[22:30] know your analytics warehouse for example and so this simple question has turned into like three days of work potentially with three different people um and our analyst just has to sit there and wait and there's not a huge amount she can do um and in that time. Obviously, a a competitor can come along and completely snap up that customer.
[22:48] So, what I'd like to do is show you a very simple way uh to do this on data bricks and you know how much how much fast this is. Switch over here.
[23:05] Awesome. So, this is a genie space. I'm sure you've seen probably a few hundred of these throughout the conference this week. Um this is one that I created uh using it's all synthetic data but it's realistic. So it's simulating a you know a home loan book of business. Um we've got customers thousands of customers you know home loans rates risk risk scores
[23:21] all that good stuff. But all of this is actually governed by unity catalog. So the important thing is that when I come in here and I ask a question of the data it will use my permissions. So it will only ever return data um that I'm actually allowed to see. I'll zoom in
[23:36] for you a little bit there as well. a little bit bigger. All right. So, I'm going to ask that same question here in Genie. And so, it's of course not just guessing. It's going to map my my English question uh pull apart the
[23:53] different pieces and then it will map that to the data I'm specifically allowed to see. Uh and then it will convert that to SQL and yeah, answer the question hopefully. The other thing I like about Genie as well uh if you expand this you can see the exact logic it's following as well
[24:11] but you know you can see step by step exactly what Genie is trying to do here you know filter the data group the results and so on and so on.
[24:30] All right. So here I've got my um my results. So yeah, I can see the the banded by risk category, high risk. There's 256 people of a balance at risk of 154 million Australian dollars. So you can see how much more powerful this is for an analyst who may not actually be able to write SQL. I can still see
[24:45] the SQL here, of course. Uh but I don't need to worry about the detail, which is nice. And this is a really good start, but it's not really driving a decision. It's just giving me a number. So of course, I can ask more interesting questions on top of that. So, here's one for example.
[25:01] Um, what I actually want to see is like which 10 customers specifically are the the highest value customers in the group and they're the people that I would like the retention team to actually get on the phone to and and and try and prevent them from churning. Now, this is a bit a bit more involved
[25:17] query. It has to hit a few more data sources. You know, it's um it's going to do some ranking and so on. But yeah, it's quite a bit more involved. But again, I can expand this and see the exact steps that Gene is taking. So I can see here that it's, you know, it's it's categorizing high risk at 0.7.
[25:34] Of course, I could change that if I like. You know, it's it's limiting to loans that are of a fixed product type and the rate expires in the next 30 days. So let's see. I can see the result. And we've got our customers. So here I can see specifically which customers. I
[25:50] can see the loan ID, how much um, you know, how much is outstanding on the loan as well as when the rate expires as well. And I can also rank those by retention risk score. Now, this gets really interesting when you start to think about machine learning as well. So, you know, I can
[26:07] see a customer, I can see their score, but it's not telling me why they're likely to churn. And if we use machine learning and explanability, we can actually start to look at, you know, why why this customer specifically might leave. What are the specific characteristics that define them as as
[26:22] high risk? And then you know you'll have some customers that are price sensitive so their retention team can offer them a specific offer. Some may not be price sensitive but uh you know they just um for example you know like a human reach out and speaking to the you know the customer service team. So once we know
[26:38] exactly why the likes leave we can send the team in with a much higher chance of retaining them. All right.
[26:55] Yeah. So you've seen here go the next slide. Yeah. So you've seen here that the data already had the answer. Um it was just locked away in multiple systems and you know potentially a free day a multi-day queue. And what data bricks and genie has done is turned that into a 30 secondond to a minute conversation with a specific decision that the team can
[27:11] act on right now. And so to summarize twothirds of enterprises are experimenting with AI but fewer than one in five have shipped it. And the difference isn't talent or tools or scale. It's really a discipline to pick
[27:27] the right first thing. So Matt showed you a great example of the bank that did exactly that. And so what I would suggest if I was in in your shoes, if you haven't done this already on Monday morning when you get back into the office, I would take all of your AI
[27:43] use cases and I would pass them through that filter. Ask those two questions. which ones are actually likely to drive the highest value and which ones are you ready to deliver on today and that will give you your your grid and then what you want to do is start with that top left corner where the value and readiness overlap and that's where you
[27:59] should spend your time and so Matt's Matt's story was about finding a different way to compete um not on scale but on how fast you can turn data into decisions and so you just watched an example of how easy it is to do that with with Genie and data bricks now without needing to build anything
[28:15] um so yeah your data already has the answer, but really the only thing left to do is just decide what to ask it. Thanks a lot for your time.
Learn more about the Databricks Data and AI platform.
The information provided herein is for general informational purposes only and may not reflect the most current product capabilities or configurations.