KPMG's SQL Server to Databricks audit transformation in 3 months
Summary
- KPMG UK migrated its entire SQL Server audit estate to Databricks in three months, enabling 200+ auditors to access governed data and AI without compromising audit-grade quality or data lineage.
- The migration leveraged Lakebridge for risk assessment, Genie Code for AI-assisted SQL translation that reduced refactoring effort by 60%, and Unity Catalog for lineage and access control, resulting in 400 scripts successfully migrated.
- A three-month training program built organizational confidence, and the governed foundation now powers AI audit products like Financial Report Analyzer, with Databricks SQL serverless compute handling extreme seasonal demand fluctuations.
KPMG's SQL Server to Databricks audit transformation in 3 months

KPMG migrated a mature SQL Server audit estate to Databricks in 3 months, enabling 200+ auditors to leverage governed data and AI without sacrificing audit-grade quality or lineage. The transformation required more than technology: it required rethinking data architecture, governance, and organizational capability.
Learn KPMG's blueprint for enterprise migrations: using Lakebridge for risk assessment, Genie Code for AI-assisted SQL translation maintaining 60% less refactoring effort, and Unity Catalog for lineage and access control. Discover how Databricks SQL serverless compute handles extreme seasonal demand in audit workflows, how a three-month training program built organizational confidence, and how the governed foundation now powers AI products like Financial Report Analyzer. this video reveals the architecture and organizational patterns that transform traditional audit stacks into AI-enabled platforms.
🤝
Chapters
00:00KPMG UK audit transformation story02:03Audit's data and governance requirements04:12Data strategy for complex, global audits07:28Four key principles for the transformation09:21SQL Server constraints and Databricks lakehouse architecture11:10Migration approach: AI-assisted refactoring at scale12:34Lakebridge and Genie Code acceleration15:38Technology choices: Databricks SQL and Unity Catalog16:42One platform for data analytics and AI19:22People and capability: training 200 auditors21:53Learning pathways and certifications24:22Outcomes: 400 scripts migrated, 60% reduction26:17AI-powered audit products and Financial Report Analyzer27:23Three lessons learned from transformation28:59Next phase: Altrix to Lake Flow Designer migration
FAQs
How did KPMG migrate SQL Server audit scripts to Databricks so quickly?
KPMG used Lakebridge for risk assessment and migration prioritization, and Genie Code for AI-assisted SQL translation that preserved audit logic while reducing manual refactoring effort by 60%. This combination allowed the team to migrate 400 scripts in just three months.
How does Databricks handle seasonal demand in audit workflows?
Audit workloads have extreme seasonality due to regulatory reporting deadlines. Databricks SQL serverless compute automatically scales up during peak periods and down during quieter times, eliminating the need to pre-provision infrastructure for seasonal spikes that would sit idle the rest of the year.
What AI products did KPMG's Databricks migration enable?
The governed lakehouse foundation built on Unity Catalog enabled KPMG to develop AI-powered audit products, including the Financial Report Analyzer. These products were not feasible on the legacy SQL Server estate due to its lack of AI integration capabilities and fragmented data lineage.
What are the three lessons KPMG learned from the audit transformation?
KPMG identified three lessons: invest in the right technology foundation before building AI capabilities, prioritize people and organizational capability through structured training programs, and treat governance as a prerequisite rather than an afterthought—principles the team is carrying into the next migration phase from Altrix to Lake Flow Designer.
Full transcript
[00:09] Good morning everybody. Come on, we can do better. I know. Like and I was Elena. Hey, come on. There we go. Was a long night last night. Everybody enjoyed chain smokers. You know, unfortunately, we got you the first session here this morning. Hopefully, we
[00:24] get some energy on, you know, an audit topic. My goodness. Right. an exciting audit topic, I'll tell you. Anyways, good morning everybody. Um, you know, uh, happy to be here. Um, my name is Kevin Cole Niatis. Uh, I'm a KPMG
[00:39] partner. Uh, we're going to talk a little bit about the story. We'll do little intros in a sec. Uh, but, um, really what we're going to talk about, uh, is, uh, KPMG UK's audit AI future and the transformation, uh, that has
[00:55] happened, uh, in the UK. um and some fantastic uh fantastic work that was done like a forward looking statement. There we go. So, with me today on stage uh my fantastic colleague Ellie Tucker from
[01:12] Data Bricks. Hi everybody. Um I'm Ellie Tucker. Uh I'm a five-year Bickster and I've been looking after the relationship between Data Bricks and KPMG for the last four years. Um, and I was working with the team who were involved in this transformation program. I'm from the UK
[01:29] and two of the other speakers are also from the UK, but obviously uh flights are super expensive because of the World Cup and Greta's also about to go on Matt leave. So, we've got some embedded videos because they are the subject matter experts, but uh myself and Kevin will uh give you the overview of the
[01:45] trans transformation program and the lessons that have been learned throughout it. Perfect. Thanks, Ali. All right. So, a little bit on the agenda for this morning. Um, so we're going to take you through uh data and AI at a high level from an audit perspective uh and what it means to us.
[02:03] Uh, and then clearly we're going to get into the transformation as a whole. Um, and go through, you know, different areas talking about the migration challenges and and things that we we work through. Talk a little bit about the govern lakehouse uh and the power of
[02:18] AI. talk a little bit about the centricity around people and why that really matters as you're doing this. Um, and then we'll talk a little bit about the ultimate outcomes uh and the value uh generated uh to our audit practice and then really sort of setting the
[02:35] stage for the path forward as we move along uh in the audit space. All right, so a little bit of an overview uh from an audit perspective. When I say audit, this is our external audit practice in the UK. If you can imagine, we do hundreds of thousands of
[02:53] audits, literally globally, thousands of audits within the UK practice uh as a whole. And those audits are complex. They range in nature, right? We can we have very large group audits that have multiple components. the subsidiaries
[03:10] across the globe where you got a group that you know may have 150 subsidiaries in 150 countries across the world right down to really small audits where they're less than 200 hours to complete but they all have a common factor they all have data that we leverage in order
[03:26] to execute those audits and we're governed in our space as you can imagine there's standards we have to follow the auditing standards we have regulators we have local regulators that
[03:41] monitor everything we do. So data is really important. Data quality is, you know, at the forefront of everything that we we do. Um, and looking at the problem, we've got to figure out how we
[03:56] can embed AI in everything that we do and take advantage of that to enhance our audit quality. We can look at all the data uh going forward. So you got to set up properly. So from a data strategy perspective we think about it in order to execute the
[04:12] audit and you guys can imagine as in the external auditors we get tons of information tons of data right whether it's transactional data from our clients that we have to look at general ledger data subleddger data etc. We also have
[04:28] all the unstructured data that comes in. So anything from contracts to board minutes right to you know various other sources invoices etc. We've got to figure out how we're going to harness all that information and you know
[04:44] leverage that within the audit. So our data strategy is really important because we need to enhance the audit quality. We need to make sure that data is governed properly within our organization because when we think about the various audits, every audit's got to be in its own silo. Only certain people
[05:00] should have access to that data in order to perform the audit for the individuals on those files. What's also really important as part of the strategy is we need the right metadata. We need the right context, right? Without that is very difficult to
[05:17] audit. it's very difficult to scale as well. So you need the right pipelines etc in order to deal with like and I'm talking vast amounts of data transactional data some of the smaller files maybe it's you know few thousand records where we our multinationals and
[05:33] that you're talking billions billions and billions of records so it's got to have the quality speed clarity and precision that needs to be in place right for us to get to where we need to be. So just a few things and to talk about a little bit why it matters kind
[05:50] of hinted to it but when you think about it we want to connect our data across our ecosystem. So we have a CLA ecosystem that has different pieces to it. So there's information sort of spread within that CLA ecosystem. So you
[06:05] got to figure out a way how you govern the data and bring that together without moving all that data around. Um, so that that's one of the key pieces in in looking at the strategy, empowering our auditors, looking at ways that we can democratize the data for our auditors is very
[06:22] important, right? Trying to get away from the traditional approaches to analytics that a lot of folks have have used. How do you democratize it? How do you use a capability like Genie as an example to open up you know that real time analysis and in querying of the
[06:39] data to understand what risks exist what anomalies maybe you identify also thinking about um the context like I said if you're going to execute certain audit procedures you need to have the right context in order to do that otherwise you're going to be
[06:55] chasing refining and it's going to take tons of time so you got to set that up properly And then finally, you got to have the right processes, pipelines, structure in place to get to speed when we're pulling that data in to potentially allow you for um continuous auditing potentially
[07:13] one day, right? Uh and driving insights through our for our clients. So really four key things or principles I would say that we looked at when we did it. So we wanted to simplify right
[07:28] one platform one stack right well structured that allows repeatability right when we're bringing this data in uh and we can execute on the audit stacking is standardized so shared patterns right if we get certain
[07:44] information structured in the right way and common data models to execute the various procedures that we need to do to satisfy our audit requirements. I talked a little bit about democratizing the data, right? Being able to get the data in so our auditors can interact with
[07:59] that data uh in an appropriate manner for the for the audit. And then finally, always the ultimate goal to enable AI within our audits, right? And fully embed it to allow us to execute.
[08:16] So, hi. In this section, I will speak about the transformation story in a simple way. Why change was needed? Why now? and where are we heading? The main message is that this is not technology change for its own sake. It is about building a scalable governed foundation for audit analytics today and for more AI enabled
[08:33] capabilities in the future. And it sets up why the next step matters. The trigger was a mix of practical need and strategic opportunity, not just a technical refresh. The on-remise stack was becoming harder to scale, particularly around seasonal peaks when
[08:48] audit demand increased sharply. At the same time, users wanted advanced analytics and AI to become part of real workflows rather than separate experiments. There was also more focus on moving towards stronger evidence, traceability, and assurance in a more AI
[09:04] enabled environment. So, the goal was clear. Bring structured data, analytics, and AI together on one governed platform while preserving the rigor audit depends on. Our starting point was strong, but increasingly constrained. We had a mature SQL server estate with more than
[09:21] 400 scripts and stored procedures supporting a large number of global audits. It was reliable, proven, and valuable. But the architecture was built for a different operating model. Workloads could vary from thousands to billions of rows, which made capacity planning manual and seasonal. It also
[09:36] limited how widely we could open up advanced analytics or introduce AI and machine learning at scale across the audit population. The destination is a datab bricks lakehouse on cloud. One platform for storage, processing, governance,
[09:52] analytics, and AI. Elastic serverless computes let us scale with audit demand instead of planning around fixed capacity. Unity catalog gives us lineage, access control, and audit logs builtin. And delta sharing supports secure data exchange. Most importantly,
[10:07] the same trusted data foundation can serve analysts, engineers, and auditors. That means the platform becomes easier to govern, easier to use, and easier to extend as new analytics and AI needs emerge. This final view for this section shows
[10:22] how the architecture comes together. Data enters from audited entities and source systems, lands in cloud storage, and then moves through standard pipelines, transformation, and orchestration. In the middle, teams can explore, process, and warehouse data with data bricks intelligence and unity
[10:37] catalog providing governance, lineage, and control. On the right, curated outputs flow into dashboards, applications, and other tools that people already use. The key point is balance. Teams get flexibility to solve different audit problems while data quality, access, lineage, and
[10:53] auditability remain anchored in the platform. Hi, this is Greta. I am going to talk about the approach that we followed in this migration and the data bricks technologies that were used to deliver a govern platform that enables more AI
[11:10] innovation. Now that you've heard why KPMG UK needed to modernize and what the vision looks like, I want to take you inside the how. How do you actually migrate over 400 SQL scripts in 3 months without breaking
[11:27] audit grade quality? The answer honestly is that you let AI do what it's good at and you keep engineers doing what they're good at. Here's the principle that we operated on.
[11:43] AI did the heavy lifting, but engineers remained the decision makers. This wasn't a press a button and hope for the best migration. We designed a workflow where AI accelerates the mechanical work, the
[11:59] line by line SQL conversion, the syntax translation, the boiler plate, but every output went through human review. The result approximately 60% less refactoring effort compared to a
[12:16] traditional fully manual approach. So we took a twofphase approach for this migration. First we used lake bridge to scan the entire estate that gave us a datadriven view of
[12:34] complexity, effort and risk across every script. Instead of engineers manually triaging hundreds of scripts, we had a risk classified backlog within hours. That alone probably saved us three weeks of
[12:50] planning time. Phase two is where it gets interesting. That's where Genie code and data bricks hosted LLMs like cloud became part of this developers workflow. I will zoom into genie code as this was
[13:06] proved to be one of the biggest accelerators for this and the subsequent migrations of the DNA transformation. Genie code lives directly inside the databicks workspace in the notebook in the SQL editor right where engineers and
[13:22] analysts are already working for migration. It understands both sides of the translation. It reads TSQL idioms and proposes idiomatic data SQL equivalents that take advantage of Delta
[13:38] Lake and Sparks distributed execution. But it's more than just syntax translation. Genie code understands intents. The engineer's job shifts from writing
[13:55] the conversion to validating it, reviewing the proposed logic, checking edge cases, confirming the output matches, and because it is integrated into the workspace, you're testing against real data straight away.
[14:12] What's critical for audit work is that test coverage was maintained throughout. We weren't trading speed for correctness. Every converted script was humanly validated against original output before
[14:28] the signoff. Now, let me give you a real example. We had an 800 plus line TSQL store procedure. one of those monolithic scripts that does so many things nested cities branching logic etc. and nobody
[14:47] wants to touch the AI tools decomposing decomposed into modular testable steps. The end result wasn't just translated code. It was better code, more readable,
[15:02] more auditable, easier to maintain. And that's the most interesting part of an AI assisted migration. you can actually come out on the other side with much higher quality code than
[15:17] what you started with. Not just equivalent code on a different platform. Now let me quickly cover the platform that makes all of this possible and importantly makes it auditable.
[15:38] The tech stack to back up the audit platform was deliberately chosen on the basis that it actually solves a specific audit requirements. The core landing place for migrated workloads was data brick SQL. We also chose Delta Lake for storage,
[15:54] Unity catalog for governance, Delta sharing for secure data exchange with audited entities. On the migration side, Lakebridge and Genie code as I just described. Now, why data recycle? Well,
[16:10] well, audit workloads are incredibly spiked. January through April, teams are running intensive analytics across massive client data sets. The rest of the year, demand drop significantly.
[16:26] SQL serverless autoscales to meet that demand without operational overhead. So, KPMG only pays for what they use. No more provisioning for peak and wasting budget on idle clusters.
[16:42] But here's what makes this strategic, not just economical. The same serverless compute that runs our SQL analytics also powers AI workloads, feature engineering, model inference, vector search queries, it all runs on
[17:00] the same elastic infrastructure. So when KPMG team started deploying AIdriven audit solutions, they didn't need a separate AI platform with its own capacity planning. The foundation is
[17:16] already there. One platform, one compute model, one bill. That's what lets a team of this size move fast without building a second infrastructure stack for AI.
[17:36] And finally, governance in audit. This isn't optional. It's existential. Unity catalog gives us a single governance layer across all assets, tables, machine learning models, AI functions, vector searches, indexes, not just traditional data.
[17:52] So the same lineage that tells a regulator this number came from these source records through these transformations can also tell them this AI model was trained on these governed data sets accessed through these permissions and
[18:09] produced this particular output. That's the traceability chain that regulated industries need before they'll trust AI in production. Delta sharing gives us governed data exchange with audited entities.
[18:26] Every access logged, every permission is explicit and full lineage is preserved. And Ela Lake underneath gives us time travel and versioning. So we can always reproduce exactly what the data looked like when an AI model made a decision.
[18:45] Now why does any of this matter? Because you can have the best model in the world, but if it's drawing from ungoverned inconsistent data, it will produce ungoverned inconsistent answers.
[19:04] High quality auditable data as default inputs. That's what makes AI trustworthy in a profession where trust is everything. Thank you.
[19:22] So, Greta has explained how the technology was transformed, but as I'm sure you're all aware, uh it it's the people who are going to interact with that technology that you really need to bring on the journey with you. And uh KPMG UK did this migration last summer. Uh any of you from the UK might know uh
[19:40] the peak season for audit uh is September, October through to April. So they did a condensed migration over the summer, but they were going straight into that peak season of seasonality. Therefore, they had 200
[19:57] 200 200 auditors who needed um to be trained on this new technology stack. They'd been using the on-rem SQL uh for for many many years. Therefore, um this was as much about culture and training
[20:12] as it was about the technology uh migration itself. Therefore, over a three-month period, KPMG took the 200 users and they took them through a um tailored training program. This was both
[20:27] um aimed at data engineering on data bricks. So learning the best practices, learning how to do SQL within data bricks, also understanding uh Unity catalog and how data is governed within data bricks. That was also a shift in in
[20:45] thinking. Uh they also implemented new operational processes and agile ways of working. What this really was was how the organization did their standard work on this new platform. uh they had a whole pillar building that capability and making sure
[21:03] that uh that all of the team were working in exactly the same way. Um and this was done at the same time as the migration. So uh there was our rolling program of the training enablement and um it wasn't just our kind of standard
[21:20] data bricks training. It was bookended by uh audit on data bricks and quality within KPMG. What is expected? What is expected with how you use the technology for the
[21:35] outcomes? Uh can I have a raise of hands of people who have done data bricks training and know these pathways? Okay, a bit of a mix. Um so data brick uh so KPMG took uh their cohort of uh
[21:53] users for this and they did a learning needs assessment so they understood the kind of level of skills that uh team had how they were used to using the um previous technology stack what was transferable and where the gaps were.
[22:09] They did a learning needs assessment. It's a little bit like top trumps for any of you who played chop trumps. It's kind of got got me. So I've got these skills. These are my gaps. Therefore, we're going to focus on these areas. And over that three-month period, we did rolling instructor uh instructorled
[22:26] training. So uh all of the data bricks content is the same if you do self-paced or if you do um uh formal training with a trainer. However, it gives the ability to ask questions as you go and make sure that people uh actually got got to where
[22:44] they needed to be by completing that training. They also did a uh round of certifications. So once the uh engineers had done their training, they would get certified um and everybody knew they had
[22:59] the full level of skill sets. But not everybody needs the same skills. So those who were doing uh lighter weight engagement with data bricks, the data analysts had a different pathway where they just focused on uh SQL and
[23:15] Genie. Um and we didn't just do the training. We then had a series of office hours throughout the buildup to the users being live in in the in the platform. And we kept those office hours
[23:31] running for the first 3 months so that there was always somebody from data bricks and KPMG who could help the teams who were doing the live audits. Technology alone doesn't deliver transformation. People and capability
[23:47] building sat at the center of the program. This is a direct quote from the blog blog that KPMG have done. So you can read more around how they built that training muscle and skills skill sets in that blog.
[24:06] All right, just to go through some of the outcomes and the business value that we have seen and what we expect to achieve as we go forward here. Um so just sort of setting the stage and and just want to call out some key metrics I think that are important. um you know I mean thinking about you know
[24:22] 400 plus SQL scripts that got converted uh refactored at the same time I mean 60% reduction in refactoring time did all this in 3 months it's a pretty short window right when you think about uh everything that needed to be done we
[24:39] had 200 plus users on boarded I mean Ellie alluded to the whole program to get those folks up and and different skills um and the most important part we didn't cut corners right? We needed to make sure that this was strong, the quality was there, right, for the
[24:54] outcomes we needed or we do need for our audits. So that really important. Um, and so I mean just sort of going beyond the numbers and thinking about what's the outcome from a business perspective for our audits as a whole. Well, we now
[25:12] have the proper structures in place. We have the data that's properly governed through Unity, right? We've developed the right lineage. We built in the right context to allow us to execute at scale from an audit perspective, right? So,
[25:28] which is really important. Um, we've set up the platform to be able to scale to bring it up, bring it down. Like I said, going back to what I started with our 200 hour audits versus our multinationals with 150 components across the globe. The volume of
[25:44] transactions varies, but it doesn't matter. We've got the right foundation pieces in place. We've got our engineers that are focused on the audit logic. They're not sitting there trying to, you know, tweak knobs and fix infrastructure and all that good stuff. We have a solid platform. Um, and
[26:01] the the foundation flexes because we're in a regulated industry. Like I said, rules requirements change all the time, but we have the flexibility built in the platform to allow us to to, you know, make those changes as we need. And you know the last point on there is about an
[26:17] AI enabled platform and we've been able to do that to allow our auditors you know to leverage that data and that democratized data really important. So and and this is an example of a
[26:32] product that was built uh by the UK firm which is the financial report analyzer. Um, and what this product is, it looks at our clients financial statements, analyzes the disclosures in the statements and uh will tell you whether
[26:50] or not one all the disclosures are there, are the disclosures accurate, etc. Right? But case in point, you got to build the right foundations in order to deliver that type of product. And that was the first product uh that came out of there um leveraging Genie uh or
[27:07] sorry Genie and Genai within that that tool. So just some lessons learned um sort of three key points I want to talk about real quick. This was a transformation not a translation. We're just taking our
[27:23] SQL scripts and you know just converting them. Uh this was all about you know we talked about the different pieces setting up the foundation relooking at the pipelines refactoring you know the code to get it in the right spot setting
[27:38] up the right governance around this data right within unity um figuring out what information need to be in there. So structuring it right is really important. So you know flowing down to that second point um
[27:53] there was a lot of planning involved at the beginning because if you don't plan properly right you're going to fall into the trap and you know that 3mon window may turn into 18 months and you can't afford it right so do your proper planning understand the existing processes what needs to be reworked
[28:10] re-engineered understand the data coming in so you can get you know the right context the right governance structures in place etc. So properly plan I think is the message there and trust is earned through validation. That's a really really important part uh point especially in our industry and I'm sure
[28:25] for anyone else in here working with your organizations it's all about the quality of the data and how do we do that? We ran parallel runs we reconcile data make sure the data was complete right the integrity of the data is there. It's fundamental to our audits, right? It has to be right. There's no if
[28:42] ends or buts. We're built on trust, so that matters the most. So, a little bit on the path forward and I think I mean a lot of the stuff I talked about today and Ellie talked about and that I think we've heard a lot
[28:59] through the conference the last couple of days, right? Talking about you know governance around data, the context around data, etc. And that really sets you up right for success into the future. So just some sort of thoughts on on the value going forward. Um I mean I
[29:18] talked about the financial reporting analyzer Genie enabled uh product. We've set the foundation to allow for our people to leverage that Genie capability to be able to interact with the data. democratize the data which is absolutely
[29:35] critical uh as we move forward. Over to Ellie for the couple more. So last year at Summit uh one of my favorite announcements was Lake Flow Designer. And what I love about working
[29:50] with KPMG is they never stand still. You would think you're just coming off the back of a threemonth uh SQL migration. you might take a breather, but they came to the summit last year and they were inspired by LakeFlow Designer. Um, and
[30:06] as a result, we've now mobilized exactly the same premise. We've now got three months in order to migrate all of their altrix workloads into Blake flow designer. Uh, we've kicked that project off. Um, and exactly the same process we
[30:22] used with the SQL migration, we're doing this AI first. The amazing thing I I I love is um the speed at which AI enabled migration has come even in a year. So those mig the migration of the workflows
[30:38] from Ultrix it's pretty complex. It's not something that we had done before. It's not a um set of users that typically have come into data bricks. LakeFlow Designer allows those users to be in data bricks to do the drag and drop transformations and we're using
[30:56] Genie code in order to do the migration. Uh at the moment we're getting really good uh refactoring um success in accuracy for that migration. Again, KPMG are thinking about this with people in mind and we're doing training for that
[31:13] user group. Uh I think typically Ultrix users are very used to using it and um therefore it is a a cultural move to using late flow designer within data bricks.
[31:30] G mentioned the financial report analyzer and the UK team have set an ambition that they want to go much much faster and their road map of AI use cases and so uh they already prioritize those use cases and decide which use cases they they put their resources on
[31:47] in order to uh get them into production as quickly as possible. and financial report analyzer took some time to build and they set the ambition that the next use case should take half the time and each time reduce the time in order to
[32:03] get production grade uh solutions in place. they've now done a factory approach. So using full stack of AI within data bricks but also using tools like claw code in order to um uh
[32:22] in order to develop the AI use cases much much quicker and deliver that value back to their business. Um it's really exciting to see uh those move through that life cycle.
[32:40] Perfect. And just to sort of finish up uh and this is sort of one quote um that's up in the blog. Yep. That you know you can have a look at uh online and I think it's really important. So I I'm just going to read it quick. Most AI agents fail at data quality not reasoning. The lakehouse
[32:56] makes audit grade data the default input. Okay. So at the end of the day it's not a model and infrastructure setup problem. It's usually a data quality problem that you start with. You can tweak the models afterwards, but it's really having a look at the data
[33:12] and the quality that you've got in there. And that is it.
Learn more about the Databricks Data and AI platform.
The information provided herein is for general informational purposes only and may not reflect the most current product capabilities or configurations.