Automating Data Engineering with Agentic AI and Databricks Agent Bricks
Summary
- KPMG built Nexa, an end-to-end agentic AI platform powered by Databricks Agent Bricks that orchestrates the entire data software development lifecycle, eliminating the context loss and rework caused by multiple handoffs between business analysts, data engineers, quality teams, and product owners.
- Persona-driven agents act as information architects, data engineers, and QA validators with humans kept in the loop at critical decision points, achieving 50–60% time reduction in the requirements and data development phases in production at KPMG.
- Unity Catalog provides version control and time-travel audit trails for every data product generated by Nexa's agents, enabling self-service data product delivery with full lineage and enterprise governance on the Databricks Data and AI platform.
Automating Data Engineering with Agentic AI and Databricks Agent Bricks

Data engineering lifecycles remain fragmented across requirements gathering, development, testing, and deployment. Multiple handoffs create context loss, rework, and delays. KPMG solved this by building Nexa, an end-to-end agentic AI platform powered by Databricks Agent Bricks that orchestrates the entire data SDLC as a unified system.
Discover how persona-driven agents handle requirements interpretation, code generation, quality management, and business validation with humans in the loop. Learn how Nexa achieves 50-60% time reduction in requirements and data development phases, maintains version control and audit trails in Unity Catalog, and enables self-service data product delivery. See the live system handling production data migrations and enterprise transformations.
🤝
Chapters
00:00Introduction: Elevating from data platform to agentic AI01:44Problem: Fragmented data SDLC and context loss04:09Nexa solution: Four core capabilities and architecture06:33Persona-driven orchestration: Information architects to QA08:29Building with Databricks: Agent Bricks and Unity Catalog11:07Unified platform: Version control and time travel in Databricks13:35Live demo: Nexa UI and end-to-end workflows16:58Business validation and automated quality enforcement18:50Roadmap and enterprise transformation strategy
FAQs
What is KPMG's Nexa platform and what problem does it solve?
Nexa is an end-to-end agentic AI platform built by KPMG on Databricks Agent Bricks that orchestrates the entire data engineering lifecycle from requirements gathering through testing and deployment as a single unified system. It addresses the context loss and rework that accumulates when requirements pass through multiple human handoffs across business analysts, data engineers, quality teams, and product owners.
How do persona-driven agents work in the Nexa platform?
Nexa uses specialized agents modeled after human roles—information architects, data engineers, and QA managers—where each agent handles a defined phase of the data SDLC with specific inputs and outputs. Humans remain in the loop for business validation and approval steps so the system augments rather than fully replaces human judgment, particularly for decisions with downstream business consequences.
What time savings does Nexa achieve for data engineering teams at KPMG?
The Nexa platform achieves 50–60% time reduction in the requirements gathering and data development phases by automating repetitive interpretation, code generation, and quality validation tasks that previously required manual effort at each handoff. This reduction was measured in production at KPMG, where the platform is actively running enterprise data migrations and transformations.
How does Unity Catalog support agentic data engineering in the Nexa platform?
Unity Catalog provides the version control, time-travel, and audit trail capabilities that give Nexa's agentic workflows full governance and reproducibility across every data product they create or modify. This means enterprise teams can trace the lineage of any output, roll back to a previous state, and demonstrate compliance with audit requirements even when the data transformation was performed by an AI agent rather than a human engineer.
Full transcript
[00:08] Good morning everybody. Uh very happy to be here. Uh we are here today uh just to set a little bit of the background. We were here exactly a year ago. Uh we were talking about our front office transformation journey. With that, we rolled out our global front office
[00:25] global data store platform where we talked about bringing all of the data into a single data hub. And here we are exactly a year later. We were like, "Okay, we need to now elevate our game." With all of the new and cool stuff
[00:41] happening with AI around us, it's now time to build an end-to-end agentic AI platform. So, before I jump in, I'll do like a quick intro. So, I'm going to I'm Bindu Beerur. I head the data engineering and analytics delivery excellence
[00:57] organization inside of KPMG uh what is called as Digital Nexus. Uh so, with me I have uh Sameer Baggi who's our AI engineer. Uh and we have Ganga Demata who's our uh product uh owner, I would
[01:12] say, for this product. And uh we are like really excited to present this to a story here. I know 20 minutes is not going to do enough of justice to us, but we'll do our best. And we'll save the best for the last, which is going to be a sneak peek and demo of this product,
[01:28] which actually is in production running today. So, I'm I'm sure the like, you know, you've all heard about, like, you know, AI and agents and what we are doing out there in in the real world, but one thing is very, very clear, right? Our
[01:44] data engineering life cycles on how you do like your data SDLC is still fragmented. There are multiple hand-offs which happens bit across like different human personas, right from collecting your business requirements, your data
[02:00] mappings, handoff to like data engineers, then to our quality teams, then to like the product teams. Like there is like so much of context which gets lost and which leads to a lot of like, you know, rework and you know,
[02:16] having to go back and do things again and missing something entirely. So, that's when, you know, though we had like engineers using AI co-pilots like to do development much faster, but we were not able to get like that full end-to-end experience.
[02:32] So, that's when we thought about, okay, you know, necessity is the mother of invention, why don't we use a AI to like orchestrate our end-to-end landscape of what we do with our data SDLC, which now is called as AI LD, you know, LDC,
[02:49] right? So, that that's what is is really the genesis of this product, what we call as Nexa. So, I'm sure you've all seen like this left-to-right picture when when you do a data delivery transformation or a program, right? Like
[03:05] there is the manual process of getting requirements from your business leads, your business owners, handoff to our, you know, our analyst and data analyst teams who, you know, collect this in Excel spreadsheets and documents, which is like all over the place, and then it
[03:22] gets handed over to the, you know, to the actual technical teams to say that, okay, now I need to look at this, write my notebooks, take it through your medallion architecture, and then finally when you go into like UAT, you realize like you missed a big piece of the
[03:37] requirement, right? So, this was like this fragmented kind of the life cycle and chain had like so many like, you know, steps in there which needed to be looked at end-to-end. So, that's when, next slide
[03:53] the the team here, right? Like they said, let's become like these forward deployed engineers and look at building a product which gives you like a single eco-ecosystem of doing everything what happens in a data life cycle.
[04:09] So, what exactly is this Nexa, right? So, the the kind of the four big facets of it is doing your intelligent requirements interpretation and gathering, uh autonomous engineering acceleration, which is the code generation. I'm sure
[04:25] you've all used it for coding, you know, wipe coding and coding applications where like there's so much of strides made. Uh the persona driven driven orchestration. I mean, like we all know the human in the loop is still extremely important in what we do. And last but
[04:41] not the least, right? Like this continuous quality and generation, right? Like you you you cannot like just trust and let the agents do what they do. It's like how do you measure your quality, how do you preemptively catch like errors in in the data pipelines, and and like, you know, report it and
[04:57] correct it, right? So, that at end end-to-end is what Nexa as a platform is all about. So, with that, I'm going to turn it over to Ganga to talk through what exactly is the solution and how it works. Hi. Uh this is Ganga Dimeta. Uh I I am the
[05:14] product owner for the Nexa. So, with what we thought about is like if we have the requirements somewhere sitting in the SharePoint versus the code sitting in the Databricks or any other database, it is very hard to analyze whether the
[05:29] requirements are met or not. So, you write the code once and you deploy it and it is running in production, but we don't know if the requirements are being met every day or not. So, the first thing what we did is we take the requirements and digitize them
[05:45] in a format which sits beside the data. The business requirements are sitting beside the data now. And with that, you get a wonderful ability later on. I'll explain how we do that, but keep that in mind that we we will enable
[06:01] the magic to happen later on, right? So, now once the mappings are the once the business requirements are digitized, the next AI will take it from there. It It categorizes personas, right? Like the data analysis persona, the engineer
[06:16] persona, and the QA persona, and moves ahead from there. So, this is the new model of how we are working today with more generic components everywhere.
[06:33] So, here are the basic personas that we have built on. So, the information architecture persona. So, when a requirement is given to us, the first thing we do is the information architects come into picture. They decide what needs to be done, where the data should sit, uh logically separations, all that. The The
[06:49] information architecture persona will take care of that. Once the output of that persona comes out, the the respective human in the loop approves the output of it or reiterates it again. And that output is given to the data modeling persona, which creates a data
[07:04] model out of the requirements. So, and again, that output is given to the data analyst persona who works on the business requirements, data profiling, and everything. And from there, the output is given to the data engineering persona, which takes that as an input
[07:21] and writes the code required, and puts it into the gate wherever we want. And then, the QA persona will take it from there, and it does the a bunch of QA tests on the tables. And the lead persona is the one who oversees all of these and like like an
[07:38] admin persona. Uh uh go back. Also here there is one point here, right? Like we we write the code, we put it into the production, it is running every day, we don't know whether it is it is confining to the
[07:56] requirement or not. Now that requirement sitting in the data bricks beside the data, we can run those validations after every ETL load. And say whether the data is confined to the requirements or not. If it is not confined, it can flag. It can flag and
[08:11] say okay, today's load some of the data is not confined to the requirement. And any anyone can query the requirements and the data together. And with that I'll hand over to my colleague Sameer here. Yeah. Thank you, Ganga. So
[08:29] Nexa was built by KPMG but powered by Databricks. As you guys can see, we've built so many different persona agents that automate a bunch of tasks, but how we built the engineering behind the scenes, we're using a lot of the Databricks inbuilt services, starting with Unity
[08:45] Catalog. Everything is our back enforced. So these personas, how we're defining um these personas are portal based. So when users log into Nexa, they're automatically put into security group based on what persona they're defined.
[09:00] They have access to execute these tasks. So these agents are running based on user executions. So if an engineer had to generate code, then they click the generate code, the AI does the work in the back end, and then the human in the loop comes back into the picture with
[09:16] the actual engineers going and making changes to the code so they can review it and they can approve it. We have our table and column descriptions, which are which essentially become the ontology in the semantic registry of our platform. This is the biggest input that we get from
[09:32] our users is once we do have these source to target mapping documents. Right? We also get the business descriptions and the column and the table descriptions so that the AI understands behind the scenes what data we're actually working
[09:47] with, what are the lineages, what are the relationships, and that's once all of that is defined in our data analyst persona, then we go into the engineering portion where we have our context retrieval layer, which uses vector search indexes to actually uh vectorize
[10:03] the datasets that we have our mapping document and requirements, and then we use our relevant um metadata discovery through our internal rag that we've built. And then we also have our persona execution layer, which uses the uh knowledge assistance from Agent Breaks.
[10:18] We also have our MCPs once our code has been generated for execution on the Databricks platform. We take the code back, the MCP has uh executes it, sees if there's any error messages, any um flags that the code has generated,
[10:35] brings it back to our Nexa product, and then we have this feedback loop system that goes back and forth between the user, the AI, Databricks, and uh so on and so forth. As Ganga mentioned, everything is sitting inside Databricks right now. Before, we had our requirements coming in as Excel
[10:50] documents, we didn't have any control over them, changes happen all the time, and since we work in this delivery life cycle that's been established in the market for decades now, any changes in the requirements means you go all the way back to the very beginning, make changes there again, and it's just a very fragmented process. With Nexa,
[11:07] everything is sitting inside Databricks. From the data from the requirements to the code to your unit tests, everything is there. So, any changes, we have version history enabled, so we'll be able to time travel back, see who made the changes, what the changes were made, and now going between the processes uh
[11:25] DA requirement changes to the engineering code changes to the QA finding some kind of of kind of issues is all automated, and it's just a click of a button. Before this process used to take us end to end from requirements to code to production several months. Now,
[11:41] this is happening in the span of seconds and with human in the loop approvals, everything uh orchestrated end to end. The final point that we want to make is that ETL has been standard, right? We build pipelines, we
[11:57] do code, put into production. With Nexa uh our goal is to automate ETL with human in the loop. So, we're not taking humans away, we're empowering the humans to come in with AI kind of uh taking on majority of the tasks, but the T, the transformation,
[12:14] writing the logic, the code is all Nexa driven. Internally in KPMG we have our internal data movement frameworks that take care of the extraction and the load. But uh as you can see, this is the impact that we've seen with Nexa so far. Um these are conservative numbers. We've it's in production and we're seeing
[12:30] about the requirements gathering phase because the metadata retrieval to convert business layman language into technical pseudo code that engineers understand is about 50% of the work that Nexa's taking on. The data development portion of the side where Nexa actually writes the
[12:47] code, we've seen it being 60% of the time with 40% of the human in the loop standard and QA is about 40% and when we go into production and have these tasks repeat for day zero processes, day one processes, seen Nexa's taking about 40%
[13:02] of that effort. So, overall on our first iteration of this being in production, these are the numbers, but these agents are meant to get better and once we start seeing observable patterns that we can start integrating back into our Nexa product, then we assume these numbers are going to shoot up as uh
[13:19] as they were meant to be. So, I'm going to quickly give you guys a glimpse of what Nexa looks like. I know we're running a little low on time, but uh uh
[13:35] So, this is our Nexa product, right? It's driven for user interface. So, we want users to feel comfortable with this UI. It's easy to explain, easy to flow, and we've defined the features here. So, Nexa is able to help you understand your data
[13:51] first. And then we want to generate precision sequel. So, this is what the AI is doing here as well. And then we want you to validate your data with confidence.
[14:06] And here we have our personas. So, every role one studio. We've defined We crafted a studio that allows all of our users to come in, but as I said, there's our back and forth. So, people are put in different security groups. When they log in, they come into Nexa, and they know exactly what they have access to execute, what they have access to approve, what they have access to um
[14:24] review, and also edit. So, starting off with the data analyst persona, then goes into the engineering, QA, tech lead, modeler, architect. I'm going to go into the studio now.
[14:47] So, this is what our studio looks like. As you can see, this is a project room. So, based on the project that you're working on, you're assigned to, I can go click on this customer care 360, and you'll see the interfaces. In KPMG, we define these as interfaces. It's nothing but a source system to a target system, and the tables associated with
[15:04] that transformation process. And if I click into one of them, here we'll be able to see the mapping documents, right? So, to for reference, this is what the business requirements look like. For every downstream profile record,
[15:21] blah blah blah blah blah. As an engineer, when I look at this, I don't understand what it's talking about. So, we go down to this generate relevant metadata after giving the proper table descriptions and column descriptions of what these uh what the ontology is, what the
[15:37] semantic registry is. Then we click on metadata, and metadata actually takes that business requirement and converts that into technical pseudo code. So, you can see here um Once this technical pseudo code is done,
[15:52] then we go ahead and take that pseudo code as an input to our engineering agents, and then they convert into SQL. And then the data quality validation happens. You'll be able to see if they passed or not, and I'll give you a better example here.
[16:25] Of all the columns, we run automated uh unit tests on them to see if there's any flags. If there's any flags, for example, this is 100% null. So, something to the engineers to look at. They have this portal to see uh and then once the QA raises as an issue or a bug, then we go back and make the necessary changes as required. And then um
[16:42] this is how the table SQL is going to be looking like as well. And as you can see, there's several iterations on this version, which means that every time an issue was raised, the engineers went back, regenerated the code with the noticed findings and whatnot. Ganga, do you want to speak on
[16:58] the business validation side of things? This is the magic we are talking about, the business validation. Now we know the requirements sitting beside the data, and we can run this business validation anytime or even after each ETL load. And you can capture the issues even before
[17:15] the client captures it. If there are any flags that runs and shows up after every ETL load, we can we can stop sending the data to the client. Say that okay, this is a high severe observation, we can stop it here because the data doesn't fit to the business requirement.
[17:31] So, that is the business validation. So, yeah, maybe I I know we are like running a little bit short on time. So, just wanted to provide like some factual figures, right? So, this actually is now getting used in a live system. We are
[17:46] implementing onboarding a new member firm onto our Salesforce platform. So, that is going through an end-to-end data migration. So, that entire migration experience is now sitting inside of Nexa. And as like, you know, you saw
[18:01] Sameer and Ganga walking through those different personas in like the previous world, so there were tools like Alteryx and like, you know, your Microsoft Excel and like so people were doing things outside the the realm of, you know, what they typically operate. And having to
[18:17] work with each of those teams and getting like answers was so difficult. Now, Nexa is forcing that experience almost like that unified experience where you're working on a single platform and behind the scenes with like what we have heard this week, like all of this is getting collected into the
[18:34] Unity catalog metadata, right? So, your your corpus is getting like lot more richer, the context is getting richer, and in in the process, right? You know, I think our next part of the road map is to roll it out to other end-to-end like massive transformation programs, right?
[18:50] It could be some kind of a data conversion or a data migration effort, or you just want to hydrate a brand new lake. So, Nexa is going to become the new way of how we do our SDLC process, right? So, I think that is a change in behavior. So, we I mean, this team has
[19:06] done the hard work of, uh, you know, bringing in some of our pilot users across like the different personas, taking them, onboarding them, taking them through like some training videos, right? But like the experience and the feedback what we're generating, so the next version of this product is
[19:23] going to be like, you know, like almost do a mass rollout to like across our engineering community, right? So I we're really excited about this. I know we're running short of time. Uh if there are any questions, we can meet you all at the KPMG booth and and answer them. But if not, thank you so
[19:39] much for coming. Thank you all.
Learn more about the Databricks Data and AI platform.
The information provided herein is for general informational purposes only and may not reflect the most current product capabilities or configurations.