Data Products for AI: Building Trustworthy Agents with Knowledge Engineering
Summary
- KPMG demonstrates how data products have evolved from simple datasets and dashboards into managed business assets with defined SLAs, contracts, metadata, and governance designed for consumption by both humans and AI agents.
- Their five-layer AI data foundation pipeline—covering source data, product curation, knowledge engineering, semantic translation, and vector search context—provides the structured foundation that prevents agent hallucinations and model drift.
- The KPMG Nexa platform uses a data analyst persona to automate requirement gathering and generates ontologies and knowledge graphs so agents can understand table relationships, dependencies, and cross-domain data connections.
Data Products for AI: Building Trustworthy Agents with Knowledge Engineering

KPMG demonstrates how to shift from data-centric thinking to product-centric delivery that prepares enterprises for AI agents. Data products are no longer just datasets or dashboards, but managed business assets with defined SLAs, contracts, metadata, and governance built for consumption by humans and agents alike.
Explore the five-layer AI data foundation pipeline: source data, product curation, knowledge engineering, semantic translation, and context for vector search. Learn how KPMG's Nexa platform uses data analyst personas to automate requirement gathering, how ontologies and knowledge graphs enable agents to understand table relationships and dependencies, and how dynamic metadata prevents hallucinations and model drift. See live demonstrations of semantic layer functionality, cross-domain data unification, and how metadata context flows from business requirements through knowledge engineering to agent decision-making.
🤝
Chapters
00:00Introduction and Speakers from KPMG00:53What Are Data Products and Why They Matter02:14Unifying Data Estate: Architecture and Personas04:40Building Blocks and Fundamental Elements05:27Data Products as Foundation for AI-Ready Enterprise07:51KPMG Use Cases: Resource, Pricing, Information Exchange10:17Knowledge Engineering: Ontology, Graphs, Semantics11:52Agents as Consumers: The Metadata Challenge13:27AI Data Foundation Pipeline: Five-Layer Architecture14:31Demo: Nexa and Data Analyst Persona18:06Closing: Context Is King and Best Practices
FAQs
What are data products and why do they matter for AI agents?
Data products are managed business assets with defined SLAs, contracts, metadata, and governance—distinct from raw datasets or dashboards—designed for consumption by both humans and AI agents. This video argues that without well-defined data products as a foundation, AI agents lack the context they need to make reliable decisions, leading to hallucinations and model drift.
What is the five-layer AI data foundation pipeline that KPMG describes?
KPMG's pipeline consists of five layers: source data ingestion, product curation, knowledge engineering, semantic translation, and context for vector search. Each layer adds structure and meaning that agents can consume, with the knowledge engineering layer being particularly important for encoding ontologies and table relationships that agents need to navigate enterprise data.
How does knowledge engineering prevent AI hallucinations in enterprise agents?
By building ontologies and knowledge graphs that explicitly encode table relationships and dependencies, knowledge engineering gives agents a structured map of the data estate rather than requiring them to infer connections from raw schema metadata. This video explains that dynamic metadata generation keeps agents grounded in current business logic and prevents them from generating plausible-but-incorrect answers.
What is KPMG's Nexa platform and what problem does it solve?
Nexa is KPMG's internal AI data platform that uses a data analyst persona to automate requirement gathering and translate business questions into semantic queries. This video demonstrates how Nexa enables cross-domain data unification and provides a semantic layer through which agents and analysts can access consistent, governed data across different business domains.
Full transcript
[00:08] Good afternoon everybody. Uh I know it's pretty late on a Wednesday afternoon. Thank you everybody for uh coming. Uh um the topic we have today uh I know it's not a very like sexy topic. But, you know, I wish we had the
[00:23] phrase something about genies or contacts engineering or something like that that would have got it made a lot more catchier. Uh but uh let me quickly go ahead and uh introduce uh myself and my co-presenter here. Uh my name is uh Bindu Deodhar. Uh I'm
[00:38] part of uh KPMG US. I head the data engineering and analytics delivery excellence org. Uh it's a internal org inside of KPMG. We support all of the data and analytics uh delivery services.
[00:53] Uh so, with me I have uh Sameer Baghi, who's a senior data and AI engineer. Uh so, what we're going to do do today is um we will talk at a high level about, you know, what are data products. Uh it's
[01:09] it's make a very like strong comeback, I would say, uh in in the last uh few years, thanks to AI. Uh and uh we'll talk to you like few of our use cases on where we have deployed data products at scale inside of KPMG.
[01:24] And and how now how that like, you know, this discussion is elevating and shifting towards AI adoption, right? You know, the other thing we were discussing is some of our thought process in what we were doing when we put this deck together, it was really good to see Ali yesterday
[01:42] on the stage talk a lot about this, like the the foundation of building a data ecosystem, having your data products defined, having your contacts defined, building your business semantic, you know, all of that, right? So, uh it it's
[01:57] great like the you know, Databricks as a vendor is is really like listening to all their customers and and that's what here we are here to talk about. All right. So, I'm I'm sure like you have seen a variation of this what we kind of as the
[02:14] layered cake slide. This is kind of like the holy grail of any data and digital estate what we see during the age of AI. Right? Everybody has to like unify the data estate. I know like today data does
[02:29] live in like silos, but is there a way like where you can connect your data together so you could have like a unified view of your entire data estate. Right? And and like the the middle layer is where all of the magic is now
[02:45] starting to really happen when you see all like the these like top data and AI vendors, they are like really scrambling for that space to become the leaders in that, right? Like how do you build that curated semantic layer through which you can empower your
[03:02] AI agents, right? Like inside of KPMG like the the dialogue has shifted, right? Like we offer our data to like different personas. We we have our business users, what we refer to as like our functional users, right? We have
[03:18] a tax user and audit user and advisory user and and then we also have like I'm part of the KPMG US member firm, but we have other the global member firms, too. So, we are unifying the data and segregating
[03:34] the data where like we could now say that as a member firm user, you could come and access your data in a very safe and secure way, right? And and last but not the least, the agents are on, right? Like I'm sure you all have like the agent sprawl inside of your organization. So, now how do you enable
[03:52] that trusted data for your agents? Like so that's the number one question we're all trying to solve and we are like taking you know one small bite at a time. Next slide. All right. So this might sound a little professorial
[04:09] like you know the what exactly is a data product, right? I mean I'm sure like back in the 2015 and 16 there was a lot of talk about data products. I'm sure you heard of data mesh. So it like when you look at the Gartner charts, you know that there was
[04:24] like that hype and then it went through the trough of disillusionment and then people stopped talking about it. But it it is really back with like full force. So what is very very fundamental for anybody is to now start building that
[04:40] data product, right? And data product as it is right it's not just like a data set. It's not a dashboard. There's like some fundamental building blocks of what will it take to enable a data product, right? So I mean in terms of like the why data
[04:56] product you can see it there like you want to break down your silos, accelerate innovation for agentic adoptions, drive tangible business value, right? Like so if you want your agents to perform end-to-end tasks and workflows, so you you need to know like your data
[05:12] products what you're providing them is you know meeting the SLAs, has got the right contracts and like you know that data has like the right metadata and trust behind it, right? So and and like your like readiness to prepare for the
[05:27] AI driven future is really like dependent on defining your data products. Next slide. All right. So again elements of a data product I'm I'm not going to like drain this slide too much, but at a high level right
[05:42] everything starts with a business. Like what are your basic business definitions you're trying to get out there. What are your technical definitions technical definitions your what you want your data analyst your engineers to collect and and maintain, right? The ground up when you start ingesting data
[05:58] into your legs, how do you set that? So having your enterprise standards defined through your governance teams is is extremely critical and it's hyper important that you know, people follow those repeatable patterns so that what's coming out on
[06:14] the other side is is a tangible data product. So like I mean like any other system, right? You have the inputs and outputs coming out of each side you know, in terms of like defining the the SLAs your timings and like you know,
[06:30] the whole design patterns what you're using and and then finally, you know, like the content wise what's going there and what's coming out on the other end. Next slide. All right. So just a little bit of the hierarchies, right? And and
[06:45] typically being in a data team and data ecosystem, we are very familiar with the like from this left to right like diagrams, right? You you have your source systems, you have your systems of records and now you start like curating and building your data products. So I
[07:01] like you know, the farther left you go like your data and products what you're building is more aligned to the source aligned like data domains, right? like that's what we call. And and then you know, you know, you'll have to co-mingle some of the datas across these different
[07:19] data domains. So you might have to join the product to a customer, you know, to a partner so that like somebody can go and do all the 360° analysis across data, right? So and and then to your farthest right, you have those consumers, those agents what we were
[07:35] talking about which you need to do. So on this slide I'll touch a little bit about you know, like some of our use cases what we have gone through. Like, you know, I think we've tried to be bold of adopting this data product methodology. Uh we we had a few pilots we ran. We We
[07:51] have something called internally as an integrated value chain. Uh so, where we This is about running the business of KPMG, right? Like, what what how exactly business operates inside of KPMG. How do you like we build something called as a
[08:06] a resource management product, right? Like, I mean, KPMG is all about consulting and resources. So, you we built a product about like, you know, what does it take to build a resource product, which is what are the skills, you know, what what what is the level of different um
[08:21] uh you know, like demographics of them. All of that was bundled into that resource management product. Then, we built another product for pricing. Right? Like, so the uh what it goes into uh you know, like bill our customers, right? Bill our clients. So, that kind of became our pricing product. So,
[08:39] it was definitely a mindset shift, I would say. Uh when when we built these data products and put it out there, people were like, "Okay, how do I use this, right? What does this mean?" I mean, they they were kind of happy to see that, you know, we were taking the right steps and like the right amount of
[08:54] input from these business before we were building these data products, but there was a little bit of that adoption curve we had to go through, right? Do I access this through a Power BI? Or do I just do self-service and run some queries on it, right? And and then this also led to our
[09:10] enabling of some of the Genies which like Sameer is going to now talk about, right? And then there was another bold program which we launched. Uh we called it as information exchange, right? So, through that what we did was uh we went after I mean, the the kind of
[09:28] the tagline for information exchange was AI-ready data. So, what we wanted to do was bring like data from, you know, we were like actually parsing it out of our lot of our contracts and engagement documents. And the idea was to build
[09:44] like a permissibility engine on top of this data, harness the metadata out of it. And the idea was so that we have our teams on the fields, they they know like how could they cross-sell and upsell to like different clients, right? We we
[10:00] have a tax client, can we offer some you know advisory services to them? And vice versa. So, we started building those kind of data products. And with that, like what we found out very quickly is just the data product was not enough. Like there was something missing in
[10:17] terms of providing additional context and semantics in like defining the taxonomies and ontologies and everything like that. So, that's where the discussion changed to doing more with like knowledge engineering and things like that. So, with that I'm going to turn it over to Sameer
[10:34] to talk through some of this stuff. Yeah. So, thank you, Bindu. As Bindu mentioned, building data products has been something that the industry has been working towards and has been kind of structured and standardized. We have these established definitions of what a data
[10:49] product is, but what do we do with data products now? And to answer that question, one of the prime examples that we have been geared towards in KPMG is Genie. So, once these data products have been created, as Bindu mentioned in the previous hierarchy that okay, products, customers, all of these are {{}quote}
[11:05] general purpose data products. So, these general purpose data products can be plugged right into Genie and have you create a conversational chatbot. You provide it the relevant instructions. I'm sure we're all very familiar with this Genie process, but have that
[11:21] having that defined set of a data product, giving it the proper classification of what are the table descriptions, what are the column descriptions, all coming together to create that semantic layer on top of a data product enables this genie to perform and give you proper insights without hallucinations or without any
[11:37] kind of model drifts. So, I do want to take it to the next level though. But, until now data products we're using across our uh across our enterprise, but we have a new customer in our new consumer in mind.
[11:52] And the new consumer is going to be agents. The problem is agents don't understand data like humans do. We have to really gather it with the relevant information related to what the metadata is speaking towards. What does the tables that we uh create these data
[12:07] products for? What do the tables mean? What do their columns mean? So, that business metadata is something that we are trying to define. And as Ali also said in the keynote yesterday, speaks to the ontology of where um the
[12:23] data product sits. So, what are the needs of the agent? Coming back again, the metadata, the descriptions, and the relevant information related to the data product. So, knowledge engineering is the process of structuring the organization's
[12:38] collective knowledge so that AI systems can interpret it accurately and use it effectively. What does that mean? We have three layers of this knowledge engineering starting off with the ontology. This would be the business uh dictionary, key concepts, etc., etc. Coming down to
[12:54] the knowledge graphs. What are the relationships? What is this the lineage for uh the data products and uh a bit data products and scope. And finally, the semantic layer. How This is the transition translation engine that connects your technical data infrastructure
[13:10] to your business vocabulary. We defined an AI data foundation pipeline. And this is something that we've used to build our uh data products and the AI layer on top of the data products starting off with our source
[13:27] data sets, where this could be raw data coming from your source systems, internal data, whatever it may be, and then to the data product curation. So, as Bindu mentioned, curating the data product with the relevant uh key requirements for what it takes to
[13:43] curate these. Going on to the next phase, which would be the knowledge engineering phase, where we give it the pro- provided the proper taxonomy, the ontology, coming up with the relationships behind the scenes, as well as defining the semantic layer, all the way coming down
[13:58] to the context layer. Now, this is getting more into the vector search indexes and whatnot, providing the embeddings, the vectorizations, metadata lineage, APIs, retrieval logics, etc., etc., to providing this information for the agentic layer. So, the process of
[14:15] this pipeline, the ideal goal is to create an agentic AI-ready data that is easily understood by your agentic layer to help you build personas, agents, chatbots, etc., etc. So, I do want to talk through the last 5
[14:31] minutes with a quick demo that we've put together. So, in KPMG, we built this product called Nexa, and the pro- the the ideal purpose of this product is to come together orchestrated end-to-end
[14:49] delivery for data engineering. And I do want to focus on one of the personas here. And the persona is called the Oh, shoot.
[15:05] But, we have a data analyst persona. The data analyst persona, what it's meant to be doing is uh
[15:26] Yep. So, the data analyst persona what we were meant to be doing is when we gather business requirements in a data engineering cycle, the data analyst persona mimics what an analyst would do. Gather the requirements and turn them into technical requirements so the engineers can convert them into code. In our data analyst persona, we've built this agent behind the scenes that
[15:42] actually takes care of that process. But, what does that mean, right? When we have the source data coming in, the source data to be transformed and going into our target systems with the target streams require metadata context to be able to map out what the necessary joins
[15:59] are with other tables, what the necessary relationships are with the other tables. And through that process, we put together an ontology behind the scenes where we talk about the table descriptions and column descriptions. Here, as you can see, we've described what a
[16:14] table is, we've described the definitions of it, any core purposes, additional considerations, and what not. Through this Through this process, we're giving the agents enough information to understand what the tables they're working with, as well as what columns
[16:29] are involved, and we have the same process through our column descriptions. And see through this, the column definitions, the descriptions. This way, when we do vectorize this information and give it back to our agents, our agents are trained with the necessary
[16:47] uh ontology. Now that they have the understanding of the ontology, the next step in process would be the knowledge graphs. The knowledge graphs, given that we've given the primary keys, the descriptions, and how these tables correlate to each other, the knowledge graphs behind the scenes are able to map
[17:03] out the relationships between these tables. So, when the technical descriptions are given to the end user for saying, "This is how table A maps with table B. This is how you can join with each other." This process actually takes in the data product that this case would be the source data sets behind the
[17:18] scenes to be able to map out for our target systems. And here, we have this thing called the relevant metadata that now takes a business requirement from our analyst, which would be talking about um
[17:34] which would be talking about how the business wants to two tables to be translating to each other or the transformation between each other. It takes the business description and gives you the source tables, the source columns required based on the knowledge engineering we do behind the
[17:50] scenes. So, this is a live process of KPMG. Bindu, if you want to Yeah. Thank you, Sameer. So, um I know it we probably have like 2 more minutes left. I want to go ahead and close. Um so, from from a KPMG standpoint, you know, I
[18:06] I think we have like some really good white papers on, you know, how do the whole concept of knowledge engineering. I know it's a still nascent field and it's evolving, I would say. Uh but if if you do like stop by at our KPMG booth, our advisory team can give
[18:21] you like further guidance on it. But like I said at the beginning, it's good to see like vendors like Databricks are starting to invest a lot more time, focus, and energy on building that ontology layer, right? And and what exactly constitutes your context? Uh you
[18:37] know, like the way you say context is the king, right? So, how do you build it? How do you kind of put the right guardrails around it where that, you know, you don't get into like your token maxing or get like the wrong results. So, I I feel like this is a area where things are going to evolve quite a bit,
[18:54] right? And I also wanted to like take a moment and thank um Hexaware, who's like one of our uh premier partners, uh who we have worked with to run and build a lot of these data products behind the scenes, right? Like so, this is uh I mean, it's it's I I do want to close by
[19:10] saying that it it is hard work. You you have to do the grunt work of uh you know sometimes we we get extremely you know focused and say that I'm going to just ingest the data. I'm going to like put it out there and generate a report. But make sure you take the time to you know
[19:26] whatever is your way of you know like capturing metadata, your whatever templates you're using, right? Like do that additional work of gathering the technical metadata, the business metadata. Make sure that is all getting logged into Unity Catalog. And
[19:42] make sure you like define your like SLAs and contracts when when you're defining those those data products, right? And you will see the the kind of the the shift in the mindset with all of your users too that the trust level and the confidence will will start seeing to like ramp up, I
[19:59] would say. So with that, I know we are out of time. Thank you so much for coming and appreciate your time.
Learn more about the Databricks Data and AI platform.
The information provided herein is for general informational purposes only and may not reflect the most current product capabilities or configurations.