Enterprise AI in Manufacturing: Building Agentic Systems and Knowledge Platforms with Databricks
Summary
- Cummins built an agentic supply chain system on the Databricks Data and AI platform that automates demand planning and supplier communication across 135 global manufacturing facilities, using MLflow observability and human-in-the-loop approval to ensure transparency and trust.
- Corning deployed an enterprise RAG solution for manufacturing documentation in PTC Windchill, processing 90,000 documents for 2,000 users through Delta Lake ingestion pipelines, context engineering, and a multi-tool agent design called Windchill GPT.
- Both cases demonstrate that production-grade agentic AI in manufacturing requires governance, cost management, and transparency by design — not as afterthoughts — including patterns for managing hallucination risk and controlling inference costs at scale.
Enterprise AI in Manufacturing: Building Agentic Systems and Knowledge Platforms with Databricks

Manufacturing enterprises face two critical challenges: coordinating complex global supply chains with thousands of active purchase orders and enabling rapid access to mission-critical knowledge across distributed systems. this video features industry leaders from Cummins and Corning discussing how they deployed agentic AI systems on Databricks to solve these problems at scale.
You'll learn how Cummins built an agentic supply chain system using MLflow for observability and human-in-the-loop approval, handling demand planning and supplier communication. Corning shares their enterprise RAG solution for manufacturing documentation in PTC Windchill, covering context engineering, metadata governance, Delta Lake pipelines, and multi-agent deployment patterns. Both cases demonstrate production-grade AI systems managing governance, cost, and transparency.
🤝
Chapters
00:00Precision and Progress: Manufacturing Excellence with AI03:09Supply Chain Optimization: The Planning Problem at Scale06:01Agentic Solution Architecture with MLflow Observability08:02Design Principles for Production Agentic Systems10:58Building Transparency and Trust in AI Systems12:16Demo: Supplier Communication and Human Review15:04Business Impact: From Hours to Minutes per PO18:25Key Takeaways and Lessons Learned20:32Corning's Manufacturing Knowledge Challenge23:00Context Engineering and Private Data Grounding26:04Enterprise AI System Architecture at Corning30:43Document Intelligence Pipeline with Delta Lake34:07Cost Optimization and Deployment Patterns37:59Windchill GPT Agent with Multi-Tool Design39:53Scaling to 90,000 Documents and 2,000 Users
FAQs
What agentic supply chain system did Cummins build on Databricks?
Cummins built an agentic AI system that processes planning signals and executes supply chain decisions across its 135 global manufacturing facilities and more than 70,000 global employees. The system uses MLflow for observability and incorporates human-in-the-loop approval steps, handling demand planning and supplier communication while maintaining full transparency into agent actions.
What business impact did Cummins's agentic system deliver?
This video states that the Cummins agentic system reduced processing time from hours to minutes per purchase order, a significant improvement given the scale of their global supply chain operations. Reduced manual effort in supplier communication and faster decision-making were highlighted as key outcomes alongside the observability and governance benefits.
How did Corning build its enterprise RAG solution for manufacturing documentation?
Corning built an enterprise RAG solution called Windchill GPT for accessing manufacturing documentation stored in PTC Windchill, with a Delta Lake-backed document intelligence pipeline ingesting 90,000 documents. The system uses context engineering, metadata governance, and a multi-tool agent design to serve 2,000 users with accurate, grounded answers to manufacturing questions.
What design principles did Cummins apply to its production agentic system?
Cummins designed its agentic system around transparency and trust, ensuring that every agent action is observable through MLflow tracing and that critical decisions require human review before execution. The team emphasized building governance, cost management, and explainability into the architecture from the start rather than adding them after deployment.
Full transcript
[00:07] Okay, welcome. Uh we just had a question uh that there was a little bit of confusion. Precision and Progress is the name of this session. It's a combined session today with uh speakers from Cummins and Corning. So, um I'm from Cummins and we're going to start uh discussing our solution first and then Corning's will Corning will go after us.
[00:25] Um so, uh welcome here. Um we asked for the hottest room Databricks had and I think they came through and delivered on that. So, hopefully everybody can stay awake after lunch. Um we're we're here to talk about from planning signals to executed decisions today uh building a supply chain agentic system.
[00:42] Uh my name's Tom Anderson. I'm the uh leader of our AI and analytics foundational systems area that we call the digital core. Um here with a uh a few colleagues from our team, Srini Gopinidi, who's our AI and analytics solutions leader, and Prashant Kolur, who's our AI solution architect on this
[00:58] project. So, we got a pretty packed agenda. Uh we got about 20 minutes to discuss our ours and we're going to uh I'm going to hopefully set the scope um of the the business problem, turn it over to Srini to talk a little bit about the workflow and how we went about solving it. Uh and
[01:15] then Prashant is going to step in and talk about architecture, show a quick demo, and we'll end with the business value and lessons learned. So, just a little bit about Cummins. Um we are 107-year-old company. We're headquartered in Columbus, Indiana. Uh the company was founded uh by uh
[01:32] Clessie Cummins in 1919, who built uh some diesel engines for marine and on-highway usage. That's still a very big part of our business today. And so, if you've ever bought anything online or gone to the grocery store for for groceries, pretty good chance that Cummins was part of the supply chain
[01:47] that that made that possible for you. Um over the years, we've evolved into power generation business as well, powering some of the most critical uh of and infrastructure in the world. Large distribution centers, stadiums, airports, hospitals, just to just to
[02:02] name a few. And we do all of that by concentrating on the lowest emissions possible so that we can achieve our destination zero goals. So just to set the the scope within our supply chain business, we have around 70,000 global employees, but over half
[02:20] of those are part of our supply chain business. So that can tell you kind of how large the problem starts to be. We have 135 manufacturing facilities around the world with an additional 57 logistics and distribution centers. So approaching 200 locations, we build 1.3
[02:35] million engines per year. And then we have a large global network of suppliers that's approaching 40,000. So as you start to see, you know, 40,000 suppliers shipping millions of parts to hundreds of locations in almost 200 countries around the world every day,
[02:53] there's a there's a pretty large area that we can go after for optimization. So I'll turn it over to Srini now and he can tell you a little more about the business problem. Yeah. Thanks. Thanks, Tom. Hopefully can you guys can you hear me okay? So So
[03:09] um this particular case is about from a supply chain side the ideal state what we define in the supply chain is about having the right part at the right time at the right place. Right? So we built an agentic workflow which helps us to achieve that particular ideal to an extent.
[03:25] So if you look at the as Tom mentioned, we make we build engines. Um we buy thousands of parts from thousands of suppliers. So at any point in time we have got a lot of active purchase orders on many of them. And then when the ERP runs uh after the after initial purchase orders,
[03:41] it will recommend changes about change the dates. It will say some some purchase order it will say to bring it early, some it will say to bring it late, some it will recommend to cancellation. So that's a typical life of any planner. what happens is that
[03:56] What happens is that like when it recommends changes, uh typically in any manufacturing industry, a planner has to review what the recommendation is, evaluate whether it is worth because sometimes the recommendation is just moving for a couple of days. And uh and then look at all the historical supplier position,
[04:11] inventories, and other aspects of it and decide whether he should take action on it. That's the first part of the puzzle. Then he should communicate in an email with the supplier saying that he is the move or the recommendation action is okay. Then when the supplier responds, then they'll take an action. This is kind of
[04:26] a day in anybody in the supply chain, anybody has seen any manufacturing entity. Like this is a part of any planners on any of the manufacturing plants role. This typically involves, as you can see, a lot of context switching at it it be estimated takes around an hour and two for each purchase order. In simple and an action level for one
[04:43] purchase order, it looks simple. But the challenge goes like in one plant itself, we have around more than 100,000 active purchase orders because of the thousands of parts and then each one we purchase it by different dates. Right? So at this magnitude and the given the given the
[04:58] number of plants what we said, this is a big problem for us in terms of the time and then the we planners don't have the capacity to manage all of them actively. They prioritize and do their best at what they can. And it has other implications in the supply chain which we'll talk about. So that's the problem statement
[05:13] about so many purchase orders having planners having manually doing it and prioritizing it. So we thought Currently like until the agentic workflow, we have other things like so obviously have dashboards before. And uh it helps us to prioritize which one to purchase order to act upon.
[05:30] That's the first part of it. We even looked at the LLMs to see whether they can add some more intelligence to it. But the extent they can help us is to give us some kind of prioritization. Therefore, the planners can look at that particular purchase order which can communicate on one. But it never actually helps them in an in an entirety.
[05:45] Where the agentic workflow what you can see I will show you a demo as well. It takes the entire work state because we can carry the state forward in this case and carry the context forward with with the kind of an architecture what we've used. Right, so at a high level, this is what we have
[06:01] done. What the agent will do is that it will try as all the information and follow the exact business rules what we can just currently follow. It will look at all the recommendations from Oracle and then look at all the apply all the business rules about which one gets filtered and communicated with the
[06:17] supplier. And it even sends an email to the supplier about that this particular purchase order we want to change by a date, for example, or cancellation. And then when the supplier responds, it he can respond in a simple English and they it'll understand it. And it'll package all that information and bring it to the
[06:32] planner to say, this was the context in which we sent a communication to the supplier, and then the supplier responded to this one, and then this is the context of it. Then the planner, what the planner will do is just look at all that context of it in that in that application and then just approve it to say, yeah, it did the it did the
[06:48] interpretation correctly of the email and then he approves it and then that action gets executed in our ERP. That's the entire workflow. We'll show it the architecture as well as the demo of that particular one today and then we'll wrap it up. Yeah. Prashant.
[07:14] Thanks, Shini, for helping me with the business problem. Now, well, I'll talk about how we have implemented this agentic system like technically, architecturally. So, this is like the complete architecture we have like two agents, one data table for the memory tracing and the control tower for our
[07:29] UI. So, there are two agents. The outbound agent is responsible for understanding the demand and it picks the right POs and sends it to the supplier and updates the database. In the inbound agents, once the supplier responds, it will look at what supplier
[07:46] had respond, it interprets, understands it, and make a decision and updates the database. So, every action is done by the inbound and outbound agents with a human in the loop. So, the there might be changes where how the integration works with the
[08:02] agents. There is a delta table which has the memory which shares it, and the control tower is the UI. I want to talk about the four basic principles that we have like we have used to build this agentic system. The first one is like separating reasoning
[08:18] from execution. So, in lot of agentic systems like we throw all the context and let LLM do the work. But, in our case, we have kept the reasoning and the execution separate. We haven't given the LLM to execute any queries on the database or create any use any database
[08:35] operations. The second one is like one base framework and many agents. So, we have designed framework in such a way that like if you want to add one more agent tomorrow, you just add one more prompt and add tools to it. You need not to design the complete architecture completely. The third one is like graph
[08:51] state. Every decision at every tool is not done randomly. Say for example, we have five tools, the decision of the fifth tool is done based on the previous tool decisions. It's not taken randomly. The final one
[09:07] the best one is like observability. So, we are using MLflow to maintain the observability like every tool input, every tool output is traced using MLflow. Now, I'll talk about how we have scaled this system to a very enterprise level.
[09:24] I think these are the good practices like we have followed which will help us to scale the system to implement it for multiple plans. The first one is like state graph routes dynamically. So, think of like a graph like we don't want to build an if else conditions or like a
[09:39] state fixed graph. We want to build a architecture in such a way that graph is generated at run time based on the decisions or based on the tool actions. It's not like something a fixed one. So, if supplier replied it, you have to interpret it. If it is confidence low,
[09:55] you have to escalate it. So, it have to take the decision based on the state what it has. The second one is like the fallback rate. In production systems, there might be scenarios where the LLM might be silently failing it, but you don't know whether it's failing
[10:10] it. So, you we have maintained a fallback rate, like if that fallback if there is a failing, then you maintain a threshold. If that failing reaches that threshold, then you go and update your prompt. You know when to update your prompt instead of guessing that, "Okay, my LLM is not doing good."
[10:26] and just chase it. And then the last one is like recover silently, not loudly. So, we have seen some scenarios where like you want the agent wants to access a database and due to some lock issues, it is able to face it. It should not fail throwing an error at
[10:42] the planner's face. It should retry. We have to implement a retry mechanism for some time and then we have to make it work and reach the final goal. The next one is like how we build transparency. The this is one of the interesting things and important things
[10:58] we need to build. Like lot of businesses are very afraid because the doesn't understand like what happens it. So, this is These are the three principles that we have used to bring the trust and transparency into the systems. The first one is like continuous evaluation framework. So, we have implemented
[11:14] evaluation evaluators on the agentic system while evaluating it and even during running, so that it keeps track of whether it's doing good, whether it's doing bad, if it's not doing it, what is it not doing good, like we'll be able to interpret it. The second thing is like
[11:29] the hallucination control verifier. So, after you run the agent, there would be an output trace for that complete agent. We just want to verify it with an verify agent whether it has met that goal or not. In that way, we know that whether the agent is doing its job or not. The
[11:44] final is like the closed-loop feedback. This is one of my favorites. Lot of times when we built agentic systems, it's that like businesses give a feedback to it and we leave it on the table. Okay, let's implement later in like a ad hoc functions. But, we have implemented in
[12:01] such a way that that feedbacks goes back into the LLM back and the agentic system and the next version of your agent is much more better than the old one. Here is the demo that I will be talking about like how actually it happens. This
[12:16] is a control tower applications where you can see the item numbers, the suppliers, and what's the demand that's happening, and all this stuff. So, here we are just showing for a demo like how it is processing it actually, but in real time we can use it as a job.
[12:32] So, once you submit it, it has processed it, and you can see that it has run success and this. And you can also see to that trace what it has done at each step like, okay, first step, second step, and third step. And not only this, you can also see in each tool what is
[12:49] the input to the tool, what is the output to the tool here. You can see the input and output of each and every tool what it has happened as part of your experimentation. This is one of the important thing that you have to establish as part of your agentic systems.
[13:06] And this is the email finally it was sent to the supplier. And now the supplier responds to the email. You can respond it like in the table or like in a free form text like any however you respond, the agent will be able to process it.
[13:22] So, now he has replied to it. So, once we have replied to it, we are processing those email right now. You can see all the steps that he had has taken to process that agent.
[13:43] Once we have processed this agent, the next step is like human in the loop. I would say this is one more important step as part of your agentic systems to understand like to improve the final decisions. Like you here this is the human in the loop page where you can see what area is the supplier responded and what is the agent decision that
[13:59] you it was present here and then he comes here approves it and submit decisions. And finally, it goes and submits into the ERP system. So a single agentic system which was able to pick your parts, process it, send emails and responded
[14:15] everything from end to end. So this is the dashboard what we are tracking how many POs were processed every day. Uh like this is like a daily dashboard what we want to see right now and we have started implementing it for one plant and we are planning to expand it
[14:30] for other plants even. These were the different tools that we have used as part of the complete agentic system. I'll not talk about all of them but we have used
[14:46] we have registered our models using UC models and we are using MLflow for our observability. So that's all about the implementation. Now I would like to hand over it to Srini to talk about the business impact how it was created in the company. Thank you.
[15:04] Now that you have seen the application about what the agent does, right? So we are excited about this one for three different reasons. One is that the the change for the planner. As you can see earlier what the planners has to do whatever the agent has done earlier about understanding which purchase order what to act on. The email
[15:19] communication and then acting on all of that was to be in the plan itself. So, that particular one changed into now the planners instead of doing all the manual triaging and then simply they'll actually focus on a prioritized version about what to act upon and then they can actually spend their time in a more
[15:35] strategic actions rather than working on these activities about looking at part by part and looking at each purchase order and and stuff. Right? And as the reason why I said why I said that about this application ideal, right? As you can see earlier what happens is that in the supply chain we have got a lot of
[15:51] these parts and then at any point in time the demand for one part increases, the other one decreases. And then which actually creates one part you'll run into shortage. That means that you will and then the other part will have an excess. Typically the way what happens is that planners prioritize shortages because it will affect the build of the plant.
[16:06] And so you they spend time in handling all the shortages, right? So, while it so happens that in there are two parts, one for which the demand increase, the other one demand decrease. While the planners are working on the shortages, it it is possible that supplier have already shipped the parts which are which are there on the excess,
[16:22] right? So, typically that will result in an excess inventory. So, in the way now you are no no longer constrained by the the capacities of the what the planners have. And then they can actually the agent can work on all the exceptions at the parallel time. This is actually a huge benefit for us in such a way that because we we don't
[16:37] have the capacity to do to handle all the exceptions which which which of these thousands of purchase orders have created at any point in time. So, this will eventually establish a good sync between the supply and demand flow. That's actually a great benefit for us. In addition to that
[16:53] uh not necessary the planner time as we talked about from hours to minutes. We just as you can see the final control tower, he just has to look at it and then just review the communication what happened and just have to approve. The key thing what he's looking at is that he just whether the agent have interpreted the response from the supplier correctly.
[17:08] In case also the agent also has cleared to the supplier in case it has a communication which is out of normal, something he responded which is out of accept reject, he said something he doesn't understand. It won't even come to this stage. It even ask the creator of the plan to say that I don't understand this about the supplier have said. Please take an action and then
[17:24] look at it. Right. So, and right now the current process is very ad hoc. Each planner does it differently. They process so we gives us a lot more other benefits about uh auditability, having an established version. We can address establish patterns about which supplier, how they are doing. And then we can actually
[17:40] eventually help us to track these things better. As as Prashant mentioned, we are in the pilot phase. We have proved the pattern. We are in the process of expanding. There are We are excited to have several more use cases and a lot of excitement from the from the various other functions as well. Right. So, we have other cases we
[17:55] are expanding on. I'll not go too much detail into the other cases, but there are a lot of opportunities what we are forcing with this. And before we wrap up, the few point is that as you can see, it's not going to replace any of the planners. It's because the planners doesn't have the capacity itself today
[18:10] to handle it. We are actually automating and coordinating. Supply chain is a huge coordination problem. It helps the It augments the supply chain planners. And then it takes away some of the work which otherwise would not be taken. And we And then you have you have seen some of the aspects where they establish the transparency and trust. Right.
[18:25] And to wrap up, I I would like to I know we passed through a lot of content. We would like to leave you guys with five points. One is that is about us We recommend you guys to start small. That We learned a We learned a lot in this and we thought of something and then we we experimented a
[18:42] lot. And then we Where we are, it came through a lot of iterations as well. So, I would We would always recommend to start small with a small use case and expand it. And And Prashant mentioned about using the LLMs for reasoning, not for execution. I think I think one of the aspect we have not Traditionally, the
[18:57] business have not seen a system which is probabilistic in nature. Most of the IT systems are deterministic. This is an This is an experience for the business where the output is probabilistic. So, I think all of this journey, establishing the trust and having the having the observability with an ML flow, they all actually helped us to gain the trust
[19:12] from the users as well. Right, so on the observability piece it was it was kind of a mandatory for us because without the ML flow observability, when something goes off and the agent does something out of ordinary, we will not be having ability to track why it happened, what we had that our debugging
[19:28] would have been much worse. So, observability we will not be able to build it and operate it without observability. And the last human in the loop, like I mentioned, the the probabilistic nature of the output, we have not seen much of the failure cases, but imagine it can actually read one of the supplier email
[19:43] incorrectly and then potentially might act propose something. Therefore, we are currently in the 100% human in the loop where every action planner approves it before it actually operates in the article. In future, as we mature on it, maybe we'll we'll see whether it it changes, but currently we are in the 100% human in the loop. And the reason
[20:00] why we are excited about this is it helps us to coordinate and helps us improve our supply chain uh better and improve at scale. Right, so that's all we have. THANK YOU VERY MUCH. GOOD AFTERNOON, EVERYONE. THANK YOU FOR JOINING US.
[20:15] UH MY name is Yiannis Papavasileiou. I'm a machine learning engineer in Science Science and Technology organization at Corning, and I'm here with Dennis Kamotzki, uh principal software engineer. And today we'll share how we uh developed a a system to use AI to unlock
[20:32] manufacturing uh knowledge access uh in for manufacturing teams at Corning. Uh Corning has deep uh expertise in manufacturing for different areas, and this knowledge is embedded in different document systems and workflows. Uh so,
[20:47] the engineers and operators need to be able to access this information quickly, and traditional search methods that were used in the past are not uh suitable anymore uh for our speed. Uh and uh speed and efficiency. So, in this
[21:04] presentation, we'll go through the problem that uh we were facing initially. Uh Dennis will go through the solution that we built the team built for the enterprise and then I'll cover the agent that we built at the end as well as some
[21:19] lessons that we learned. But first let's go quickly talk about who we are. Corning is one of the world's leading innovators in material science inventing life-changing products and technologies
[21:35] since 1851. This year marks the 175 years of Corning history. Corning is best in the world in three core technologies, four manufacturing and engineering platforms and five
[21:50] market access platforms which act like business business units. And these are the five business units. First we have display technologies which provides precision glass for the world's most advanced displays.
[22:05] Optical communication is a business that uh provides optical fiber, cable and wireless technologies as well as connectivity solutions that carry the information at the speed of light. Automotive technologies encompasses
[22:22] emissions control as well as glass solutions for advanced interior displays or external windows. Mobile consumer electronics enhances a wide range of devices from damage resistant and scratch resistant
[22:39] glass for mobile devices and wave guides for augmented reality. And then finally life sciences business offers trusted products that accelerate drug discovery and development that deliver to save lives.
[23:00] So as you can see Corning has a vast different set of businesses with a lot of expertise in different areas. So, to be able to continue innovating, we need to be able to ground LLMs with our private data. The real advantage starts when
[23:15] we give the model access to the knowledge that only our organization has, internal documentation, procedures, drawings, as well as policies that that were approved in the latest revisions for those. This is what we call here context engineering. This is
[23:32] not just asking a generic LLM for an answer. We need to ground this the LLM response with the right sources, the right permissions, and the revisions that all the governance that uh is on top of these documents.
[23:53] The opportunity is not just simply bigger models or better prompts. The opportunity is to build AI system that that trust that reads what matters, trust what's authoritative, and powers knowledge-specific company-specific knowledge.
[24:08] So, the takeaway here is simple. Industrial AI starts when the model can use knowledge that the public internet never had and it was never trained on.
[24:28] So, this brings us to the Windchill PTC challenge. PTC Windchill is a hub, it's a centralized hub that keeps all the design documentation, specifications, and procedures, and it enables collaboration for different teams, version control of this documentation, and regulatory
[24:44] compliance. It is deeply integrated in Corning's operations, specifically the optical communications business, and it supports both document control and design and On the right, you see screenshot older screenshot of the
[25:00] search capability at Windchill. The search is very powerful, but it's also complex. The user needs to know how to construct structured queries, so they can they can perform their search more efficiently. They they can start with the simple keyword or pattern search,
[25:16] but they also have the ability to create metadata filter based on metadata. Filter each document type has different types of metadata, so they kind of to how data is organized to be able to construct a query. If they don't do that, they could end up very well with
[25:32] hundreds or even thousands of documents as a result. So, it can take time to find the exactly document the exact document that they need to reference. So, for this, we need rag solution that enterprise AI system with
[25:47] rag that can help power this. And then for that, I'll pass it on to Dennis. Um My name is Dennis. I work in Corning IT. My team is
[26:04] AI products. So, our mission is to work with people like Yiannis and with other Corning divisions to enable them to power their use cases with AI in our AI and data intelligence platform, which is of
[26:19] course Databricks in Corning. So, we started working on information retrieval and augmented generation use cases as early as 2025. In fact, Windchill was not even the first use case we we looked at.
[26:36] Um as soon as Corning was an early adopter of large language models. So, as soon as the word got out, you know, people from the divisions came to us and they said, "Well, you have this AI. We have documents. Can you help us enable question answering, enable search,
[26:52] enable information retrieval with these models. For example, you know, optical fiber division needed a searchable technical library for their process engineers. Um our you know, business operations in
[27:07] Corning International in India wanted reliable answers out of their standard operating procedures. Um Corning Life Sciences wanted to build like a mega chatbot that answers about all kinds of documents that Corning Life Sciences has from marketing to quality
[27:23] to research documents. So, we started looking at these and initially started building them individually case by case, right? But then we quickly identified some common needs, common themes across all of these use cases.
[27:39] Um So, one, going back to previous slide that Yiannis showed with that search in Windchill, um it's it's search, but it's not just pure vector search. It's really a heavily metadata-driven search. So, the need is to capture and maintain
[27:56] that metadata through all stages of your data engineering, so it ends up in the vector index, and then you can filter on it um as part of that uh rag experience. Um We also have obviously some manufacturing, right? So, it's not just
[28:12] regular, you know, books. It's uh highly structured, you know, very large documents, complex with multiple languages, um heavy with diagrams and images. So, it's multilingual and multimodal retrieval.
[28:33] And on top of um these oops Oops, sorry. On top of these uh functional requirements, we have some serious non-functional requirements. So, Corning has um is very serious about information security. So, all of these document sets from different divisions, they come with a set of um permissions
[28:51] and restrictions that apply to them. So, it's almost like document um security domains. Um and we need to maintain those permissions. Unity Catalog is great for that. Um but how to do that, you know, consistently with all intermediate
[29:07] artifacts generated by data engineering pipelines, um we had to solve that problem repeatedly. And of course, you know, maintaining um op um maintainability and and operational consistency of these pipelines. You you,
[29:22] you know, you have a lot of documents. What happens if it fails? You don't want to rerun and reload this whole process. It's very expensive to rerun and reload. So, the pipeline has to be restartable and has to be sort of transparent and observable at every stage.
[29:39] And of course, cost, right? So, you you have complex documents, you have LLMs, you can give a document to an LLM to extract content from it, but do you want to? It's expensive. Can you um do it make these decisions dynamically
[29:55] to manage cost of extraction and managing that content? So, what we ended up building is not, you know, many different pipelines on a case-by-case basis, but sort of, you know, IT centralized um document intelligence solution for
[30:11] Corning. You know, now it is based on Databricks Knowledge Assistant. Um back in the day, Knowledge Assistant was not even there. It's It's complementary. Um and and it's it's it's a hybrid of open source and Databricks, and it uses,
[30:27] you know, all the best practices from Databricks in terms of how you build data engineering pipelines. So, what's the general idea? What did we build? Um
[30:43] like I said, we have um document um domains with with, you know, security restrictions by domain, which means that if you deploy it as a Databricks workflow, let's say, you can't have uh a single workflow for
[30:59] all Corning documents. You know, each workflow needs to have its own service principal, its own permissions, its own groups, its own users. So, you can't have one pipeline. What do you do? Um we we we built a template which then our CI/CD process can
[31:16] essentially clone, and so we deploy dozens of these pipelines, like instances of this data engineering pipeline, um one for each of these security domains within Corning.
[31:32] Uh but from the perspective of the actual business user, it is a single pipeline, because that's the one that they have permission to access. Make Some of them have permission to access several, but primarily they work with one for their business unit. What do they have to do? Do they have to write code? How can they self-serve? You
[31:49] know, there's a lot of business users with documents. Um you know, now with uh Databricks knowledge assistant, it's something similar. They can provision themselves through UI. We took a little bit different approach. Um we came up with this idea of a manifest, which is a single file, single
[32:05] configuration file, um which business users simply upload together with their data um to a designated location, and a pipeline picks it up from that point and just runs with it. Um this manifest uses um progressive
[32:23] disclosure of complexity principles. So, as a business user, at a minimum, all you have to do is configure the name of your data set. You configure name of your data set, you upload your files, you're done. Um for more complex use cases where business users want to configure
[32:40] algorithms, embedding models, they understand what what they're doing, they can add and override those configurations. The pipeline is completely transparent. It every uh stage of the pipeline, content extraction, chunking, embedding,
[32:55] vectorization, produces its outputs in the form of Delta tables. So, this is very helpful, for example, for people who want to develop uh an agent, right? Which doesn't just look at the vector index itself or search result itself. It joins with other tables.
[33:17] And in the end, um we can consume that vector index in multiple different ways. We can um use it directly as an MCP server in our AI portal, right? Or we can build a Databricks knowledge assistant out of it using the agent bricks, or build a a highly customized line graph-based deployment, for which
[33:33] we provide SDKs so that it integrates with Corning um governance and operations, such as token rotations, etc.
[33:49] One important aspect of this is cost. Like I said, you don't want to feed every document to an LLM. We do have AI parse document function from Databricks, great, but it feeds everything to LLM, comes with some restrictions. Um, you can choose your LLM, you um
[34:07] you know, there are restrictions uh on file sizes, etc. So, in many cases, let's say you have an image, uh but is it really a diagram, or um is it just a scan of text? Right? We had PDF files, very large ones, w- where every page is simply a page of
[34:24] scanned text. Do you really need to feed that to an LLM? It seems very wasteful. Right? You'd be better implementing some open source based solution, you know, Tesseract, Classic OCR. Those types of solutions also support multiple languages, like 100 languages,
[34:41] um which is very important for us. We have multilingual content. So, we decided to you know, build a metadata driven solution where um either, you know, in the source Delta table which lists what document you need to ingest and you
[34:58] plug that table name in your manifest file and you're done. Or in the manifest file itself, you can add a SQL transformation to add a metadata column. But, in the end, you add a metadata column. And the metadata column says, "This is the extraction type. Uh use the basic one using open source
[35:14] or you use AI parse document from Databricks or some more complex uh contributions from our AP team who's here. Thank you guys. Um So, in the end, we can control that on a row by row basis, on a document by document basis. And so, we can keep our
[35:30] cost of extraction under control. We store all that in a Delta table with a unified contract of how we represent that nested document structure, which is very similar to the output of AI parse document from Databricks.
[35:45] Um and that contract allows us to implement something like semantic chunking, so chunking which pays attention to the structure of the document and not just operates as blobs of text.
[36:01] Similarly, there is a common theme, right? So, we want to give options, right? Because by giving options, we increase adoption, we increase the number of use cases we can put on this pipeline. We have multiple deployment patterns, um
[36:17] the simplest one, but it it it became available only recently, right? Is the Databricks knowledge assistant. You can configure one based on the existing vector index. That's what we do. Turnkey solution. Benefit that when they when Agent Bricks enhances
[36:32] this functionality, you consume the enhancements right away. For our retrieval use cases where the user kind of needs to control the additional filter additional metadata filters. Like I said, we're all about flowing metadata
[36:49] with document content uh side by side. Where the users want to use that metadata, um we deploy our own version of um chat agent, which has what they call uh self-query
[37:05] capability, which essentially looks at your um user input and tries to extract those metadata fields and then filter on them before running the search so that the search space is constrained. And of course, for people like Yiannis with very complex requirements, tools,
[37:22] you know, seek SQL uh integration, we provide an SDK that allows them to build line graph-based custom agents and deploy and it's all it it it's custom, but it's built on top of Databricks uh frameworks like Databricks Agent Framework in this case.
[37:42] All right. Thank you, Dennis. And uh based on all that, now let's go back to the uh Windchill GPT uh agent that we built. Uh the key design uh decision is here the the list of tools that we built for the agent. Uh the agent has a
[37:59] all these tools available and will based on the user query and its information uh the prompt, it will decide what tool to use. First is the rag tool, uh which is powered by the solution that Dennis and his team have built. Uh it can give us information about uh you know, related
[38:16] documents that based on the query from the user. Then we have the SQL tool, which is more like metadata type of uh questions, you know, who is the owner of this document, when was it published, things like that. Uh we have the context tool, which is more like uh internal helping tool for the agent
[38:31] itself to be able to uh generate and filter uh queries in either the vector search or the SQL based on the facilities. We have so many different facilities, spelling can always
[38:48] may not always be correct, so the agent needs to know that uh before doing any filtering. Uh and then finally the content tool, it's a new addition. Uh given a document number, it will return the full content. So, the the structure that then is provided where you have all the steps uh clearly
[39:04] defined, it's very easy to add more tools based on that structure and retrieve the whole document. And this can help us uh generate create troubleshooting agents where they can respond to the full content of the
[39:20] document, how do we perform troubleshooting for different operations, and so forth. Uh one thing to add here is that, you know, the user has natural language interaction with the agent, but then the agent will respond with uh data that um based on the query and the data that
[39:36] was used to to that, it will also attach a link to it. So, it will link to the original document in Windchill so that we keep track of the traceability, how the agent responded, what document was it used, and then governance as well for the user so that
[39:53] they they they link back to the original document. And that's because we don't really want to replace Windchill, we want to enhance it with better capabilities and then the agent search. So, going back to now that we have deployed this, we have
[40:11] uh indexed roughly 90,000 documents uh from different facilities in the optical communications business and we have enabled about 2000 global users. The initial set of users that we targeted was technology and engineering
[40:27] groups where they have very visible need for technical documentation search and then now with the addition of the troubleshooting agent we're targeting operations where a larger group of people will will be using it for
[40:47] questions they have. They're on the line, maybe they need to solve some issue, they can quickly get an answer. The the feedback we have gotten so far is that our initial assumption that it will improve the efficiency is true, that it has helped a lot and makes a big difference and then also
[41:04] reduces the training requirements. Users don't need to go through extensive set of instructions how to construct these search. The structure search in They can just interact with the agent. New users as well can benefit from I
[41:19] will let Dennis close with a few notes. Yeah, so to conclude it is possible to build an enterprise-wide self-serve no code solution for building AI powered question answering
[41:37] and information retrieval that scales. This is where those best practices from Databricks, you know, fully incremental load, you know, change data feed being able to tune, own and really take ownership of those Spark
[41:52] clusters which perform all these operations, UDFs, etc. are highly optimized. It's possible, that's how, you know, Windchill has more than you know, by this time it's probably over 100,000 documents. It keeps growing, right? Um
[42:08] and it's all processed incrementally and in reasonable amount of time. Um, we also control cost cost by using open source and not relying on LLMs for everything. And we were you know, we had longer runway, so we were early adopters.
[42:24] That's what allowed us to onboard more use cases on this. Right now we have maybe so I like I said, we started maybe in 2025 and right now we have about a dozen use cases in production similar to WinChill running on the same platform. Um, that makes it almost one production
[42:41] use case per month, but we have even more times more in pre-production. Um, so this shared pipeline design, you know, we everything people talk about you know, reusing code, reusing architecture, we we we see the benefits of this.
[42:57] You know, when we make a scalability enhancement for WinChill or a security enhancement for another project, uh, we commit into the common code base. We have disciplined ML Ops, DevOps discipline rolling out code updates to
[43:12] all these pipelines at the same time, so they benefit from each other and they benefit from our um, improvements. Our future is probably convergence with Agent Bricks. As more and more Agent Bricks features become live, we will leverage our pluggable architecture.
[43:27] Like for example, now they have you know, AI prep search function which is in beta right now, right? That's like replacement for chunking, so we can plug and play that function as long as it um, consumes approximately the same uh, logical structure. We can plug and
[43:42] play that that function. So so we see that convergence and we see benefits of these features um, automatically rolled out to our divisions in the future. Thank you for your attention. Thank you.
Learn more about the Databricks Data and AI platform.
The information provided herein is for general informational purposes only and may not reflect the most current product capabilities or configurations.