Accelerating Investment Decisions: AI-Powered Loan Data Harmonization with Claude Code
Summary
- Prophecy and Waterfall Asset Management built an AI-powered tape cracking system using Claude Code that achieves 93% first-pass accuracy on complex loan data, automating the schema mapping and validation work that previously required hours of manual effort per incoming loan tape.
- The three-step workflow covers automated harmonization with confidence-flagged columns and reasoning, templated and ad-hoc analysis on clean data, and self-healing surveillance that monitors data quality and suggests corrections to incoming datasets.
- Waterfall Asset Management, which manages $13 billion AUM in private credit and structured credit strategies, reduced review workload significantly and enabled self-service analytics for portfolio managers making investment decisions on the Databricks Data and AI platform.
Accelerating Investment Decisions: AI-Powered Loan Data Harmonization with Claude Code

Structured finance demands rapid data onboarding and harmonization. Loan tapes arrive as inconsistent spreadsheets from multiple originating institutions, and manual data cleaning consumes hours before portfolio managers can run their models and pricing analysis. Waterfall Asset Management and Prophecy built an AI-powered harmonization agent that automates schema mapping and data validation, reducing this bottleneck dramatically.
Discover how Claude Code powers an intelligent tape cracking system with deterministic confidence scoring and reasoning for every mapping. Learn the three-step workflow: automated harmonization with confidence-flagged columns, templated and ad-hoc analysis on clean data, and self-healing surveillance that monitors data quality and suggests fixes to incoming datasets. In production, Waterfall achieved 93% first-pass accuracy on complex loan data, transforming investment workflow efficiency on Databricks.
🤝
Chapters
00:00Introduction: Prophecy, Waterfall Asset Management, and Structured Finance01:12Prophecy Platform and AI Data Prep Overview02:01Problem: Manual Data Wrangling in Structured Finance05:12Old Workflow Bottlenecks and Spreadsheet Limitations06:32Solution: AI-Powered Tape Cracking with Confidence and Reasoning09:12Analysis, Charting, and Data Quality on Clean Data10:18Surveillance and Self-Healing Data Pipelines13:15Architecture: Prophecy, Claude Code, and Databricks14:03Results: Review Workload Reduction Across Multiple Tapes15:23Impact: Self-Service Analytics and Faster Investment Decisions17:02Future Challenges: Context Management and Large Language Models18:08Live Demo: Tape Cracking, Mapping, and Confidence Scoring25:21Demo: Analysis Templates, Ad-Hoc Queries, and Stratifications29:10Demo: Unstructured Data and Prospectus Processing30:13Demo: Surveillance Setup and Automated Scheduling
FAQs
What is tape cracking and why is it a bottleneck in structured finance?
Tape cracking is the process of mapping and cleaning loan tape spreadsheets that arrive from multiple originating institutions in inconsistent formats before portfolio managers can run pricing models and analysis. In the private credit and structured credit space, data is not available through standard vendor platforms, making manual schema reconciliation a significant bottleneck before each investment decision.
How does Claude Code power the Prophecy loan data harmonization system?
Claude Code acts as the intelligence layer in the Prophecy platform, performing automated schema mapping with deterministic confidence scoring and providing reasoning for every column-level decision. This allows analysts to review only the flagged mappings rather than manually reconciling every field, and the system achieved 93% first-pass accuracy on complex structured finance loan data in production at Waterfall Asset Management.
What is self-healing surveillance in the context of structured finance data pipelines?
Self-healing surveillance is a Prophecy capability that continuously monitors incoming datasets for data quality deviations and automatically suggests corrections when new loan tapes deviate from expected patterns. This shifts the workflow from reactive data cleaning after problems occur to proactive pipeline maintenance, reducing the burden on analysts between investment cycles.
What results did Waterfall Asset Management achieve with the AI harmonization system?
Waterfall Asset Management achieved 93% first-pass accuracy on complex loan data harmonization, dramatically reducing the hours previously spent on manual data wrangling before portfolio analysis. The system also enabled self-service analytics so investment teams can run faster decisions on clean, governed data within the Databricks Data and AI platform without waiting for data engineers.
Full transcript
[00:07] Thank you so much for coming in. Uh my name is Maciek. I'm one of the founders of Prophecy and we are going to talk today about how a waterfall is using Prophecy to accelerate investment decisions with our product that's specialized in structured finance with Cloud Code. Uh with me is Shahzad. Shahzad, do you want
[00:24] to introduce yourself and and tell us a little bit about Waterfall? Sure. Uh my name is Shahzad Nabi. I'm the CTO at Waterfall. Uh with the firm for almost 5 years. Uh Waterfall Asset Management is uh we operate in the private credit and structured credit space where we have
[00:40] various uh strategies in the alts business uh from structured credit, private credit, asset backed finance, whole loans um and ABS space. So, we are 13 billion AUM. Uh so, our investment strategies are all in the alts and
[00:56] private side of the house. We obviously work in the public domain as well. Um and uh and we we had some interesting challenges which we'll talk about next uh and how we engage Prophecy on some of this. Cool.
[01:12] Um a little bit about Prophecy. So, we are an AI data prep and analysis platform, as we already said, powered by Cloud Code. A series B. Um I'm sure you guys are going to be in a good company. Our customers large from the range from the largest enterprises like
[01:27] JPMC, HSBC, Microsoft to, you know, companies and smaller shops, all tied together by solving data problems. Um a little bit about myself. Um I started in AI space. First built my startup in um pharma. We were
[01:44] automating a lot of data cleaning for pharmaceutical companies. Um and now leading product and engineering in Prophecy. Again, the tying thread is for AI, you you really need a specialized context um and then very clean data. And that's exactly what we do in
[02:01] in prophecy. So, Shahzad, would you mind telling us a little bit about the problem that we are trying to solve? So, before I talk about the problem, one of the things is to lay down the context. I when I joined Waterfall from a large company, the assumption was that, oh, everything will be available like a
[02:16] technology platform. Usually, most banks and everything has them. When I landed, I realized some of the industry that we operate within the private space is very complex, regulated in a different way, which means the access to data is not available through
[02:31] any proprietary technology that's out there or through vendors or through other through other software services like SAS. So, one of the things we do is we look at deals. What we call deals is when we're financing a company, let's
[02:46] say we're issuing some sort of debt, we're buying loans, we're buying whole loans, we're buying basically private assets. Before we can look at it, we need to look at the data that comes with it. And usually, that means you're engaging with those counterparties directly. There is
[03:02] no other services in the middle. So, that creates complexity because there is no normalization You can't just go to Bloomberg or Moody's or Reuters and get public data. So, one of the things that we do is in our investment process, we have embedded engineers or data
[03:19] analysts, as we call them, which are effectively part of the desk teams, but their job is to what we call crack tapes. So, they get the data through data rooms. These are all in any format that you have, for instance, spreadsheets, CSV files, PDFs, whatever you can think of in any format. There is no dictionary
[03:35] attached with it usually when you receive the data. And their job is within sometimes 30 minutes, 5 minutes, 10 minutes to be able to get the data into a space where PMs and traders can actually start making some decisions. You can pick a sleeve of loans that you
[03:50] want to bid on, and then you need to price it. You run You need to run your own models and that process takes time. Sometimes you need to iterate a lot and usually sometimes when you're late in the market, you miss the you miss price, you miss the deal and then you obviously there's an opportunity cost associated
[04:07] with it. So, this diagram kind of depicts the workflow that we had. Um we would have someone who's looking at this data. Their job is to normalize as best as possible, which means sometimes they're seeing the same things coming from the same vendor but in different formats. So, they have to always be
[04:24] manually involved. That data bubbles up with some insights and analysis to let's say PM. PM will ask questions saying, "Oh, what about this other data?" And then there may be 500 columns that come with let's say as part of the data tape um and you have only extracted 30
[04:40] columns because that's what you needed but now suddenly you have to go back and extract maybe 20, 30 more columns. And that's that's a very cyclical process and very manual. Um we used to use an old technology that would basically very file-based. Uh you you'll be surprised even larger
[04:56] banks like JP Morgans of the world, they also use very old technology where they're using Excel CSV-based platforms to solve this problem. So, so that was that was basically the problem that we have and we wanted to solve this problem
[05:12] so that PMs can make decisions because on the other side we have a Databricks platform where we have a lot of data which is cohesive on one platform. We have Agentic AI rolled out and you have an experience where portfolio managers are looking at your performance and the risk profile in a very
[05:28] uh new technology stack and then when you're looking at a deal on how you are going to price and bid, suddenly you are stuck with the old school. And and that was our problem and that's when we engaged with Prophecy team um and we figured out like can we figure out a solution where we can make a tailor tool
[05:45] that can bring us closer to data and AI stack and on Databricks. And this is where we started the journey and then we we we basically came up with a solution that hopefully solved most of these problems. You can go to the next one.
[06:00] Um but before we did that, one of the things PMs and we tried is can we use Claude on its own? And we ran into that issue early on where um it works in isolation. So maybe one person can build a product that works for them, but then the other
[06:17] person on another desk cannot use that because obviously their learnings are very isolated to their sessions in um and we need to build a platform where you can actually interact with the data in a meaningful way, maybe make changes to it. And that is something that we learned. We built custom products in
[06:32] other spaces. This is an area that we didn't want as a 13 billion AUM shop, our goal is to work um and add value for our business. We don't want to build tech products in-house if we can uh avoid that. So hence we came to Prophecy team because we already use Prophecy in
[06:48] other parts of the organization. It's like can we solve this problem collectively? And this is where I think it led to a good solution after that. Okay. Awesome. Um so what have we actually built, right? We're going to talk a little bit about it and and actually show it to you guys and then uh
[07:04] would love Shahzad to also speak a little bit about, you know, how well did it work um in a little bit as well. So it it does start all with uh tape cracking as as Shahzad has mentioned, which really to structured finance is a very specific problem, but it applies generally to the whole data industry,
[07:21] right? You're trying to bring bring in this messy data, whether on here it's loan data, on in any other industry it could be, you know, health care data or or or data in any other format. It's going to be sometimes structured, semi-structured, unstructured data, and the user should be able to map it using
[07:38] AI and then validate what AI has developed. Um in our product specifically, what we've built is this ability for AI to give you confidence for all of the mappings that it generates, and also tell you its reasoning. The confidence is really important because at the end of the day,
[07:53] AI is going to hallucinate some parts. Now, how do you know which parts are hallucinated? So, here there is um deterministic confidence. It's going to be either low or high or high. If it's high, the user doesn't have to look at the mappings at all. It can just trust that AI has produced the right result. If it's low,
[08:09] that means that AI didn't have the right context. Maybe data quality tests have failed. Now, the user definitely has to look at that. As the user is fixing all of those mappings, there is an interface for them that allows them to very quickly iterate on it, and the AI system is going to get better and better and
[08:25] better. And we're going to show you some of the results as well that we've done on on some synthetic um benchmarks, too. And on the other side, there is the reasoning as well, right? So, for every single mapping, when you're actually reviewing it, AI is going to tell you, "Hey, this is the explanation of why I
[08:40] did what I did." So, as an example here, total spend was computed using a reference that already existed for a different uh for a different mapping, and there was some additional rescaling done on the value itself. So, very easy to understand that. As soon as the tape cracking process is
[08:57] done, now we have the clean data. Now, we're getting into the more interesting part, which is well, now we can actually start um analyzing that data, right? And determining whether we actually want to invest in this asset or um you know, perform a specific outcome. So, how is that going to work? Well, now
[09:12] we have our kind of gold uh clean data layer. Um, a Prophecy gives you an interface where you can start asking questions. Uh we're going to produce an analysis, um a summaries from AI, a set of charts. But, again, the confidence is going to be playing an important role here.
[09:28] Um and now you're starting to see a theme across everything that we are talking about. We want to make sure that the results that you're getting from AI are as as accurate as possible. At the end of the day here we're going to be talking about we're talking about finance, right? Every single number, {{}comma} matters, a digit position.
[09:46] Also, what prophecy is going to do is for every single chart, for every single number that you get on the other side from AI, we are going to one, show you a visualization of all the code that was produced, right? So, even if AI generated millions of lines of code, now you have this higher level visual
[10:01] abstraction available that you can very easily review. And then on the other side we're going to tell you how confident the system was in in what AI produced. So, again, you can go in, iterate on the output, and AI is going to be going going to get better and better and better of that over time
[10:18] based on that feedback. Now that we have the analysis produced based on the clean data, we are also going to be able to set up surveillance on top of our workflows, right? So, now we've got the data, we've cleaned it, we've analyzed it, we are confident,
[10:33] hey, let's invest in this particular set of assets. The next step is surveillance. We're going to be getting new data coming on weekly, monthly basis from our data vendors, data issuers. Sometimes that data is going to be similar, is going to follow the same
[10:49] schema. And sometimes that data is going to have different inconsistencies, right? Maybe the vendor added one more column, it modified the schema. What happens next? Well, AI should be able to monitor the data quality on top of all the data that's coming in, inform you
[11:05] when the data is potentially mismatching, but also be able to automatically suggest fixes to the data coming in, so your pipeline, your data pipeline can be can be iterated on. And at the end of the day, now that you have this scheduled, working, and self-healing system, you'll be getting
[11:21] those notifications, alerts from AI with the breakdown of the updates on on on the portfolio that you're now I'm surveying and tracking. So, all of those three steps we're also going to show you directly how they look like in the demo. But, zooming out a little bit, now you're starting to live
[11:38] in this new world where previously as what Shahzad was talking about, you're starting from a data analyst that has to crack the tape very quickly. They get this messy spreadsheet, they need to you know, onboard it onto a data format. And what is the investment in doing?
[11:54] Waiting. They're not able to do anything until that data is clean. On the other side, once they do have the data, they're of course going to be asking questions, doing analysis. There is going to be some that they can do on spreadsheets, but again, a lot of more complex analysis or modeling, they're going to be blocked. There is some
[12:09] vendors that they can adopt. A lot of the times they need to go back to their analyst team and ask them questions, redo analysis potentially, right? Um so now, what's the change? How is the new world looking like? You have all of the users working on the same platform,
[12:25] on the same source of truth, and collaborating with AI to get their work done. The analyst is going to be able to, you know, now in the matter of seconds onboard a tape. A portfolio manager is asking questions to produce standard analysis. As some of them are going to
[12:41] be templated, a lot of them are going to be ad hoc questions. They can bring in their unstructured and structured documents and AI can parse them. It can actually generate new documents as well, and they can easily review and fix them together. And you can produce new data cuts based
[12:58] on that, right? Which are all of your filters, cohorts, exclusions, etc. And all of that happens on again, collaborative platform all at the same time. So, instead of having spreadsheets sends over emails, lost context, you're all collaborating on your you know, data bricks
[13:15] kind of governed data platform. And that's kind of how the architecture looks like on the other side, right? So, you have a data studio that sits on the intelligence layer. The data studio has all the visual AI components, the pipelines, the charts.
[13:31] Intelligence is what Cloud Code powers directly with a set of agents and tools. Uh, it allows the very fast iteration on that output. Uh, we do of course need all the right connectors to to to various systems and transformation capabilities, but at the end of the day
[13:47] the the core transformations runs directly on top of um, on top of Databricks. And this is some of the results that we saw at least on at least on synthetic tapes. Um, so this
[14:03] specifically again we're talking about a tape cracking use case which is very specific structured finance problem, but very applicable to other industries. Here we got a bunch of tapes from different issuers. So this is not the same file seen for the nth time. This is completely new different data files, but
[14:19] being onboarded on the same canonical data model. And we can see that with the first tape, of course the user wants to review everything, kind of go mapping by mapping by mapping. Here we have 145 columns, so they have to review all of the columns. Fine, but now what happens with the second tape? What happens with
[14:35] the third tape? That number dramatically goes down all the way down to fifth tape where now it's just four mappings that the user reviews. Which means that out of, you know, spending your whatever an hour or two hours on reviewing a full spreadsheet worth of data, now you're just reviewing two free columns that AI
[14:52] wasn't very confident in on the on the fifth tape, right? So that's a humongous work reduction um, produced with very high confidence on the other side from AI. Uh, the the medium and low confidence mappings that you're seeing right here on the fifth tape, right? This free one,
[15:08] those are the ones that the user actually now has to be uh, has to be reviewing. And of course with more and more tapes or more and more files, the results get uh, the results get only better. Um, how does that uh, so those were the
[15:23] results a little bit more on that synthetic side though, but but how are you seeing the results impact waterfall? What future opportunities do you see? I think I think it does two ways. One is it brings everything on one cohesive platform. So, PMs can self-serve. They
[15:40] don't need to go to a data engineer just to say, "Okay, do one portion and come back to me." They are basically interacting with the data in real time. They can ask questions. They can interact it. And then because this is all on top of Databricks, it moves into the next cycle in the workflow where
[15:56] they can actually start looking at their cash flow models, run a run our scenario analysis, look at price price the deals, go back and forth. And I think that's creating opportunity for us, which means we don't need to spend too much time waiting for two or three people to do one job. And sometimes you
[16:12] can imagine this is a sporadic kind of workflow where suddenly there is flurry of activity in the market. Suddenly 10 people are bidding for the deal and then they're all bottom lined on two people on the desk. So, that that's removed, which means everyone can do the job themselves. And also now we can extend this to other
[16:29] parts of the business. We are starting to build waterfall knowledge into the same ecosystem where we can bring our IC memos, investment committee memos that we have, our commentary on our own macro views, and we can start bringing that into the same ecosystem where people can actually extend and ask connected
[16:46] questions. Okay. And you were telling me that there is still some challenges that you're seeing with AI, right? I think I I just found it really interesting with context. Can you tell us a little bit more about that? So, those are like future things that can be still solved in the AI world. I think if you if you talk to anyone
[17:02] right now, everyone thinks AI will solve all the problems in the world. And it's not there yet. One of the biggest issues we are seeing is that um, how do you engage your data with with LLMs? Um, you can ask a question, but then large data, when you're looking at
[17:18] terabytes of data, looking at it, um how do you synthesize that and give it to the model in the right context so that it can evaluate and give you the right answer? And sometimes you need to look at connected data. If you're looking at, let's say, market data or your holdings data, which could be in gigabytes and terabytes just for that
[17:35] given sector, and then you want to connect with that with the alternative data, you're looking at large amount of data. And the question is, how do you pass all of that to LLMs? I think those are the next challenges we have so that we can actually get to the right answer very quickly as opposed to it looking at part
[17:51] of the information and giving you an incorrect answer. Yeah, so so there's a lot of work to be done in the context management space. Yep. Um okay, awesome. So that was it in terms of the slides. Let's get into some fun exciting demos. So everything that we were showing right now on uh slide
[18:08] where we're going to show you in actual software. Um so we're going to start from uh some messy data in a tape, uh just a spreadsheet that we're going to throw it AI and ask it to harmonize for us. Harmonize meaning uh clean, match it into the canonical data model, solve
[18:23] some data quality issues, and we'll see how AI greatly automates that. We'll do some stratifications on top of the data with AI. Stratifications meaning just cuts on top of the data, specific filters that we're going to run through both a template and a a set of ad hoc queries. We're going to talk a little
[18:39] bit about this unstructured data processing as well with a prospectus that comes with tapes, which is kind of like this almost like a research paper that tells you about the performance of the tape and and the the kind of its summary. And then we're going to set up that that surveillance on top of that as well.
[18:54] Uh which is going to be the last step. Uh so let's deep dive into it right away. Um so this is prophecy. Let's see if it's visible. I hope it is, but if it's not, scream. Um so I zoomed it in. I hope the interface is not going to break. Um this is, you know, every
[19:10] product right now needs to start from a chat, so we're also starting from a chat, but there's going to be a lot of very specific skills, tools, and and and interfaces built for this. Now, I do have here um a tape that I'm going to be onboarding. Uh so, this particular tape can be seen
[19:27] right here. Let's just pop it up. Just to prove that it's an actual file. Um yep, my Excel is breaking. Great. But, but this is how the tape looks like. Um there is a there is quite a few
[19:44] records, lots of columns. There is a lot of data issues as well that we would need to go through, but um just trust me, there is, you know, that it's a it's a messy tape um that normally it would take a few hours for a user to onboard. Uh so, let's go ahead and let's just tell the system to crack it. I'll I'll
[20:01] use the industry terminology for it, but of course, you can, you know, just say harmonize, clean it, etc. And as I'm providing that tape to the system, we're automatically going to identify that this is this will require us to run this um cleaning harmonization
[20:16] process. Uh we're going to be able to, of course, review some basics like CD actual uh data format, uh the schema uh right here. And then really importantly so, choose the data model, right? So, this data model is our data layout. This is the expected format that we want the
[20:32] data to be in. Um I'm going to deep dive into it a little bit in a second, but there is a essentially a set of columns, descriptions, uh semantically information tags, etc. that are provided uh on top of this on top of this um a data layout, right? So, this is the layout that we want all the data to be
[20:48] in. Um you're going to have many different data layouts. For this particular use case, this is a auto uh securitization use case, so that's the specific layout that I'm using right here. And I'm just going to kick off the agent and let it start at the work.
[21:04] While this is happening, just one more word around the data layouts. In this particular case, we are using a single table layout. That's just because that's the very common use case in structured finance. But but if you look at some other data layouts, of course they can they can be arbitrarily complex. So
[21:21] as an example here, there's a data layout of multiple different tables with primary foreign keys relationships. You can actually pop out any one of those data layouts. I'll zoom this in a little bit and see the kind of um metadata in it, but also a set of data
[21:37] quality tests that are defined for it, right? And those are going to be very important. The AI will have to be using them and evolving them. Uh let's see how our process is running. Uh perfect. So now we are on the actual mapping screen. Uh couple of important things. On the
[21:54] left, you can see the AI agent itself, right? That's what you're going to be talking to doing a lot of the work with. On the right, there is a very specific um harmonization interface. You have a set of source columns that we're going to be mapping from. Uh this is all the columns that were in our um
[22:11] that were in our in the Excel source that we provided. And then there is a set of targets. This is the that the data layout that we've defined. Uh now that AI has already did done some work, uh which took a minute or so, we can already start seeing those mappings being populated, right? So we see that
[22:27] asset number for instance is used by three downstream columns. I hope that that's visible. On the other side, um you know, maybe some more complex. Some of the columns are not going to be populated in this case, which is totally fine. Uh there is going to be some that are a little bit more interesting than just one pass-throughs, right? As an
[22:44] example, we have a loan age, right? So we see that loan age as a column is coming from a reporting period ending date and an original first payment date columns, right? And AI came with came up with an expression. Uh this is where we can see that AI confidence, right? This is what we were
[23:00] talking about. Uh it can be either low or high. In this particular case, AI has seen similar columns. It's created this loan age many, many times, so it's very confident in what it did, and we've deterministically determined that as well. Um data quality in this particular test
[23:15] case is passing as well, and we can see, of course, the actual data, right? So, original first payment date, some set of dates, reporting periods in a date in a different format, and we can see the loan age in in months being generated. Uh there's, of course, AI explanation,
[23:30] fine, uh data quality results, and really importantly so, the memory, right? So, this is the part where that you will see evolving the more you use the product itself. For every single one of those mappings, uh the memory is being stored as the users are refining it over time.
[23:47] Now, on the top, we can see the progress that AI has made, right? So, out of uh 88 mappings in total that it needed to generate, 87 are very high confidence. Awesome. So, those are the mappings that you can basically ignore. Uh they're the ones that AI is very
[24:02] confident in. It's seen the mappings like this in the past. You've developed code that's similar to this that it was able to adopt. Great. I can just look at all of them and basically accept everything. On the other side, there is one mapping that seems to be failing for us. Okay,
[24:18] let's look at what's happening here. So, it seems like the data quality tests are failing for this. Uh it's actually a new column that I added to this data model, so it wasn't aware of this mapping before. And what's happening here is that it's expecting a format for target columns
[24:33] that's different than the data than the data that's coming in. So, the data that's coming in says bureau for our credit score type, which should be actually classified as an other source. Uh AI already determined the potential fix for us in code, so the only thing that I really have to do here is just
[24:48] apply that fix. Um that was a fix that was hallucinated by AI, but in this case, that hallucination was correct. That's which is fine. Uh perfect. So, now I have no more mappings to review. I can just accept this mapping as well. And I have basically completed all of my
[25:05] work. So now I took this messy spreadsheet, onboarded it onto a clean clean schema, reviewed my mappings, which basically took no work from me, and I'm completely done. AI took, you know, 3 minutes to to generate this map to generate this mapping. This mapping
[25:21] under the hood is backed by, you know, code sitting on DBT sitting in the DBT format on your GitHub. So if you want to review it, etc., you absolutely can. But now we can get into the more fun part. So now that our data is ready, we can actually start doing analysis on top of that data. I have here a couple of
[25:38] pre-baked templated analysis, and you can very easily create those analysis by yourself. Strats, peer comparison, post close surveillance. But you can of course build more templates or start asking any ad hoc questions on top of that. I'm going to go ahead and start with a strats template. So this template is
[25:55] going to tell me a little bit about the summary of the principal balance. There is a demographic analysis that needs to be performed. So this is the right template. I'm going to go ahead and start asking AI to now produce this analysis for me based based on the new data that's coming in.
[26:12] This process usually takes a minute or two, but I have it already done right here. So I'm just going to switch to this tab. And this is how this analysis looks like for this particular for this particular tape. So we can see that AI has actually, you know, deep dived into the
[26:28] data itself. There's a set of skills that are specific here for structured finance, but you can really easily build them for any other industry. That already tells you the important parameters for this loan, right? So there is a credit quality, yield profile, duration, all the important information that the investor would be
[26:44] wanting to look at, portfolio manager would be wanting to look at in the first pass. There is a set of, you know, composition summary statistics. Charts that are generated. Now, this part is easy that now that our data is clean.
[26:59] Right? So, we can be very confident in in all of those results. But, we can go a step further and we can of course start asking AI any questions or to create further analysis on the product itself as well. Uh so, in this particular case, I'm going to go in and ask it to perform a data cut. Data cut very again very
[27:16] specific terminology to structured finance, but but it's essentially just um processing this data in in such a way that you're focused on a subset of loans that matches the criteria. Um so, so filtering essentially, right? So, I'm going to focus on very specific FICO less than 750 cut and so that I
[27:33] understand the impact on my loan data set and whether it's still worth uh investing in it. As you fire any one of those prompts, all of this data uh that it's operating on is of course on Databricks. So, there's going to be a set of tools that we're leveraging to talk to Databricks,
[27:49] to your semantic layer, Unity Catalog, pull a lot of information from there. And the other part is of course the code that's stored natively uh on your Git. So, you can actually see all of the those AI operations uh tools, etc. being executed. Uh so, it's reading the files
[28:05] um I assume natively code, running a set of queries on top of Databricks, and very shortly we should be able to see the uh impact analysis itself. As part of the run again, it leverages a set of skills as well that help it be more defined and fine-tuned towards the uh
[28:21] finance industry. Uh perfect. So, now we can see that um you know, what's the impact on the other side of our cut. Uh if we perform this cut, we're going to take away a lot of the contracts. Uh so, actually over 40,000 loan contracts are taken away.
[28:38] Uh so, we remove a lot of the pool, uh but it does improve the credit quality significantly of um of the loan data set itself, right? So, it it looks like it's a potentially a pretty good buy opportunity if we are able to perform this data cut. And we can keep going and asking it more questions. It can we can
[28:55] actually ask it to modify this analysis as well very easily. Awesome. So, we did clean the data, produce the analysis on the other side. Of course, anyone can now see it, interact with it, collaborate, ask questions. Uh on the other side, we can
[29:10] of course start working with unstructured data as well. Um so, just very quickly showing that not going to deep dive into into it a lot, but with prophecy, here I have a set of um PDFs that we've provided as well. So, for this particular tape, um there is a full
[29:27] prospectus that we've given it. Uh this is the prospectus that we've given it. Uh so, that the agent actually can extract information from this, compare it with the data that we're given the agent from the Excel files, and actually ask it to develop other assets. In
[29:42] structured finance, the collateral analysis or prospectus creation is like a common thing that that that is done. So, here the agent is actually producing I asked it to create a collateral analysis note on a particular loan, so it actually created this markdown document that you can now go and edit,
[29:58] modify uh very easily, and and and and collaborate on. Um the very last step that we were discussing in the on the slides is now surveillance, right? So, we have this pipeline that's created under the hood.
[30:13] This is this visualization of code that we were talking about, right? So, for every single um chart that we've seen on the visual um analysis, there is a pipeline that actually backs it up. We can very easily understand what's happening here, so we have some geographic heat map, original
[30:29] mileage, vehicle segmentation. All of that is just a code on the other side, right? I assume there's a bunch of technical folks in the audience. Let me just show that as well very quickly, right? So, this This code that's get generated. Uh we have all of those, you know, summary finals, credit scores,
[30:45] etc. SQL is generated. All of those are just DBT models that Prophecy on the other side is visualizing for you. And then we can set up the actual surveillance with agents going in, potentially fixing pipelines, sending you reports over email depending
[31:02] on what you what you've actually built. And setting that up is super easy. Now we can just go in and say, "Hey, this is our monthly report update." Set it to monthly, send it to specific users.
[31:20] Like my investor team, for instance. And save it. Everything is here stored on Git, so I just need to publish it to to actually start for it to start running. And now it's going to be actually executed and running on your on your system itself. So this is Prophecy itself. Now we've went from, you know, a messy tape,
[31:36] cleaned it up, had the agent generate the mappings, create the analysis, and and and we were able to schedule it.
[31:54] Cool. And with that, I think there are two very important outcomes, kind of takeaways. One, you know, AI works very much so. It automates a lot of the steps, but it does require really good tooling and and human in the loop to learn about the context. And you know, we've shown that it works. So feel free
[32:11] to go in, try it out, reach out as well to us. So that is there anything else that you would like to add as well? I think you covered it. All good. Okay, well, thank you so much for coming in.
Learn more about the Databricks Data and AI platform.
The information provided herein is for general informational purposes only and may not reflect the most current product capabilities or configurations.