Skip to main content

Built on Databricks: Payment Routing Optimization at Enterprise Scale with PROM and Lakebase

Summary

  • CMSPI is a payments intelligence advisor that processes approximately 185 billion transactions and built a Databricks-native payment routing optimization platform that replaces more than 30 hours of manual spreadsheet analysis and generates 14% incremental savings for merchants.
  • The platform uses a medallion architecture normalizing payments from over 300 report definitions, MLflow for feature table governance, the PROM constraint solver processing billions of routing combinations in 30–40 minutes, and Lakebase managing transactional product state as a Databricks-native Postgres layer.
  • The architecture is presented as a reusable pattern for governed Databricks-native applications, combining Delta for analytics, SQL warehouses for interactivity, and an agentic layer with Genie and a Knowledge Assistant for explainability.

Built on Databricks: Payment Routing Optimization at Enterprise Scale with PROM and Lakebase

Watch: Built on Databricks: Payment Routing Optimization at Enterprise Scale with PROM and Lakebase
Payment routing optimization at enterprise scale requires more than spreadsheets. CMSPI built a Databricks-native automation platform that replaced 30+ hours of manual analysis and generates 14% incremental savings for merchants. The challenge: payments involve interconnected variables across networks, incentives, transaction types, and processor capabilities that create combinatorial complexity traditional tools cannot solve.
Learn how CMSPI built this production system: medallion data architecture normalizing payments from 300+ report definitions, feature tables governed with MLflow, PROM constraint solver processing billions of combinations in 30 to 40 minutes, and Lakebase managing transactional state. Discover a reusable pattern for governed data apps combining Delta for analytics, SQL warehouses for interactivity, and agentic layers with Genie for explainability.
🤝

Chapters

FAQs

What is CMSPI and what does their payment optimization platform do?

CMSPI is a payments intelligence advisor that processes approximately 185 billion transactions and helps merchants optimize payment routing decisions across networks, incentives, transaction types, and processor capabilities. Their Databricks-native platform replaced a manual spreadsheet-based process that took more than 30 hours per analysis cycle with an automated, governed optimization system.

What is PROM and how does it work on Databricks?

PROM stands for Priority Ranking Optimization Model and is a constraint solver that processes billions of payment routing combinations in 30 to 40 minutes. It runs on the Databricks Data and AI platform, using normalized feature tables governed with MLflow as input and producing explainable routing recommendations with comparison outputs for merchants.

Why did CMSPI choose Lakebase for product state management instead of the lakehouse?

The payment optimization application requires a transactional database layer for managing interactive product state — including user configurations, incentive settings, and merchant-specific parameters — which is better suited to a Postgres database than an analytics lakehouse. CMSPI chose Lakebase because it is Databricks-native, enabling faster product velocity without introducing a separate infrastructure stack.

What results has CMSPI achieved with the Databricks payment optimization platform?

The platform generates 14% incremental savings on payment routing costs for merchants compared to their previous approach. It also eliminates more than 30 hours of manual analysis per optimization cycle, replacing spreadsheet-based workflows with an automated, governed, and explainable production system.

Full transcript

[00:09] Okay. Hello everybody and thank you so much for coming to our session. Uh I'm Jordan Pierce from CMSPI and I'm joined today by the amazing Sanjay Shock from data bricks. Thanks Jordan. Really glad to be here. Thanks all for joining today's session. Uh today we actually going to do u our
[00:25] talk a bit differently because I'm going to ask Jordan some questions. So we're going to keep it a very conversation style rather than us just talking through the slides. Yeah. So so today's session is going to be effectively three talks in one. It's
[00:40] going to be one part payments, one part architecture and one part cautionary tale into anyone who's tried to stretch a spreadsheet into doing a platform's job. Yeah, I see a few smiles in the room actually. So um the thread that's going
[00:57] to hold the talk uh together today is actually CMSPI had a messy domain data problem. Uh they're an interactive product state problem and a high-scale optimization problem as well and on top of it there was a big challenge of
[01:14] governance. So uh the interesting part is they were able to solve all of it by adopting a data bricks architecture. So uh let me set up where we are going from here. So this is the shape of the talk today. Uh we're going to see the uh of course
[01:30] who CMSBI are and we're going to look into the business reason why CMSBI needed a simulation right and the payments complexity is outrunning static analysis. We're going to see why is that and we're going to look into the dynamic
[01:47] data fabric because you can't really make sense out of the simulation if the data I mean the inputs to the simulation isn't correct. So this needs to be trusted, governed, pre-processed and normalized. So then we're going to look
[02:02] into prom. Uh this is the the first simulation engine inside CMSPI's digital um payment digital twin. So um I'm really excited to know more about that actually. Um and and then we're going to look into a bit of the architecture
[02:18] itself and uh why the app needs a product layer and why the product layer has to be um a Postgress database and why did CSPI choose lakebase for it. Um
[02:33] yeah closing on the broader pattern with data bricks. Over to you Jon. Yeah. And and the main thing we want to highlight on this slide or underline at least is this bottom line down here which is this today is about a product that is CMSPI
[02:49] specific but the architecture pattern is something that you guys can all reuse for your own governed data bricks native applications. Amazing. Before all that I'd like to know who are CMSPI and what do they do?
[03:06] Yeah. Who is CMSPI? Does everybody know who CMSPI is in this room? No, that's a shame. So, that kind of ruined that uh hype moment. But CMSBI, we are a payments intelligence advisor. Um, what that means in plain terms is
[03:22] that we help merchants, very large merchants, understand and optimize one of their biggest and uh well, least transparent cost and performance levers and that is their payments. How do we do it? Well, three components. Firstly, the
[03:41] well our dynamic data fabric and that is our normalized view of payments reporting data across merchants, issuers, networks, card types, regions, etc. And that gives us this marketwide source of truth of payments.
[03:59] Second, it's our payments intelligence engine, effectively our insight layer. And that's where you can think about where we apply modeling, machine learning, benchmarking, any kind of analytics and simulation as well. And that is used to turn these, you know,
[04:16] raw complex payments information into prioritized actions. And thirdly, well, we are seniorled payments experts with over 20 years of payments experience. And that means that we know that from the payments intelligence engine when we
[04:32] identify opportunities of how to make your payments more productive it's not enough just to find those opportunities we need to implement them and that means that we have worked extensively across different departments across many organizations. Amazing. Um why don't we give the
[04:48] audience um you know the scale okay so that you know that's going to change everything afterwards honestly. Yeah. So we've now got a data lakehouse that contains over 185 billion transactions. Yeah, it's quite a lot of volume and that organically grows around
[05:05] 1.5 to two billion transactions every single week. How do we get it? Well, that's because we engage with around one in four of the Fortune 500 merchants.
[05:21] Yeah. Yeah. So these are really good u lakehouse stats. So um I'd like to double click more on the complexity of payments itself. So Jonah, could you help me understand what makes payments complex? Yeah. So payments is complex. I I guess the the first part what makes it so
[05:37] complex is that it's not just a single metric really. It's a system of interacting variables. Every single transaction can vary by channel by authentication methods like 3DS. You can have te to tokenization, network
[05:52] tokenization for example, different card types, different regions, they all affect what the payment structure looks like. Um, of course there's loads of volume in payments. Every person makes money. For example, the global consumer uh index, I think that's the correct
[06:09] one, isn't it? Uh, and that comes to around 50.7 trillion dollars of spend in 2025. So obviously there's a lot of volume going through payments in the first place. And I don't need to say this in front of everybody in this room at least, but you know technology keeps
[06:24] evolving and that means that the payments ecosystem also keeps evolving. You've seen that through the rise of e-commerce over the past 15 years. Everyone buys everything online now, don't they? And now you've got payment methods like cler so buy now pay later
[06:40] sort of schemes as well going. You've got cryptocurrency, so stable coins indexed against the the US dollar for example, and most recently agentic commerce, and that's probably not going to be the last time you hear about that this week as well. So, yeah, there's a
[06:57] lot of interacting variables inside of here. So it's it means that static analysis of looking across these cross dependencies it's a multiple well a multivariate different system which means that there's multiple possible outcomes to any one scenario.
[07:13] Exactly. Yeah. And also the static analysis can't uh really explain a multivariable scenario. Right. So that's the headline especially with multiple outcomes and and also like you know for a problem of this kind you can have multiple dashboards to uh you know spin
[07:29] it up and try to solving it and that's the reason why CMSPI pushed towards a digital twin architecture. Yeah. So this is our strategic shift really to move from that static analysis to a more dynamic inference kind of analysis. So of course static analysis
[07:46] is useful bs it you know it's backwards looking it looks at a point in time the digital twin however is a control plane for configuring simulating and explaining different payment scenarios and that means that you can change the
[08:02] kind of questions that you ask for example you can say what will happen if I change X what will happen if I change my routting strategy to X what if I want to bring a new payment method in a different country, what will be the impact? So that's why we've built the
[08:18] digital twin in the first place to see how these variables interact with each other and look at those different metrics of performance across them. How do we build it? Well, we start with our dynamic data fabric mentioned just before which is this normalized transaction data across different
[08:34] suppliers across different regions and merchants. It gives us all of that context. Next, the first operating component of the digital twin, which is prom. I know we're just calling it prom at the moment and it's uh just an abbreviation. We will eventually get back to it and I will explain a bit more
[08:49] but for now just know that it's in the pin debit routing space in the US which is about minimizing the rooting cost for pin debit transactions but we want to enhance that further and we're just finishing off the uh production readiness of this now of integrating
[09:05] approvals which means looking at transaction performance as well so you can ask the question of what's what's the best way to root my transactions to achieve the minimum cost or best transaction performance. or a combination of the two of them which will make that a lot more powerful too.
[09:22] And then obviously as we continue to build further and further metrics such as like fraud rates or conversion rates these are different preferences that a merchant can you know tear towards. Of course then there might be completely different scenarios that we'll ingrain inside of the payments digital twin as well. I mentioned some before around if
[09:39] you brought in different payment methods, what would happen if I brought cler instead or if I introduced mir card in Russia for example or any other different metric or volume started to increase or oil prices started to increase what would be the effect on to my cost of my transaction performance
[09:54] for example. And finally as is all the hype at the moment is obviously the agentic layers. they will in our case at least provide a massive benefit to explainability and exploration of data but hopefully we'll also be able to generate
[10:10] different scenarios from the agents themselves as well. So, I guess the the key point of why we've built this in the first place is that we we want merchants to be able to, you know, test different scenarios of their different strategic changes
[10:26] without ever having to run that live AB test in production. Yeah, this is fantastic. But um if you look at a digital twin as well, uh it's only as good as the data that's underneath. So before we deep dive into
[10:42] the twin architecture and prom itself, we definitely have to talk about the data foundations. So this is the dynamic data fab fabric. Um and I'd really like to understand how it's set up within CSBI. And this is the
[10:59] key part here is before we do any extract any um payment optimization or do decisions on payment that we have to normalize the payment meaning in the first place. Is that right Jordan Ren? Yeah the the meaning is the the important part here. So getting to that
[11:17] normalized meaning is it's pretty complicated actually. One of the the core parts of the or the core issue in the first place is there is no regulated reporting standard for how payment should be reported. Of course, there is h a regulated standard for how the
[11:34] transaction message should be passed through the payment supply chain. The ISO 8583 code for example. If anyone is an enjoyer of ISO codes, no one. Okay. If uh hopefully some people can come and pass me back their favorite ISCO codes after this, but mine is ISO 3103 which
[11:52] is the ISO code for um brewing the best cup of tea as any English man should know. So I digressed of course but uh yeah because of our independent positioning in the market that means that we can get data from merchants, processes,
[12:09] gateways, networks, issuers and that really provides us this uh like marketwide source of truth as I mentioned before. But to to do that that means that we have to normalize each one of these different reports that comes from each of these parties differently.
[12:26] They all come with slightly different structures, different formats, different schemas and sometimes have you know different levels of quality for specific areas of data. Yeah. Yeah. Amazing. Yeah. And also like you know just by looking at this this uh is
[12:41] a classic Malian architecture right and also the work is not the pipeline itself it's about the semantics. So I would actually like to understand a bit more uh Jordan. So could you explain how this um architecture is especially on the medallion pattern?
[12:57] Yeah. So I mean this is not a bit everybody in here hopefully has a lakehouse and they they know the bronze, silver, gold uh enrichment sort of process. Suppose for us the start point is is passing from raw data. So that can be unstructured data or structured data
[13:13] for example but putting that into the uh bronze tables. The silver stage is the biggest step for us, that enrichment process where we have to take different attributes from within the transaction to be able to categorize it. And it's really about finding that meaning in the
[13:29] context of that particular supplier's data. So for example, you might have a free description from one supplier and that will be a long descriptive piece of text. Another one might just be a product code and you have to use a lookup and one might not even have any description at all. and you have to use the other attributes within the data to
[13:45] try and work out what that means. So that can be the the complicated part of that. We still use reax in 2026. Yes, three tiers for reax. Uh and it's a obviously a complicated reax system, but it is helped out nowadays with the two
[14:01] new agentic functions from data bricks, AI classify and AI query. If you haven't given those a go, I would totally recommend it. They are really, really powerful, especially when it comes to this semantic layer of knowledge as well. After that enrichment process in silver, we move to the gold table which is really just all of our normalized
[14:18] output data and that is really ready then to be used within any of our decisioning components. I suppose the the main thing that I would emphasize on this though is that it's this is not really well it is it's data pipelining but the the thing that
[14:34] makes this CMSPI is one of at least our USPs is that this ability to bring all of this data from these different suppliers across all the different report definitions makes it um well it's that that standardized truth across all
[14:50] that supply chain and when we have this across about 300 over 300 different report definitions that have gone into this across 30 different direct supplier integrations that leads to obviously a massive amount of volume but that volume is not just you know volume to us that is our accumulated
[15:07] semantic knowledge that is true yeah and also if you tie it forward um I think you need to like talk a bit more on that as well because that's the dependency the people might actually skip yeah yeah because like prom or the payments intelligence engine It
[15:24] literally cannot work until we have these canonical dimensions, these normalized pieces of of data in these semantic expressions. Basically no fabric, no solver. Exactly. Yeah. And also like you know once normalized so there's um data that
[15:42] be it becomes as available as a govern input for um all the modeling downstream. Yeah. Yeah. So over to you. Yeah. So once we've got our a lovely delta outputs, these are standardized canonical
[15:58] dimensions, you can build well what we do anyway is we build feature tables from this. So that could be things like merchant features. That could be issuer features, network features and bin features for example and bin being the bike bike. Come on bank identification
[16:15] number which is typically the first eight digits of any uh card number the full pan full card number. And that's a particularly important feature because that tells you which of the networks a transaction can be routed to. Um
[16:31] obviously in this case the the these features can be used as inputs. So inputs into prom for example but it also serves as the basis for training in impress into the approvals integration that we talked about before. And for future digital twin components it serves
[16:47] as the same exact foundation. Yeah. And also the MLflow here uh is a governance wrapper. Right. For those who um don't know what MLflow is or new to MLflow, it is a product that you use to actually
[17:04] version assets. If anyone who has done data science modeling, you could you have you could you could version assets and the model itself and log everything in MLflow, right? And also track metrics. So um I'm really curious to know why this matters here, Jordan.
[17:20] Yeah. So you don't want just to run your models in isolated notebooks as like one-off jobs, do you? You want to have trust in the outputs of your models. So you want to know that they've moved through a controlled life cycle to get
[17:35] into production and through that feature pipelining, development through to model evaluation, validation, and then into deployment finally of course and then finally into monitoring as well. Um why it's so important is because when you are then serving these
[17:51] results to influence a a decision of any kind, you really can't avoid governance. So as soon as a a recommendation or a result affects maybe like a routting decision that you start to send volume
[18:07] to a a different entity instead, you need to know what data did that model see. Is it biased or not? um what model produced the result, how was it being monitored? This is critical for being
[18:23] able to trust the the outputs. So that makes sense because you know what we've seen so far is you know understood how CMSI normalizes the data and then you're using MLflow as a governance wrapper and to also you know
[18:41] manage the model life cycle and everything etc. So I would really like to learn more about the decision engine now. So what's uh what is prom? You know if you could talk a bit more about this payment intelligence engine that'll be amazing. Yeah. So prom um well before we just get
[18:59] into all about how it works and what it does. I think it's probably important. Firstly we understand what it's for and what's the business problem it's actually solving for. So Prom is involved in the debit you the US debit card market space. Any one transaction
[19:16] can be rooted to multiple different networks and those different well you can have global networks vying for competition and regional networks all vying for competition that could be like Vera and Mastercard for example for those global networks and pulses Excel etc. for these local regional networks.
[19:33] You might have different incentives or constraints that could be applied and there might be um well different well I suppose incentives different buckets and different processor capabilities of actually how they can implement that structure too and this really affects
[19:49] what the the merchant can actually do. Um yeah think so makes sense. Yeah. Um and also like you know you mentioned a bit about the um you know the configuring the constraint optimization problem uh and everything
[20:04] right so I'm really curious to understand that all of the groups that you mentioned here they interact as well right so once you find let's say a cheapest choice for let's say any client that might not be a global one globally
[20:20] right cheapest one as well so um yeah I'm curious to understand how you're going to deal with that kind of scenario Yeah, that's so that's exactly right. All of these different aspects inside of here, the constraints, the incentives, these transaction amount buckets capabilities, the availability of the
[20:36] networks of which transactions can go to different places, the commercial rates that could apply, different fees that could apply, they all interact with each other. So that means that just finding on a particular individual transaction if it goes to the minimum cost route that doesn't
[20:53] necessarily mean that that is the minimum cost for the whole portfolio decision. So you need to understand it's the whole portfolio. Make sense? Yeah. And also like you know uh we we saw like you know why the old approach uh and
[21:09] looking back at the Excel you know with manual estimates it's definitely not going to cut it right it's not going to help you solve this kind of a problem with this complexity sorry yeah so Jordan can you help me
[21:27] understand why Excel uh is probably not the right choice yeah this is this is the cautionary tale, I suppose. Um, but of course we'll know Excel, for example, has a limit already of about a million records in any one file. That means that you
[21:43] automatically, when we're thinking of this volume, that you're going to have to aggregate that data in the first place. That means that you're also going to have to look at potentially simplified assumptions. And that means that you're going to have to potentially do single pieces of analysis or piece of
[21:58] rooting analysis in one go. And that means that you end up with a likely view of what the savings could be for changing your routing strategy in prom. Of course, this is applying now at the transactional level across billions of transactions.
[22:14] Um, and that's really looking at now what networks are available for every single transaction, how the incentives will play out with each other as they interact with one another, and understanding then what the capabilities of the processor is to implement any of those. um incentive structures or
[22:32] implement that routing strategy that ends up being delivered. Yeah. So the the combinational space then in this particular case is looking at the number of transactions with the number of network routes that are available, the incentives that are in play, the rates that could apply and then the capabilities of the processor.
[22:49] Yeah. And also the optimization u it has to look over multiple combinations, right? It has to look at the whole scenario rather than just looking at one row at a time. So that's a line between Excel and a solver. Yeah. So yeah, so Excel estimates the,
[23:07] you know, SQL returns some rows, but PROM returns an optimized decision under constraints and that's what turns it into more of a decisioning product rather than just a you know a workbook trial and error kind of result. Amazing. Yeah, I mean that makes sense
[23:23] because you know the platform earns its keep. So data bricks gives scale optimization prompt gives the math that's the brain you know behind uh this optimization and we're also going to look at lakebased right so lakebase gives the product state for you and uh
[23:38] so the users can quickly interact with the prompt by applying the configurations so um that actually gives CMSPI the ability to um actually tackle a different set of problems altogether right so you're actually dealing with a different class of problem
[23:54] yeah for sure I I mean, Excel would give us a bit of a a best estimate um or e estimates, but PROM gives us a a reproducible optimization workflow that allows us to continually evolve that as
[24:09] well. Okay, so I've mentioned it about a hundred times and I've never actually said what it stands for. So, this is the moment. Um, PROM, what is it? It actually stands for priority ranking optimization model and its job is to
[24:27] find the lowest achievable cost for any userdefined configuration. Configuration here is the keyword prom isn't running in a vacuum. It takes the scenario that the user has defined for different incentive structures, the different
[24:43] segments that they want to route to different transaction amount buckets for example and it takes those the available networks that could be rooted to in those for all those different transactions and then it evaluates that scenario and that could be you know billions upon billions of combinations
[24:59] of where that transaction could be rooted to instead. And it does this all in about 30 to 40 minutes for very large retailers anyway. Yeah. Yeah. So, so are you saying the output is uh not a savings number? Well, partly partly contains a savings
[25:15] number. Um, but the main thing that is it brings out is it brings out this optimized routting order. It then needs to be able to explain how it came to this optimized routting order. So it shows you the movements of where that volume's come from, the cost difference of where each
[25:32] of those when it moved from one network to another for example and then you're able to compare the results of different scenarios with each other and that leads to a very explainable result set. So that's what makes you know prom being the first operating component of the
[25:49] digital twin. It gives you these configurable inputs, simulated outcomes and explainable decisions. Yeah. And also to uh matter to all humans, you know, who are interacting with the system, uh the solver has to be wrapped into a product,
[26:05] right? Uh so that people can configure, they can run, they can compare, they can also repeat the same simulation that they've done before. So uh without having to actually think like an optimization engineer.
[26:22] So over to you. If you could talk a bit more about how this would look like a user journey, that would be amazing. Yeah. So this is the basically the product abstraction of that constraint optimization problem. The user can now go in and configure a different scenario. They can add any of the
[26:37] incentives that they want to apply to a particular situation. Any any other constraints inside of there. They can then run that scenario and see how it um yeah changes anything. seen the review the results compared to the baseline of where it was before. They can then add
[26:52] further scenarios, add new incentives, etc. And then they could just repeat this over and over again as many times as they want to to be able to keep comparing those different result sets until they find their best scenario. And I think that last step there is probably
[27:08] the, you know, the whole game inside of this in the old world of the Excel, my cautionary tale coming back to haunt us again. um you'd basically have to build a new workbook every single time to each time you had a different scenario that you wanted to implement.
[27:24] But now in this more product abstracted way now it just turns into I have a new scenario a new governed scenario run inside of prom. Yeah, this is very interesting because the product is actually translating
[27:40] um a user's intent into the optimization rules, right? Yeah. So this basically means that the user can come into here and they they don't need to think ah this is a big constrained combinatorial big optimization problem. They can just
[27:56] think in the things that they know best the the business rules. They get to think, okay, here are my incentives. Here are the rates I think might be applied or here's some custom rates that I want to try and apply. These are the buckets that I'm looking to implement with the processor. And then PROM takes
[28:12] all of those different, you know, u rules, decomposes them, and then evaluates them. Makes sense. So let's make this a bit concrete, right? So for for a merchant who's actually coming and using the system uh let's look at it there in
[28:30] their point of view in a merchant specific point of view because what matters is incentive. Yeah. So I would love to learn more about how all of this ties to the incentive. Yeah. Sure. So merch the incentives is where you know routing becomes very
[28:45] merchant specific. Merchants could have one or many different incentives in instead. So they could for example have a a minimum volume requirement to send a certain amount of transactions every single month or year for example to a particular network. They might have a a
[29:02] rebate but they need to achieve a certain volume threshold for example. They might have a minimum requirement to send or or a minimum requirement to send a portion of transactions to eligible networks to a particular when the network is eligible for example or they
[29:17] might have to send all transactions to that network if it's eligible on that particular transaction. Um yeah but just thinking about these as you know those types of incentives they they're not things that you can think about after the fact. You can't just get
[29:34] the to the minimum cost per transaction to start with and then come back to these. They completely change the the optimization itself. They're part of the solver itself. So they need to be considered right from the beginning. Exactly. And incentives are just one
[29:49] part of the constraint set. Right. So you have you have rates, you have buckets, you have availabilities. You know, I think that's the next layer that we're going to just do a deep dive about. Yeah. So this slide is really about making the app the optimum actionable.
[30:04] So you know the the rates here they they could be you know switch fees or pro fees or interchange fees for example and they could be published pass through or custom rates which could be also defined you know negotiated custom rates for example they define the the economics
[30:20] for example of the transaction. Then you've got the transaction value buckets and that's really you know a processor might to only be able to implement like eight different buckets for example it's it's quite complicated if you said I wanted to route every single amount difference to a different route
[30:36] potentially so zero to one one to two one to etc. So usually it might be about eight different buckets or they might be set particular buckets as well. And that means as well that you know the the optimization here can't always just be about the uh theoretical optimum that
[30:53] you know 0 to1 one to2 for example it needs to be about what is actually possible to implement and then you've got the the network availability as I mentioned before not every single uh transaction can be rooted to every single network. So that means prom needs to know what are the generally or
[31:09] genuinely available um networks that can be rooted to before it goes and tries to allocate any volume to those networks. Yeah, makes sense. Yeah. So once we configure everything, you know, prompt runs and the real value is uh actually
[31:26] the outputs, right? So not just like any other answer, you um should be able to explain and compare these results. So shall we have a quick look at how the output looks like from prom? Yeah. So you know these the user needs
[31:42] like clear answers and and clear results but they also need to know how did we get those results. So of course you need to be able to see that what's changed what volumes change from one place to another the cost differences between each one but you need to be able to compare the different scenarios that
[31:58] you've you've tested as well. Um, I suppose that the main part of this that makes it so important is that that comparison of results because when you eventually get to a decision of which is the the optimum, you come to a conclusion of a recommendation and that recommendation needs to be explained not
[32:15] just to the merchant and the different departments across the merchant but it also needs to be explained to the supplier and the processor themselves. Yeah, that's amazing. you know uh just by looking at the output uh you could actually see that how there's a shifted from you need to have the analysts who
[32:32] build the workbook in the room to somewhat a selfserve kind of setup where you know the merchants can come and ask a question give their criteria and get the output that they want yeah and I I probably should just highlight on this on the slide here you
[32:48] can see at the bottom that these uh result data sets that's a routing uh example routting strategy. For example, you can see there there's about eight different segments. It shows about reg for regulated transactions different amount buckets zero and one 0 to 10 for example. Those
[33:04] are eight different structures that could be sent to. There's 11 different networks that could be rooted to. If you think about that now, that's 11 different places that a network could have been rooted to in any one of those places. So that means that there's 11 factorial choices per segment. So 11
[33:22] factorial is around quick math. Anybody? Nobody knows that. Okay. 40 millionish around there. So that means that there's 40 million for segment one, 40 million for segment two, 40 million for segment three. That's 40 million times by 8.
[33:41] I can't I can't quite get to that number. So it's a very big number of course and this extrapolates or you know exponentially increases the amount of these different segments that you increase. And in fact, this is a simplified example anyway because there's around 42 different network subtypes. So instead of 11. Yeah. Wow.
[33:57] Yes. A lot. Yeah. You can see see how complex it is, right? So okay, this is cool. So we have seen uh the the data layer. We have seen the complexity that you know solving and and what prom looks like and what the outputs are from prom, right? So let's take a quick look at the value itself.
[34:14] You know what you guys have got. So you've saved like 30 hours of manual work and the uh like 30 to 40 minutes is the solver runtime right you know for the scenario results and you've also seen 14% incremental savings so once you run prom you're able to identify uh like
[34:31] you know 14% more savings for on on the clients which leading to about 7.4 4 million annual savings for all your clients. That's really really impressive. Yeah, it's a good result. So, we're happy. Thank you. But one thing Jordan I think uh we uh we
[34:46] haven't really touched on and this is also an underrated one. It's not in this slide actually. Yeah, confidence. You know, users need to know that the results that they achieve here and what they're comparing is trustworthy. um they've obviously had to think about
[35:03] the different inputs that go into there, the constraints, the outputs, the comparison, but that history of how they got to that result is the important thing there for them. Yeah, makes sense. Yeah. So, we've seen a govern product uh
[35:18] gives you know everyone a common language, right? So, the value is actually three things here. So, you get faster analysis, you get better negotiation and of course repeatable governance. So we we have we have seen again the data foundation we got it right now we have also seen the
[35:33] workflows that are involved around it. Now let's look at the product interactive side of things where um CMSPI went ahead with lakebase right so this is an architectural decision uh you know that made the product experience work for CMSPI so um this is you know
[35:52] question to you Jordan like you know this is a data bricks native app where should the product state actually live yeah not within the lakehouse okay makes sense yeah so So the that is true. So lakebase owns the user and
[36:08] the contextual state um and lakehouse owns the analytical data right. So you have all your um app level interaction from scenarios drafts edits user configuration and let's say what the user is selecting around incentives that
[36:25] we saw. All of these need to be managed by something like a Postgress one, right? You know, so which follows CCRUD create yeah um update delete and then you also have the uh lakehouse side of things which is an old app analytical
[36:40] platform if which involves larger table scans doing complex analysis. these uh typically in the data tables fully governed in UC you use SQL warehouses for these kind analytics um and also the workflows that orchestrate the
[36:56] prominence prom itself right so that is a really important distinction here because the architecture principle that led CMSBI not to se choose separate platforms so that you know they can achieve both of these uh rather they
[37:11] just went ahead using data bricks and achieve both OLAP and OOLTP in a single place. Yeah. And so the question for us wasn't you know what's the best or what's the best OOLTB database in the world for example. It was what is the
[37:29] best OLTB database that meets our specific scenario which is to reduce the amount of friction as possible to introduce this transactional state layer. Yeah. So um this is a quick matrix of
[37:45] what CMSPI explored and uh you know this is just not a question into uh which OLTP is the best right Jordan this is more of why was lakebase the choice for them because it made things easy it actually fit their architectural pattern and this is
[38:02] something they would really wanted to go ahead with it's not just about the architectural tidiness right it's about how fast CMSPI can build using data bricks
[38:18] Yeah. So, um, if you could help me tell about the product velocity choice, John. Yeah. Uh, I'm looking at time, so I'm going to speed through this a bit quickly, but, uh, you know, why we chose Lakes Base in the first place here. One of probably the main important reasons was this aspect of being able to have governance and and sync of the data between Lakehouse to Delta from Delta
[38:35] tables into the Postgress um, SQL tables for example there. Um, it allows us to build new proof of concepts way faster in in terms of the the UX workflows. If anyone's used the branching and restore features inside of the uh Postgress
[38:51] database here, the lakebase, that's such a powerful feature. You're trying to build different concepts but without having to amend any of your production data that can just be spun up. Obviously, it has the standard sort of Postgress compatibility that brings object relational mapping tools and a
[39:07] familiar sort of uh environment for web developers to to work in as well. Yeah. So, um yeah, let's let's actually summarize all of that and consolidate in a single slide, right? If you could walk us through what an end toend architecture currently in database looks
[39:23] like. Yeah. So, this is the main architecture slide of the presentation. So this is the one to take your photos of to start with. Um but this is a user signs into a web application on inside of the browser. They authenticate using entra.
[39:39] That means that the access and role-based access is managed by groups that syncs through to unity catalog. Of course, they draft and make a scenario inside of the web application which then syncs that or sends that transaction down into lakebase that then triggers an
[39:54] Azure API which sends a notification to sync the rec the new scenario into delta which means it goes into a delta table inside of there. When it adds that new row inside of the delta table that means that it then triggers the prom workflow. The problem workflow then contains the
[40:10] solver of course. The solver then runs and outputs those to delta tables as all the results. And those results are then served back to the users using DBSQL of course and you've obviously got delta sharing there. And the idea of delta sharing would be to be able to share the
[40:25] outputs and results from the solver to external merchants or potentially even to share incentives or other structures of information back to us as well to combine the solver. Yeah, I mean that's a reusable pattern right there, right? I know for anyone
[40:41] who's interested. So if um if you also look at the agentic extended agentic workflow, we are adding uh quite a lot of things again. So you have Genie, you have knowledge assistant which actually interacts with the structured and the unstructured um knowledge payments
[40:56] knowledge respectively and that powers the web application right. So the um the clients can come and ask question and like you know trying to understand how the optimization works. You have you have all these supported knowledge that you provide um through Genie and helping
[41:13] the uh clients get get the right answers on you know top of the data. So um why don't we pull it back in one message Jordan over to you for your final slide and then we can wrap it up. Yes. Um so data bricks gives us a
[41:28] flexible innovation stack. Um it's it's helped us produce our dynamic data fabric which is our normalization of that messy payment semantic data. It brought us into our first multivariate constrained optimization solver which is prom of course and latebase then is
[41:45] brought on to handle that transactional app state to make it a user centric product. Cool. Oh, we missed out workflows in the agent layer. But, uh, hopefully everybody recently just heard about that, so they know what it is. Uh, but thank you so much. Uh, I'm sorry
[42:01] that we've run over, but if you have any questions and you want to ask further about anything, we'll be around and you can take any questions. Thank you. Thank you.

Learn more about the Databricks Data and AI platform.

The information provided herein is for general informational purposes only and may not reflect the most current product capabilities or configurations.