Databricks AI/BI for Regulatory Reporting: The PowerBI Alternative
Summary
- CareSource, a nonprofit managing government healthcare for 2 million+ members across 12 states, rejected a lift-and-shift to PowerBI Paginated Reports and chose to keep regulatory reporting inside Databricks, eliminating duplicated governance, additional data movement, and extra licensing costs.
- The Databricks AI/BI solution delivered measurable results: reports now run in minutes instead of hours, the organization achieved $120,000 in cost avoidance, and adoption reached 98% without any user retraining.
- An AI-powered report accelerator cut implementation time per report from 52 to 23 hours, while Unity Catalog provided single-pane governance and a metadata-driven configuration approach handled the regulatory variations required across 12 states.
Databricks AI/BI for Regulatory Reporting: The PowerBI Alternative

Regulatory healthcare reporting faces a perfect storm: complex state mandates, strict audit requirements, and accelerating submission volumes. CareSource initially planned a standard lift-and-shift from legacy SSRS to PowerBI Paginated Reports. But that approach meant duplicated governance, additional data movement, new licensing costs, and split infrastructure. This case study reveals why a healthcare provider with 2 million members chose to keep reporting inside Databricks instead.
Learn how Databricks AI/BI eliminated the multi-tool stack while delivering measurable results: reports running in minutes instead of hours, 120,000 dollars in cost avoidance, and 98 percent user adoption without retraining. Discover the architecture principles enabling regulatory compliance: single governance through Unity Catalog, metadata-driven configuration for regulatory variations, and self-service access with observability. See how an AI-powered accelerator reduced report implementation from 52 to 23 hours per report.
🤝
Chapters
00:00Welcome and Introduction to AI/BI01:13CareSource: Managing Regulatory Reporting at Scale03:16Legacy SSRS to Modern Databricks Lakehouse04:37The Three Forces: State Mandates, Governance, Volume08:08Why Lift-and-Shift to PowerBI Fails11:10Rethinking Architecture: Reporting in Databricks12:17AI/BI Solution: Single Platform, Single Governance15:08Metadata-Driven Design: Core Models and Configuration22:51Impact: From Hours to Minutes, Cost Savings, User Adoption25:11Adoption Strategy: User Experience and Muscle Memory27:55Implementation: AI-Powered Report Accelerator32:46Future: Genie, Databricks Apps, and Roadmap
FAQs
Why did CareSource choose Databricks AI/BI over PowerBI for regulatory reporting?
CareSource evaluated a lift-and-shift from legacy SSRS to PowerBI Paginated Reports but found it would require duplicated governance, additional data movement, new licensing costs, and split infrastructure. Keeping reporting inside Databricks eliminated all of these issues while delivering better performance and a unified governance model.
What results did CareSource achieve with Databricks AI/BI?
CareSource reduced report run times from hours to minutes, achieved $120,000 in cost avoidance, and reached 98% user adoption without retraining. An AI-powered report accelerator also cut the time to implement each new report from 52 hours to 23 hours.
How does Databricks AI/BI handle regulatory reporting variations across multiple states?
CareSource implemented a metadata-driven configuration approach that centralizes core data models and applies state-specific regulatory variations through configuration rather than duplicating logic per state. Unity Catalog enforces governance and access control across all reports within the same platform.
What is CareSource's future roadmap after the AI/BI implementation?
CareSource plans to extend the platform with Genie for natural language querying of regulatory data and Databricks Apps for expanded self-service access, building on the governed foundation established through Unity Catalog and the AI/BI reporting layer.
Full transcript
[00:07] Well, thank you everybody for joining. We really appreciate everybody being here. It's very end of the conference, so thanks for sticking around for us. We're excited to cover uh this topic on AIBI on top of data bicks today. Um so I'm Cam Cross. I'm a partner with West Monroe. If you haven't heard of West Monroe, we're a mid-market consulting
[00:23] firm. Uh we are a data bicks gold partner. I've worked personally with data bricks since about 2018 and have worked with Caresource for about that long as well. So, a long-standing partnership. Um, today we're excited to talk a little bit about uh AIBI. Uh, I
[00:40] don't know about you guys, but I've seen a lot of like really cool talks this week. A lot of uh ontology and agents and AI. I haven't seen a whole lot on BI and kind of traditional reporting. So, I think this will be a good session. We'll get into uh the AIBI topic a little bit
[00:56] here. Um I'll just call out Roger uh and his team have had done an awesome job kind of reimagining you know reporting in his organization. Uh regulatory reporting is a really hard thing to solve for especially across multiple markets. So excited to tee up Roger and introduce uh the talk here.
[01:13] Thanks Kim. So a quick introduction here's what we're going to cover today. Three parts. First, the introduction to care source and why we're regulatory reporting in healthcare is uniquely painful to us. The second, the AIB solution itself, the
[01:31] architecture, the design principles, and the impact that it had. And third, what we learned along the way, the things that worked and the things that didn't, where we're taking it to the next level. So, let's get into it.
[01:49] Caresource is a national nonprofit organization residing in 12 states. Currently we have a little over two million members in our environment. The um focus of our program is on government
[02:07] healthcare. So, we are in Medicaid, we are in the dual eligible special needs program, we are in the federal marketplace, and as of this year, we are in the Triricare demo for our military families and veterans. The word nonprofit
[02:23] matters here because it shapes everything about how we make decisions. Our mission is to make things better, to make a lasting difference for our members lives. both in their health and their
[02:39] well-being. And that's not just a poster for us. It's a mission. The lens for that makes sure that everything flows through. Technology decisions have to go through it because every dollar we
[02:55] save on not duplicating structure goes back towards serving our members. So much of what we do is government sponsored. We live in a regulated environment as those of you that work with the government do.
[03:16] Here's our story. About six years ago, we were pivoting off a legacy SQL server environment. Um the standard stack that you're familiar with, SSIS for ETL movement, SSRS for reporting. And we were doing loads in daily, weekly
[03:34] and monthly cadences based on the ability of our infrastructure to support that activity. We made the switch to Microsoft Azure data bricks and
[03:49] we are in the process of loading now 365 days a year. We load the entire warehouse nightly. We are um currently sitting at 447 teabytes of
[04:06] information in our production environment. We use the controlm tool to coordinate Azure data factory pipelines to kick off Azure data bricks notebooks. and our environment follows the typical
[04:22] medallion model that they recommend from bronze to gold standard. So that's the foundation, a modern lakehouse governed, refreshed daily and keep that picture in mind as we go
[04:37] forward. Now there are three forces at play. The first is statemandated complexity. We deal with many state agencies and each one has its own unique cadences,
[04:52] its own quirks. Approximately 415 regulatory or contractual reports exist in our environment. As of last Friday, our states are being more prescriptive about what we deliver, when we deliver,
[05:10] and how we deliver it. Some are even moving to fixed temperature that we're not allowed to change. They have inbuilt formulas that we have to populate data into. So the same standard report
[05:25] can mean many different things or have many different variations. There's no such thing as build it once in our environment. The second force is our audit and governance control. Because it's regulatory reporting, we have to be able
[05:41] to track lineage all the way from source to delivery. And we have to be able to recreate that activity, not just the asis today, but the asis that exists um as of a point in time.
[05:56] So if a regulator comes to us and says, I need to see what that data looked like 18 months ago, saying I think it was this is not going to fly.
[06:12] The third force is the volume and the frequency that is continuing to climb. We are getting more reports more often with less time to deliver.
[06:28] The government moves slowly until it doesn't. So what we are seeing is year-over-year submissions are growing. The time frames in which to deliver are shortening. We're getting requests in as little as seven days from request to delivery.
[06:45] Now, when you have that kind of environment, your architecture either bends or it breaks. Now, when we face this problem, most health care organizations follow
[07:01] the same migration path and so did we at first. Where we started was the aging infrastructure, hundreds of pageionated reports, tribal knowledge built into the stored procedures.
[07:17] And I want to be very clear that legacy world actually was worse than it looks like on the slide. Our code base was unique to each report. We duplicated code for the submission and for the
[07:34] monitoring version of the report. We had hard-coded values that were buried deep inside stored procedures. We had queries that would run for hours only to time out at the very end without delivering anything.
[07:52] And if your users are like ours, they would take that data and they would immediately dump it to Excel and manipulate it. So where were we headed? We were going to follow the standard playbook. We were
[08:08] going to lift and shift from SSRS to PowerBI pageentated reporting. We were going to ensure that the look and feel that our business users were used to remained so as to reduce the abrasion and the risk
[08:25] to our tooling. But here's the thing we kept running into the path created brand new problems. Lift and shift actually
[08:40] created problems or broke things that we had already solved in our modern data platform. Think about what Kesource was about to inherit. Extra data movement, extra security, extra governance,
[08:58] latency on an additional refresh to get to our reporting layer. And remember that traceability, that auditability from source, every data hop instills risk.
[09:14] We have to be able to trace back to source. So with added data movement, we had another layer that we had to be able to track and follow. We were adding another layer of governance, another layer of quality
[09:30] testing, another layer of model building. and we actually were going to have to have additional security entitlements just to make sure that things stayed in sync. In the long term, it meant two stacks to operate within. And that's two points of failure that
[09:46] were available to us. Most concerning, our source of truth would now be split between our data layer and our presentation layer.
[10:03] In a world where strict reproducibility is a requirement, that's dangerous.
[10:23] We also risk instituting gaps between our governed objects and our reporting environment, our display layer. Not to mention the additional total cost of ownership that this was going to cause for us. We would have needed to increase our PowerBI fabric capacity to
[10:40] handle the new workload. We would have had two testing environments, two areas to govern and costs per run in the PowerBI environment. The other thing that we discovered is
[10:55] that the PowerBI doesn't really have a way to deploy in an automated fashion through pipelines yet. Meaning that activity would be manual.
[11:10] And that brings us to the moment that changed our architecture. Somebody on the team, and I really really wish that I could take credit for this, asked a deceptively simple question. What if reporting never left data bricks?
[11:27] That's it. Sometimes the simplest question can have the most profound impact on changing the perspective. We had been anchored on which BI tool do we move to? So much so that we never seriously
[11:43] considered or asked whether we needed to do that at all. The data already existed in data bricks. It was already governed in data bricks. Why were we so determined to carry it
[11:59] somewhere else just to display it? That question is what took us into part two. So let's talk about the architecture rethink and what it means to keep
[12:17] reporting on site of data bricks. Three things changed. First, reports now run where the data lives. No more shipping data out to a reporting layer. Second, Unity catalog governs
[12:34] everything. We don't need to worry about additional layers or two types of governance support. And third, AIBI replaces pageentated reports as our reporting layer entirely. So picture the before data flowed from
[12:51] our sources into the datab brick lakehouse then out to fabric and powerbi pageionated and then to our end users or our regulators. two distinct sets of controls, two refresh cycles,
[13:09] more data movement, more models to maintain. Now let's think about the after the data sources are in the data bricks lakehouse. They flow into AIBI and they go straight to our analysts or our
[13:24] regulators. We can use our existing warehouse capacity. We keep one set of entitlement and we retain familiar business functionality to our users. In the end,
[13:44] there were secondary benefits to doing this as well. With AIBI on a serverless warehouse, we have scalability. We can scale capacity as we need to. So that data latency challenge goes away.
[14:02] We can adjust that as needed with a few clicks inside of the administration layer. We also gain the ability to tag workloads. So for those that are concerned about PinOps or cost monitoring by tagging the workload, we can see where our expenses
[14:19] are actually coming from in our reporting layer. And with the DABS, the datab brick asset bundles deployment pipeline capability, we had a way to put our reporting in a repo and automatically deploy that through our Azure DevOps pipelines into
[14:35] our production environment and our taxonomy folders. We'll do a little bit more conversation on that in a minute. So, we walked through the Medallion architecture. We talked about the challenges that we
[14:51] had in the legacy environment with the sprawl and the duplicated code. So we wanted to change a few things when we moved into the AIBI platform. Our core models are the foundation. Remember when I talked about there were multiple
[15:08] variations of the same code to run the same report. We now have one set of code that runs both the submission and the monitoring side of the activity. We now use metadata tables for configuration of our reports. All of
[15:23] those parametric values that used to go into a piece of code itself now goes into a metadata table and those are called as needed. What that does for us is eliminates the full software development life cycle we would need to make a change. In the legacy world, if
[15:41] we wanted to add an additional CPT code, for example, to a report, it had to go through a full stack development cycle. With our new structure, all we have to do is do an insert statement on our metadata table, and we have that value in production within minutes.
[15:59] We want to change the configuration, not the code base. The additional benefit is it lets us reuse code. So if we have that similar regulatory report, we can create a separate configuration for it to use for
[16:14] a separate state. Unity provides the governance for all of this. It also helps us as developers because it lets us see the data freshness in the environment through the catalog tool. It gives us traceability to the objects
[16:31] that exist. So if we have to make a change to a report, we can see all of the users that have been utilizing it to let them know. We have the ability to see which dashboards are impacted if for
[16:46] example we need to change a source object. That Unity catalog experience helps us to move faster.
[17:01] We are also using custom code that we wrote in Python for what we call our metadata extract tool. So the same views that underpin our regulatory reports can be used to call from our metadata extract tool to create an extract to a data bricks volume for the purposes of
[17:18] sending that outbound to an internal network storage location or to an SFTP folder to go outbound to an external organization. Underneath all of this, there were several core principles and I find it
[17:34] useful to look at these from both a business and an IT lens. The first principle is simplicity. We have a single platform both for querying and reporting our data out. Now we went from two interfaces down to one. We unified
[17:52] submittal reports and monitoring reports into the same object and AIBI on two separate tabs. For the business, that means they have one tool to go to. With the embedded search and column
[18:08] reordering capability that exists inside of Azure data bricks, there is less need for them to excfiltrate data. Now, I will be honest, they're still going to exfiltrate data. they're going to push it to Excel when they want to, but that doesn't say we didn't give them
[18:25] the opportunity to stay within the tool. They can view data for a good portion of the day. That load that we run daily really only impacts reporting for two hours out of the day. So, 22 hours out of the day, our business users can
[18:40] access their reporting as they need to. Lets them work when they want to work. And in the event we have contention, if the warehouse is overloaded, as I said before, within a few clicks,
[18:56] we can scale that up, remove that contention, and get the business owners back to functionality for it. This meant a reduction in our code base. We no longer had multiple sets of code to maintain.
[19:14] We have the same code in the same AIBI using the same parameters to deliver both a submitt and a monitoring version of a report. So there's less work for us on a maintenance side.
[19:30] The second principle is self-service. We want to enable our users to be able to access the data when they need it on their own without it interference. In the world before, request would end up in an IT work queue and it would end
[19:46] up being a blocker to getting things through. Now with the self-service capability, the user goes out to a folder where all of their data or all of their reporting is stored and they can run it as they want. With the inbuilt
[20:04] capability of AIBI, they have the opportunity to schedule that to run every morning so that it's ready for them when they come in. They also have the ability to schedule extracts out of the tooling as they need.
[20:23] From an IT side, that means that we no longer have a work cue of existing reporting that we are waiting to deliver to someone. We are no longer the roadblock. The third principle we have is governed distribution. And this is where people sometimes get a bit nervous. If you give
[20:40] the business users the control, don't you lose the control? And we found that the opposite is true. We actually have better controls letting them run it. Due to the observability capabilities inside of data bricks, the
[20:57] system tables, we can see every execution, who ran it, when they ran it, and what the results of that were. So we have better governance, not less.
[21:15] That leads naturally into a single pane versus multi-pane comparison. By keeping governance on the report and the reporting is where the platform pays off the most. Picture a multiplatform
[21:31] environment, permission sprawl between two systems, running sync jobs back and forth between data bricks and whatever BI tool or pageionated tool you choose. You end up having duplicate policies that have
[21:46] to be kept in lock step. And inevitably, no matter how tight your controls are, you're going to end up with drift. And when you have drift in a regulatory environment that adds risk and that is dangerous. Now the single platform model we have
[22:03] unity catalog as the only front door. One policy applied everywhere. One audit trail in one place. We have a cleaner experience for our end users. We have fewer controls that are
[22:19] necessary. Everything is maintained via a single intra ID group with fewer processes. We have lower discrepancies because
[22:35] there is only one environment and a much lower barrier to entry for our business. So let's talk about what actually changed. I want to focus on the tangible benefits. And here I'll tell you upfront
[22:51] that these are conservative estimates. Your mileage may vary. But even with conservative estimates, the numbers are pretty striking. We didn't have to have any retraining done because we were using some new tool or some new
[23:08] technology. Our end users simply needed to know data brick SQL which they already did. So three areas. The first is runtime. Our heaviest reports on our aging infrastructure would run for hours on
[23:23] data bricks in AIBI. They can run in minutes. And again, we have the ability to scale that as needed to help improve runtime performance. Those places where our onprim infrastructure tapped out, AIBI is able
[23:40] to perform well. All told, we have saved approximately 3,700 hours by updating queries in code optimization and the benefits of using Azure data bricks serverless compute.
[23:57] Item number two is the cost and the complexity. We eliminated the PowerBI pageentated capacity growth we were going to need entirely. We removed duplicate governance requirements. We have one platform to run, one bill to
[24:15] pay. To us, that's a minimum of $120,000 a year in savings that we don't have to put out. And again, that's a conservative figure. The final impact area is adoption. And
[24:32] this is the one that I'm most proud of. Business users now access directly in data bricks the reporting that they used to get off of SSRS. They are more active and more engaged than they were ever before. 98% of our
[24:50] legacy SSRS users are using PowerBI I'm sorry using AIBI reporting for regulatory activities. We didn't leave people behind. We brought him along with us.
[25:11] And that's the perfect bridge into part three because adoption didn't just happen, we had to earn it. So what did we learn? Let's start with the adoption challenge because moving SSRS users onto data bricks meant
[25:26] changing how they work and that's never trivial. Think about what our users were used to an SSRS reporting environment. You click a folder, you find your report, you have familiar parameter values inside that
[25:44] you use. Decades of muscle memory. For many of our analysts and business users, this is the only tool that they were familiar with their entire career. You don't overwrite that muscle memory
[26:00] overnight. And you shouldn't try to bulldoze through it either. So what we built for them was a datab bricks business interface that respected that muscle memory. The same tasks that they were used to with a new flow and a
[26:16] slightly new pattern. The environment is fast to query, fast to analyze, and more concretely where SSRS gave them two or more reports named differently and potentially in different folder structures. Today they get a
[26:32] single item or a single taxonomy folder that they can favorite which allows them to see all of the reporting that they have access to. So the moment that they log in, the report is right there front and center for them to use.
[26:48] We kept the parameters similar. The extract capabilities are the same, the ones that they relied on anyway, like extracting to Excel. We deliberately minimized the what changed surface so that they could focus on the benefits and not the disruption.
[27:07] And we're not done yet. Looking ahead, we're looking to leverage attribute-based access control and the Genie chat or now Genie1 functionality so that we can give users a single place to land. They can type in plain language
[27:23] what they're looking for and it will automatically start filtering down to those objects that they have access to to be able to run. We're making it easier, not harder as we go forward.
[27:39] Now, a huge part of why we could move so fast in this conversion is the accelerator our partners, West Monroe, built for us. Cam. Awesome. Yeah. So, we're going to talk a
[27:55] little bit about like how we implemented these reports quickly as well. So, Roger uh mentioned that there's obviously a lot of kind of technical reasons and functional reasons why AIB is a strong choice. One of the, you know, delivery um components that's really was
[28:11] interesting to us is it's uh it's all JSON behind the scenes. That's all code. So when this program started it was a large team. I think 20 plus um individuals were building out PowerBI reports. So we kind of said like hey uh everybody's moving to you know using AI
[28:27] for development. AIBI is JSON under the hood. Can we use AI and speed up this process right? But the car source environment is very challenging. Uh as Roger mentioned we're moving from a old data model that nobody really liked to this new indust industry standard data
[28:43] model. a data model was much more complex. So it's not a onetoone, it was not a lift and shift type of conversion. It's much more complex than that. So what we had to work through was kind of figuring out how do we do that mapping from the old fields into the new fields. One thing that Roger and his team had
[28:58] done really well is they documented a lot of the legacy code base and we had a lot of documentation as well from the you know new industry standard data model. So we were able to build essentially a mapping between the two using AI not to completely aug or automate the process but to augment the
[29:14] developer. So if the developer had a field in the old SSRS report they maybe had five options for that same field in the new data model. We wanted to give those developers the options to pick kind of the right uh piece. I think the um essentially what we built together
[29:31] was we ended up calling Intellio hopper which is an agentic transformation tool. We were using lang chain right there in data bricks notebooks to speed up the mapping process and then generate the JSON for the AIPI report. Uh and Roger's team was nice enough to let us do kind
[29:47] of a time study as part of the effort to say hey like can how how fast can we really go on this? We're going to share a few of the results here in a minute, but um I think big picture we saw just huge efficiencies and we shifted from essentially what was a 20 person team, 20 plus person team down to a fivep
[30:02] person team that was really using AI to speed up the process. So huge efficiency savings not just in the platform and the cost savings but also kind of fundamentally in how that work was delivered. So 400 reports um originally would have taken you know about twice as long to implement across the uh PowerBI
[30:20] plan. uh big picture maybe just kind of one thing to hit on that we learned from this process in using AF for this delivery um is that really you really have to focus on that last mile so in this case especially regulatory reporting we couldn't just you know fire this into chat GPD and hope for the best
[30:36] right you really had to make sure this was right so we worked to not just you know create an output and you know have the QA's test it right we worked on building a workflow where the developer could experience and kind of see the different steps in the process could make some their some of their own decisions on what was right and wrong.
[30:53] Um, and we found that that last mile really focusing on helping that developer experience was really helpful to speed up the development effort and it also helped our business owners in understanding because it it spit out pseudo code. It was it was an understandable English of the activities
[31:08] that were happening so we could get through that review cycle faster. So I'm going to let this quote stand for the whole journey here. The last clause
[31:24] is the entire thesis of this talk. We didn't get faster despite staying on data bricks. We got faster because we stayed on data bricks. The fresher data, the tighter governance, the speed at which we are able to deliver are all
[31:40] downstream of one decision. Don't move your reporting away from your data. So as uh Cam said uh while reporting went from hours to minutes, we also did a time study and we saw improvements on
[31:56] our development as well. Uh we uh across five roles did a comparison and you can read the roles that are on the screen. I'm not going to read them to you. Our base case doing this the oldfashioned way through the layers was about 52
[32:12] hours per report using the AI hopper tool. We were able to bring that figure down to 30 I'm sorry 23 hours per report. So we more than half our effort and we did
[32:29] it while strengthening the review process not weakening it. So where do we go from here? We have three different directions and we're generally excited about all of them. First, AIBI unlocked. When we started, AIBI was in its infancy. Um, they've
[32:46] added a lot of new capabilities and functionalities and we are excited to try all of them. We are looking to expand our footprint even further. Uh, we're looking to build more semantic and metric view models to help fuel that
[33:01] expansion. Second, we're looking to utilize Genie at scale. We're rolling Genie out across the entirety of Cares Source right now. This pushes self-service reporting even further by making conversational data access available to our end users. The
[33:18] goal is simple for us. We want our users to be able to use natural language to ask questions of their data. And third, we're looking at data bricks apps. We imagine a
[33:35] environment where we can build bespoke appbased activities or workflows to replace reporting as workflow that we have today. We're thinking about adding cues and alerts, a true single pane of glass
[33:51] for our business users. If you have an opportunity, go to the data bricks expo and go to the health and the life sciences display. They have working examples of what we're thinking about building for our organization.
[34:13] So, let me bring it home. We started with a problem that every healthcare organization in the room recognizes. regulatory reporting that's complex, heavily audited, growing in volume, and shrinking in time for delivery. We almost solved it the standard way with a lift and shift to a separate reporting
[34:29] tool. And then one simple question changed the entire perspective and sent us down a completely different path. The results, our heaviest reports went from hours to minutes. 3,700 hours of development time saved, at least $120,000
[34:45] in averted cost, and 98% of our users active on our system. We have the same data, the same platform, the same skill set used in the way it was meant to be used. If you take
[35:01] away one transferable lesson from all of this, it's that this the instinct to move your reporting away from your data layer is worth challenging hard. The data already lives in data bricks.
[35:17] The governance already lives in data bricks. Sometimes the best architectural decision you can make is not what tool to add but where you can subtract. I thank you for your time today. Cam and I are available to answer any questions
[35:32] that you may have. Appreciate your time.
Learn more about the Databricks Data and AI platform.
The information provided herein is for general informational purposes only and may not reflect the most current product capabilities or configurations.