Large-Scale Data Platform Modernization: Johns Hopkins and Databricks on AWS GovCloud
Summary
- Johns Hopkins Health Plans migrated its legacy SQL Server data warehouse to Databricks on AWS GovCloud, reducing full-load pipeline runtimes from 22+ hours to under one hour while maintaining FedRAMP compliance for sensitive armed forces member data.
- A metadata-driven ingestion framework built with CitiusTech reduced transformation logic by 60% and enables reusable, modular pipelines across the medallion architecture's bronze, silver, and gold layers, replacing a legacy environment with 400+ hard-coded transformations.
- Unity Catalog provides fine-grained governance and PHI masking across claims, membership, and provider data for a multi-line payer operating Medicaid, commercial, and military programs under strict healthcare regulations.
Large-Scale Data Platform Modernization: Johns Hopkins and Databricks on AWS GovCloud

Johns Hopkins Health Plans is modernizing its legacy SQL Server data warehouse into a cloud-native analytics platform on Databricks and AWS GovCloud. This multi-year transformation reduced full-load pipeline runtimes from 22+ hours to under an hour while maintaining FedRAMP compliance and governance controls critical for armed forces data.
Learn how Johns Hopkins and CitiusTech built a metadata-driven ingestion framework across the medallion architecture, reducing transformation logic by 60% and enabling reusable, modular pipelines. The platform unifies diverse healthcare data sources, implements standardized data quality checks, and leverages Unity Catalog for fine-grained governance and PHI masking across claims, membership, and provider data.
🤝
Chapters
00:00Johns Hopkins Health Plans Modernization02:00Payer Business Context and Data Challenges03:37Legacy EDW Limitations and Modernization Decision04:58Why Databricks: Unified Platform Architecture07:05Metadata-Driven Ingestion Framework08:23Medallion Architecture: Bronze, Silver, Gold Layers09:26Semantic Layer with Databricks Metric Store10:30Unity Catalog: Governance and FedRAMP Compliance14:19Data Quality Engine and Reference Standardization15:21Legacy System Complexity: 400+ Transformations16:24Plug-and-Play Architecture and Development Savings17:45Transformation Rule Management and Schema Evolution19:36End-to-End Pipeline and Observability21:24Future Vision: Data Product Enterprise
FAQs
Why did Johns Hopkins Health Plans choose Databricks on AWS GovCloud?
Johns Hopkins Health Plans selected Databricks on AWS GovCloud because it provides FedRAMP compliance for handling sensitive armed forces member data alongside a unified platform that supports analytics, data engineering, and conversational AI workloads. The platform's fine-grained governance through Unity Catalog and support for PHI masking were critical requirements for a regulated healthcare payer.
What is a metadata-driven ingestion framework and how did Johns Hopkins build one?
A metadata-driven ingestion framework uses configuration tables rather than hard-coded logic to control how data is ingested, transformed, and validated, making pipelines reusable across different data sources. Johns Hopkins and CitiusTech built this framework on Databricks to reduce transformation logic by 60% and support modular pipeline development, replacing a legacy environment with over 400 individual transformations.
How did the Databricks migration improve pipeline performance for Johns Hopkins?
The migration from a legacy SQL Server data warehouse to Databricks on AWS GovCloud reduced full-load pipeline runtimes from over 22 hours to under one hour. This improvement was enabled by the medallion architecture's modular design, distributed Spark compute, and the metadata-driven framework eliminating redundant transformation logic from the legacy system.
How does Johns Hopkins handle PHI data governance on Databricks?
Unity Catalog provides fine-grained access controls and PHI masking capabilities that Johns Hopkins uses to govern sensitive claims, membership, and provider data across multiple lines of business including Medicaid, commercial, and military programs. Governance controls are integrated into the medallion architecture pipelines so that data quality and compliance requirements are enforced at each layer of the platform.
Full transcript
[00:08] All right. Welcome everyone to the large-scale member and claims data platform modernization um with Databricks on AWS GovCloud. I am Tyler Sipe. I'm the director of data solutions and the data governance co-chair at Johns Hopkins Health Plan. Um then I'll kick it to Riz. Hi guys. Thank you all for joining
[00:24] today. My name is Riz Qazi. I'm part of the program leadership team at CTS Tech Healthcare Consulting supporting Johns Hopkins data modernization initiative. Um so just to set the agenda before we dive in. Um So today in that next 35 to 40 minutes,
[00:39] we'll leave little bit of time for the Q&A, but we're going to go through the Johns Hopkins journey. Uh we'll talk about the current challenges, current landscape, and then really look into, you know, from why Johns Hopkins decided to go with Databricks as a tech stack on AWS GovCloud. And then future uh future
[00:55] vision and then end strategy. And I don't like to use the word end goal because transformation is a journey that's constantly evolving and you need to continue to be on that path. Um with that, let's dive right into it. Uh Well, actually uh if you go back a little bit. So
[01:11] today, just to set the expectation also, we'll also look at the overall architecture, but we're not going to dive too deep into it. But the core of the engine for this platform is really that scalable and modularized component. So we will focus on metadata-driven ingestion framework. We'll also look at the data quality engine that we built.
[01:27] And really um the scalable uh standard reference structures that we have embedded into the pipelines so that we can enrich the data. And then towards the end, we'll also talk about the data quality in general, how it drives the insights and healthcare outcomes. So from taking the legacy ETLs all the way
[01:44] through the transform data, uh curated governed data, and insights built on AI for BI, you know, leveraging Genie and other other conversational agents and solutions. So that's that's our agenda for today. Uh with that, Tyler, let's kick it off. So, before we go into the overview, just
[02:00] to get a lay of the the land of the room a little bit, how many folks are Data Bricks customers currently? Wow. Okay. How many are payers or in the payer space in health care? Okay. So, at Johns Hopkins, we're obviously
[02:16] we are a payer and just to give a little bit of a background, we have four lines of business, five, but one is very small. So, we have a Medicaid line of business, that's our largest. We have an EHP commercial line of business. We have a Medicare Advantage line of business, and then we have, which puts us in the wonderful
[02:31] as you see, AWS GovCloud, our TRICARE DoD USFHP line of business, which keeps me very busy wondering what we can and cannot do on Data Bricks, right? But, they've been a great partner for that. Um So, like many health care orgs, as you all know, right? Our data ecosystem has
[02:48] evolved over not 1 year, 5 years, 10 years, over decades, right? Our current data model of the data warehouse, um and I have to be careful over these next things. My team is sitting in the audience in the front here. So, I have to go over our current struggles, but I have to be careful. Um you know, you can tell that it was
[03:05] built in mind that we would never change a claim system, that we would never change a provider credentialing system, we would never change any system. So, it's got a little bit of rigidness to it, right? Um At the end of the day, it served us well. It's does it has done what we
[03:21] needed it to do, but we squeezed the juice out of that lemon. And that's where the modernization has come in and our move to Data Bricks. So, with the increased demand I think we're all probably facing around analytics, regulatory reporting, and then the pressure of being able to do advanced
[03:37] analytics and AI capabilities, right? It just wasn't cutting it to scale anymore. Um and that's where we brought in Cityus Tech and Riz and our partnership to go on a 3-year modernization journey. We're currently going through year two.
[03:53] Um so, we'll talk about some of the things in in in this session today, some of the accelerators we worked on with City National Tech and my teams to get us to build out the foundational layers, right? This is probably one of the only sessions that you guys will see and have that is not heavy AI focused. Um that
[04:10] was important to me because to do all those capabilities, right? The foundational layers have to be there. If you don't take the time to do that, um it it's it's just going to cause inherent problems. Now, the other thing that I want to equip you guys with that I hope that you can leave here if you're in the middle of a journey
[04:26] similar to us or newer to it, right? As you get to those capabilities, you have to show incremental value, right? And and along the way, right? You guys will see some of the things that we're going to talk about, you know, your CFOs of the world, right? What did I pay for? Right? These are the things that we're
[04:41] going to go over as you get to the advanced AI capabilities that don't which don't happen, going to tell you in 1 to 2 months, right? Or 3 months or 6 months. You can do something in 6 months, right? But but it's going to give you guys, hopefully, the information and the the you know, tool set there to be able to to to give that
[04:58] message, right? That you guys made the right solution and the right decision. So, why Databricks? That answer, honestly, is pretty simple. You know, it's it's the one-stop shop for us for a unified platform for engineering, governance, and advanced
[05:13] analytics. Um it's it does everything that we need to it to do, and it gave me the ability to take our, you know, tech stack, which on premises is not very lean, and we were able to lean that out. So, when I say that we're moving to the Databricks platform, we are all in,
[05:28] right? We do not have a lot of secondary tooling, um you know, outside of some CDC injection, which will probably be replaced with new Databricks functionality that wasn't available when we started our implementation, right? That's always the big and the fun of it is when you start your first year,
[05:44] there's going to be something, I promise you, within that first year that you needed when you started, right? Um Um so so we're working through those those those uh you know, various things now. And you know, for our our um let me go.
[06:01] The end goal vision is really simple. I'm not going to read everything on this slide, but right, we at the end of the day as as data folks on my team knows, we're customer service, right? Our job is to solve business problems and to enable data for our business users,
[06:17] right? It doesn't matter what we think, right? And and and how good we think we're doing. It matters that our customer feedback is positive, right? And and and those are consumers of data. So at the end of the day, we're moving to a data product-centric enterprise, right? That those products are solving
[06:32] real business problems instead of just serving up data in semantic models that may or may not be used, right? And one of the big thing that is happening throughout that is the governance of it and the business engagement. We're you know, we're not just building, we're building with purpose. Um so I think that's
[06:49] the end goal for us and and what Databricks gives us is that unified platform. Um from there, I will kick it over to Riz. Thank you.
[07:05] All right. So as Tyler put the context in place, right? We There There are legacy challenges with tightly coupled pipelines and sometimes the performance can can be a concern. Sometimes overall cost and how integrate integrated you are with as you try to bring in additional sources as requirements
[07:20] change, this can add a lot of complexity from code maintenance and portability perspective. So from a overall architecture perspective, how we approach this modernization was we needed to do something different, right? So we didn't look at it as one big migration. We We kind of looked at it as smaller chunks of easy solution
[07:36] tenants that we can build to scale. Um so first um really from the solution tenant's perspective, if you look at it the ingestion side of it, right? So, from an ingestion perspective, we built the metadata ingestion framework. Um it's really based on
[07:52] patterns, so you're bringing in files, you're bringing in from SQL Server, DB2, Oracle, whatever the sources may be. As Tyler mentioned, we are currently using AWS DMS for CDC, and in the future we'll try to look at the Databricks capability as well, so that we can completely be independent of the cloud um uh
[08:07] provider. Really in the In the data repository side, we use the best practices when it comes to the medallion architecture. So, you guys are Databricks consumers, you already are familiar with it. Uh when it comes to the raw layer, you want to keep the source fidelity. So, you want to maintain your source structures, but at
[08:23] the same time, silver is where you have a lot of the business transformation and um the logic, and you drive standardization and curation across your layers. And then you can curate the business-ready data for the gold layer. So, we built that medallion architecture, and we leverage a lot of the foundational
[08:39] models, but enhance them from enterprise-grade, you know, data design patterns for our gold layer. At the end of the day, the whole point of uh you know, re- going on to the Databricks was the unified platform, but at the same time, improving the performance. I think Tyler, you hinted
[08:54] it right. Uh today it takes about 22 hours to complete the pipeline end-to-end. With with And we'll talk about the results, too, but I I just want to mention it that within Databricks, the first injection from claims and membership, we brought it down to within an hour. So, that was a huge achievement, and uh we were able to achieve that because we looked at it
[09:11] every individual component, and made it modular, so it didn't matter what the sources were. We weren't building hundreds of pipelines bringing in various different sources. We had one pipeline, and everything was driven off the configurations in metadata tables, rather than code. So, if you have to
[09:26] bring in an a additional source, we're not opening up writing notebooks or writing creating another pipeline to bring that data in. We're configuring those entries in the metadata table, and we'll talk about that a little bit today. Um and then at the same time from a semantic layer perspective, you know, taking in the existing semantic layers
[09:42] but leveraging Databricks metric use to build in the semantic models at the use case level, it gives you a lot more flexibility as you scale up and looks at across the business domains. You have a lot of business use cases for claims life cycles or performance provider benchmarking. You can really build your semantic models catered to those and
[09:58] that can serve up your genie spaces and in the future genetic workflows. So, that that really also help as we re-imagine that. From a data delivery perspective, we we are trying to be as much native as possible within Databricks. Databricks has come a long way in terms of the visualizations, but
[10:13] at the same time they have Lakehouse Connect that that would be launched in GovCloud in Q3. So, we'll be looking at that to be able to use that for not only member portal integrations you through the APIs but also Power BI to be able to get that. But, you want to keep the semantic modeling within Databricks. What that allows is that at least you
[10:30] have all of your Unity Catalog governance and uh uh audibility, everything within Databricks. It doesn't matter what interfaces are using it, what products are consuming the data, they're doing it in a consistent manner. And then as I talked about Unity Catalog, so that allowed us the
[10:45] foundation across our our pipelines from access management perspective, user governance. So, we created entitlements uh as Sridhar talked about um Johns Hopkins deal with a lot of armed uh forces data, Department of Defense. So, it is FedRAMP compliant. So, for in order to achieve that, within Unity
[11:01] Catalog you can ensure that PHI masking functions can be implemented with entitlements. So, based on who you are in the business area, what uh use cases you're consuming Databricks for, you have uh appropriate access across the layers uh so that we are able to use that within uh Unity Catalog as well.
[11:18] So, here's the overall architecture. Um I know it can be a little bit overwhelming, so we'll go a little bit left to right and we're not going to spend too much time on it because we want to focus on data ingestion framework and um the data quality engine and uh the standardization. But, on the left-hand side you have
[11:34] multitude of sources, right? In any pair you you guys already know you have hundreds of sources for claims, membership provider, labs data, pharmacy data. Just like that at Johns Hopkins we just listed some of those. We have multiple adjudication systems that are handling different lines of businesses at Johns Hop- Hopkins. So by building
[11:51] this metadata ingestion framework, we can codify the configuration as part of the metadata rather than building separate pipelines for each of these ingestions uh sources. The the core of the you know, the meat of the process is in the middle where we are really processing through the
[12:07] metadata driven framework. Um and what that means is that we're not you know, when you're coming in bringing claims data, it doesn't matter how you want to drive the inpatient processing for one source system versus another. We want to unify that uh business rules transformations within the silver layer
[12:24] across the source system. So that curation and standardization um that happens from uh metadata perspective, it allows us to be able to just simply say, it doesn't matter what the source is, what is the source list defined that, what are the key attributes, how I'm going to drive this execution block, and
[12:39] uh it's really PySpark based, so it's what is the execution order and dependency, and then it kicks through it, right? So if you if tomorrow the requirements changes and recently it happens I want to use a use case real use case. Uh our data governance council came back with a different rule set to
[12:54] drive the place of service at the header level. And in traditional modeling and from a development perspective, I I know our architects are laughing here, but when that happens traditionally, then you have to really go back, go through the SDLC, and you have to open up code, figure out where you need to change the SQL logic or the Python code to figure
[13:11] that out. But for us, these were metadata entries. So all we had to do was just go into that after the design and just going and making sure that we took those requirements, we added and updated those entries, run up run the job again, and we were able to validate and see the place of orders derived accurately.
[13:28] Really has metadata tables No, these are Delta Delta live tables within Databricks. So, if yeah. So, we created a schema for that within within Databricks and again, everything is within parquet behind the scenes. But, in our framework, it goes through dynamically figures out everything is
[13:44] object-oriented. So, it just looks at it, okay, here's my primary key, business key. This is the logic I'm implementing. Here's the execution order. This is where I am from a context perspective. So, the pipeline gets that context as it kicks off and then it runs through those metadata entries configurations.
[14:01] Um, but that really is the power and we'll talk a little bit more about metadata and the value and we're always seeing in terms of the development savings through this framework. Um, and then, you know, we'll a little bit talk about the data quality engine. That's another thing, right? You can have the data across all the layers
[14:19] across all the source systems, but if it's not trusted or validated within the healthcare space, the business is not really going to rely on the data. So, how do we ensure that? We've built a CTS stack is also built an accelerator for data quality engine that we've implemented at Johns Hopkins with the collaboration with the Johns Hopkins
[14:34] team. And in that engine, it's also a rules-driven piece, so that you can define the rules that you want to implement, validations you want to read out at the bronze level, at the silver level, depending on because at the source level, you're not doing a lot of transformations. So, your validations
[14:50] are more around if the data is coming in is null, is it in the right format? So, more standard checks. But, then when you go through silver, there are a lot of real business use cases that you can validate. If the enrollment is coming in, but you have claims with the date of service where that member wasn't active, that is a data quality issue because it
[15:06] shouldn't happen in the source system. How do we surface those up? So, Tyler's going to talk a little bit about the data quality engine and how it notifies the leadership of these data anomalies and really ranks all of the data quality through the each of the layers. And then also, the key piece about the
[15:21] standardization reference data. So, today a lot of the reference data at Johns Hopkins comes from various different source systems. So, as you know, faster has reference data coming in, they are also ingesting it from Optum, but we are trying to standardize and add categorization and enriching that reference data. So, we'll talk
[15:37] about that architecture a little bit later as well. But, that also allows us to be able to drive semantics business semantic definitions and from a business perspective, if you think about it, you're not going to be asking what are my claims total dollars? That's a very base question, right? That's not what drives insights or
[15:54] analytics. So, what they really care about it for this particular case disease, you know, what are my population health and what are those numbers, right? So, they're going to go through that conversational piece and that's where Genie power is going to come in and we'll briefly talk about it towards the end as well. But, this
[16:09] reference data standardization and categorization will enrich all of the data in gold there. Um I'm going to continue to march towards. All right, so let's dive into it, right? Um I'm not going to repeat myself, but I just wanted to highlight key few points
[16:24] here. We had 400 plus transformations in the legacy EDW today. Um probably more, but more than 20 percent of the stored procedures were written up to handle a lot of this complex logic. What all of that does is that not only from code maintenance
[16:40] perspective, as things change, you have to go and go through the SDLC, but also from a performance perspective. So, imagine you're loading up claims and you have these 100 columns that you're building out fact tables and you know, detail tables, you really have to go through and call out these procedures to drive and curate the data. You have to
[16:57] drive financial amounts. You have to look at not covered amounts. You have to look at place of service as we talked about earlier. A lot of this would add time to your whole pipeline in terms of when the data comes into the system and it gets to the point where the data business areas can consume it in the way they want. Um so, we had to do something
[17:14] about that. We couldn't just do lift and shift and just move from war EDW platform over to Data Bricks. We had to really rethink about how we wanted to modernize it. At the end of the day, again, tomorrow if you have another source comes in, we don't want to be thinking and building new pipelines as I talked about it
[17:29] before. So, this also would enable that where it's just becomes an entry. So, there are some stats that I'll get to that, but at the end of the day, very high level, you know, it's all reusable. It's plug-and-play architecture, right? So, at the end of the day, you're just defining your sources, you're defining your business rules, you're defining the
[17:45] execution orders, and you're done. Um from an audibility and compliance perspective, everything is uh traced and auditable with batch IDs, and every transaction and the transformation that happens is tracked. Um from an operational efficiency perspective, not
[18:01] only from a maintenance, but we were able to save 60% of the development savings. We're already seeing the fruits of our labor um as we bring in more and more transformations. Uh from a development perspective, it's a really quick turnaround. Within We're talking about sprints to bring in an another source rather than an entire program in
[18:16] human really think through a new source system and bring that in. Um there is a transformation rule management that's the core behind this metadata engine framework. In that transformation rule management, as I talked about it, you you you configure all of these aspects of it, what validations you're applying, what are
[18:32] the target columns. And the way we build it, if let's suppose tomorrow and I'll use the places service example again. If we had to add another additional attribute in our uh gold air tables in terms of, you know, we want to be able to drive these things for the business, these KPIs, maybe the claims turnaround time.
[18:49] We don't have to go and update the schemas. We don't have to deploy databases again, right? In the traditional uh sense. In this scenario, the framework itself will figure out does this column exist? If it doesn't exist, it's going to expand the table definitions. And it's going to calculate the value based on metadata entry, and
[19:05] then it's going to expand your table. So, you're it's almost like a touchless system to continue to maintain your period tables. And again, it's built completely natively within the Databricks and it's by Spark. It's SQL oriented and Python based.
[19:20] And last point, DQ is integrated with MDIF. So, as those rules are called, we're also calling the separate engine from a data quality perspective and that engine takes care of validations, business rules validations, and other dedupe checks. And then it spits out a report and there is a dashboard on top
[19:36] of it. So, we'll talk a little bit more about that as well. So, here's the architecture. Again, not going to dive too much into it, but very high level if you can look at it from the sources as the data comes in and lands on S3. We really, you know, bring in through
[19:51] those configuration tables. We have this one pipeline. So, you have this one process that brings the data into the bronze and looks at those rules and bring it in and then it there's one process that really looks at silver curation and business logic transformation across all these sources,
[20:07] right? So, we are not creating separate pipelines and notebooks. We are able to save not only to build those hundreds of notebooks, but also all these transformations are in one place. So, in the future, as you're maintaining the architecture, as you're expanding on it, there's whole visibility in terms of what really happens because nothing is
[20:22] hardcoded in the board. That was the key point for us. We didn't want to hardcode anything. Whether it's standard functions we are calling to curate and cleanse the data, whether it's business transformation rules that we are we are calling, all of these are through those standardized notebooks and with fallout schemas and audit checks and um
[20:39] um a call out to the dashboards and there's a notification and event going back to the developers, to the ops, and even to the business leadership as the pipelines complete. And at the end of the day, you know, our curated gold there really have the trusted data, hopefully with the validation and risk scores that we also
[20:54] put out through this pipeline. So, because it's 1 minute 30 second, wanted to show that end-to-end flow towards the end because we have built this genie claim space. And Tyler, you can talk about then what's next from a vision perspective where really, you
[21:09] know, this is an example of where we use the semantic modeling and curated views to build on top of gold and added additional facts and dimensions to surface up the Genie space. Uh Genie today can answer 50 of those business use cases that we have collected from the business and it's in um testing
[21:24] validation phase with our issue consumption team. Um Dara, you want to talk about it what's next and how do we take it forward from a journey perspective beyond the Genie claims room and conversational agents? Yeah, so you know, one one of the biggest things that we're really trying to do and and I call it shift ownership
[21:40] but it's shared ownership, right? As we're building out the data products, we're engaged with the business very very intimately, right? So, we're not just rogue building, you know, uh a claim self-self-self-service tool because as payers we all actually need that, right? A member and eligibility
[21:56] self-self-service tool and Genie space. So, we're intimately working with the needs of each group. So, is it a slower rollout? Sure, but it also avoids us from building things that aren't going to be used, right? Which I know that we all do regularly, right? We We the whole
[22:11] build it and they will come just does not work, right? Um so, that's the vision is, right, to to from a Databricks perspective is the same vision as us is to democratize is to democratize data, right? So, so we use the word data data-driven organization a
[22:29] lot, right? And then you get back from executives, what does that mean, right? Well, what it means is different for each group, right? So, so every single group has a different one to need and we're here to enable that. So, again, back to the payment integrity example,
[22:45] health services, all of those groups, right? We're We're We're tailoring the data needs to them. And at the end of the day, the biggest thing for us is to deliver trusted data that is trans-trans-transparent to the organization, right? So, they're not We're really treating data as in
[23:00] We're treating data as a true asset, right? So, you know, in closing, what you see here is we're going to continue with CDS stack on the path that we have. You know, we've built the foundational layer in that year and a half now, where now is where we're really excelling in the the the actual data data product
[23:17] space and delivering value to the organization, right? And and we're starting to to realize and we're rolling out now actually end of the month more and more adoption of, you know, the business and the more people we get on, the closer we get to actually deprecating the on-prem data warehouse.
[23:34] I know we have time, but more than happy to stay back and take further questions and dive in deeper. But, thank you so much for being here. Appreciate it.
Learn more about the Databricks Data and AI platform.
The information provided herein is for general informational purposes only and may not reflect the most current product capabilities or configurations.