Skip to main content

Modernizing for AI: Legacy to Lakehouse with Databricks

Summary

  • McKinsey's Nebula platform migrated from a centralized lakehouse to a federated data mesh on Databricks in three phases, achieving a 40% cost reduction, 20–30% ETL speedup, and a 50-point NPS improvement.
  • The modernization shifted from custom centralized components that had become a bottleneck after 10 years to a hub-spoke governance architecture that balances domain autonomy with enterprise-wide data governance for healthcare analytics.
  • McKinsey's seven principles for AI-ready data platforms position platform modernization as a prerequisite to AI adoption, emphasizing treating data as a product, sharing meaning not just data, and securing trust by default.

Modernizing for AI: Legacy to Lakehouse with Databricks

Watch: Modernizing for AI: Legacy to Lakehouse with Databricks
Legacy data platforms become bottlenecks when organizations scale. McKinsey's Nebula, built over 10 years, reached an inflection point where custom components and centralized architecture were slowing innovation and increasing operational complexity. The modernization wasn't just a technical upgrade, it was a strategic shift from a centralized lakehouse to a federated data mesh that balances governance with domain autonomy.
See how Nebula migrated to Databricks across three phases: establishing foundations with hub-spoke governance, piloting with parallel coexistence, then optimizing post-migration. The results: 40% cost reduction, 20-30% ETL speedup, and 50-point NPS improvement. Learn the seven principles for building AI-ready foundations, treating data as a product, sharing meaning not just data, securing trust by default, and measuring behavior at scale. Most importantly, discover why modernizing the platform must come before AI adoption.

Chapters

FAQs

What is McKinsey Nebula and why did it need to modernize?

Nebula is McKinsey's proprietary analytics platform built over more than 10 years to support healthcare analytics engagements for payers, providers, and health systems globally. After a decade of evolution, its centralized architecture and custom components became a bottleneck that slowed innovation, increased operational complexity, and created barriers to AI adoption.

What results did McKinsey achieve by migrating Nebula to Databricks?

The migration delivered a 40% cost reduction, a 20–30% speedup in ETL pipelines, and a 50-point improvement in Net Promoter Score. The three-phase approach — establishing foundations with hub-spoke governance, running parallel pilots, and optimizing post-migration — allowed the team to achieve these results while minimizing disruption to active analytics workloads.

What is a data mesh and how does it differ from a centralized lakehouse?

A data mesh distributes data ownership to domain teams, giving each business unit autonomy over their own data products while sharing common governance standards, in contrast to a centralized lakehouse where a single platform team manages all data. McKinsey shifted Nebula from centralized to federated data mesh to eliminate bottlenecks and enable faster domain-level innovation.

Why must data platform modernization come before AI adoption?

According to this video, without a well-governed and modern data foundation, the platform and data become roadblocks rather than enablers when organizations try to scale AI. McKinsey's seven principles for AI readiness — including treating data as a product and measuring platform behavior at scale — are positioned as prerequisites that must be in place before AI use cases can be successfully deployed.

Full transcript

[00:07] All right, good afternoon. Almost good evening. And thank you for joining us. Um we realize this may be one of the last session that you might have today. So we're pretty standing between uh back day of uh sessions um and well-deserved happy hour drink.
[00:23] So we try to do our best to make this as uh worth it as possible. So today we'd like to share with you our experience in leading large-scale analytics um platform organization, the journey of our Nebula platform, which is
[00:38] making the healthcare analytic platform. Um and perhaps more importantly how we evolve from a lakehouse architecture to a data mesh platform.
[00:55] We will walk you through why we made that decision, how we made that decision, um how we approached it. Um the benefit that we achieved. What went well, what didn't go as well because you know like any transformation some things as um much as you want to
[01:11] plan them don't go as well. So we'll look at uh kind of uh what we would change if we were to do this again. Um and towards the end we will also give you a bit of a glimpse on what we're currently working on, which is the rewiring of Nebula for the analytic age.
[01:26] Um we'll share some key practices that is uh helping us to make sure that the platform and the data um do not become roadblocks for scaling AI but actually key enablers for scaling AI at an organization. My name is Pierre-Arnaud, the CIO uh of
[01:43] our healthcare analytics practice. And I'm very lucky today because we have Nico Schippers with our uh product manager. Uh and we have in the room some of the team members. So if you have questions at At end of the presentation, we'll be happy to answer some of those questions that we might not have been able to
[01:59] address during the during the session. All right, before deep diving into the migration itself, um let me take a step back and give you some context on Nebula, which I think is quite important. So, Nebula is McKinsey's proprietary FK intelligence
[02:15] platform. Um it has been developed um and refined over more than 10 years. It supports our analytics work and analytics engagement. Um which we're doing for payers, providers, and FK system globally.
[02:31] Um it is also the platform that we use internally for building and skilling our analytics assets. So, in many ways, Nebula is our playground, it's our laboratory, and our foundation for creating impact through data. Although, given the sensitivity of the
[02:47] FK data, you could probably consider this as the most carefully governed playground that exists out there. There are three core requirements that have shaped the way that we have developed Nebula. The first one is security because even
[03:04] though the data that we have are pseudo anonymized, they remain highly sensitive. Um they're subject to many regulations, HIPAA in the US, GDPR for Europe. Um so, keeping data safe and in a
[03:20] controlled requirement and sorry, in a controlled environment is our number one requirement. The second one is scale. The volume of data that we have, and I know it's all subjective, but the volume of data that we have is quite significant. So, we also needed a platform that could scale and that would allow us to process those
[03:36] big data in uh a fast and efficient way. And thirdly is isolation, which is very specific to the type of work that we're doing, but I'm pretty sure that you have this type of requirement as well in other industries like banking, to name just one, which is like every data set needs to be
[03:51] handled separately because those data set belong to, you know, clients or um country-specific health care system. So, they all need We need to have the controls and the way so that we control who is accessing it, what analytics can be performed.
[04:07] Um and that principle is standard or is central, sorry, to the hub and spoke model that we have designed, which Nico will will talk to you a bit more uh later on. So, as I mentioned, we developed Nebula over uh 10 years and if just go back in
[04:22] the history, our whole journey really started in 2013 where um we actually set up Nebula, you know, back then it was on-premise SQL- based analytics solution. Um that gives us and the team, you know, a high HIPAA-compliant space to work with
[04:39] complex health care data and deliver the analytics. In 2016, we we took a leap of faith and um and went to the cloud. We moved to AWS. Back then, we were working with Hortonworks and later on Cloudera. Um that gives us the scale, uh the scale
[04:55] to actually handle more and more um larger volume of data. Um and also being able to uh leverage the platform performance for different use cases. Over the following years,
[05:10] Nebula kept on evolving. In 2018, we uh achieved HITRUST and ISO 27001 certification, which is critical for the industry that we're in. Um and by 2020, we actually moved fully to
[05:26] AWS managed service. So, we actually shifted from Hortonworks Cloudera to EMR. That helped simplify operations. Um it improved the overall performance. Um It reduced some of the burden of running everything ourselves.
[05:41] But by 2023, after nearly a decade, we really started having like um legacy pain. Um the platform had become powerful, but also increasingly um, it increased in complexity. Um, we needed to have more and more individuals to actually run the platform.
[05:57] Um, it also slowed down our pace of innovation. Um, user satisfaction was under pressure. And maintaining the required the the environment required to have a dedicated team. So, in other words, Nebula was still delivering value, but it was, um,
[06:14] uh, impacting our way to actually innovate. So, we needed to open a new chapter of, um, of Nebula. And that's when we decided to actually migrate to, um, Databricks.
[06:32] Our legacy platform was also originally built much more around centralized lakehouse model. And let me maybe do a quick, uh, show of hands like, uh, first, who is using Databricks in the room? Can just show Okay. Okay, that's what I was expecting. Quite a few of you. Who is now using
[06:47] Databricks more as a data warehouse solution? A few hands, right? Who is using more as a data lake or data lakehouse? All right, a bit more hands. And who is using it more as a data mesh platform?
[07:05] All right, that's what I was expecting. So, um, quite a few. Anybody using like more like data fabrics, which is quite advanced? Okay, no one. Okay, that's what I thought. That's normal. Um, so, our legacy platform was really geared towards, uh, data lakehouse. Um,
[07:22] it was originally built around centralized, uh, system. At that time, that was absolutely right foundation for what we were doing. However, um, as the platform grew, the same centralization that was core to what we were doing, um,
[07:38] you know, starting to create some friction. Uh, Uh, we had more and more data sources, more and more uh different use cases that we need to have. Our domains were also getting more and more uh smart about what they were doing. Uh, they had better technical capabilities, they had different use
[07:54] cases, they wanted to use different tools, and they would basically wanted to have more autonomy. Um, as a result, um, we needed to change something, and that's when we decided to actually move to a more data mesh um concept.
[08:15] Nico. Thank you, Pierre-Arnaud, and welcome to our session. When we first built Nebula on Garth Cloud,
[08:31] we were ahead of where the market was at the time. Many cloud-native capabilities we take for granted today simply wasn't available back then. And we had to build a lot of these ourselves. Um,
[08:46] it had the benefit that it gave us the control and the flexibility where we needed early on. Uh, but it also meant that we accumulated a number of custom components during this process. And by 2023,
[09:04] cloud providers were beginning to offer the same natively. At the same time, our business needs kept growing and expanding. And, um, to meet those needs,
[09:20] we uh needed specialized tools um, to bring it into the platform. So, we added EMR, Spark, Druid, Redshift, and others. And each of these tools solved their own
[09:36] unique problem at the time. But, with each one that we added, we added another layer of complexity. And over time, Nebula became a collection of powerful, loosely connected capabilities.
[09:53] Um we created operational overheads, performance inconsistencies, um integration challenges, and rising maintenance costs. Uh the platform was still doing important work,
[10:08] uh but it was becoming harder to scale, uh harder to evolve, and harder to simplify. We reached an important inflection point.
[10:23] And we had to ask ourselves a practical question. Do we continue rebuilding and maintaining more of it ourselves, or do we partner with an industry leader uh that can help us scale faster, innovate faster, um and remain
[10:39] secure for the long term? For the modernization effort, uh we had a few clear objectives. Uh first of all, uh we wanted to shift our team's energies away from infrastructure and
[10:56] the support work that we were doing, and back into insights and client impact. Uh a lot of our legacy platform required hands-on effort to maintain it, to operate it, and to troubleshoot it. We wanted to simplify it. We wanted to automate it,
[11:12] um as much as possible. Uh so that our data teams could spend more time on analytics and innovation. Secondly, we wanted to strengthen our security and our government governance um without slowing down our business.
[11:27] Uh there or that was the hub and spoke model, um that became very important to us. It gave us a way to apply consistent controls and policy policy centrally without still allowing
[11:42] whilst still allowing individuals, uh, to operate with flexibility that they needed. And for us this became the foundation to federated governance across our mesh. Third, we wanted to make
[11:59] the user experience better, and that means simplifying access, improving onboarding, and making data easier to find, and giving teams more standard dies toolkits, uh, so that they could serve themselves instead of waiting on platform teams,
[12:16] um, for every request. And then finally we wanted to create a platform that we could we could support new client delivery needs on, um, and models and use cases that were emerging in the market. Whether it was as a product, um, or as a service, delta sharing APIs,
[12:34] or future AI enabled services. Uh, the goal was to move a platform that mainly supported internal analytics to one that can support a more flexible, scalable, and client-facing delivery model. Modernization
[12:50] wasn't just about replacing technology for us. It was about creating a platform that was simpler to operate, easier to govern, and better for our users, and ready for the next generation of healthcare analytics.
[13:11] What was the solution? Well, after reviewing several options, we selected Databricks as a cornerstone of our modernized Nebula platform because of its compatibility, openness, and flexibility. It's native support out of the box for
[13:26] SQL, Python, Spark, and their medallion architecture aligned naturally with how our teams already worked, and making the transition practical and easy to adopt. Architecturally,
[13:42] the modernization would allow us to consolidate a fragmented ecosystem of loosely connected components into a single integrated platform. From a security and governance standpoint, we worked with Databricks to deploy our hub and spoke architecture that
[13:58] strengthened our compliance and data protection, while supporting our data mesh principles. Common platform capabilities, such as governance, security, compliance, reusable data methods, and shared tooling, are now centrally managed as a hub.
[14:22] And this balances control with consistencies and compliance matter the most, with the domain level autonomy, where speed, ownership, and context matter most. From a user experience standpoint, Databricks provides a unified workspace with streamlined access,
[14:39] onboarding, and data discovery. It enables self-serve analytics for us through standardized reusable tools, and it also accelerates our time to insights dramatically through the automated ingestion, integration, and workflows.
[14:55] And finally, from a delivery model standpoint, Databricks supports secure cross-company collaboration for for us through their Delta Sharing capabilities, enabling governed read-only data to be exchanged
[15:11] safely across organizations without copying data or exposing the underlying storage. Thank you, Bear. Thank you, So, what did all of this deliver?
[15:29] Um I think the first one was cost. We were able to actually decrease our cost by 40% from 2020 to 2025. Uh that reduction came from the consolidation that Nico mentioned in terms of the infrastructure, the
[15:45] simplification of the tooling landscape. Um as well as having much more capability to actually manage cost. And I'll get back to that in terms of the operation excellence that goes with any transformation. Um compute cost and storage decreased by
[16:00] 32% also because we, as I will mention later on, we also reduced significantly our data footprint. Um We also reduced the third-party licensing system. But more importantly, I think the
[16:16] financial impact is of course very important, but I think what was key for us was actually two things. One, we were able to decrease the time that it takes for standard ETL to be performed, reducing them by 20 to 30% faster. So, on some of our ETL which run for many,
[16:33] many hours, that was quite significant. Um and more importantly, the user satisfaction really improved. Our net promoter score increased by 50 points. 50 points, right? 50 to which actually 70 uh in terms of net promoter score,
[16:49] which was quite actually important. We got a drop after that of 5% after like end of the hype and the happiness of of users. I think that's a normal drop. Um but since that we remain at that at that very high level. So, in summary, the Nebula modernization
[17:04] was not just about technical infrastructure shift, but this was really about like enabling the business, enabling the environment, and users to actually achieve new things and be able to innovate. So, how did we do that? We had three approach strategy manage that that
[17:21] transformation. The first one was foundations and alignment. Um we first set up the direction that we wanted to actually take.
[17:36] So, we clarified the strategic intent with our stakeholders, our users. We align on the data mesh vision. We define a federated governance model that balanced domain autonomy with the actual enterprise wide standard. We also designed the target
[17:52] architecture, identified first priority domains, and appointed domains data owners to embed accountability closer to the source. So, ensuring that as we giving more freedom, we also hand the responsibility to the domain.
[18:09] We did a lot of user research. Trying to identify what are the pain points for the user, identifying who are the users. Typically, you would have lightweight analytics users, you'll have the builders, you'll have the data engineer and data science. They each all have their own
[18:24] pain point and preferences. That was also important because that was a good way to actually make sure that users would feel heard, that they were part of the journey, that they were part of the transformation, and not something which was imposed top-down to the users.
[18:46] The second phase was the pilot and platform enablement. So, we actually run both environments in parallel. We had our legacy environment, and then we had our new environment. For the new environment, we selected domains that were kind of opinion leaders. Uh we selected pilots that we run, and
[19:03] then we tested. So, we tested the new environment on those pilots. Um those pilots were important for two reasons. First, they allowed us to validate the technical uh concept. Secondly, and more importantly, they also helped us build the confidence and buying from the users uh who ultimately
[19:20] need to adopt the new model. From there, we expanded to broader um deployment. We scaled the mesh across more domains while keeping uh the legacy where it was running in parallel. And that coexistence first um allowed us to
[19:36] de-risk the overall project. It gives more time for domain that were a bit lagging behind all the domains. They were not all uh at the same level in terms of um requirement, but uh at the same level of complexity as well as uh competencies.
[19:52] Um and then at some point, we reached the level where we all felt confident that yes, we could finalize this migration and we had to push a bit uh some of the domains giving like strict uh target date in terms of a By that moment, we're going to cut off the environment to force the last domain to actually migrate uh over to the new
[20:08] platform. The third phase, which is super important, is uh scale and continuous improvement. So, at that stage, we focused on shifting from retiring um legacy component and optimizing performance and
[20:24] cost, automatic governance, setting up dashboard, uh putting up cost control in place. And and that um phase actually lasted 12 months after the migration uh was actually completed. So, that's something I recommend not to skip cuz that's what
[20:41] it is allowing you to make sure that your team is uh uh operating with the right elements, the right dashboard, the right capabilities, um and then can move on to innovation again. So, having a post cool down period, make sure that your team is
[20:57] operationally sound, that you recover from that migration effort, um and that you take the time to put all your operational process in place. Once that done, then you can start innovating again. So, some key lessons and I already started discussing about that, but all
[21:13] the key lessons were first one, you cannot build and run at the same time. We tried it, doesn't work. So, you need to have a team that is dedicated to the migration. And you set up your team just like you setting up a product team. You need to have a product owner, you need to have
[21:29] tech leads, you need to have client success manager who are going to coach, train, explain why this migration is necessary, accompany the different users. Second, the coexistence, the hybrid phase, being able to run both your previous environment and the new
[21:44] environment all together. That is good way to de-risk a project. Although, you need to make sure that it's not an excuse for some domains to, you know, slow down their migration, right? So, you keep the same tempo, you keep a hard target, but you run those two elements or those two environments
[22:00] in parallel. Thirdly, more domain autonomy does not mean less governance. I would argue it actually requires more governance because you're giving more power to the users, you also need to make sure that there is right governance in place to make sure that you're not risking or
[22:17] increasing the risk for your organization. We also learned the value of starting small. So, high impact pilots helped us to prove the model, build confidence, create momentum before scaling to a bigger
[22:32] audience or bigger projects. Um one of my favorite is make sure that you're managing the platform as a product. You treat the platform as a product in its own right with users, with personas,
[22:48] a roadmap, a feedback loop. We look closely at the need of the data engineer, we close closely at the need of the business analyst, make sure we understand them and that we develop a platform where we can actually serve each of those needs.
[23:06] The product management mindset was critical because it created engagement and ownership. People were not just being asked to migrate um to the new platform. They could see how the new platform was solving problem and what would benefit for them.
[23:22] And beyond that, the success required strong uh leadership alignment because we had to make some difficult decision. We had to convince some people. Uh we had to force some people to actually do this migration. Uh so having strong um ownership and buying for your from your
[23:37] senior leadership is key. Like don't move if you don't have it, basically. So in short, we did not just modernize the technology. We built an operating system, an operating model around data as a product.
[23:53] So what now? So the move to Databricks and the federated data mesh put Nebula in a much stronger position to innovate with AI. Not something that we necessarily anticipated, but we kind of lucky we had the right timing to actually do that migration before AI became really a thing. So we now have a more scalable
[24:08] foundation. We have stronger governance. We have clear data ownership. Um and analytics closer to the domains uh that understand the business context. That foundation has helped us to very quickly uh run our first AI pilot and to
[24:25] get some success already. However, pilot are not the same thing as running things at scale. So to move from experimentation to real impact and spirit towards agentic AI, we realized Nebula needed yet another revolution. And the next step is about
[24:40] rewiring Nebula so that data governance models, application, agents can work together reliably, safely, and at scale.
[24:55] A house is only as uh strong as its foundation. And the same is true for AI. We're seeing many organizations that want to jump on board on the trend of AI for the fear of you missing out. Um very frequently, however, their foundation is just is just not right.
[25:12] Right? So, they're not going to be able to achieve what they want to do at scale. More importantly, they're going to take some risk because they don't have the right level of governance for their platform. So, we're sharing here a survey that we that we shared. This was with the in the telecom sector, but
[25:28] I would argue like it it is applicable to to many industries. Um and what that survey showed is that the challenges about moving to AI and adopting AI at scale is about building the foundation that allows user to adopt AI safely, efficiently, and organization to work in
[25:45] a in a reliable way. So, there are a couple of of hurdles for that. The first one is data. Is the right data available? Trusted, well-defined, and usable. Um the second one is the technology platform.
[26:01] Can the organization deploy, monitor, scale AI, and enable workflow across users and teams? The third one is security and trust. So, especially in a health care, but increasingly in every sector, as model
[26:16] become more and more capable, we've seen that through Metos, there's more and more risk that are being taken. So, actually reinforcing your security is paramount at this moment. The fourth one is cost control. Um we start seeing this at first, you know, AI was kind of a
[26:33] free meal. Right? Everybody could use it. Now that we start receiving the building, right? People start saying, "Well, I don't know. I've consumed all my AI budget for the year." There's plenty of story that says that. So, setting up the control so that you know um you make sure that you're controlling your cost. You also give the information
[26:51] to the user so that they know how much are they, you know, actually consuming in terms of tokens so they can make the right choice in terms of what they want to identify, what they don't want to identify. So, to move on from AI pilot to AI scale, we need to strengthen the
[27:06] foundation. Um and we've identified for ourselves seven key principles that we're following for Nebula. The first one is to continue treating data ingestion like product. So, data should enter the platform through consistent reusable patterns. Um
[27:24] so, it can be governed once and reused many times. The second one is about sharing meaning, not just data. So, we need shared definition, metadata, taxonomy, um ontology so that analysts, humans, as
[27:41] well as machine connect machines can actually leverage your data. So, we're spending we're we're lucky because we had a good data model already, we had good metadata. Now, we're investing everything to building the ontology. The ontology is which is consumed by humans and by um by machines.
[27:57] The third one is use one data foundation for analytics and AI. It does not make sense to create two different set of data, right? It should be the same data that is being accessed either through for dashboarding, machine learning, or generative AI.
[28:13] The fourth is to build the trust into the platform uh by default. So, security, access, control, um lineage, um all of that needs to be embedded and automated, not added later on.
[28:30] The fifth is about ex- about exposing capabilities through stable interfaces. So, team needs to have reliable API, MCPs, model endpoints, interfaces so that they can build without constantly reworking um the plumbing on the how they're getting access to the data.
[28:50] The sixth principle is to make behavior visible and measurable. So, we need to monitor data quality, model performance, uh latency, cost, and usage so that we can uh uh ensure that we have a controlled environment from which we can actually uh work.
[29:08] So, the message is actually for me simple. It's like scaling AI is not just about having better model. I would argue, you know, don't focus too much on AI if you're not ready yet, right? Focus on the platform. Focus on building the foundation that will allow you to scale AI at scale within your organization.
[29:29] To close, we wanted to leave you with a few resources to go deeper on some of the topics that we've covered. Platform modernization, AI transformation, and the foundation for agentic AI at scale. The our main lessons from our Nebula journey is that modernization is not a one-time migration. You've seen this is
[29:46] like a more than 10 years journey to actually build the platform. We went through already three large migration. Um we hope that we're not going to get no migration in a couple of years. I think we have a stable environment. Um but take this opportunity to strengthen
[30:02] your governance, strengthen your data product, improve the user experience, treat your data as a product, treat your platform as a product.

Learn more about the Databricks Data and AI platform.

The information provided herein is for general informational purposes only and may not reflect the most current product capabilities or configurations.