Enterprise Lakehouse in One Year: Building for Optionality and Scale
Summary
- Hospital for Special Surgery built a data lakehouse from the ground up on the Databricks Data and AI platform in 12 months, ingesting 60 source systems, creating 70,000 tables with a shared data model, and deploying 130 applications.
- Chief Data and Analytics Officer Dr. Michael Uohara deliberately chose enterprise-first architecture over use-case-driven development, building a comprehensive foundation that creates long-term optionality for AI rather than chasing quick wins from individual projects.
- The approach demonstrates how governance by data rather than by use case — with high decision velocity and shared engineering principles — enables downstream AI capabilities including Databricks Genie and applications that transform patient care and operations.
Enterprise Lakehouse in One Year: Building for Optionality and Scale

Dr. Michael Uohara, Chief Data & Analytics Officer at Hospital for Special Surgery, shares how enterprise modernization requires principled engineering, not use-case-driven development. During his first year, HSS built a data lakehouse from the ground up on Databricks, intentionally prioritizing enterprise data assets over quick wins to create long-term optionality.
Learn how principled engineering enabled HSS to ingest 60 source systems, create 70,000 tables with a shared data model, and build 130 applications in 12 months while maintaining enterprise governance. See why building a strong foundation first, not being driven by immediate use cases, positions organizations for AI capabilities like Genie and delivers applications that transform patient care, financial reporting, and operations.
🤝
Chapters
00:00Introduction and Background: From Physician to Data Leader01:36The Problem: Theory vs Practice in Lakehouse Architecture03:10One Year of Building: HSS's Intentional Approach03:58Three Key Principles to Remember05:32Principled Engineering: Building with Humility and Team Execution07:58Tale of Two Cities: Use Case Driven vs Enterprise First09:34The Leadership Role: Enabling Wholesale Data Ingestion11:09Governance by Data, Not by Use Case12:30The Execution Challenge: Decision Velocity and Shared Principles13:52Technical Leadership: Embedded in the Weeds14:57Results: 12 Months of Enterprise Data Architecture18:25Key Takeaways and Closing Thoughts
FAQs
How did Hospital for Special Surgery build a full lakehouse in just 12 months?
HSS achieved their 12-month build by prioritizing principled engineering and wholesale data ingestion across all 60 source systems rather than building incrementally around individual use cases. Dr. Uohara's approach required high decision velocity and a team aligned on shared principles, avoiding the rework that use-case-driven builds typically generate.
What is the difference between enterprise-first and use-case-driven lakehouse development?
Use-case-driven development ingests only the data needed for specific projects, which over time produces a repository of disconnected small datasets rather than a true enterprise data asset. Enterprise-first development, as HSS practiced on Databricks, builds a comprehensive shared data model from the start that creates optionality for any future use case including AI.
What AI capabilities did HSS's lakehouse foundation unlock?
By building a strong enterprise data foundation with 70,000 tables across 60 source systems on Databricks, HSS positioned themselves to implement AI capabilities like Databricks Genie for natural language querying and to deploy 130 applications serving patient care, financial reporting, and operations within the same 12-month period.
Why does Dr. Uohara argue that use-case-driven data development doesn't work?
Dr. Uohara observed across multiple organizations that a collection of many use-case-driven data assets never adds up to a cohesive enterprise data platform — the analogy he uses is that in data, unlike in personal finance, a lot of small things do not necessarily accumulate into one big thing. Building an enterprise lakehouse requires intentional, upfront architecture rather than incremental addition.
Full transcript
[00:09] Please welcome to the stage Dr. Uhara.
[00:28] These are some forward-looking statements and a legal disclaimer. I did not want to write these. The lawyers told me I had to. Um this is who I am. I think um everybody heard and I don't love this slide cuz I don't want to introduce myself as Dr. Uhara or you know, a
[00:47] technologist. I'm just a data person and my team can attest that that's the way I I introduce myself. Um this, you know, recent pivot in my career is the second best thing. I I I tell everyone this is the second best thing. I was a physician by trade.
[01:04] This is the next best way that I get to affect change in our industry. Um I have had a career in in big tech and big consulting and then most recently moved to HSS about a year ago. Um HSS is a specialty hospital. Um I
[01:20] guess that's the tagline, number one in orthopedics and musculoskeletal care. Uh but we provide wonderful care across four states. We're an integrated health system in the tri-state area and also have practices in Florida. Um
[01:36] when I joined HSS though, uh a little bit of reflection, um I I did so with the context that the way we've been talking about building data warehouses and lake houses, you know, for the past better part of the last 5-10 years,
[01:53] you know, these common repositories of big data, enterprise data assets, um you know, a shared ontology across an organization, you know, the the perfect picture really fell apart when we put it into practice.
[02:08] You know, and in practice, uh when we tried to implement these things for customers when I was at Microsoft or uh Booz Allen or when I even tried to do these things uh in a previous career, we ended up being driven largely by use cases. So, a use case would come in and
[02:24] somebody would complain really loud and we would build our lakehouse or build our data warehousing on a collection of use cases. And over the course of well, extrapolate a year or a couple of years, you had you know, what is essentially
[02:40] uh a large data repository of a lot of small data. And I also like to say that our parents lied to us. Uh a lot of little things don't necessarily add up to being a big thing. You know, it may be true in um
[02:55] uh in finance, you know, saving a lot of pennies, one day you're going to you know, we tell our kids this, but the truth is that's not the way data works. Um and so, when I came to HSS, it was with the intentionality that I was going to build this data warehouse, this lakehouse
[03:10] architecture and I was going to try to do it the right way. Um you know, I really uh wanted and when I contacted the Databricks team when we first started this project about a year ago. In fact, today marks our 1-year anniversary of standing up our lakehouse
[03:27] from when we started the program uh till to now. Um it was that I want the data purists um data warehouse. We're going to build enterprise data assets. We're not going to be use case driven. We're going to go into the basement and we're going to try
[03:42] to build the right way instead of being driven by, you know, corporate needs or corporate politics or whoever was complaining the loudest or you know, the the pressing issue of today. And I'm really going to try to convince you, you know, of of three things. Don't
[03:58] read the slides. I also banned PowerPoint in my organization, too, so that I can tell you a little bit about what I how much I prepared to give this PowerPoint. But, you know, what I'm really going to try to convince you today is that there is a, you know, long-term value that's created by
[04:14] building way beyond the asks. Um because what that does, right? It gives you the power of optionality. So, when you build in this in this in this way, you give yourself forward-leaning optionality. And then, you know,
[04:29] that sets you up. That's a fertile soil for that nice enterprise architecture, but it doesn't get you all the way there. Right? At the end of the day, there needs to be an execution arm of this. We need an all, you know, uh leave it to the technical person to talk
[04:44] a little bit about change management at the end of this. So, apologies for those of you who are technical and have to hear me talk about that. Um but really do making this real really requires focused execution. It requires a little bit of changing a philosophy. It requires us to enable our
[05:00] teams to uh to build quickly and to stand up the enterprise architecture because in many ways, you're building against the clock. There is only so much time that you have from the kickoff of this project of the kickoff of building out a lakehouse until those use cases really do start
[05:16] knocking at your door and really start to change your architecture fundamentally. And what you want to do is to have the foundation be extremely strong before you start inviting use cases. Right? And so, that's what I'm really going to hope to focus um us on today and and convince you. And the first two
[05:32] things are extremely important. I don't have a ton of slides, so I'm going to spend most of the time talking about this. Um you know, principled engineering, what does that mean? Um you know, what you do when you start to be use case driven uh from a lakehouse
[05:48] architecture is you start to actually fundamentally erode the value proposition that is a lakehouse. Right? When you do this, you inevitably are a collection of use cases, which are collection of governance paradigms,
[06:05] which have very limited usability. Right? And therefore also cannot predict where this market is going and what the needs are going to be for tomorrow. And so when we do this, when I think about principled engineering, when I use
[06:20] that term, I'm really talking about approaching engineering of a lakehouse with the humility to understand that we do not know what the future holds. I don't know and I'm a very poor predictor of the future even though I think I'm pretty good at it. I'm a very
[06:37] poor predictor of the future. I don't know what the use cases are that are going to come in and be the hot topic of today. I think all of us in this room, if we are in data leadership or IT leadership, we probably would not have predicted the AI boom that came to our
[06:55] our doorstep. And that's exactly the point. The organizations that are winning today are the ones that understood fundamentally they could not predict the future. They were not the ones that predicted the future accurately.
[07:11] Let me say that again. The organizations that are winning today were the ones that understood that they could not predict the future accurately, not the ones that predicted the future. So, why build in this capacity is because we
[07:26] are hoping, like all of the presentations that you saw, you know, to enable AI in our organizations, to enable Genie 1, right, for all these products that you hear, you know, heard mentioned, they're nothing without principled engineering, they're nothing without a fertile ground
[07:42] and an entire repository of corporate information in one common place. You don't want to be reaching across into other systems because integration burden is a real thing. All right? You want to have this common large common repository. And so when we
[07:58] built the lake house, you know, you can see sort of the tale of two cities here on each side. One is, you know, you sort of have a use case come in. Uh you know, it's for a clinical use case where, you know, a hospital will provide clinical care or maybe it's a finance use case or maybe it's
[08:13] operations. And to get that use case done, we need to ingest a couple of data sets, but we don't need to ingest the entire thing. So we ingest part of it. Because why? Cuz it's fast. It's quick. And we can get the job done. We model the data in that way, right? We have a team of data modelers who model the data very use
[08:30] case specific. And at the end of it all, at the end of the project, the team gets a product, a deliverable, a dashboard, an analysis, whatever it is. It doesn't matter. But then you have to ask yourself the the level of effort that it took to preserve and to to prepare that use
[08:46] case one, did it lower did it lower the potential energy for use case two, and three, and four? And the answer, if you're honest with yourself, is no. It didn't. And use case two comes along and you do the same exact thing. And maybe there's
[09:02] 1 or 5% overlap. And use case three comes along and so forth. And over the course, again, you extrapolate, you know, many years down the road, you have big data in maybe in theory, but you really have a
[09:18] collection of small data. And in order for us to actually have enterprise data assets, we need to have a shared common ontology. We need to have a shared common data model. And so the other end of the spectrum is what we've we've tried to do. Um you know, I
[09:34] my job as a leader, right, is to protect my engineering teams so they can go out and accomplish these bigger engineering pipelines to go and ingest do wholesale ingestion from our source systems. And so what that looks like is we go after system A and we get a
[09:49] essentially a mirror copy of system A in the lakehouse. We have a reusable pipeline. We have evergreen data. The understanding and the theory is that all data is good data at this structure. All data. If it's the logging, if it's the permissioning, if it's the um if it's a
[10:05] clinical, you know, prescriptions table, if it's a diagnosis table, it doesn't matter to us. It's all everything. We ingest it all wholesale. We're metadata driven on how we model data in the first, you know, your bronze, silver, gold medallion architecture. And so we model data
[10:20] appropriately without any use cases in mind. And then we move on to the next system and the next system and the next system. You know, and then through throughout the course of that for us, use cases do come along. But we were able to successfully sort of hold the business off just a little bit
[10:37] for 3 to 6 months or you know, whatever the time period to get critical mass and data gravity. So that we were not use case driven on how we were preparing our data model so that we were preparing an enterprise data model. One that was reusable.
[10:53] Right? Repeatable. Governance structures in place that were repeatable across business units cuz we govern the data, not the use case. And you know, also ultimately this can potentiate this future that we talk
[11:09] about. And that to me, Genie 1 or any one of these systems that you see lowering uh you know, lowering the the barriers to individuals querying the data or lowering the barriers to people using the data are all about democratizing access, right? It's about pushing this
[11:26] data expertise all the way to the perfect edge so so the end user can do the the same query that a you know, a data access developer or a SQL developer can or a data engineer can. And you can't get there if you don't have the right governance structure.
[11:42] You can't get there if you don't have reusable data assets because there's no way for you to predict what that question is going to be or what that query is going to be. So, self-service doesn't exist in in practice until you architect your lakehouse in the right way.
[11:58] All right. Um and so, all that is great, right? Unless you don't have a team to ultimately deliver this. You don't have a team to ultimately architect in this way. You don't have speed to deliver. It all sounds really really nice in practice.
[12:13] The architecture can look beautiful on a Visio diagram, but when it comes to actually having a team deliver this architecture, I think this is probably where most organizations fall apart. Right? I talked about like in theory how we've thought about lakehouses,
[12:30] ingesting these systems, and having common repositories of information. I don't even think in practice that's where it fell apart. I actually think this is where it falls apart. You don't have decision velocity, so things are slow, and so then you get pressured by more use cases to do use case driven.
[12:47] You know, you don't have shared principles and foundations. And so, what I you know, when I came to HSS, I mean that was exactly what my my principal mission was is to convince the data engineering team about this operating model. And so, we, you
[13:03] know, make decisions very quickly. I think I've gained either a bad or a good reputation depending on who you ask for me for making uh decisions pretty quickly, but the truth is I can make a lot of decisions, but I cannot make all of them. I need my teams, my engineering teams to make
[13:19] really good decisions on my behalf or on the organization's behalf. There's not enough me's to go around, and so you need to have decision velocity. You need your engineers and that's that's not my my engineering leader for the whole organization. That's his engineers making those good quality decisions and
[13:37] you don't do that. You don't have decision velocity unless you have a shared principle. How we ingest data, what we ingest data for, how we model data appropriately, the standards that we have, how we do governance at HSS or any one of your organizations. All right?
[13:52] And then ultimately the very end of this is, right, getting leadership embedded and and you know, into the weeds. I think I I'm one of one of our data Data Bricks reps told me that I'm the most in the weeds leader that he's ever seen and now I can't tell whether that
[14:07] was a good thing or a bad thing. So it's still still jury's still out. But you know, having leaders embedded at the edge, able to go into the technical nitty-gritty, able to answer questions quickly. And this is not how we've built IT shops or data shops in the past.
[14:24] This is not. This is unfortunate, but this is the truth. And with the tools that we have today, we can write code, we can build pipelines at scale. But ultimately a decision has to be made and when the decision point or when the teams are, you know, the the hierarchy
[14:40] of the organization is what's blocking you from developing well, then we have a major major problem. So I'm going to go forward and talk a little bit about what this becomes. Uh this is not a highly technical talk and so this is a slide that um I ran a
[14:57] query, I think it was last week or maybe the week before in preparation for this about what we were able to accomplish in about 12 months, a little less than 12 months. Uh we've ingested, you know, nearly 60 source systems. I put 57 plus. I keep asking People keep asking, "Why do you put 57 plus?" Because there was actually
[15:13] two or three systems that were in flight at the time and I knew it was going to be 60 by the time this conversation happened. So about 60 source systems. Um you're moving from left to right, you can sort of see uh your medallion architecture starting to be built, right? About 200 plus schemas.
[15:29] Um, 70,000 tables or so, about 16,000 of them, a little greater than 16,000 of them in production. And what I'm most proud about is now you can see the architecture coming to fruition at the far left. There have been 130 applications built, 10% of
[15:45] those are in production today. A vast majority of those were built in the past 3 months. And to me, you know, when you ingest data at this scale, when you start to have shared engineering practices, when you have shared data assets and
[16:01] repositories, it's that thing that we ultimately care about on the far left, the application science of the data, rather than like the tactical engineering of how you model the data, you know, model the data perfectly. And so, those applications span everything from, you know, how we
[16:16] schedule exam rooms at HSS, built entirely on Databricks. There's a real-time nature of that, and then there's an ETL process of that. We have both the source system on the lakehouse in addition uh to the reporting mechanism. So, they span things like our financial
[16:33] reporting platform lives on the lakehouse. So, we actually report to our auditors and our um you know, our bond holders on Databricks. And a year ago, that was on spreadsheets. You know, that that spans to our unified intake.
[16:48] So, how anybody gets access as an employee to to shared services at HSS is supported um on the lakehouse, um to our ITSM platform that we're building out on that on that very same platform. You cannot get there unless you do the
[17:05] hard work up front. And I think a lot of folks have been really excited about the application science, but they didn't understand for the first 6 months what the heck are we doing? Cuz they're just building a lot of data assets. They're just doing a lot of data modeling and no end users there. So, at this juncture, where we stand is,
[17:21] you know, we probably have uh in the queue 30-plus applications that are beginning to go into the production uh environment in the next 6 to 12 months. Those are either takeouts of SAS systems, those are either, you know, source systems that did not exist
[17:37] that are going to production. Um those are uh applications that will touch patient lives. Those are applications that will touch providers and nurses. Right? That's the ultimate stuff that we care about. And we're able to now spin up applications. We're able now to build
[17:53] with that sort of velocity because of what you see across the bottom, right? A pretty massive data repository that was built out. 300 TB of structured data elements is quite a massive amount. This is These are Delta tables.
[18:10] Um and that's, you know, kudos to the engineering team that we have, but is that true enterprise first approach. Um and I think this is the last slide before, you know, I've been told to ask you mention to you guys that I'm going to be
[18:25] taking questions cuz I didn't want to bore you with I think the the devil is always in the details and you don't get to the details unless you get to questions, right? Um so, these are the key takeaways. Don't read the slide. I I think Yeah, just don't read them at all. Like, the truth is that you want to build
[18:41] enterprise data assets because enterprise data assets create value. Use cases are limited by their value creation. Long-term assets will create that value for you, you know, the those use cases that you cannot imagine in the future. Data engineer with humility
[18:58] instead of this idea that I can create and model the perfect data set that's going to predict the future. Engineer with the idea that the folks that are going to win the future are not the ones that predict the future, but understand that they cannot.
[19:13] Uh and then lastly, like, you know, I think talking about enabling your teams, if you have teams that you work for you that do data engineering, data modeling, products folks, you know, you got to make it easy. Uh make the right way very easy. Um I don't know half of the decisions
[19:28] that my teams are making every day. I have a team of approximately 100 folks or so, give or take, that are doing engineering, both contractors and full-time employees. And every single one of them, I know for a fact feels like they can make decisions, and they feel like they know exactly what is best
[19:45] for each HSS in terms of data engineering, in terms of data modeling, about how to create the best data assets. Um and, you know, that is I want to make one thing clear before I close. None of this is saying that the use cases are not important.
[20:01] I am, as a physician, care deeply about what we can do with data. In fact, again, I said this is the second best thing that I could be doing. The first best is being with patients. This is the second best thing. I'm not saying that I don't care about use cases. I'm saying that when you
[20:17] build a house on a shaky foundation, when you build a house so that you can have a nice kitchen instead of an entire home, when you build this lake house with that very that one very finite use case that you can get done pretty quickly,
[20:34] but unfortunately is not going to help the second and the third use case, you're not going to ultimately gain and win the future when it comes to all these novel tools that we're talking about. So, really this was the long-term game. I'm a long-term investor in my mind,
[20:50] right? I'm investing in in HSS into the future and trying to create a long-term asset. And so, with that, I'll close and thank you guys all for for for joining. I think we have about 15 minutes uh to take questions. So, there you go.
Learn more about the Databricks Data and AI platform.
The information provided herein is for general informational purposes only and may not reflect the most current product capabilities or configurations.