Build Enterprise Data Mesh with Databricks: Pharma Case Study and Best Practices
Summary
- BeOne Medicines built an enterprise data mesh on the Databricks Data and AI platform in under a year, treating data strategy as decision strategy to address fragmentation across internal systems, CROs, investigator sites, and external partners in an oncology research environment.
- The medallion architecture with federated governance through Unity Catalog creates AI-ready data products that power dashboarding, machine learning, and Genie from day one, with semantics and data quality embedded at ingestion rather than applied retroactively.
- Practical use cases including portfolio decision-making and global medical translation demonstrate how BeOne's data mesh accelerates R&D outcomes, with lessons learned around scaling cross-domain data products and breaking common myths about data mesh complexity.
Build Enterprise Data Mesh with Databricks: Pharma Case Study and Best Practices

Enterprise data organizations struggle with fragmentation, silos, and incompatible sources. In regulated industries like pharma, where decisions directly impact patient outcomes, data quality and trust become critical bottlenecks. BeOne Medicines faced this challenge across internal systems, CROs, investigator sites, and external partners, preventing fast, confident decision-making in R&D.
Watch how BeOne reframed data strategy as decision strategy and built an enterprise data mesh on Databricks in under a year. Discover the medallion architecture pattern, federated governance with Unity Catalog, and data product design that powers dashboarding, AI, ML, and Genie. Learn how embedding semantics and data quality from day one creates AI-ready data, enables cross-domain scaling, and turns fragmented sources into decision-ready intelligence.
🤝
Chapters
00:00Introduction02:02About BeOne Medicines Company04:29The Data Fragmentation Challenge07:34Data Strategy as Decision Strategy10:07Enterprise Data Mesh Architecture12:14Medallion Architecture and Federated Governance14:09Designing AI-Ready Data and Data Products16:32Business Value and Use Cases18:56Portfolio Decision-Making Use Case20:03Global Medical Translation Use Case22:13Lessons Learned and Myths to Break25:14Key Takeaways and Future Path
FAQs
What is BeOne Medicines and what data problem did they solve?
BeOne Medicines is a mission-focused oncology and hematology company with approximately 12,000 employees that has tripled in size over the past four to five years, creating complex data governance challenges across distributed global teams. The company faced data fragmentation across internal systems, clinical research organizations, investigator sites, and external partners that prevented fast, confident decision-making in R&D.
What does 'data strategy as decision strategy' mean at BeOne?
BeOne Medicines reframed their data initiative not as a technology project but as a means to accelerate specific business decisions in oncology research, including portfolio prioritization and regulatory submissions. This framing helped secure executive sponsorship and ensured that every data product was designed with a concrete downstream decision in mind rather than as a generic data asset.
How did BeOne build an enterprise data mesh in under a year?
BeOne built their data mesh on the Databricks Data and AI platform using a medallion architecture with federated governance through Unity Catalog, embedding semantics and data quality standards from the initial ingestion layer rather than treating them as post-processing steps. This approach created AI-ready data products quickly, enabling dashboarding, ML, and Genie capabilities within the first year.
What are the key lessons BeOne learned about enterprise data mesh?
This video covers BeOne's lessons around scaling cross-domain data products, including the importance of starting with clear use cases rather than comprehensive infrastructure, and challenging myths that data mesh requires large teams or long timelines. The presenters emphasize that a focused, value-oriented approach anchored in semantic alignment and data quality from day one can deliver a production data mesh in under a year.
Full transcript
[00:09] Good morning. Good afternoon, everyone. Uh thanks for your taking the time uh to attend this session. I know anything that follows a keynote is a tough act to follow, but uh we'll try our best to make this a worthwhile uh use of your time. Um So, uh
[00:26] Uh my name is Chirayu Desai. I have been in the data and analytics space for 20-plus years. Uh I was only 10 when I got started, but just just kidding. Um and uh I carry some of the same battle scars that many of you
[00:42] carry, right? Uh converting data into uh usable data products, insights, so on and so forth. So, uh today we are going to talk about our journey at B1. I want to thank the Databricks team for this opportunity.
[00:57] Um uh our journey at B1 is a young journey. We started less than a year ago uh with the goal to create a uh robust data foundation that can accelerate uh decision-making at B1. Um and we'll talk about the approach uh
[01:14] and some of the key decisions we made along the way. This is not a very detailed technical session, so we are not going to go through the you know, uh uh to the guts of our engineering frameworks and things like that. We are going to talk more about the strategy and the approach we took.
[01:30] Okay? Um I'll be joined by Subhranshu, my partner in crime here, uh and together we'll for the next uh uh 25 uh uh minutes, we will try to uh cover our journey and uh you know, what
[01:46] we have accomplished. Uh this is the customary forward-looking statement. Um you know, we are sharing the uh our experience for information purposes. Your mileage might vary, as they say it. Um And so, with that, let's dive in. How
[02:02] many in the room have heard of B1 Medicines? With the show of hands. Okay, a lot of hands from B1 itself. Uh but, uh so, we are sort of the best kept secret in the world of oncology and hematology. We are a
[02:18] uh a mission-focused company focused on curing many types of cancers. Uh specifically focused on uh oncology and hematology space. Uh we are a young company that has grown very rapidly, especially in the last 4
[02:34] to 5 years. We have tripled in size. Um we are a very globally distributed company with uh offices all over the world. Uh about 12,000 uh employees and growing. And so, we have all the complexity and
[02:50] all the challenges of a global scale of cross-functional collaboration uh and all the regulations, right, that our that life sciences has to deal with. And, you know, it is in this context that we have to enable
[03:05] uh uh you know, uh improved decision-making, improved outcomes for business through use of data and analytics. So, um so, that's a little bit about B1. Um before we dive into the presentation, I just wanted to start
[03:22] with what we would what I would like to you to take away at the end of this 25 minutes. Uh and this is going to be focused on four themes that we will talk about during the presentation, right? The first one is partnership is critical to
[03:37] unlocking value, and I'm talking about business uh and data uh teams partnership. Consolidation accelerates simplifying the architecture uh and trying to bank on a integrated platform is critical.
[03:53] Uh ready by default. It cannot be an afterthought in this day and age. You have to start with the mindset of making the data AI ready from day one. And then, you know, federated execution delivers, you know,
[04:09] the days of having one centralized team doing everything are gone. You know, the the boundary between technology and data and businesses are blurring. And we really need to focus on how do we enable value by leveraging whatever resources skills we have across the company.
[04:29] So, before we go into what we did, let me frame what the challenge was, what the need was for B1. So, these are very well-known stats that should resonate with everyone, right? But in context of B1, as I said, it is a company that grew
[04:46] very rapidly in last four to five years. And so, if you think about it, every business function did what they needed to do from a data and analytics perspective to achieve that growth. And that resulted in a lot of fragmentation, a lot of silos, a lot of duplication, because there was not a
[05:03] coordinated enterprise approach to data. And while we had a lot of data, we had an enterprise data warehouse where we had a lot of data, that was not resulting into insights that were trusted by the business. And so, really,
[05:19] you know, the common assumption all along has been like, let's get get data in one central place and then everything else will follow. That's required, but that's not necessarily adequate, right? Because you see multiple copies of data, you see multiple KPIs derived from those multiple copies of data, and then things
[05:36] don't match. And so, people download those data from the reports into an Excel and reconcile and create a third version of the truth, and that goes on. And that's not really accelerated uh accelerating the the uh outcomes or the decisions that the company has to make.
[05:56] This problem is much more amplified in R&D, right, which is the engine that drives a pharma company. Uh in R&D, data comes from dozens of internal systems, external partners, CROs, investigators, so on and so forth. It comes in different formats, it comes at different cadences, it has different
[06:13] definitions. Uh and you know, a lot of the decisions that are made based on this data are high-stake decisions. Right, it is hard to uh I wouldn't say they are irreversible, but it's hard to reverse them. There is a material impact if you make a wrong
[06:29] decision about committing to a portfolio or committing to a study or how you design a study. And so, the accuracy, the trust, and the readiness of data to make a decision becomes really, really critical. And not having that really delays
[06:45] getting the uh most importantly, the drug in the hands of the patients. Right, at the end of the day, we want to put our therapies in hands of the and we want to put our therapies in hands of the patient as soon as possible. And we are talking about diseases like cancer, right, where every day matters. Uh and
[07:02] so, as we step back, right, we and and thought about, you know, what do we do differently here? What do we need to focus on in order to make these decisions more data-driven, more trustable, more confident for the
[07:17] leadership. Uh we understood that you know, the key outcome is accelerating the critical decisions, right? It's not about getting more data, it's not about, you know, uh sophistication of a platform, but it's really focus on the decisions that matter the most, and then try to get the
[07:34] data ready to enable that. Um And so, our data strategy essentially became a strategy to enable decisions. The critical decisions that the company has to make. And so, our data strategy is really a decision strategy.
[07:51] And we are a small team at B1. Right? And focusing on the critical decision really allowed us to focus on what matters because we cannot boil the ocean. And we really have to focus on what's really, really important. And in order to then enable those decisions, right?
[08:08] We need to make sure that the data is trustable, usable, and ready to make those decisions. And that's where you have to change your mindset from you know, go from pipelines to data products. Right? Data products are those trustable, curated, usable sets of data that the
[08:25] business can use uh to make decisions. Shared semantics become very important. Right? Data is bits and bytes unless you put a meaning to it. And making sure that everyone is interpreting that data consistently, we have a a common definition for KPIs and
[08:43] metric that are used to make the decision is really, really important. And that takes work. That takes partnership. That takes working with the domain experts and the data teams together to create that. And then lastly, it's not about technology. Right? It's not about, as I
[08:59] said, how sophisticated your platform is or at the end of the day, it is about are you enabling that decision? And if so, how quickly? Right? So, that's that's really what grounded us on what we needed to focus on and and sort of build the foundation that we
[09:16] needed to build. So, with that, we'll switch to what are we building? Right? How are we enabling that? I will hand it over to Subhranshu to kind of cover that. Yeah. Thank you. Thanks, Chirayu. Uh so, I'm Subhranshu. The way you see
[09:32] my role is I'm a information architect, but I'm a forward-deployed information architect. So, I work with R&D team. Although I come from enterprise and IT background, I had the experience of working with many pharma companies in the R&D space. So, I bring in that R&D
[09:50] domain knowledge, and that's why I was selected for this role so that I could bring that technology domain knowledge together and partner truly with the R&D teams to understand their business challenges in detail and structure the
[10:07] enterprise information architecture and the data mesh. Now, what we're building essentially is a enterprise data mesh, which is built on four key principles. We looked at what are the key decisions we need to make,
[10:22] and what are the data products we need to build, and for those data products, what are the data sources we need to ingest. So, our enterprise data mesh is built around this principle of integrating data from multiple disparate sources and building them into data products.
[10:38] And those data products are organized into data domain. So, we followed a domain-driven data modeling principle. We brought those data domains to answer those important question and prioritized based
[10:54] on the top questions that we need to answer, and R&D was the first domain which had the most important set of questions that we need to enable and decision-making process that we need to help our R&D teams with. We then started
[11:11] looking at governance because while you build the data products, it's important to have the context. Every data product should have the business domain context at the table and the field level, and that's something that technology teams are not the best
[11:27] teams to put those definitions. So, we had to partner with our domain SMEs, our business users, and we have to work with them to get those definitions into it. So, that's why it's called more of a federated governance. So, our
[11:43] data SMEs and business SMEs are who are governing the data products. And finally, all of these data products are registered in a data marketplace, have the check gates around data quality, around metadata readiness. Without
[11:58] those, we don't certify those data products and register them in the data marketplace. So, at the end of the day, if we are creating a data product, it has to be reusable and has to be used for many decisions across the R&D organization, and that's what we made.
[12:14] Now, for the builders and the developers in the room, on the right-hand side, you could see the architecture, which is pretty standard what Databricks also recommends, a medallion architecture, where we have the bronze layer, we ingest the data once, and we we that's
[12:29] why the enterprise team governs the bronze layer, so that we are not duplicating data ingestion. The silver and the gold layer are where our analytics ready data is created. Now, most important part here is on top
[12:45] of that layer sits the governance, which is which is the real context in the days of AI. And that enabled us to have all different type of use cases, dashboarding, AI, BI, ML, AI. All those
[13:04] sort of use cases to sit on top of our data lake. And then all our users interact also directly with those modules. Now, the good part is when we started the journey, we had a mix of technology in play, where Databricks was the core, but there was a
[13:21] different ingestion framework, different data quality framework. We're evaluating multiple technologies for metrics, but in the very first few use cases, we realized that the real speed comes with you know eliminating the tool sprawl. If there
[13:36] are multiple tools, there are multiple moving parts, and then you you end up integrating tools and spending a lot of time. So, we are essentially all this that you see on the right is built on Databricks and the different components of Databricks.
[13:53] Now, on this day and age, we talk about AI-ready data, and for us, what that meant was two primary takeaway I'll I'll say from this slide. The first one is you model your data in a way that it is
[14:09] analytics-ready. So, our gold layer, we do all the pre-science, metric calculation, and summarization, so that it's very easy to get all the data needed for decisions in one single table or at most two or three tables. You could get all the decisions. So, when
[14:24] I'm talking about study-related metrics, a study can have 150 metrics, but the way our gold layer is designed, all those 150 metrics are in one single study table, which is sourced from at least eight different sources, and goes through multiple data model. In
[14:41] silver layer, the study has around eight tables, but in the gold layer, it is only one table. So, it's essentially making it very user-friendly and ready, and now your AI can easily use that one table and answer your questions accurately, right?
[14:57] And to enable that, when we started, we had enterprise data catalog, which was not on Unity Catalog, because of course we were just starting our journey on Databricks. But we built the connector between the enterprise data catalog and the Unity Catalog, so that when in Unity Catalog a
[15:14] steward is writing a glossary of term, because there was a glossary of term which was already created in our data catalog, we just have to map those glossary to the Unity catalog elements so that we could reuse and synchronize our enterprise data catalog to our Unity
[15:30] catalog. Now, to do all this, when we started our journey, we started with our data strategy because that's the most important pillar. Then we looked into the governance because
[15:45] guardrails are very important while when you do things at the enterprise scale. We made the platform decisions and then we created the engineering capability or the muscle behind it. And then we realized there are certain accelerators
[16:00] which can be used by multiple different business functions in our organization. So, we built those accelerators. So, these five pillar in total which helped us build a template in the first few use cases and then we are able to reuse this template across different
[16:16] functions and enable our rapid growth of the data lake as the foundation. Now, all that is the foundation, we also want to quickly talk about some of the use cases and the early business value that we are able to deliver. Now, as you
[16:32] see, we are still very young and early into this journey. We have spent around 10 months. So, it's still we are building a lot. It's no way we are done with this and I don't think any data program truly ever gets done because if
[16:47] you're done, you're not innovating, your business is not changing, your questions are not changing and that's that's a bad sign. A data program should always be on. And from that perspective, where we are today, we started 10 months ago. We are able to activate eight
[17:04] different domains. We started with clinical, but then we used that template to onboard seven other domains and we are able to support 15 plus use cases to make those decisions faster in those use cases.
[17:20] And then we have built a certain number of silver and gold layer data product to answer those questions. And you could see the data is sourced from 70 different sources into the data frame. We coming from different background and experiences within life sciences
[17:36] consider this as a very good achievement and throughout B1 we had a very talented team and good support from Databricks to achieve that. And this can be done in a short duration of time because we template ties everything. We
[17:52] reduce the tool sprawl and the standards were followed. Because we are a small company, it was easier for us to get people together and follow the same template. Now some of the use cases that you can see we started with R&D resource
[18:08] planning. We then enabled spend analytics for finance, inventory tracking for manufacturing. And then we are working now on many other interesting use cases around you know, site financial metrics reporting for leadership team. And then
[18:25] also a very interesting use case was around our global medical translation where we enabled translation across and there is a case study I'll talk more detail about that use case. So there are many such use cases as I said
[18:40] 15 use cases but we picked two of them which were very I'd say important to us and we had made a lot of progress in those areas. So we have picked those two use cases. So first of all was the portfolio level
[18:56] decision making at the R&D team where before the data mesh they used to have multiple silo data marts and the R&D analyst has to go to multiple data marts to pull different type of information to pull together the PPT which can be presented to the R&D leads and see you
[19:14] know decision makers in R&D. With the enterprise data mesh we were able to pull together around 30 different sources, create those data products needed across R&D decision making process and enable the portfolio level decision making. It's still I would say
[19:32] we have enabled the early set of use cases and now they see the power in it and there are more and more demand for this data within R&D. Uh where Databricks really helped was with the right side of Unity Catalog, the
[19:47] governance and the Lake Flow. We were able to accelerate our journey and make it available to R&D teams in then quick 4 to 6 months of starting our journey. Another use case which I'd like to touch upon is the global medical translation.
[20:03] So as a company we had to translate our CSRs into multiple languages for making them and you know submission ready. And these are not just like any other general translation. These are high fidelity translations which are submission ready packages and need that
[20:21] amount of accuracy which can be taken and submitted to the regulatory authorities. So for that we were able to pull together end-to-end solution using Databricks platform. We used Databricks Lake Flow to keep all our
[20:38] Lakehouse to store all our context behind how we translate and what are the semantics around our documents that we are translating. We used the AI playground to experiment with different models and then we built
[20:54] end-to-end forward deployed app using Lake Base and Databricks apps using which people could upload a document and do the translation. The great part about this The was that you could see the metrics on top, the accuracy percentage
[21:10] and the productivity it was able to deliver just by using the power of the platform and the semantics that we are able to give to the platform to do this translation. Now, I covered a lot of our use cases. I
[21:27] would ask Chirayu to talk about the lessons and what comes next. Thank you, Subhranshu. Um So, that was a rapid fire on, you know, what we built and some of the use cases we have enabled.
[21:42] I especially like Subhranshu how you characterized your role, like forward deployed architect. Uh we've heard a lot about forward deployed engineer, but the forward deployed architect, the architects who are embedded in the business, who are both technology and domain
[21:58] fluent, who can work with the business and translate into data solutions is a critical pillar of our success. We are following the same uh sort of recipe in other domains that we activate. And so, uh if it catches on, forward
[22:13] deployed architect, you heard it here first. Just remember that. Uh So, uh lessons and what comes next. Uh any journey is is fraught with learnings and uh you know, uh myths that we had to kind of address as
[22:30] we went along. Um so, these were the four myths we had to break to move fast, right? I I don't think this is a surprise to uh majority of the people in this room. Uh the first myth is governance slows down teams, right? The word governance
[22:45] is has sometimes considered has a negative connotation attached to it. And so, what we had to be very careful about is not introduce governance as a gate in our process or a check and control point.
[23:01] We pivoted governance to be part of the process itself, right? As we work with the business to design our data products and create the shared semantics and define the domain ownership, it very much became a part of our life cycle and we kind of try to make it seamless.
[23:19] So that at the end of the day, right, when you create that data asset or a data product, it has a well-defined owner, well-defined semantics, data quality checks all built in from day one. And I believe that, you know, AI is a really, uh, you know, what data governance was
[23:35] missing was a killer use case. The advent of AI and, you know, Genie AI, BI and the fact that you want AI to answer your question is that killer use case where you can now get the business and other stakeholders to say, "If you want to If you want to AI to answer your
[23:51] questions reliably, this is something that we need to We need to put the context on the data. We need to make sure that AI understands your data as well as you do." And so, I think that has, uh, helped us in this journey. Um, a data foundation is infrastructure. I
[24:07] think we have to look at it as a business capability. Right, Databricks started as a data platform, but it is becoming a data intelligence platform. And so, a lot more personas interact with that data platform. It's not just the data engineers, it's not just your platform engineers. You have your data
[24:24] stewards interacting with the with, uh, your data platform. You have your data consumers now interacting with with your data platform directly through, you know, your AI, BI dashboards or Genie spaces or whatever else. And so, you really have to design it as a business capability that works
[24:41] for all those personas. And with the right governance, right? Because as you democratize more, you also need to make sure that you are governing it, uh, properly to make sure that the usage is within the guardrails of, you know, regulatory compliance as well as, um,
[24:57] uh, you know your enterprise policies. Uh AI readiness is a separate program. I think what emerged for us is that good data practices and embedding semantics, data quality within our data products
[25:14] made our data AI ready. It's very hard to do it after the fact and to go back and add those things. And so one of the critical learnings for us is make it part of your design process itself and think of it as your data is not going to be just consumed by humans, it's going
[25:30] to be consumed by AI. And so that context has to be understandable by machine. Now, AI is really great at understanding unstructured data uh and and so it doesn't have to be your complicated graphs and things like that, but you really need to have those
[25:47] definitions figured out. You really need to have that context attached to your data products so that it can be used by AI. And then lastly, big foundations take 18 to 24 months. Um our journey is a proof, right? We achieved what
[26:03] Subhranshu mentioned in about 10 months. Uh and I think that's really by reducing not only the sprawl of the tool, right? Like you know, bank batting on an integrated platform like Databricks. Uh in my past life, I have you know, we all have integrated many
[26:19] many tools and you spend a lot of time integrating them. But you're not spending time delivering the data products. So here, we decided to bank on a uh integrated platform like Databricks and really reduce the complexity of the architecture. Simplify the architecture,
[26:35] focus on what really matters. Uh B1 is a young company. We don't have a lot of the legacy that you see in other companies, so we were lucky that way. But it really important to sort of start decoupling that so that you can focus on what's really important.
[26:51] Um and uh Uh you know, the other thing is focusing on the right things, right? The focus on critical decisions enabled us to focus on the right things. Because if you try to boil the ocean and bring all the sources into the data lake, we would have we
[27:08] would not have been able to deliver what the business really needed. So, and create this momentum that we have created. So, those were the key lessons learned for us. And so, as we end this presentation, uh you know, fourth
[27:23] as I said, I promised for everyone to take home. Uh partnership is the unlock. We couldn't have done this alone. Uh today in the room, right? I have my business partner sitting in this room, you know, uh cheering us telling this story. I think that kind of collaboration is really, really
[27:38] important. And data it's not just about technology, it's also about infusing that uh enterprise domain knowledge into your data, and that only comes through through that partnership. So, that's really, really important. And one of the things that really helps is focus on
[27:55] what matters to business. Right? Uh and that will go that is what gets them excited and create a co-ownership, co-accountability towards that outcome. Uh consolidate the technology, as I said, uh simplify the architecture where you can,
[28:10] uh so that you can focus on what matters. Uh Uh don't make AI-ready an afterthought, make it part of uh your design uh and your practices up front, uh and I think it will emerge uh you know, uh and then, you know,
[28:27] federating the execution, as I said, um you know, every organization has multiple data teams, uh and you know, the key is to align everyone on an enterprise approach, and sort of leverage all that capacity, all that talent you have to build
[28:44] towards the outcomes. Um you know, I I think uh you know, centralizing or complete decentralization is not the answer. You really have to figure out how you would federate, you know, set the right standards, the right approach, the right guardrails,
[28:59] use the platform to enforce some of those guardrails, and uh let the let the teams go and do what they need to do uh to really achieve the outcomes. Thank you.
Learn more about the Databricks Data and AI platform.
The information provided herein is for general informational purposes only and may not reflect the most current product capabilities or configurations.