Skip to main content

Real-Time Clinical Data Sharing: Pfizer and Enphase Modernize with Databricks Delta Sharing

Summary

  • Pfizer and Enphase replaced overnight batch clinical data ingestion with a near real-time architecture using Delta Sharing and Unity Catalog, reducing data delivery latency from 24 hours to 3-5 minutes across 90,000 tables and multiple clinical studies.
  • Delta Sharing eliminates data copy between Enphase's UDMS clinical data management system and Pfizer's analytics environment, while Delta Lake manages schema evolution through protocol amendments and provides ACID guarantees and time travel for GXP compliance.
  • Serverless compute reduces operational costs, Auto Loader simplifies ingestion, and immutable audit logs support the regulatory requirements of a GXP-validated, 21 CFR Part 11-certified clinical trial environment.

Real-Time Clinical Data Sharing: Pfizer and Enphase Modernize with Databricks Delta Sharing

Watch: Real-Time Clinical Data Sharing: Pfizer and Enphase Modernize with Databricks Delta Sharing
Clinical trials race against time. Every delay in data access extends the wait for patients seeking breakthrough treatments. Pfizer transformed clinical data sharing from overnight batch jobs to near real-time delivery by partnering with Enphase and Databricks to build a modern, compliant architecture handling 90,000 tables across multiple studies.
Learn how Delta Sharing eliminates data copy, Delta Lake manages schema evolution through protocol amendments, and serverless compute reduces costs while maintaining GXP compliance. Discover zero-downtime deployments across clouds, Auto Loader ingestion, Unity Catalog governance, and immutable audit logs enabling real-time clinical insights with reduced latency from 24 hours to 3-5 minutes.
🤝

Chapters

FAQs

What is Enphase UDMS and how does it integrate with Databricks?

Enphase is a SaaS provider for large pharmaceutical companies supporting clinical trials; their UDMS is a GXP-validated, 21 CFR Part 11-certified clinical data management platform built for real-time data collection from global clinical sites. UDMS integrates with Databricks through Delta Sharing, enabling Pfizer to access clinical trial data in near real time without creating data copies.

How did Pfizer reduce clinical data latency from 24 hours to 3-5 minutes?

Pfizer and Enphase replaced a slow batch ingestion architecture with a real-time cross-cloud data sharing model using Delta Sharing and Unity Catalog on the Databricks Data and AI platform. By eliminating orchestration complexity and data copying, the new architecture delivers clinical trial updates to analysts within 3-5 minutes of data capture at clinical sites.

How does the architecture handle schema evolution in clinical trials?

Clinical trial data evolves as protocols are amended, which previously created rigid schema management challenges. The new architecture uses Delta Lake's schema evolution capabilities and protocol amendments to handle these changes gracefully, while ACID guarantees and time travel ensure historical data remains accessible and auditable for regulatory purposes.

Why did Pfizer and Enphase choose Delta Sharing as the foundation?

Delta Sharing was chosen because it enables real-time, governed data sharing without the need to copy data between systems, eliminating the orchestration complexity and redundancy of traditional batch pipelines. This approach also supports zero-downtime deployments across clouds and maintains the immutable audit logs required for GXP compliance in pharmaceutical clinical trials.

Full transcript

[00:07] Good afternoon everybody. My name is Jennifer Roller. Clinical trials is a race against time. Every delay that we get every delay in accessing data means a delay in knowing if a treatment actually works. Today, we're going to show you how
[00:22] Pfizer and Enphase change that equation. We will walk you through how we move from a slow batch ingestion to a real-time cross-cloud data sharing using Unity Catalog and Delta Sharing.
[00:43] At Pfizer, we truly operate at a global scale. As one of the leading biopharmaceutical company, we're running clinical trials across several therapeutic areas. At any given time, we have hundreds of studies that are currently in flight. Every
[00:59] single one of them producing constantly evolving data. And our teams need the ability to act on them the moment it's ready. Good afternoon everyone. Thanks you for joining and uh my name is uh Jayaprakash
[01:15] Rao. I'm one of the co-founders and CTO for Enphase. Uh we are a SaaS-based solution for uh uh large pharmaceutical companies supporting clinical trials. We have two products. Uh today we is focuses on uh
[01:31] our real-time uh CDMS called UDMS uh and also we are manufacturers of developers of RedCap Cloud uh So, Enphase is a uh UDMS is a real-time enterprise uh clinical data management
[01:47] platform purpose-built for global pharma. Our flagship UDMS uh is a GXP validated 21 CFR part 11 certified, uh uh captures and manages and uh reports uh
[02:02] real time uh uh data collection from various uh global clinical sites. Uh The design from the day one, I mean, has been always say instant data visibility and faster data driven decisions, I mean, downstream all the way to the
[02:18] submission. So, here is what we are going to cover today. So, we'll start with the what was the friction, I mean, what made our old uh solutions, I mean, painful.
[02:34] Then we'll talk about, I mean, partnership and why we have chosen Databricks and uh you know, to solve the problem. In section three, we will go through the architecture and both where we started and where we are now. And we'll go through the journey, I mean, where in
[02:50] and uh including the lessons, I mean, we have learned, I mean, the and you know, if you were to start over. This has been a three-year journey almost. So, section four is all before and after and what actually changed operationally,
[03:06] how Databricks actually helped us out in reaching that goal. And uh and we will actually share with you where we are uh all this is heading, I mean, to the next phase. So, So, let's start with uh the friction. Jennifer, uh take us where we were
[03:21] before and uh where we are uh you know, and give us a snapshot. Okay. We're going to um walk through the the journey that we'd have uh we've done. Clinical trials require modern speed. In clinical trials, speed isn't just a
[03:38] metric, it's a patient's life. From the moment a patient enrolls, we are constantly asking ourselves, is it working? Is it safe? The answer to those questions leaves in the data, but under our legacy system, our legacy
[03:55] legacy architecture, we assert That legacy architecture, don't get me wrong, it has served us well, right? But with that, our clinical data was always a day behind, it's always a step behind.
[04:11] First first friction, 24-hour lag delay. Historically, we were stuck with a 24-hour delay. We run our data maybe maybe daily, maybe at most four hours a day. If Imagine this, if a clinician logs in in the
[04:28] morning, right, puts in a vital update, our analysts are kept in the dark until the next morning. When patient safety is on the line, a 24-hour blind spot seems like an eternity.
[04:44] Rigid data systems, our second friction. As we all know, clinical trials don't stay still. Study protocols get amended, it gets changed, new data points get added. And when a trial evolves, the underlying
[05:00] structure changes right along with it. In our old setup, there is no graceful way to resolve this when things are changing. It would just cascade downstream and break the pipeline.
[05:16] You simply cannot build shareable, trustworthy data sets when your architecture breaks when real world changes. And finally, right, we had friction on fragmented collaboration. When we're sharing data to our partners, it was
[05:33] slow if not manual. It The workflow was literally you export the data, save it to a file, SFTP it, and when our partners our collaborator receive it, the data was already a stale
[05:49] data snapshot. If you look at this particular diagram, it looks very logical, right? So, you have the clinical site, um you have your legacy EDC, it goes to your Pfizer analytics, right? But, for those of us
[06:04] that actually maintains their architecture, it was like a daily crisis. Starting on the left, when our clinical sites enters the data, as soon as the data is entered, it stopped. Why did it stop? Because it has to wait for that fixed nightly
[06:20] schedule. That Think about waiting uh a full 24 hours to receive a critical email. By the time it reached our analytics team, it was Yes, it was accurate, but only for yesterday. We were using our brightest minds
[06:38] by Our brightest minds are being forced to look at old data, and they're acting like historians, and they instead of telling us what's happening, they're telling us what used to be true. But, the absolute worst part about this
[06:55] particular architecture, it was it was very fragile. With these rigid systems, I essentially call them pipelines made of glass. Whenever there's a change, right, those pipelines break in silence. What does
[07:10] that mean? We were making decision on incomplete data. Right? It's like a silent bomb ticking, and it's getting detonated further downstream. Protocol amendments, whenever we do that, if anybody of you has done
[07:26] clinical trial, we do a lot of protocol amendments, and this required extended downtime, sometimes for hours, if not days. At the end of the day, right, this old legacy setup that we currently have here, built a brick wall between us and
[07:43] our strategic goals. It built a brick wall between the patient and the possible cure. I'm going to hand it over to Jaya. He's going to go through the architecture. Okay. So, Jen just mentioned about all the issues
[07:59] what we had with our older applications and uh served faithfully for many years uh uh but as we look at uh you know, we wanted to build in a real-time uh clinical data management system.
[08:14] We wanted to fundamentally change the architecture. So, essentially what uh you know, through collaboration with uh Pfizer and Enphase and Databricks uh we built a a three-party application over here uh essentially if you look at the on the
[08:30] left over here uh uh is Enphase uh UDMS captures and manages all the clinical data and at the source uh Databricks provides the governance and orchestration and data intelligence layer
[08:45] and uh Pfizer GP DIP uh uh uh is the primary consumer of this uh governed real-time data. So, what does Enphase UDMS do? Uh it is not just a form builder. It's a unified system designed for full
[09:03] life cycle of clinical data uh from site level capturing I mean from site level eCRF and non- non-CRF from various central labs and all and uh going through any medical coding and safety and risk-based monitoring
[09:20] randomizations and all the way through you know, which uh you would expect a modern GSP uh validated 21 CFR CFR part 11 system would do. And uh Databricks and specifically the component I mean we are talking about is
[09:37] Delta Sharing over here. So, to transfer the data from the, you know, be on from Enphase Red Cap Cloud or to Pfizer, the key technology edition what we have chosen over here is Databricks Delta Sharing.
[09:53] And you know, most approaches to clinical data sharing evolve copying data from, you know, extracting and transforming and eventually moving it over to, you know, the sponsor systems over here. So, Delta Sharing, I mean,
[10:11] and inverts this model. I mean, there's nothing like this concept of copying and making, I mean, all this. So, there is only one copy of data which is actually shared in between a a CDMS vendor like Enphase and Pfizer.
[10:27] So, the data stays stays with where it lives, I mean, and in this case in our cloud over here. The recipients, I mean, can access in through a open protocol, I mean, you know, I think it's actually Databricks just announced another solution called Open Sharing, I
[10:43] mean, but say Delta Sharing also was actually a precursor to that. And regardless of whether you leave it on AWS or Azure GCP or any other clouds also here. So, so no data movement, no pipelines to
[10:59] maintain. This is actually the fundamental core value of Delta Sharing over here. And Pfizer has its uh there is a Unity Catalog and it is a governance layer and it actually provides all the uh
[11:15] provides the governance across all data assets uh and essentially the ability to, you know, have permissions at not only at table, but down to row and column level. And this is quite
[11:30] important over here. And there are other features which uh Delta sharing offers, something called serverless compute. It scales from zero to, you know, uh uh uh infinite capacity and and it has all the notification layers, it has the
[11:47] web hooks and everything and retry logic. So, we eliminated using uh the the lake flow and Delta sharing Unity Catalog entire orchestration layer here. So, data simply from most of them
[12:02] all the way from in-phase UDMS to Pfizer, uh uh you know, in a in a very simple and uh orchestrated way, okay? So, we also transitioned everything over to what is called as a lake flow. And now everything is orchestrated
[12:19] within Databricks, uh notifications are built in and failed jobs are visible. And we can re-trigger anything I mean uh in in place of it. So, that is actually the say at a high level what is uh how we have modernized over uh
[12:36] clinical data sharing I mean using uh using uh the partnership with Databricks, okay?
[12:59] Okay, so real-time is actually the the key here. And here is what we have built. I mean, we have five layer layers of the solution and each with a specific job. On the left is coming in is the in-phase UDMS, uh
[13:14] it has uh the ability to do the data capture, uh the validation, and uh and it pretty much uh uh real-time across various different clouds over here. So, essentially you are, you know, sites are able to enter the data and you know,
[13:31] everything gets uh you know, you go through the data management and eventually ends up in in our uh ClickHouse I mean which is our uh large data warehouse over here. So, within from ClickHouse I mean we move the data into uh the data is
[13:48] actually stored on AWS S3. And uh that's where the the Databricks Databricks actually picks up. So, data lake I mean what we have is actually we use something called uh delta tables over here I mean which is a
[14:04] key fundamental thing in the delta sharing I mean essentially you are uh keeping a version of the data at any time and you can always roll back at any time to to a given version here, okay? So, we are essentially in the
[14:21] in the uh Databricks data lake uh it allows us to do what is called as acid guarantee, the integrity of data, time travel at any time, and uh schema evolution. So, this is schema evolution Jen mentioned because uh clinical trial
[14:38] tables are not uh you know, with protocol amendments I mean the you may drop or add columns into your tables. So, so we need to actually maintain all this stuff. So, all this is actually main managed I mean within the lakehouse itself. And uh then there are jobs I mean which orchestrate uh
[14:55] they they do the conventional ETL, the reporting, and everything and and uh finally we have the the delta sharing I mean which we talked about. It is the open protocol, allows you to share data with the uh no data movement. I mean this is
[15:12] actually the key. So, data stays within the our cloud but we are able to share it with any uh sponsored partners. So, you know, in this case uh uh Pfizer access the data. If If they are able to share it downstream with their
[15:28] own CROs and uh their own partners having our share. And the the key outcome of this entire thing starting I mean from the left when the data enters into our big house uh warehouse it is near real time all the way to
[15:46] the sponsor. I mean, this is all happening in near real time. When I say near real time, you're looking at uh less than 3 to 5 minutes of delay here. Okay? So, this is actually the reporting ready data ability to consume and uh
[16:03] further down you can you know is uh statistical computing I mean, you know, they are able to actually uh consume and start actually running the analysis at the time. So, So, this is actually the benefit what we have. So, So, the consumers are here what you're
[16:18] seeing I mean, which is actually the layer four and five are primarily uh Pfizer's uh uh downstream systems which are about 26 of them and Pfizer's partners including CROs as we and potentially regulators and uh
[16:33] and getting all this uh near time uh uh near real time safety signals. So, So, this is the advantage and what we have over here. Okay? So, four outcomes, near real time access. So, essentially we are
[16:49] eliminating the batch delays. Zero down time. And protocols changes deploy instantly and they are available downstream for systems. Cross cloud sharing I mean. So, sponsors are may not be on AWS all the time. So,
[17:05] delta sharing allows us to actually share the data with any cloud. And uh open protocol which is a key thing I mean, so there is no vendor lock-in or uh and and you are able to actually transfer out to anyone. Okay?
[17:33] Okay, this is another view of the same thing. And you know, this actually highlights on the Delta sharing and the Unity catalog and essentially the data flow, full data flow. So, three lines top to bottom. And you know, we talked about having the
[17:50] clinical sites, I mean they you know, like any other thing they enter their uh CRF data into the InPhase CDMS and it flows into our ClickHouse, which is a columnar OLAP engine and uh so, every row is timestamped and soft
[18:07] deleted only I mean we don't actually allow any deletes out there. So, this is our GXP constraint over here. We the warehouses fully I mean compliant. So, the in terms of the process, essentially the which is the second one,
[18:24] so, the there is a uh ingestion config table. I mean that actually drives drives the pipeline. So, every time we introduce a change into our ClickHouse, I mean essentially the updates them in this ingestion config
[18:40] table and it allows them in which studies actually which data gets moved over and what cadence, you know, so and along again no code changes or nothing. So, there is a Databricks component solution
[18:57] called Auto Loader, uh which actually detects them in these are changes, incremental changes happening coming into the into the into the Delta Lake and uh so, it automatically pulls the data and it looks at
[19:13] any schema evolution and eventually uh they the data lands in within a short duration near real time into uh onto the AWS S3 and Unity Catalog. So, we also look at Unity
[19:29] Catalog over here to the right. It uh essentially row and column we talked about it permission set at level. And again, Unity Catalog is also uh immutable they they maintain immutable audit logs which are
[19:45] very important in our GXP or 21 CFR compliant environments over here. And so, from so I basically I think from Unity Catalog, I mean you have Delta Sharing which is critical here.
[20:03] And essentially you are from distributing the data and using pre-signed S3 URLs over here. So, as I say, you know, the entire process is zero data copy I mean and zero egress so everything stays within the cloud over
[20:18] here, okay? And essentially the data gets actually mapped over to the Pfizer's environment. So, everything what we have on the InPhase cloud on our S3 is automatically mapped so you're no data
[20:33] copying just actually the metadata is getting copied all the time. And essentially the that that uh from there I mean Pfizer systems are able to actually pull the data using uh multiple things JDBC or including SQL or
[20:49] or even using any of Power BI tools or any other tools I mean what you have here, okay? So, net net the result is a 24-hour batch delays are down to less than 5 minutes over here. This is actually at volume I mean we are
[21:06] talking about not uh phase three trials. I mean, some of them are mega trials. I mean, with 40,000 subjects and everything. And this is actually at scale. I mean, what we are talking about. I mean, you know, the and uh you know, this is a very unique achievement. I think uh working with uh
[21:23] Databricks engineering, we were able to actually uh uh you know, architect, I mean, a solution at scale. We are talking about about 90,000 tables are getting synchronized over uh multiple hundreds of uh studies. I mean, so that that's the kind of scale what we are talking
[21:39] about over here. Okay? Next one. Okay. So, this is a a quite you know, we started this journey with uh Pfizer and uh Databricks more like in 2022 or uh you know, so
[21:55] uh there was a lot of uh lessons learned. I think I mean, 4 years ago, I mean, you know, look at the success. I mean, when we look at uh Databricks today, I mean, you know, I think I mean, they evolved over time. And it was a great partnership I mean, over here. Uh you know, uh
[22:11] on the journey, I mean, we invented, I mean, we have matured, I mean, Unity Catalog. We have created, I mean, several things together. And uh so, you know, initially, we managed some access controls outside of Databricks. And uh and then we went to migrate all the
[22:27] governance into Unity Catalog. I mean, you know, it was more significantly complex. And you know, luckily, our partner actually stepped in. And and uh so, if I were to start, I mean, a lot of the groundwork already has been done. Uh but, you know, I think I mean, you
[22:43] know, if I were to advise you guys of anyone, I mean, you know, uh wanting to architect a solution, Unity Catalog should be your first place. And it is actually uh you know, first edition, I mean, as you enter into this uh so- a solution like
[22:59] this. The second thing, I mean, we talked about this schema evolution. Now, clinical trials schemas, they are living things. I mean they every PPC could introduce a change and they change constantly. So, new fields having drops field, you
[23:15] know, so you know, drop index or add a new column, drop a new column having protocol amendments. So, so we had to actually make sure I mean all the downstream systems are uh as you you know, instead of assuming the schemas are static
[23:31] uh and uh so we need to we had to you know, initially they were breaking repeatedly. Now, we build the schemas on on read patterns. So, Delta Lake evolution uh policies having from the start I mean so
[23:47] schema drift is actually now a pretty much a a old thing having so it it is a an important uh issue what we have solved over here. The third one is actually main main thing. So, there are a couple of ways of uh
[24:02] using I mean a SQL warehouse and what is called as a server serverless computer I mean in uh in Databricks uh uh you know, I think we initially over provisioned having some of our uh SQL warehouses and uh essentially
[24:19] incurring having you know, we were guestimating having you know, here is what what we think I mean a synchronization of a 90,000 tables would actually cost and you know, we were calculating and trying to optimize using different calculators and everything.
[24:35] And for peak loads and all, but honestly the best way to do I mean these days if I were to do this one here, you know, let serverless do its job and uh you know, essentially having you configure and uh and get this uh uh you know, right from
[24:52] from day one. Not not later. I mean there's no reason to double guess and you know, go through that pain also here. So, our you know, I think we after we moved out to our serverless our cost profile actually dropped our cost actually dropped.
[25:09] And you know, we have seen you know, the proof is in the pudding when you see your monthly AWS or cloud bills actually going down. Okay? So, the the last one is uh you know, I think I mean you know, sponsors should not anticipate I mean everyone uses
[25:26] a single cloud. I mean you know, I think people are using I mean different things. Uh it all depends on their corporate policies. Uh uh you know, AWS is fine, Azure is fine, GCP. You know, as a technology companies I mean and even sponsors I mean need to
[25:42] anticipate that uh some of your own consumers of data downstream would be on using multiple clouds over here. So, a solution I mean need to be planned I mean right from get-go. Uh that uh you know, we need to anticipate there would be cross cloud
[25:59] and uh so factor into the those additions I mean right from from from the early days and often 20 a solution. You know, you need to start thinking about validation, testing, and everything also as part of the solution.
[26:18] Listen it. John John Okay, so um in this slide I'm going to go before and after. What is the transformation that's been done? Very quickly cuz um Daya already reiterated this, right? So, before there was significant delay between data capture and access. Now, that data is
[26:36] now available very fast. It's available shortly after it's entered inside. That changes how teams work in Pfizer. Before every customer, there's a different pipeline. Custom pipeline for each partner. Now, there is no data
[26:53] movement. We're not sending files anymore. Our partners basically access the data in place. Before we have very complex monitoring. Just imagine we have so many different systems here. Now, we have everything in one place. If something fails, I can see
[27:10] it, I can investigate it, and then I can re-trigger it. And the good part is you can always ask Genie what went wrong and how to how to resolve it. Right? That's the new thing now. Whenever our support team calls, "Hey, I have an issue." I said, "Ask Genie. Ask Genie why it failed and
[27:27] what Ask him what are the possible solutions." Now, another thing is before whenever there's a protocol changes, that always equates to a downtime. Now, we have zero downtime deployment. This was a significant operational burden.
[27:44] Now, it's a non-event. The last one is fragmented governance. Now, everything is in Unity Catalog. All of our tables, our access control, our lineage, we have regulatory um grade audit trail. And then in it's
[28:00] all in one system. Now, these are not incremental improvements. This is a fundamental different on how we're operating. It's a a fundamental change on how we operate our clinical data.
[28:20] Okay, I think I think we are going to run out of time, but uh you know, essentially what we have developed having this partnership if I would summarize and leave any questions over here uh uh you know, it is architecture is never destination. I mean, it's always
[28:35] need to be proper foundation. And I think what we have done over here is essentially having uh enabled having the a a, know, I think having the not just a proof of concept, but a production
[28:52] architecture for for the not just for Pfizer, but for the industry. And you know, I think I'll leave it there, and I know we are coming up on time here.
[29:07] Okay, so today we've shown you the mechanics, but let's be very clear here, right? We are not just optimizing pipelines, we're not just bridging clouds. I would like to think that we are forging the lifelines of medical discovery. An overnight batch job is not just a
[29:23] technical debt, it's an overnight wait for a patient. It's an overnight wait for a family. Every bite that stagnate, every insight that waits, extends the clock on suffering, uncertainty, and lost. We did not build
[29:40] this real-time for We did not build this real-time foundation, right? For just performance metrics alone. The true measure of our platform, of our infrastructure, isn't processing speed. It's giving back time.
[29:57] Remember, the patient is waiting, and they are not waiting for smarter data archive. They are waiting for a tomorrow that they can believe in. So, let's steer the data silos, let's close the data gap, untether our AI, and
[30:14] accelerate the heartbeat of discovery. Let's deliver the breakthrough that they desperately need. They need it today, because for them, tomorrow simply cannot wait. Thank you all for coming. Okay, thanks.

Learn more about the Databricks Data and AI platform.

The information provided herein is for general informational purposes only and may not reflect the most current product capabilities or configurations.