Clinical Data API Modernization: Lakebase and Databricks at Novo Nordisk
Summary
- Novo Nordisk replaced fragmented, self-hosted Postgres silos with Lakebase, a Postgres-compatible engine built on the Databricks lakehouse, delivering clinical trial data 30 times faster than legacy systems while maintaining pharma regulatory compliance.
- The platform standardizes ingestion from the Viva clinical trial system through automated API layers built with Databricks agent skills, and implements attribute-based access control to enforce clinical blinding requirements.
- The architecture distributes data management across five data cores aligned to Novo Nordisk's business units, with Unity Catalog providing unified governance that reduces integration time from months to days.
Clinical Data API Modernization: Lakebase and Databricks at Novo Nordisk

Novo Nordisk modernized clinical trial data integration by replacing fragmented, self-hosted Postgres silos with Lakebase, a Postgres-compatible engine built on the Databricks lakehouse. The transformation delivers data 30x faster than legacy systems while maintaining the regulatory compliance, audit trails, and data governance critical to pharma. this video walks through the architecture, migration path, and governance patterns that enable speed without sacrificing control.
Learn how Novo Nordisk standardized ingestion from their Viva clinical trial system, built automated API layers with Databricks agent skills, implemented attribute-based access control for blinding requirements, and reduced integration time from months to days. The approach generalizes across Novo's business units through five distributed data cores, unified governance with Unity Catalog, and a managed platform that balances self-service flexibility with enterprise guardrails.
🤝
Chapters
00:00Introduction and Why Clinical Data Modernization Matters01:26Novo Nordisk: 100 Years, 47 Million Patients, Pharma Scale02:29Clinical Trials, Blinding, and Regulatory Requirements04:07Current Challenges: Silos, Manual Processes, and Trust08:53Solution Requirements: Lineage, ABAC, and Data Quality09:58Platform Architecture: From Concepts to Data Cores14:47Data Ingestion Across Distributed Business Units17:39Multi-Cloud Implementation: Azure, AWS, and Fabric24:57Viva Integration and Legacy Architecture Problems30:48Lakebase: The Modern Clinical Data Solution32:41Data Standardization, Domains, and Automated APIs35:21Key Benefits: Faster Innovation and Compliance
FAQs
What is Lakebase and why did Novo Nordisk adopt it for clinical data?
Lakebase is a Postgres-compatible database engine built on the Databricks lakehouse that provides transactional capabilities alongside analytics and governance features of the broader platform. Novo Nordisk adopted it to replace fragmented, self-hosted Postgres silos that created manual processes, limited data trust, and made regulatory audits difficult in clinical trial workflows.
What is the Viva integration and why does it matter for clinical trials?
Viva is a widely used clinical trial management system in the pharmaceutical industry. Novo Nordisk standardized ingestion of Viva data through automated API layers built with Databricks agent skills, replacing a complex legacy integration architecture and reducing integration time from months to days.
How does Novo Nordisk handle clinical trial blinding requirements in their data platform?
Novo Nordisk implemented attribute-based access control within Unity Catalog to enforce blinding requirements that restrict certain team members from seeing specific trial data until designated phases are complete. This governance capability is critical in regulated pharmaceutical trials where premature disclosure of treatment assignments can compromise scientific validity.
How does the five data core architecture work at Novo Nordisk?
Novo Nordisk distributes data management across five data cores aligned to their business units, each with their own ingestion and processing workflows while sharing unified governance through Unity Catalog. This structure balances self-service flexibility for individual units with the enterprise guardrails and compliance requirements of a top-10 global pharmaceutical company.
Full transcript
[00:06] Uh we have an opportunity now to give you a little bit of introduction to how we at Novo Nordisk modernize our clinical data integration and with a special focus around the use of Lake Base and the Databricks Lakehouse. So we in this context is I'm Henrik Lyng.
[00:24] Top of the image line and then we have Jonathan Selsing over here. Uh and then we have Ahmed. We are all architects, but we are the type of architects that still code and we haven't disappeared in the ivory tower, which is super cool. Um
[00:39] and then without further ado, let's dig into it. I'll give you a little bit of introduction to why some of the control elements that we've been implementing are important and also important in relation to some of our integrations and the use of Lake Base in that regard. And then when I've
[00:55] given you a little bit of introduction, I'll do this speedy, then Jonathan will give you some insights into the platform and how we built that out. And then I will will do a showcase on the use of Lake Base and how that has modernized both our team composition and reduced
[01:11] the complexity of the solution that we're running. We are taking a very concrete example of integrating data from a solution called Viva that most of you don't know about, but in pharma it's really big and really hot. Um in pharma
[01:26] speed is of the essence and having a solid data foundation is mission critical for ensuring that we can process at speed and we don't have a lot of room for error. There's not margin for doing major failures or
[01:41] hiccup both from a regulatory compliance paid, but also from from the fact that a lot of the compounds that we are working with are ultimately lethal. So the potency of for instance insulin is is quite high. Uh Uh, so making this
[01:57] just a few numbers about Novo Nordisk, who we are. Uh, one of the top 10 farmers in the world serving approximately 47 million patients on a daily basis now. Uh, therapeutically, it's very much diabetes, obesity, and rare disease.
[02:13] We deliver approximately half of the world's insulin uh, across the globe. Uh, and we have been around for 100 years. Uh, so we recently had our 100 years uh, uh, anniversary and appreciate sort of
[02:29] the value set that we are built on. Um, the value chain in pharma is very much anchored around early research and then molecular engineering development. So this is where we do the clinical trials. And a significant part
[02:45] of how we do clinical trials is a very controlled process of different phases. And then for every phase you do what's called randomization. So you distribute uh, drugs either active or inactive to patient based on uh,
[03:01] based on on randomization. And to hold that randomization, we are using something called uh, concept of blinding, making sure that we do not know who got what compound. And that is the what forms the foundation of the statistical analysis and the
[03:17] justification for why we can actually come to the conclusions we we hold. So this blinding thing is actually quite important and Alva is going to come back to that a little bit. But then once we've developed the product, we need to make sure that we have it approved and then produced.
[03:33] Finished product supply chain, commercial strategy, uh, commercial execution, and then finally serve the patients with drugs at affordable cost. Affordable cost is becoming increasingly important uh, these days um, as we are feeling uh, amongst others in
[03:49] the American market. Uh, some of you will uh, I don't know if any of you are on GLP-1. I think it's an eighth of the grown of the adult population in in the US. Um, but the price has just reduced significantly. Uh, which of course impacts us. Underpinning this is quality
[04:07] and then a lot of IT solutions that historically have been very siloed. So, that's part of what we are trying to address. And then when we built this platform, we wanted to make sure that we had adequate controls in place
[04:22] so that we could actually do the needed qualification, validation, and control of the data all the way through the pipeline without doing too much qualification and validation ongoingly. And I think we've succeeded with that.
[04:38] So, making sure that we can actually work with speed with data under proper control and make sure that because failing on this is unacceptable within the pharma sector. Um, then discovery, clinical execution,
[04:54] submission, and manufacturing. The world that we live in historically has been about 1 million ideas gets to 10 products that get tested and hopefully out of those one make it to the market. Uh, that is one of the
[05:10] reasons that some pharmaceutical products are pricey. Uh, that's because the rate of success is actually relatively low. This has and this process historically has taken around 10 years, ballpark figures. This is now being shortened,
[05:25] uh, so the amount of ideas is exploding. Now we are more in the sort of 10 million or 100 million ideas because of AI coming into the labs and gene synthetics, uh, making sure that you can target more faster. Uh, so now it's to
[05:40] the 10 or 100 million ideas that lead to 10 that sometimes lead to one. And the 10 years, I think the ambition there is to cut that to at least five years, accelerating the deployment of new products to patients. Um we are surrounded by data, but we are
[05:57] starving a little bit on trust. And a lot of that is because historically a lot of our processes have been based on document creation. Because ultimately what we sent to the authorities was documents where we described what we did, how we did, so forth, and they
[06:13] approved based on that. And then there was a sidecar which which contained data. That is shifting as we speak, so there's a lot of acceleration around making sure that the data flow gets streamlined and automated, that the silos between the different operational
[06:28] systems is removed, that there's no more duplication between teams, so more transparency on if you build a product, I can actually reuse it. Uh you see how some of the capabilities of Databricks become valuable here. And then that
[06:44] manual validation, we are accelerating that by automation, so more automated tests, test execution, feedback loops, so forth as part of the process. And then super clear ownership of data. So that we have well-defined
[07:00] owners that can actually make critical decisions around the data. That goes for access control, but that also goes for quality parameters and indications. Um so weeks are lost in validation before we actually get to the analysis. There's a lot of rework because of
[07:17] inconsistencies in sources and alignment between different points. Delays in trial submissions and data issues, and then bottlenecks also on the manufacturing side. So this is where we need to sort of work smarter in order to
[07:32] get cheaper products to patients. Um the pressure that we are exposed to is very much driven from AI-driven uh That's back to the 10 million ideas instead of 1 million. Uh the data volumes are increasing a
[07:48] lot. Omics is basically genomics, it's proteomics, it is very large data sets that we are working with in order to find the right targets, but also to find the right treatment options that we could pursue. Uh a lot of real-world
[08:05] evidence, so making sure that we actually do create more targeted treatment for sub-populations. We are not all the same. Uh so making sure that we actually have differentiated treatment to the differentiated patient cohorts is critically important. And then more
[08:21] regulatory scrutiny, making sure that we actually have visibility on where did what come from and what happened to it under the way. Authorities are very keen to identify if there was fraudulent in the process and who modified what data. And then there's a high cost pressure
[08:37] that is hitting us. So, smarter, uh faster, cheaper. And in order to do that, sort of we need to make sure that we actually have really good data lineage. So, visibility around who did what and what did they do.
[08:53] That we have attribute-based access control. This is one of the features that we are using to drive the the blinding process. And big kudos to uh to Databricks for actually stepping up on enabling that. Um we are having a massive benefit from
[09:09] that. More focus on mastering of data and shifting that from the classical sort of super-mastered one source uh to something that's more this is a trustable reliable uh source that we are actually using and working much more
[09:25] dynamically with that. That's one of our focus. And then data quality indicators, so getting reliability in between teams and within the process. And then also data products to make sure that we have the scalability and that the different teams are not replicating work
[09:42] uh from each other. So, this is just give This gave you a little bit of introduction to what we are trying to achieve. Also, for those of you not familiar with Pharma, uh it's sort of a quick intro to what we are trying to do. And then, Jonathan, can you explain a little bit about what we've been doing?
[09:58] Yes. Yeah. So, I'll take you through sort of a whirlwind tour on some of the large sort of the large-scale concepts around how we approached our platform and sort of where it where we've landed now. I thought it was important to start
[10:14] with this this this pyramid around what it takes to achieve uh impacts with AI because I think sometimes there's can be a a disconnect with between how executives see the impact of AI and the reality of
[10:30] of of getting there. And I think we've seen it yesterday in I got his keynote talk where there was this almost like an a Marvel action type movie around the impact of AI with large sound effects and a lot of things going on.
[10:46] Um and that that might be true, but I think in order to to get there, there's a lot of hard work. Uh and there's a lot of hard work for many many people. And the larger an organization you have, the more people are going to partake in this this hard
[11:02] work. So, I think we all agree that the AI piece is is what we want to get to, but then and Databricks for us is the technology, but data is really the important piece, data and and metadata around data. I can see that I jumped a
[11:18] bit in the slides, so I'll go back here. Um and and I think with with Databricks we found a really strong technology partner that that handles many of the technology pieces. Uh but and we are making some way on the data piece, but we are by all means
[11:35] uh not there yet. And I think we have a multi-year journey before many of these things are in place. One piece is getting the data in, the next is is metadata. As Henrik mentioned, Novo is a 100-year-old company. That means we have a uh and it's a very large company that
[11:51] we that means we have every possible version of IT you can imagine everywhere in the IT chain. Some of it is uh old servers running from the '90s that keep running '90s that keeps running. Um and it's a very broad company, so you have a decentralized IT where they do a
[12:08] lot of both good and also crazy stuff around in the company. Um so, we started 3 years ago 4 years ago in our Databricks journey where we sort of said, "Let's try and dig deep. Let's incubate a new platform in one of the single business areas." So, this is the
[12:25] development space the R&D space that Henrik and Arvil are coming from. And then we built sort of a new Databricks native platform from from ground up in that area. And we did it with a few principles in mind. We know we're a large company. That means we can never let IT be a
[12:40] bottleneck. We have to make it self-service. We have to delegate responsibility to the units that use our platform. I think as Henrik was mentioning, you can never have a very fast car without good guardrails and good brakes. So, that means we have to think in
[12:55] governance from the beginning. I think that was a bit dated difficult with Databricks from the beginning. I think a lot of the Apex stuff that is coming out now is going to make this increasingly easy to make the guardrails sort of a part of. Uh but it has sort of
[13:11] been a transition. And then finally this this data product concept that I think everybody or many people work with right now where we think data is a standalone valuable asset for a company and collecting it and re-exposing it to the rest of of company provides value in itself. And I
[13:28] think some of the stuff that has come out of our users of Databricks is the increased proliferation of the same data products to different lines of the value chain that you would never think about to begin with, but due to the availability of these
[13:43] data products, you suddenly see these cross-pollination projects that make that that have a large impact. And the way we did it was that we sort of took a a hybrid mode of the of this data mesh thinking where you distribute a lot of effort in a lot of different
[13:59] places. We were sort of say we can't do too much distribution because we've tried it before and the overhead on the single units are too large, but what we can do is that we can map sort of this model to somewhat to our to our value chain that we have inside our
[14:16] company, so traditional business units. And then every business area is sort of responsible for for the activities that takes on themselves. They typically have their own IT departments in in Novo. So that means we have sort of a central
[14:31] IT department that takes care of the of the platform setup and the governance around it. And then we have sort of distributed smaller Databricks platforms that is responsible for subcomponents of of operating. And that's among other things getting the data in, ingesting all of
[14:47] the strategically important data we have across. So for us, I think so far it looks something like this. We have our product supply and quality organization, we have our early research, we have our commercial, we have our development, we have our enterprise IT functions for IT
[15:03] data products because there's surprisingly a lot of IT native data products that make sense in in in many cases. And this model has sort of is sort of a bit back and forth where we have the domains that are responsible for the ingest piece, but then really
[15:19] you only get value out of using Databricks by having your your business teams building on it, building the the Databricks apps, they're building the dashboards, building the consumer interfaces where you you you you touch your customers, and then in our case uh
[15:35] touch touch our patients. And uh those use case teams, they need to be able to leverage and use Databricks with the smallest possible IT overhead as possible because these are to a smaller and smaller degree IT functions as we can more and more easily
[15:50] write code applications, uh it becomes increasingly important that the overhead for managing this IT uh uh follows around. So, there we have this uh uh managed offering. And we call our sort of ingest capabilities, our ingest frameworks, we
[16:06] call them data cores, and we have five of them uh spread across uh the company by now. And uh that is a dedicated data engineering IT function where they do internal prioritization. I think we do sort of a dot voting type system where
[16:21] uh we have a list of all the data products that might exist, and then we say uh how difficult is it to get in, and how much how many use cases are voting to get this in, and then we get sort of a prioritized list around uh what do we get in uh and when.
[16:36] So, that means because we have this distributed setup, it's not only a central IT function, it's a distributed business function where they get the data in. Uh and then this benefits uh not just the the local domain teams, but everybody in the company. Because we've implemented Databricks in
[16:51] this full mesh solution where all workspaces and all catalogs can speak to each other. That means that if one one domain ingests an important data product, it immediately becomes available to uh the rest of the company. Uh and I think we've seen some of the
[17:07] central IT data products that least have extremely fast been consumed by uh very many uh parts of the value chain. And then in the end, we have these consumer products. So that's the gold products, the single use case teams. They work with the data product that's
[17:22] the data that had already been put in the platform. Um to build the the value cases. So how does it look like sort of from a practical perspective? So we started on Azure Databricks Azure. That was where we started our journey. Novo, as you
[17:39] know, is a large enterprise and that's that's synonymous with Microsoft. So we have a very large Microsoft presence, of course. So we started with Azure Databricks. It provided a really good integration with our intra ID. So there was easy governance and connection
[17:56] with our identity management systems. And then, as Henrik mentioned, there's a lot of guardrails and controls also when working with software. If any of you in the pharma industry, you know that we have to operate under controlled processes and this kind of stuff. And that means that we've sort of
[18:12] nobody touches anything by hand anywhere in our ecosystem. Everything is fully automated. Uh we have sort of a GitOps approach. And that's good because we can be in a lot of control. It's also kind of bad because all of this control because now we have sort of
[18:28] put the process IT process to our consumers. And we do have VPs approving pull requests in GitHub. But they hate it. So that's sort of something we have to work on internally. But we are getting there.
[18:43] By now we've sort of spread across two clouds, both AWS and Azure. Now we've sort of mirrored the setup across both clouds. And we do provide multiple options of using Databricks. There's sort of a a more self-managed offering where people
[18:59] they can bring their own Databricks almost and off they go. So they get a lot of freedom. Uh that's only for the business cases that have that have a correspondingly large IT budget. And I think that has shown to be not a very effective pattern for us.
[19:16] So I think we are transitioning more and more into a managed offering where IT takes a larger and larger responsibility for the platform components. And the model looks identical. We have this So every use case gets three catalogs, gold, silver, and
[19:33] and bronze catalogs that are attached to them. We attach them to workspaces, and we have workspaces globally across all regions for clouds, and they select at consume time where do they want to go? And where they want to go can be based on a
[19:49] number of factors, hardware availability, you name it, these kind of things. Where is your data being generated already? If you're sitting in Brazil in some of our production sites, you would probably use some of the local capabilities in Brazil. It's the same with the US, you would use
[20:05] local capabilities. And then I think what has shown over the last year at least is that the multi-cloud thing has become increasingly important to do for hardware availability because we see that many of the hyperscalers, the cloud providers, they have difficulty keeping
[20:20] up on the hardware side in in some of the capabilities and some of the capacities, especially for for newer hardware types. And that means we see sort of an increasing trend where we move workloads around to where there is availability on
[20:37] the compute side. But that being said, I think we do have a principle that you need to minimize the cloud plus region combinations you grow across because everybody hates egress and and the performance hits that come here. So this is sort of a design principle we
[20:53] have. But it's not it's not enforced because there can be, I mean, data sovereignty laws for instance around where you keep the data. So it's always a trade-off between what you can do. And then I think one of the interesting points here that I think and I think one
[21:10] of the main reasons why we went for Databricks to begin with is the sort of integration to the to the cloud layer. Because, of course, with Databricks, there's a lot of nice stuff you can do inside in the in Databricks. But, because we're a large company, many of the IT systems we interact with,
[21:25] maybe they didn't implement the Databricks API. Maybe they don't know Databricks, but they but I think by now S3 and storage accounts have sort of become the the facto API standard for many integrations. So, that means we see quite a number of use cases where we are like "It's great.
[21:42] You provide Databricks. We don't care. We have to integrate to some other system." But, what we can do, because we sort of let them store the data on blob, we can anyway, even though even though if the use case can't leverage Databricks, they store the data on blob store that
[21:58] is mounted as Databricks storage. And that means even if the local team can't use it, we can still expose this data product through the Databricks ecosystem. Um and I think this freedom of of diverting from from sort of a strict technology
[22:14] has proven very effective. Maybe just as much in in sort of a buy-in component. And because a lot of IT typically are very opinionated around they do. And we can say now, "Yes, you can do Databricks, but you don't have to use Databricks."
[22:31] And then, a lot of units are are okay, then I'll go and try it. And then, fortunately, Databricks is getting better and better because then once they're in, they're hooked. And that's of course good for for us in the in IT. We have a large sort of a consumption ecosystem around Databricks as well.
[22:48] We have a relatively large Snowflake presence as well. And I think those integrations have become better. They're not still super good. I know there are some integrations, and I think we're working actively with them. Uh but it's it's always when you are working across two large projects, there
[23:06] is always a certain amount of friction. That the friction is fortunately dropping for us, but it is still there. But we do have a wide sort of ecosystem of consumers and especially Power BI Fabric type, we have quite a number of users and there the integration between Azure Databricks and
[23:22] Fabric is uh is doing wonders for us. And I'm hoping the new Lakehouse RT is going to sort of move more workloads into into Databricks because I think traditionally a lot of our business users, they were very happy about the Power BI semantic model performance you
[23:38] would get with the with the import mode. I think hopefully now the that balance has shifted. So if you want to get a really performant dashboard, probably you would going to go for Lakehouse RT and I think for for central IT, that's a good thing because that means we can
[23:53] consolidate our our IT landscape. And we've been at it for a couple of years now. Um and and I think we've gone very far in in what we have. And I think now we are seeing full end-to-end projects that are
[24:09] basically building Databricks native applications. So applications that live their full life Databricks native, Databricks apps hosted lake base as the back end for it. And I think we are seeing sort of an early adoption, but I think what we are seeing also at the
[24:25] keynote here, I think that transition is going to increase because there is sort of this I guess that's part of the game that Databricks is playing that it's it's better on the inside and the more people are on the inside, the better it's going to keep get. And I think there's sort of a certain self-gravitating property
[24:42] around the Databricks that we hope to to leverage in our business transformation. I think we've seen that uh quite successfully. So Now I'll transition over to Arvid that will take us through how the nitty-gritty things of one of
[24:57] solutions are looking like. Thank you. Perfect. But yes, let's talk about our situation in our development area. We have a lot of systems, a lot of internal systems, a lot of sales providers. We have a lot of partners we do integration to. So it's a quite complex landscape. And I had to to show
[25:15] the real benefit of data bricks I want to go a little bit back and what we had in the past. We got to think that that's not good. Yes. So but if you look at the at the slides then you can see we have our Viva system that that is where we conducting our clinical trials. That's one of our key system. So when we
[25:30] are starting up a study, when we are creating sites, when we are start to monitor the participant that have their first visit onto the last visit and also where we document and say our all of our document beginning. There's still a lot of documents. So we are not not done with that yet. So there's a lot of documents we getting from site that we
[25:46] need to document so we have a copy of that. And there I can say there's a lot of need for for having that to should be able to be used for AI use cases. And yes, but this is the current flow that we have today in the past we also strategic partner with Viva so we have a good collaboration with them. In the past
[26:02] they didn't had a way for getting data out in a good way. They had a waste API and then you can say you needed to query that with some day time finding out what have changed. You need to find some deletes. Collect that and then you need to expose that. And if you have a lot of consumers
[26:17] that need to use the API there are also some limitation on doing that with a lot of frequent calls and stuff like that as well. That's where you need to offload the data so and that's where it's really nice that we have data bricks so we have a good way of of getting data in and that's quite performance. They can do a lot of request. Yes. The flow was that
[26:33] that we called Viva then we had a new set of application. You can say that was getting the data and all of the conversion of what to extract that was done in Viva. So that was a lot of team. So there we specified this object we allowed to get out, this fields we allowed to get out. And so it
[26:49] was really to trouble because that needs to be part of your release and that really delayed for need for new objects that we need to have this communication. Then the build application put the the data down to S3 bucket. Uh and then we need to talk to another team with AWS.
[27:05] This should do a lambda and we say for digesting data. Then we need to talk to say a Postgres database and that team as well for doing some schemas and what is coming and indexes and and stuff like that as well. And then we also needed to give data to S This is
[27:20] quest for notifying our consumers as well that the data has changed. So we are delivering an API on top of this. And I saw a lot of steps. So so it would not be ready. So so we could not do something in days. We could not do it in months. It's months and maybe also
[27:37] longer time as well. And and it was not scalable. Yes. So so the solution that will not scale for AI. So it's too slow for doing experimentation. No way to prototype. Schema management was a mess. We had no documents in there.
[27:53] Uh audit logs that was not possible to to extract in a good way. We had some process mining use cases where we want to learn and how people are using the systems so we could improve the process as well and we want to do that real time so we could change processes and see the effect on on what we're doing.
[28:09] The next one I think that's aha the IDP integration. So our security team they don't like you can say having static keys and stuff like that as well. That's not a good practice to have. Uh so and there we have tried with our new that was maintaining your set of Postgres. They tried for years to making that work. It should be easy to do but at
[28:25] least they had really really issues in the in in making that work. So so that was really really difficult. Fine grain access control like Kim that said with the blinding it's important for us that we can say because we don't have a way to do role based access attribute based access in Postgres at that time then we say it required us to create views for
[28:42] each use cases maintaining that and really really difficult to maintain. And uh the biggest thing for also for us and because we are farmer, that is we need to ensure that data integrity. And when we are calling the API, then we also don't know if indexing on Viva has been done correctly as well.
[28:57] So, we could miss that there was an object that was missing and and stuff like that as well. And also when we commit, it was difficult. So, so, so, it was not good. And then performance was not good. And the biggest thing, that is the one in the bottom, it was not a standard solution. So, we have a lot of Viva
[29:13] systems, quality, safety, regulatory, and and others. Everyone built their own solution. So, so, so, so, and that make it difficult. So, if a new AI solution needed to be data from two different Viva vaults, then it's two-way process of of learning and stuff like that as well. So, that was something that we
[29:28] want to change. Yes. And that's where I can say we Yeah, that's really partnership with Viva. We had some good talks. We would like to have real-time data, but oh, they didn't give us that. But at least they gave us uh something called direct data API. That's where we can get files from them every 15 minutes on what has changed.
[29:45] And then they're ensuring that data integrity is in place. So, we get all objects out. Uh I think that that that that that was really really nice and we don't have the need for real-time that 15 minutes. I think that's okay. Uh used in the past we have past been a lot older data. I think it's more important
[30:00] for us that we can rely on the data. We know the data is consistent and people is not getting bad data. Yes. And then we also worked on documents. And there I can say they didn't have a solution for that. So, there we needed to use the the API for for for getting documents out. But we in the past we had
[30:16] a use case every time then it was can I integrate these documents, get those out? And then we had a lot of documents around. And then who ensuring that what if I can say a document has been wrongly called or wrongly classified? Then you wish that they have a document that they shouldn't have. And that's where we wanted to build a
[30:32] standard solution that you can say on data mix I can say for for how for how to handle that. So, we ensure that that that that I can say that the documents are removed or also change parameters and then do a with downstream systems so we can say if if the metadata should be updated. Yes. So, we're sending data to
[30:48] lakehouse, getting it to lakebase. Uh and that was really really good because we really need to say a strong backbone for for doing a lot of requests on API. And it's really been nice to have part of the journey from the beginning with with lakebase. I think there's also been
[31:04] small issues and stuff but again, I think they're fixing stuff and they're coming a lot of new stuff. So, so really happy with the new things that is coming. Then we have a little bit of conflict on the right side as well because we also need to deliver our streams. Also, I think Databricks have done some good work on zero boss and stuff like that as well. So, I think they're going in the
[31:20] right direction but they are still you can say we try to push people to to go directly to Databricks, read the tables so you can say directly but there's also some that have a standard pattern how they want to do that and there we need to use streams in some use cases. Yeah. Yes. Awesome. So, this is the the new
[31:36] way of doing it. So, we wanted to have one team that that that that could that could do it end-to-end. Uh we would still like to say the back-end system to agree on what we're doing. So, so I'll I'll show that later how that look like. Yeah, but but so we still have a a configuration but instead of it was that
[31:52] you're allowed to do this and this and this, then it is the other way around. What should we not extract? And right now on all of the we we only have to say one exception list for one of the walls but all of the other ones we we're sending all data. And that's where you can say when we are getting it to bronze layer, there we are
[32:08] already enabling a like and and one level access as well. So, so we can filter stuff out so we know if data has been blinded and stuff, then then it's not there. We can only do that on Deltas right now. It would be nice to have a feature also on lakebase as well. So, so please say that to Databricks and
[32:26] then they'll have to support that. Yes. Yes. Awesome. Then you can say we have a silver data bucket and that that's where we have tried to define some data domains because in the past we did a lot of different modeling and different teams doing that and have a lot of replication. And then you start
[32:41] to see the business have have different you say how data is is not the same thing that I'm seeing. And and because we have this amount of system that we need to standardize that. So, we have tried to define different data domains so we know you can say owner with that and area that is responsible for certain set of the data. And that will make it a
[32:57] lot more better you can say and also for us to rely on the data. Then we are getting into lake base to with two tables. Um yeah. Yeah. And I'm really really happy about yesterday that the now this is you can see they're saying that you say that's not needed anymore that you can say what you have in delta set will also will be
[33:13] in lake base directly so it can be queried. So, I think that's really really nice. So, that's also a a step less that we need to do. Yes. Then we're using a fast API. Yeah, as I said we used MuleSoft in the past. And I really think I really like you can say what you can say Databricks has done you can say with
[33:30] skills you can say for for for for doing stuff get something faster on running. That's also develop you can say but we need to do that's what we make it production grade but at least I think it's a good starting point to to show something with some prototyping and see the data you can say also in in a good way. And then we have notification that's
[33:45] being sent back to a Kafka. That's it. Yes. On the second side we have you can say documents you can say with the best API. Um yeah. So, so that that's going in and we have a lot of documents so so that's there's a lot of stuff being transferred. But we have a a good uptake
[34:01] on AI use cases. We also tried we standardized the way to get data over to Databricks. The next step that is also to all with back and everything try to standardize that as well so so we can instead of each use case need to do all of this groundwork for doing it because there are so inside you need to have for for learning how to pass these documents
[34:16] in a good way so it's easy to use for AI. Yes. Awesome. Yes. So, adding a new object in the past three teams a lot of coordinations and and you can say it's we had PI planning you can say that's every second month at that time and then you can see you need to
[34:31] arrange with people on on when they could do work as well and and so so it it could take a long time. But now we are down to one team can can do everything. And then it's just we say we need to get the approval from from the VSU system to to get data out and then we can start getting the data and then
[34:47] it will flow automatically to to to to bronze. I think that's really nice because there are also users that's more tied to the system that's getting data out. And this is really giving value to to to them that because it can be also difficult in Viva to understand their data. That's where we also started to get business users in to use data bricks
[35:04] as well to get a good understanding on their data as well. So so they can see where are the data quality and stuff like that. Yes. Yes, key benefits. Yeah, faster innovation. So that was really good. One team, one pipeline as I said. Schema evolution was a lot easier. Yeah, really bronze,
[35:21] silver, gold. Really really nice. Yeah, object and document access control. So we have a lot of say media data that we also getting over. One thing that is for example our HR data. That's coming from success factors. We need to see what kind of department
[35:37] people are part of. Other media data on the employee that could be part of granting access to data. Then we also have our training records. For example with blinding, have you taken blinding training? Are you compliant with your blinding training? If not, then you should not see the data because else we can say the authorities can can come after us and
[35:53] that's where it's really nice that we are getting a lot of these media data in and then we can automate. So instead of asking people, can you please review the access then maybe they are not compliant the day after and stuff like that as well. It's not bulletproof. So so you need to have something that is automated and that's where data bricks really helping
[36:08] us with that. We have a fully managed offering now. So so we can say lake based and everything. And we are getting a reliable notification out. Um All of our legacy dependencies you can say that's starting to say away from AWS. The only thing that that is confirmed but yeah, let's see if we can
[36:24] get streams out of data bricks at some point as well. And then we also getting document auto looks out. Yes. Then just a little bit of demo. So I not showed the So we have a lot of stuff to do today. So I try to do some screenshots and it's moving a little bit
[36:39] about but it's still being done. So this one that is you can say the document extract and this is for quality system. And here you can see there's a list of objects and that we are filling out. So so it will not be transferred. They are still complaining a little bit because the access model on Viva that is still
[36:56] you download a zip file you can say from Viva and there you have the all of the objects. And then they all then you still have the object and who can have access for doing that. But at least it's nice that we can fill it out so it will not go to Delta some problems. Yeah. Then we have you can say configuration with all of the documents as well.
[37:13] And there you can say there's a lot of minor version and we are only interested in major version approved version you can say that we get getting transferred. And then you can say this system they have had a filter they wanted to you can say filter out strictly confidential data. And then that's part of the
[37:28] filter. So if somebody is wrongly classified that the document has been not confidential then it will be that to begin with. When it's being changed then it will be removed and so it's not available to see in data bricks. Yeah. So and one thing that we have to done as well also with AI we try to give
[37:44] this to our business owners please approve this and they cannot reach equal. For me it's it's quite simple it's almost like reading it but but again then we have used AI please print a really nice readable format on this and then they are happy with that. Yeah.
[38:00] And that they are really happy about the new process instead of creating documents we need to document in document tool and then they are part of the pull request so that's why we have our documentation. Yes.
[38:15] Then next week then we have created our branch where we getting all of the the tables in and we have already seen a lot of object on people in central area you I a bit of a better understanding. We have tried for 4 years ago. Please, can we get everyone to to classify your data? It seems impossible. They cannot do
[38:31] that. Uh but but but now you can say you can say by having all this data available, then they start to get see the raw data. And then now they have done it suddenly. I was quite impressed again. Now we have done it. Yes. So so that's really good because then we can start using that for actual data access and all of that as well. Yeah.
[38:46] Yes. Then uh yes. Go to script next slide. Yes. Then uh yeah, for the API layer um used the the data bricks um agent skills. I really think it's a it's a strong tool.
[39:02] Uh we have showed it a little bit internally to our legacy developer on some of the platforms uh just saying, if you have this table, how long time will it take for you to to to do this with a click ups? You can say there is we have to I think 30 different kind of objects in in
[39:17] in this one with a lot of fields. It will take months. You can say if they're doing this manually. It it's not automated. And and here you can say where we have done we have done it automated. And we have also done it in the code so it's now it can be done dynamically as well. So it can if when they are released, then all of
[39:34] this is being done automatically. We have a configuration where we specify what kind of objects you want to do, what kind of schema is is it, and also what kind of filter parameters we have. And then it's quick generating the the the the the fast API specification and then we have that available. So that has really improved our way also to deliver
[39:49] API faster as well. Yes. Perfect. Yes. This is really is is the API that is being exposed. Yeah. So this was running locally, but it is also running in production, but I didn't want to show the URLs that we
[40:05] have in production. So and and uh yeah, that's it.
Learn more about the Databricks Data and AI platform.
The information provided herein is for general informational purposes only and may not reflect the most current product capabilities or configurations.