Skip to main content

Sponsored By: Informatica | Unlocking AI-Ready Data

Summary

  • Most enterprises are not short on data but on context—AI agents are forced to reason on fragmented, stale, and inconsistent data because every system stores its own version of the truth with its own rules and definitions.
  • Informatica MDM solves this 'context gap' by creating golden records through matching, merging, and survivorship capabilities, then making those governed master records available to AI agents via MCP servers and Unity Catalog integration on the Databricks Data and AI platform.
  • Eisai, a life sciences company, uses Informatica and Databricks together to unify clinical research and drug safety data, turning fragmented enterprise data into an AI-ready foundation that governs the full pipeline from ingestion through audit trail and query monitoring.

Sponsored By: Informatica | Unlocking AI-Ready Data

Watch: Sponsored By: Informatica | Unlocking AI-Ready Data
Fragmented, inconsistent data is one of the biggest barriers to AI in life sciences. In this video, learn how to use Informatica MDM to build a unified, AI-ready data foundation on the Databricks platform. We'll cover: Golden records on Databricks How Informatica MDM unifies disparate data sources into a single, authoritative record; enabling AI in life sciences How governed master data unlocks AI/ML workloads across clinical research and drug safety; MDM for pharmaceutical research How Informatica MDM's matching, merging, and survivorship capabilities solve real pharmaceutical research challenges at scale. Walk away with a concrete blueprint for turning messy enterprise data into a competitive advantage.
Talk By: Sean Jacobs, Global Data & Analytics Technology Lead, Eisai ; Stefan Glover, Director, Strategic Alliances, Informatica

Chapters

FAQs

What is the context gap in enterprise AI?

The context gap is the problem where AI agents must reason on top of fragmented, contradictory, or stale data because enterprises lack a shared understanding across systems. This video explains that it is not a lack of data volume causing AI projects to stall in pilot—it is the absence of consistent, governed context that gives those numbers business meaning.

What is Informatica MDM and how does it create golden records?

Informatica MDM (Master Data Management) unifies disparate data sources into a single, authoritative record through matching, merging, and survivorship capabilities. The resulting golden records represent the most complete, accurate version of an entity—such as a customer, product, or clinical subject—and can be shared with AI agents and analytics systems through Unity Catalog.

How does Informatica integrate with Databricks for governed AI pipelines?

Informatica connects to the Databricks Data and AI platform through Unity Catalog integration and native data engineering connectors, allowing golden records to flow into Databricks pipelines without requiring custom ETL. The solution also exposes governed master data to AI agents via MCP servers, enabling agentic architectures to access clean, unified context at runtime.

How does Eisai use Informatica and Databricks in life sciences?

Eisai uses Informatica MDM and the Databricks Data and AI platform together to unify disparate clinical research and drug safety data into a governed, AI-ready foundation. This video presents the Eisai deployment as a concrete example of how governed master data can unlock AI and machine learning workloads across pharmaceutical research at scale.

Full transcript

[00:08] Hello, thank you for coming everyone. Um, my name is Stefan Glover. I work for the strategic alliances group uh within Informatica um for Salesforce and excited to have with me Sean Jacobs um from ASI who heads up their data and
[00:20] analytics uh technology platforms and we're excited to talk to you guys today a little bit about Informatica and a little bit about how ASI is using Informatica uh and data bricks together to help their enterprise.
[00:34] So let me start with just kind of a few questions first. So, does your organization trust its data? Can you think of a time when you
[00:46] actually did trust your data? Uh, how about trusting your data holistically across every system, every source, and every application in your company?
[00:59] Most enterprises today are drowning in data. But what they're actually starving for is context. Every system stores its own version of the truth with its own rules, its own language, and its own
[01:13] business definitions. And AI agents are trying to reason on top of all of that. Which means they're reasoning on sometimes fragments of data or maybe data that has contradictions or
[01:26] maybe snapshots of outdated data. This is what we call the context gap. So, it's not that enterprises lack data. It's that they lack a shared understanding of it. AI can't see the
[01:40] full picture because the context is scattered, stale, and inconsistent. But the good news is that this problem is solvable. But solving it requires connecting your data, connecting your
[01:54] metadata and your systems into something that AI can actually trust. So the amount of data is exploding in your enterprise and as the data practitioners you're at the very center
[02:08] of this shift moving from the role of a data custodian to agentic data architects. We're now looking at 10 sorry we're looking at 100 to 200 zetabytes of data
[02:21] being AI generated versus 10 to 20 zetabytes of human generated data. That's a 10x multiple. And now we've got AI agents that are going to amplify that
[02:33] even further. But here's the thing. AI doesn't just require more data. It requires better and more complete context. Data alone gives you numbers. Context gives those numbers meaning,
[02:47] understanding, and business value. That's the distinction that separates AI projects that stay in pilot versus those that actually get into production. And I for those that attend attended the
[03:00] keynote this morning, I love the fact that Ali actually kind of touched on this and the criticalness of context to really getting um your organization to under to the agents to understand the data that underlies it. So kind of
[03:13] really liked that uh that validation of the message. So what we did was we surveyed over 600 chief data officers and the results were pretty interesting. First off, and I don't think this will come as a surprise
[03:26] to anyone, the number one priority was AI. Shocking, I know. However, the more interesting part was hearing that governance, quality, and lack of completeness and reliability of data are
[03:38] holding them back. So, just looking at that first number as an example, if data governance has not kept pace with AI, that could mean that your agents are potentially acting on ungoverned,
[03:51] unsecured data. That alone could expose you to huge risk and potential liability. So the trust gap is real and closing it isn't optional if you want to build AI that's actually reliable. You
[04:05] don't just need trusted data, you need trusted context. And so where do you start? In our experience working with enterprise customers in every industry,
[04:17] the journey to trusted AI ready data always comes down to a handful of fundamental questions. And if you're trying to power enterprisecale AI initiatives on data bricks, you're likely going to run into all of these.
[04:31] How do I discover the data that I need to bring in to my lakehouse? How do I integrate that data from dozens of disparate enterprise sources? How do I ensure that the data is actually of high
[04:43] quality? How do I holistically govern it across clouds and on premises? How do I get a true golden record of my customer, supplier, or product? And how do I
[04:56] securely democratize sensitive data without creating compliance risk? Each of these is a real blocker that stops AI from reaching production. They are blockers to closing the context gap
[05:09] and they're all connected. Solving one in isolation is not enough. When just one of these breaks down, trust breaks down. And when trust breaks down, the whole AI use case breaks down with it.
[05:24] These are the challenges that IDMC, Informatica's intelligent data management cloud, is designed to solve. IDMC is a cloudnative microservices-based platform built for
[05:36] multicloud and hybrid data management at enterprise scale. But a huge point of relevance is the connectivity story. Over 300 unique data source connectors
[05:48] and 50,000 plus metadataware connections come right out of the box. So IDMC is really designed to support your entire data estate. And that matters for context because a real system of context
[06:02] is only as complete as the data that flows into it. If you have gaps in connectivity, you have gaps in context. And agents operating on incomplete context are unreliable.
[06:15] IDM IDMC ensures no enterprise data asset gets left behind. Every source and every system is connected, understood, and trusted.
[06:28] So the context your AI agents act on is never incomplete. And the intelligence layer behind all of this is CLA, which we'll come back to in just a moment.
[06:40] So let's loom out zoom out for a second and talk about what the full picture actually looks like architecturally. Salesforce's vision is that every AI agent, every workflow, every business
[06:52] decision should be grounded in complete and trusted understanding of your business. And that system of context is a foundational layer that makes every AI agent trustworthy enough to act
[07:05] reliably. And at the heart of that system of context is Informatica covering data integration, data catalog, data quality, governance and lineage combined with Muleoft for things like
[07:17] API and app integration and zero copy copy activation into the data bricks lighthouse with data 360. Salesforce we call ourselves customer zero internally runs on this
[07:30] architecture and the customers winning with a AI today are the ones who have chosen to invest in this kind of foundation. This architecture tells a pretty simple story. Unlock your data, govern it from
[07:43] end to end and activate it across every agent and workflow. And for those environments with data bricks, data bricks is the intelligence layer, the processing, the analytics, the scale with Informatica owning the trust layer.
[07:57] Together they give every agent, every workflow and every decision a complete and trusted understanding of your business.
[08:10] So when customers think about data governance, risk and compliance are typically top of mind. They want full visibility into where data enters the enterprise, how it's transformed, who accesses it, and how it's ultimately
[08:24] consumed in AI and analytics. Databick's Unity catalog provides a strong foundation for lakehouse governance, managing access controls, and tracking technical lineage within that
[08:36] environment. But most enterprises have far broader data landscapes spanning multiple clouds, on premises systems, and business applications. And that's where Informatica's cloud
[08:48] data governance and catalog, CDGC, extends the value. CDGC aggregates metadata across all of your sources, giving you a unified enterprisewide view beyond the
[09:00] lakehouse. It enriches Unity catalog's technical metadata with business context, glosseries, policies, and stewardship workflows and adds enterprisegrade data quality, including
[09:12] things like proactive anomaly detection. The most powerful capability is the ND end-to-end lineage. CDGC combines UD catalog's technical lineage with business lineage, giving a true 360
[09:25] degree view from ingestion to final consumption. And with AI governance for data bricks models, you can trace exactly what data trained a model and what model is being used for inferencing.
[09:38] And just recently at Informatica World, we announced that CDGC is going to support the extraction of governance tags from Unity Catalog. So you can now surface Unity catalog tags natively from within the the Informatica data catalog
[09:52] federating governance across both platforms without duplicating the effort for data stewards. So let me bring CLA back into the conversation because Clare is really
[10:04] what makes everything I've described intelligent rather than just automated. Clare is Informatica's AI engine and is really the backbone of everything that IDMC does. Clare doesn't just automate
[10:16] pipelines. It understands your data. It autoprofiles data sets. It detects data quality issues before they reach your lakehouse. It recommends transformations and automatically generates governance
[10:29] metadata. With Clare, you're doing much more than just moving data from point A to point B. You're continuously building trust into the workflow. Every time data flows through IDMC to data bricks, CLA
[10:43] is working to make sure it's of high quality, it's well-governed, and is fit for purpose, whether it's feeding your lakehouse, your agents, or your analytics.
[10:55] Cloud GPT takes all of that intelligence and makes it directly accessible through natural language. Think about the problem we all know exists in large enterprises. You have multiple copies of data and your users want to know which
[11:08] version is the source of truth. Whether the data has been certified, how recent it is, how clean it is. And today answering those questions can be difficult. Clar GPT answers all of those
[11:20] questions in one conversational interface grounded in empirical metadata and data quality scores. whether it's a primary source, a transformational layer, or the analytical tier in your
[11:32] data bricks environment. So you can have multiple data management agents all working together under a unified management and security layer all underpinned with consistent
[11:44] organizational specific semantic knowledge. So Clare GPT is not just a chatbot, it's a data intelligence assistant that knows your enterprises data estate.
[11:56] And if CLA GPT is how you find and understand your data, CLA co-pilot is how you build with it faster. Generative AI has fundamentally changed developer productivity and CLA co-pilot brings
[12:10] that to data engineering. If you need to move data from Oracle to datab bricks or build a trans transformation from bronze to silver in your medallion architecture or maybe just define a data quality rule
[12:22] in plain English, just tell Clare Co-Pilot what you need and it generates the working pipeline automatically or or mapping for you. No boilerplate code and no manual configuration is required. If
[12:35] you want to come see clear co-pilot in action, I would encourage you to come and visit our booth and we'd love to give you a deep dive. So trusted context is only as good as the data that underpins it. And nowhere
[12:49] is that more true than with master data. If your agents are reasoning across fragmented, inconsistent records from dozens of enterprise sources, they aren't reasoning from truth. That's the
[13:02] problem MDM solves. Informatica MDM consolidates fragmented data across all of your sources, customer, supplier, product, location, and produces a single highquality golden record. Lineage and
[13:17] governance are built natively into the process, so you always know where the record came from, how it was derived, and who owns it. For data bicks customers, MDM is the trust anchor at
[13:29] the data layer. Whether you're in financial services, getting a clean customer identity for risk and compliance, in life sciences, reconciling clinical and research data, or in retail, aligning product and
[13:42] supplier hierarchies, golden records are a fundamental aspect of every reliable AI use case. And we've made this native um to to data
[13:54] bricks. MDM now works with agents through MCP servers available on the data bicks marketplace. and supports data bicks genie to drive better analytics through those same MCP servers
[14:06] and the recent announcement MDM extension for datab bricks publishes those trusted golden records directly into data brick SQL so your agents are always working from the highest fidelity
[14:18] most trustworthy data in the enterprise has built a reference blueprint for datab bricks agent bricks that shows you exactly how to wire trusted governed enterprise data into your agentic
[14:31] workflows. So you're not just starting from scratch when building your agents in agent bricks, you're starting from a foundation of trust. So as enterprises start building out agents at scale, the question we hear mo
[14:45] most often is, how do I give my agents access to trusted governed enterprise data without rebuilding my entire data management stack? So, as we've talked about, IDMC gives you a complete picture
[14:58] of your data estate, whether it lives on premises, across multiple clouds, or in SAS applications. And once you have that picture, you can enrich it, cleanse it, and make it agent ready. But what we've
[15:10] done now is extend IDMC's data management capabilities into five native MCP services designed to accelerate the Agentic experience. There's an MCP to help you search and discover data assets
[15:23] across your enterprise. There's an MCP that gives you a checkoutlike experience from a data marketplace for analytical needs. You can do things like verify and enrich addresses to enhance data quality
[15:35] or access certified golden records from MDM to uniquely identify customers. These MCP services work natively within data bricks agent bricks and through genie and you can query your data state
[15:47] through natural language natively from within the datab bricks environment. So going one step deeper at the bottom you have your enterprise data sources being ingested into data bricks
[15:59] medallion architecture leveraging IDMC's set of built-in connectors. That's your bronze to silver to gold pipeline enriched and governed by Informatica. Then you've got your five Informatica
[16:11] MCP services sitting as the intelligence bridge between your governed data estate and your data bricks environment. These MCPs are what gives your agent bricks agents access to trusted data without
[16:23] having to rebuild any of those data management capabilities. And at the top it all comes together. Your agents call the MCP services directly from data bricks notebooks or agent bricks and your business users get
[16:35] the same trusted data through Genie using plain English. So the same governed highfidelity data whether you're an engineer running a pipeline a data scientists building an agent or an
[16:48] analyst asking a question and one of our la latest announcements also from Informatica world is the general availability of our MCP services
[17:00] on the datab bricks marketplace. So customers can call them directly from within their databicks environment and use them when creating agents in agent bicks or when querying with Genie. The use cases these can serve are very
[17:13] broad, but one I wanted to highlight is the MDM customer identity MCP where customers can cross reference data stored in data bricks against their customer 360 golden record in MDM to
[17:26] validate and reconcile customer identity across all sources. And we have very innovative customers like ASI in the room who are actively trying out these MCPs in their environment today. And with that, I am
[17:40] thrilled to have Shawn Jacobs from ASI here to tell their story firsthand on how ASI is using Informatica and Data Bricks together. Sean, the floor is yours.
[18:03] >> So, I want to thank Informatica for inviting me to speak today. It's a um an honor and uh I I've been looking forward to it. Um and
[18:19] So first off a little bit about ASI is that um we we specialize in you know neurological and oncological therapies and and uh drugs that we
[18:31] build. The main
[18:43] listen and to learn from them. We call this our human healthcare mission. Um, using all these tools, we help support this. This is why I go to work every day.
[19:12] what kind of tools could we use including with data bricks to make our data quality better. We did some data quality checking in data bricks. It wasn't as robust as we wanted in some cases. Um, and in other cases it worked
[19:26] fine on its own, but we we had a desire from our business to actually have significantly more data quality, but also we required a catalog. Where is your data? How do we unsilo the data
[19:39] from all the different places that we have in our company? So the basic end thing was that you know what were we going to adopt? what was the potential you know to get accurate
[19:51] reliable reporting as opposed to you know unreliable or inaccurate reporting. So we started a data modernization journey. We learned that um if the data is not trusted, validated, accessible, supported by clear lineage, it's not
[20:05] going to be of use to our internal stakeholders and then they're going to stop using it. And so we decided that data bricks combined with Informatic would be our foundational platform for the company. I don't know if everyone in here has
[20:18] probably seen if you've worked in pharma industry, data is siloed, right? This is the thing that happens constantly. Clinical teams don't necessarily want to share with the discovery teams. The discovery teams don't want to share with the commercial teams. We wanted to break
[20:31] that down and bringing all these tools in with a cat um just for the data quality was the first thing like I said, but then it was becoming the catalog, the data marketplace that we could build. How do we give access make the data findable? Right? We wanted
[20:44] to verify our data.
[20:57] accomplish all these things. Ingestion, integration with orchestration and automation. Bottom line, we needed that. We wanted to decouple architecture for storage and processing, right? What we had was other systems and solutions
[21:10] which I'm not going to mention here but you know they would stay on over the weekend incurring money. We wanted to decouple that architecture. That was data bricks was doing that for us, right? We wanted the ability to transform data while providing cleansing
[21:23] and modeling to significantly increase overall data quality. That's where Informatica started coming in, right? So, we're starting to work together with these two tools. It should natively provide governance, security, lineage, and cataloging. So, data bricks does its
[21:37] own cataloging. It does its own lineage, but we wanted more. We wanted to see it from the true source all the way to the destination. so that the end users could see the lineage, right? We also wanted to provide governance
[21:50] and cataloging. Like I've said, I'm going to keep saying cataloging over and over again and I apologize, but that was one of the most important things. We didn't know where any of our data was. We couldn't find it. If you can't find your data, you can't analyze it. And if
[22:03] it's not in, you know, a foundational database of something like data bricks, then it's not of use to anybody. So I keep telling everyone that I work with and everyone that our friends and colleagues that are even in this room,
[22:16] if it's not in your database, it's not indexed, it's not cataloged, and you don't know where it is, it's of no use to you. You can't build aic AI, you can't get to the journey of the full-long data. So ultimately, data bricks combined with Informatica covered
[22:28] all of our requirements and provided additional capabilities we not anticipated, required or planned for as well like the MCP things that we're talking about, right? We went to I was at Informatica World what three weeks
[22:40] ago I guess I saw the MCP stuff in a demo. I sat in a room like this and actually went through a whole lab with them. Immediately went back to the hotel and started hooking up the MCP servers
[22:52] into our data bricks environment. It's it's unbelievable. Um ultimately with a global roll out for our Informatica pods and our data bricks workspaces, we put them globally in our company's based in
[23:04] Japan. So we have pods in Japan. We have data bricks installations in Japan for data residency. We also
[23:16] all of our data across the company that's available in data bricks from the UK that's available in data bricks from Japan that's available in the US.
[23:29] respond to data requests from the end users in a in a highly scalable way. you know now they can go in see where the data is in the catalog request access in Informatica
[23:41] and it automatically hooks into data bricks and grants them access in data bricks so we use uh if you want to talk to me after the meeting I can give you a little bit more details on how that works but it's it's seamless and it
[23:53] takes 30 minutes for them to get full access to any data set worst case usually it could be one minute could be 30 minutes but um it's something that's amazing to me to actually not have to answer an email or a text message or
[24:07] something and log into data bricks and then go and sit and change the grants for the data to each individual person. We have it fully automated now and Informatica catalog is combined with
[24:20] the marketplace is what made that possible. So we talk a little bit about one of our main solutions. So this is the first big solution that we had to do when we uh purchased Informatica for the data
[24:33] quality. We went through um a redshift migration. We wanted to take the data out of Redshift. We wanted to get it into data bricks. We wanted to make it again verification of the data, right? F
[24:46] AIR, findable, accessible, interoperable, repeatable, right? Um and we decided that we were going to follow the recommendations of Medallion Architecture from data bricks. We had
[24:58] the bronze, the silver and the gold, which is you know the standard thing but we got a little advanced into the technology of how do we data model in there. So in our silver layer we decided that we were going to use advanced data
[25:11] vault 2.0 know um to show us a full lineage of the data, full history, right? To break down all the fact that we had 16 different data sources for claims for instance, right? and that we
[25:24] could have a history with a repeatable reusable pipeline built in Informatica that's all parameterized by the way and it inest one pipeline ingests all of our data for all of our data sources because
[25:36] we parameterized it
[25:49] of high quality and then we produced it into report ready, analytics ready data sets in the gold layer in the form of you know star schemas, data marts and data warehouses.
[26:15] because they work so well together, data bricks and Informatica, the amount of support that we have to run this right now for the entire company is about four people. And that to me is is pretty
[26:28] impressive because it was really hectic when I was doing it by myself first off, but it became much easier to do this globally. and we're we're expanding to global support and so there'll be more people working with this but in the
[26:41] complexity of having to support this the operational support is very little it I mean we need a 24/7 support so we'll have to have people you know in our global capability center in in India that will actually uh you know work to
[26:54] support this overnight when I'm sleeping but the uh the point is that this combination of of the platform powered with um you know Informatica and data bricks combined feeds our Tableau dashboards and our
[27:07] data bricks dashboards and our other various dashboard types that we have at the company all from one place. You know, we have PowerBI over there too. So, this is um really amazing. Now, the next generation of things that we're
[27:20] going to have here when you start taking a declared GPT, we're doing the same process again for another database and this time it's going to be based in Japan. We are going to fully leverage
[27:32] all of the different things that Informatica brings with Clare GPT that data bricks brings with you know the the uh the Genie code to actually build our
[27:44] schemas and our pipelines in combination together and we've already talked to one of our partners um who had provided us with a quote before we started talking about doing all of this and it has
[27:56] decreased their quote. They've came back with a quote after I showed them all the things, hooked up the MCP servers, had another analysis. I think that they the number came around 38% less in development. So they gave us a quote for
[28:08] a certain number, fairly large quote that I'm not allowed to talk about here, but um you know it was a good size quote and we dropped it by 30%. a significant cost savings by using the new tools with
[28:21] CLA GPT and Informatica combined. And it's going to be uh really fun to see them be able to type in create a pipeline from this data source. Give me
[28:33] my bronze data table. Please build it with data vault. I've already tested this by the way. It works. Um please build it with data vault 2.0 in the silver layer. give me some data quality in the middle checks that you would like
[28:46] to recommend and create a business layer in the silver as well and then produce gold with star schema. It actually created the entire pipeline for me and
[28:58] the tables with giving me all the different uh um coding for the SQL tables. It was pretty impressive end to end and it was about 3 minutes to generate the whole thing for one of our
[29:10] data sources. And if that's going to be the case and how you can develop now with you know tools like Informatica combined with data bricks with all the AI and the MCP endpoints that are being created um it's going to change the way
[29:23] we develop every one of us and I highly recommend that you look at these MCP endpoints from Informatica because it's been it's I think they're going to be adding more is what I understand and
[29:35] it's going to be able to do full functionality. You don't even have to leave data bricks now. Like we we were sending the people into marketplace to go find data. Now you can go into just the genie code and say tell me about the data that's in the catalog in
[29:48] Informatica. It'll go up and it'll bring down all the different data assets and tell you what it is. It will actually let you say well I want to check this out. When you say that you want to get access to that data, it actually sends
[30:00] whatever you routed in data in um Informatica to the data owner that this person's asking to request. this is why they want it and what they want to do and the person approves it in Informatica. The person in data bicks
[30:12] never even has to leave data bicks. They just get notified in an email that they've got access to it now and suddenly within 20 minutes they have access to that catalog or that table or that data asset that you're asking for. So I think it's uh from my perspective
[30:28] it's some of the coolest stuff I've ever seen in my career and and it's going to make my life so much easier. So what how did this really affect us all in the end of the day right? So we decreased cost by 50% increased
[30:40] performance 10fold of running the application. So the old applications in red shift that we moved over including now I just went through this exercise too with analyzing how much does it cost
[30:53] to run in informatica processing units versus data bricks units at the same time combined and our total cost of ownership is 50% less with this architecture now than it was um we migrated one of our existing platforms
[31:06] like I said from Redshift which utilize smaller capacity homegrown ETL ELT tool so we got rid of a a licensing cost that we didn't even need anymore. The system also uh often saw bottlenecks too. We used to get constantly when people were
[31:20] querying red shift simultaneously from like 10 different analysts. We get the phone call eventually. Hey, it's not scaling right. We're not doing it. It's not giving us back the response in a in a timely fashion. Um so we've gotten rid
[31:32] of that bottleneck, right? We also saw a largely increasing monthly bill for compute because we kept turning up the size of the cluster to accommodate in the old system. Obviously data bricks we all know scales scales very well um and
[31:46] that eliminated all the problem and so there's no more singlethreaded blocking waiting while uh you know other SQL queries are completing and um you know the compute there was not decoupled as
[31:59] it is so well with data bricks and then moving to our foundational architecture like I said better scaling with spark clusters unmatched performance increases I mean I could tell you that we had queries that used to run for like 6
[32:12] hours. I could get them to run in like 4 minutes. We ingested with Informatica. When we first bought it, I ingested a 90 gigabyte file in 4 minutes with Spark
[32:25] doing SQL push down and the ability to prove to this advanced architecture assisting with data organization operational support. It's absolutely dropped our operational support requirements. We now the other thing is
[32:39] we have a single pane of glass to look into all of our different data quality results. What does our data quality look like? I can go into our data quality dashboard in Informatica and see at any given time what the quality of all the
[32:52] data sets are and drill into it. So I know that the data is trusted. But more importantly, the stewards in the business side can go in and look at the same thing and see the same information and have that comfort level that we're
[33:05] making business decisions on data that I trust. And that is the ultimate goal that we had and we feel that we've accomplished that now. We also have the self-service part of it which we didn't have before. That was a big part where
[33:18] again they would you know literally call me on the phone or text me in teams or you know hit me up in Slack and say hey can you get me access to this data set well that's not really governed model right and in a farm industry you need to
[33:30] know who has access why they have access are they doing are they doing a secondary or tertiary analysis of the data what's the purpose of that analysis and do they you know have access to that data indefinitely or is it for a certain
[33:42] amount of time very often if you're using data sets in pharma from like third sources, you know, out outside um you know, things like the ADNE data set or uh you know, the uh different
[33:54] Alzheimer's drugs like Biofinder from the Swedish Bioinder studies. We have contracts that we need to notify who has access to this. This allows us to to be adhering to our contracts.
[34:08] Also, when we have those contracts expire, if we choose not to extend, we can remove access to people to be in compliance and show that we're following a compliance. any kind of accountability that we have to show whether it's
[34:22] through an audit or anything. This audit trails between data bricks and Informatica combined gives us visibility into not only who has access to what data but I mean I can see what queries they've run on the data if they've done
[34:35] anything nefarious with the queries. You know we can track everything down to the minutiae.
[34:47] accountable, interoperable. I I always forget that and and uh reproducible, right? And so that is a a a goal that we had that we've been saying for 7 years. With all of this in place now, I feel
[35:00] like we actually can say that the data that's in this platform now is is verified. Um it doesn't mean that all of our data in our company is verified. We're still,
[35:12] you know, there's always going to be the battle of unsiloing data in any company and we still have those challenges and we're working through them. Um, and again, that last thing that was the MCP integration of informatic agents into data bricks exposes the detailed
[35:26] information in catalog for data bricks, right? So, think about this. The data bricks catalog has certain amount of information. It's got some technical description areas. You can put little descriptions for each of the columns. But Informatica's catalog has a
[35:39] significantly deeper dive into the business glossery terms into um you know starting to get into ontologies um starting to get into the semantics of of the entire data sets. You can actually
[35:51] do a lot more describing and inter relationships of the data sets from a metadata standpoint than you can do in data bricks. And so with this MCP server integration now and I kept getting asked by the business they're like well we put all this stuff in Informatica but I'm
[36:04] working on data bricks all day so why can't I you know I really want to see all that information about my data the data about my data. Now with the MCP servers this is combined the Informatica
[36:16] catalog is available in the data
[36:28] descriptions. It'll tell you all the relationships. This is a game changer for our business because they were there was an argument about well if we can't really why do we enter all the information in the catalog for Informatica if we can't see it when
[36:40] we're working with the data. This solves the problem. So all these improvements to come with implementing both data bricks and Informatica allow us to move on to additional data solutions at our company with a trusted and repeated
[36:52] foundation. Right? So we're using this wash rinse repeat. Right? We're going to the next project. We're going to do the same process. We're going to use the same data modeling techniques. We're going to reuse the plat the pipelines
[37:04] that we built in Informatica because they're parameterized. So we just take that lift and shift all of that over into another project and we can reuse it. It doesn't mean that we'll only use
[37:16] that because now we don't have to worry about building so much. Um we can have AI do a lot of it for us. But this really helps us along. So with that, you know, I'm done and I would love to talk
[37:28] to any of you at the end if you have more questions. Um again, thank you to the folks from Informatica. Thank you to Data Bricks for having us here as well. And um I much appreciate everyone's attendance today and have a wonderful
[37:40] rest of the trip.

Learn more about the Databricks Data and AI platform.

The information provided herein is for general informational purposes only and may not reflect the most current product capabilities or configurations.