Databricks and Palantir Integration: Enterprise AI Operating Models
Summary
- The Databricks-Palantir partnership has grown from approximately 30 joint customers to over 150 in production—including more than 20 Fortune Global 500 companies—by streamlining interoperability between the open data platform and Palantir Foundry and AIP.
- Rich Products demonstrates a six-principle operating model that uses Databricks as the governed data foundation and federates compute and complex workflows through Palantir Foundry, preventing data duplication and platform sprawl.
- The integration is built on four technical pillars—data federation, compute push-down, secure governance through Unity Catalog, and closed-loop publishing of validated insights back to the lakehouse—enabling faster project delivery and consolidated data architecture.
Databricks and Palantir Integration: Enterprise AI Operating Models

Building enterprise AI requires more than connecting platforms. It demands a governed operating model that scales. In this video from Data and AI Summit, we explore how Databricks serves as the foundation for enterprise data and AI, and how federation with Palantir Foundry enables complex workflows while maintaining governance through Unity Catalog.
Learn how organizations like Rich Products are converging on a Databricks-first approach, leveraging data federation, compute push-down, and secure governance integration. Discover the six principles of a scalable operating model, how to prevent data duplication and platform sprawl, and how to close the loop by publishing validated insights back to the lakehouse.
📂 Technical documentation: https://palantir.com/partnerships/databricks
🤝
Chapters
00:00Opening01:12Databricks-Palantir Partnership Overview02:37Partnership Growth: 150+ Customers in Production03:40Technical Architecture: Four Pillars of Integration07:14Rich Products Operating Model09:04Safety, Scale, and Speed Framework11:07Six Principles of the Playbook12:43Cost Allocation Use Case and Closed Loop16:10Benefits: What the Model Prevents and Enables18:04Reusable Takeaways for Any Organization21:35Key Message: Governed Operating Models
FAQs
What does the Databricks and Palantir Foundry integration enable?
The integration connects the Databricks open data platform with Palantir Foundry and AIP, allowing organizations that invest in one platform to easily leverage those investments in the other. This video explains that it enables faster project delivery, consolidated data architecture, and secure data sharing between the two platforms without duplicating data.
What is the six-principle operating model demonstrated by Rich Products?
This video describes a six-principle playbook for governing a Databricks-Palantir deployment, aimed at preventing data duplication, platform sprawl, and governance gaps when the two platforms operate together. Rich Products uses these principles to define how data flows from the Databricks lakehouse through Palantir workflows and back, closing the loop with validated insights published to the lakehouse.
How does Unity Catalog provide governance across the Databricks-Palantir integration?
Unity Catalog acts as the governance layer managing data access controls, lineage, and security policies even when compute and workflows are pushed to Palantir Foundry. This video explains that maintaining Unity Catalog as the authoritative governance source prevents the data access fragmentation that commonly occurs when two enterprise platforms operate side-by-side.
How many customers are using the Databricks-Palantir integration in production?
As of the Data and AI Summit presentation featured in this video, over 150 customers are using the Databricks-Palantir integration in production, up from approximately 30 customers a year prior. This includes over 20 Fortune Global 500 companies powering more than 100 use cases across the joint partnership.
Full transcript
[00:09] It's Thursday at Data + AI Summit. We've made it to Thursday. First and foremost, I want to thank everybody for being here with us this week and today. It is not lost on us how hard it is to take a week out of your busy schedules to come to Data + AI Summit, but all of you, the attendees, are who really make
[00:25] Summit special, so thank you. Before we can jump into any of the exciting stuff, we have an obligatory statement from our legal team. Any forward-looking statement that Brandon or I make today while on stage,
[00:41] you know, while we will try to hold to it, is not uh necessarily a promise. Anything unforeseen, of course, can change these statements. And finally, please do complete your surveys after today. For every session you attend, you will
[00:56] have a survey in your Databricks event app. Your surveys and your feedback are really how it is prioritized what content will be at next year's Summit, so please help us make next year's Summit great. Fill out your survey.
[01:12] Awesome. My name is Ben Aboud. I'm the global tech lead for our Palantir partnership. I'm joined on stage by Brandon Lyle. Brandon. Hello, my name's Brandon Lyle. I lead our data and AI platforms team at Rich Products. Today, I'm going to give an introduction
[01:27] and an overview of the Palantir partnership. And then I'm going to kick it over to Brandon, who's going to walk through a customer playbook for how Databricks and Palantir is used at Rich Products today.
[01:43] So, before we get into the what of the partnership, I want to talk a little bit about the why of the partnership. It should be no surprise that Palantir and Databricks have many joint customers. What we often saw was our platforms, the open data platform and Palantir Foundry
[01:58] and AIP, deployed side by side to really power our customers' businesses. We set out to streamline interoperability between our two platforms, so that as your teams make investments in one platform, you can easily leverage those investments in the other.
[02:16] What this means for our customers is they're realizing faster project delivery. They're able to consolidate their data architecture on the open data platform, and they're able to securely share data between Databricks and Palantir Foundry and AIP.
[02:37] A year ago at Data and AI Summit, I introduced our partnership. At that point in time, we had about 30 joint customers using the partnership integration. Today, that number is over 150. This includes over 20 Fortune Global 500
[02:52] companies. These customers are powering over 100 use cases in production. And they're doing so while minimizing data duplication, ensuring secure governance between our platforms,
[03:08] and without compromising on the strengths of either platform. The result of this partnership has been incredible. Our customers are very frequently realizing at least 20% overall efficiencies gained when they're
[03:23] deploying with our partnership architecture as opposed to a a traditional deployment model. So, I gave a little introduction to the partnership and how it came to be, and an update on how the partnership has been going over the past 12 months.
[03:40] Now, I want to talk about some of the technical details about what our partnership really is. Our partnership is a technology integration between Databricks and Palantir Foundry. This integration is centered around four pillars of product integration.
[03:57] The pillars are data federation, governance integration, compute federation or compute push down, and AI and workflow integration. Now, you can seamlessly federate data from Unity Catalog into Palantir Foundry
[04:14] and AIP to power the ontology and operational applications. Before our partnership, a very common deployment model that we saw was our customers would build their medallion architecture in Databricks. Then they would ingest some or all of
[04:29] this data into Foundry. Once that data got to Foundry, maybe additional transformation occurred, and then those data assets would need to be written back to Databricks. It was really challenging to keep these data assets in sync, and you were double paying for ingest.
[04:45] With our partnership today, Foundry can seamlessly federate data from Unity Catalog. There's no data movement, you can power the ontology directly from Unity Catalog. The second pillar of our partnership product integration is governance
[05:02] integration. Today, the governance integration covers how the platforms authenticate to each other. Primarily, we use workload identity federation, meaning that workloads on Foundry can run and hit Databricks APIs
[05:17] without the need for Databricks secrets. Something that we've started to scope is user identity federation. This would mean that user credentials could be passed from Foundry to Databricks so that Unity Catalog can truly be your single point of governance for integration.
[05:34] Additionally, Foundry can accept vendored credentials from Unity Catalog so that Foundry can directly read and write to Unity Catalog. Our third pillar, maybe our most exciting pillar, is compute push down or
[05:49] compute federation. With our partnership, you can now leverage the best-in-world Databricks compute engine, regardless of which platform you're using. If you're in Palantir Foundry, whether you're using pipeline builder or code repositories, you can now directly
[06:06] hit Databricks clusters to read and write to Unity Catalog. This has been incredible. This means that Foundry can federate data from Unity Catalog, so data can stay in place at read. And then at write time, you can actually invoke a Databricks cluster to perform the write directly to Unity
[06:23] Catalog. So, from within Foundry, you can bring your own storage and bring your own compute in the form of Databricks. The fourth and final pillar of our partnership is AI and workflow integration. As your data science teams make
[06:39] investments in building, training, and deploying machine learning models in Databricks, you can now register those models directly in Palantir Foundry for use in your ontology and in your operational applications.
[06:55] So, we've given a brief overview of the partnership, an update on where things are today, and a little technical deep dive on how the integration actually works. Now, I'm going to hand it over to Brandon Lyle's to talk through how our partnership is powering use cases at Rich Products. Thanks, Ben. And thank you all for joining us today.
[07:14] I want to transition from the partnership story into the customer operating model. What did it look like for us to actually make this real within enterprise? The lens for today is not just about how the platforms work, but how we determine where they fit, how
[07:30] we govern them, and how do we create a model that was repeatable and useful for the business. A little bit about us. Rich Products is a global food manufacturing organization with a broad operating footprint. We're 100% family-owned, we have over
[07:46] 13,000 global associates, over 4,000 product codes, and we operate in over 110 countries, based and headquartered out of Buffalo, New York. We have a very wide range of product categories, as you can see here, which
[08:01] means that we deal with real complexity, not just in data, but in decisions, processes, and how work flows through the organization. So, at that scale, platform decisions cannot be case-by-case based on preference.
[08:16] They have to be governed, they have to be repeatable, and they have to be understandable enough for many teams to be able to operate within the same model. So, what is this talk about? This is not a feature tour,
[08:32] it's not a tool comparison, it's not a debate about what one platform can do versus another. This is about us sharing with you how we've decided to essentially converge on a Databricks-first governed operating model,
[08:47] and then only extend beyond that for a subset, narrower class of complex operational decision workflows where the business process required it.
[09:04] So, why does the model exist? We weren't solving for a tool or technology problem. We're actually solving for safety, scale, and speed and enterprise value realization.
[09:20] Safety. It's about finding a centralized governed platform, supported and governed integrated pathways, and full auditability. The business had to trust what came out on the other side. Scale. It's about reusable patterns,
[09:36] not solving for one use but for hundreds of use cases. Speed. Eliminating months long of architectural debates. Being able to have the team focus on creating and delivering value
[09:52] versus re-litigating the model every use case. Databricks is the foundation for our data, AI, and analytics at enterprise scale. And I'm going to talk about how we fought to preserve that while enabling composition with Foundry.
[10:13] So, we think about Databricks, it is the substrate for enterprise data, AI development, applications for internal processes and operational workflows, AI, BI dashboards, you name it. But, it also can serve as the operational layer.
[10:28] Agents, apps, orchestration, workflow execution. So, the question that we asked was, "What is the Where should these capabilities sit from a tool standpoint?" The question rather was, "What native capabilities can we
[10:45] leverage Databricks for? And is there a real justified governance reason to go beyond that?" Every extension beyond native required governed justification.
[11:07] So, once we solidified and aligned on that foundation, we turned it into a playbook. And a playbook is intentionally simple. Because if an operating model requires a manual to follow, no one's going to follow it. So, the big ideas are six. The first one
[11:23] being Databricks first. The idea around Databricks first is go native and prove the exception. Secondly, patterns over projects. I think we all can appreciate that when we think about platforms, we're more interested in patterns for versus
[11:39] one-off project-based workloads. We want to be able to solve for 100 plus use cases, not one or two. Third, trusted publication. We wanted to be able to bring data back into our enterprise platform that was
[11:56] certified and trusted, so that it become a swamp and a dumping ground for intermediary data sets from Foundry. The next three ideas, federate, don't copy. Ben talked a little bit about this at the onset. We wanted to be able to virtualize and
[12:12] federate our data into Foundry versus driving duplication of data. We wanted to enable the lightest execution path, especially as we think about compute. Enable compute first pushdown as a first path.
[12:28] Only then, if that doesn't work, and there's cases where that doesn't work, we enable local compute in Foundry. Last but not least, observe everything. If it runs, we monitor it.
[12:43] Cost, compute, lineage, outcomes. We can't govern what we can't see. So, with this playbook, we needed a use case that will pressure test it, to make it real. So, we converged and aligned on cost allocation, cost of service allocation.
[13:01] And the reason why we picked this use case is because it was high complexity. It required human judgment, collaboration across multiple teams. It was stateful. We have draft, validate, published outputs coming out of Foundry. And we had a hard deadline. We had to
[13:17] migrate off of a legacy system. So, we didn't have all the time in the world to iterate. We had to make it real. So, as you can see here, the design decision that emerged from
[13:33] this use case was this closed loop. Data comes from our ear piece system, comes into Databricks through ingestion. We then federate that data to Foundry where we build our allocation model. The output is then reviewed, validated, and
[13:50] approved. And then we publish that back into Databricks through a write-back catalog, which then is leveraged at consumption for downstream systems. And at month end, P&O data comes back into Databricks, becoming the certified trusted source of
[14:06] financial data. In this loop, Databricks appears three times. Ingestion, source data at the top, govern output coming back in the middle, and value, certified data
[14:24] at the bottom. The closed loop, the federation, the composition with Foundry made the lakehouse more valuable, not weaker.
[14:40] If we take a step back to architecture that we sort of put together was around, again, reusable patterns. So, not supporting an allocation use case only or the next use case, but the next 100 use cases beyond. So, you see here, we have many different source systems and and data sources that
[14:57] come into our enterprise data platform. We call it our EDP, which essentially is is Databricks lakehouse on Azure. We then, again, federate the data through virtual tables, leveraging VIN credentials,
[15:13] and JDBC as a fallback, to now enable teams to build data sets within Foundry. And only once those data sets been certified and validated, do we then write them back into a dedicated catalog within UC.
[15:30] Not into any other curated or source catalogs, but a dedicated catalog. From there, many different consumer archetypes now have the ability to take advantage of that certified published data. Of AI, BI, Genie, I know that's not the
[15:46] name anymore. There's a lot of different changes this week, so I may be a bit behind. Databricks apps, even non-Databricks consumer interfaces, customer experiences to drive open reusability and consumption.
[16:10] So, when we think about this model and we think about this approach, our value was realized actually in the things that we prevented. As you see here, we prevented casual data duplication. No No casual copying of data, right? We federated and
[16:25] virtualized the data to Foundry. We prevented over-engineering. The playbook made it super clear and simple. We gave engineers, developers a decision framework on how to do work, how to operate work, ways of working, so that
[16:40] they didn't have to again relitigate the approach and the model every use case. We had a repeated repeatable architecture framework. No month-long architectural debates. This was written before use case number
[16:56] two. And we prevented platform sprawl. There's intentionality in how we compose Databricks and Foundry. We didn't just add another tool to the stack, but we intentionally thought about the closed loop and ensuring that
[17:12] the lakehouse increased in value every time the loop closed. What do we enable? Again, we enable a centralized governance construct. Where composable workflows, whether they were inside of Foundry,
[17:29] outside of Foundry, inside of Databricks, outside of Databricks, still follow through a govern process. Unity Catalog became that governance spine and played the role that we wanted it to play. Approved outputs came back into
[17:46] Databricks so that we can enable broader enterprise consumption of data products. And last but not least, we enable speed to value. I mean, that's what the business wants. This is all great, but the business wants to go fast and drive value creation.
[18:04] This investment, this discipline up front enable that. So, you don't need necessarily our stack for beta. But, what you do need is a decision framework. As there's Here's a few things that
[18:19] anyone in this room can walk away with as reusable concepts and frames and mental models as you go into your organizations to think about your journey. First, again, is Databricks native. Native until you can prove
[18:37] beyond the native pattern that you need to compose outside of it. Federate in, publish out. Again, we're going to federate and virtualize data into Foundry, and we're going to publish back out certified, validated data sets.
[18:55] This is important. Define the framework before use case number two. A lot of the times, what we see organizations do is they go off and do about 10 different use cases and then figure out, "Okay, how do we now get this under control?" Invest the time to think through the framework, your governance posture, before use case
[19:12] number two. Allow use case number one, just like in our case, to put pressure on your framework. So, then again, you can move for speed, scale, and safety.
[19:27] Make governance executable. Shouldn't just be wikis, slides, documentation. It should be embedded in the operating model. It should be policies in your code. It should be just a way of working. And last but not least, close the loop. This is my favorite, because this is
[19:44] where the lakehouse again gets more powerful and more valuable. And as we sit through a lot of sessions this week, and all these great features that are being rolled out into the product, we can quickly adopt them because we're thinking about closing the loop and making the lakehouse more
[19:59] powerful by publishing certified, validated data back into Databricks. And this is a huge mental model shift for us. We didn't want value to be created outside in a different system or tool and not ever come back inside of Databricks,
[20:17] which is our central platform. So, four things you can take with you leaving this session today. Number one, Databricks is the governed foundation. Make that your central
[20:33] brain, your central engine of data, AI, analytics in your organization. Number two, when we compose, we compose on top of that foundation, not beside it. Compose on top, regardless what sits on top,
[20:48] the foundation begins to continue to be hydrated with value and benefits. Number three, governance applies uniformly. Again, regardless if it's Foundry or Databricks, Vin again shared a lot of things that they're bringing to
[21:03] bear with governance, so we can have Unity Catalog remain that governance spine and apply governance uniformly across the entire ecosystem of platforms. And last but not least, the lakehouse gets stronger every time the business
[21:19] uses the model. Every time we close the loop, every time we introduce a new new use case, the lakehouse gets stronger. Let's make it simpler. One sentence.
[21:35] If there's one sentence that I want to leave you all with as you leave today's session, the goal is not to connect platforms merely, but the goal is to build a governed operating model that gets stronger every time the business uses it.
[21:52] That's it. That is the one thing if you can ground yourself on that, you're set up for success. I'm going to kick it back over to Ben, where Ben's going to share a little bit about how the partnership is scaling with partners
[22:07] and how you all can continue to stay engaged. Thank you. Thank you, Brandon. As you guys can see, Brandon and the
[22:22] rest of Rich Products have been incredibly thoughtful about how they want to deploy Palantir and Databricks next to each other. You know, I often think about things in terms of people, product, and processes. And no product A product is only as good
[22:37] as the people and the processes around it. At Rich Products, you have incredible people, and they have defined a fantastic process for how to integrate these products together. So, Brandon, thank you for all the thoughtfulness you've put into this. Now, I would be remiss if I didn't call out some of our thought leaders and
[22:54] system integrator partners for helping us further the partnership by helping our customers integrate the two platforms together. This is just a small sample of some of my SI partners who are currently trained on the partnership.
[23:10] So, what's next? If you want to learn more about the partnership, there are a bunch of uh you know, public-facing documentation and blog posts you can check out. Rory Patterson, Palantir's Chief of Staff, recently put out a blog post in October where he covers how over 100 customers are using the partnership to
[23:26] power their business. If you want to do a little bit of a technical deeper dive, there's a great YouTube video. It's Chad Walquist, an architect from Palantir, and myself. We cover an introduction to the partnership, and we give a demo of what the integration actually looks like in
[23:41] practice. You can find it on YouTube. And then finally, one of our partners, Blueprint, recently put out a blog post to describe what they're seeing in the field for how this partnership uh is really helping to power their customers' businesses.
[23:56] If you want to get started today, there's a bunch of technical documentation online. I've got links to everything. On Palantir's website, you'll find the Databricks connector and all of the documentation under palantir.com/partnerships/databricks.
[24:14] Great first place to start. And finally, if you'd like to connect with me or my team, you can reach out to Taylor Melville, and he can help facilitate a connection. I'll give it over to Brandon to tell you to how to connect with rich products. Yes, if you would like to reach out to me, you can connect with me on LinkedIn. I will be uploading a supplementary post
[24:31] on today's session to probably go a little bit deeper. I know a lot of us are probably like, "I want more than that." So, I will leverage LinkedIn to share a little bit more about how to make this work. Awesome. Again, thank you everybody for spending the week with us at Data and AI Summit.
[24:47] It has been a pleasure talking to you guys about the Palantir and Databricks partnership. I think we've got about 15 minutes left. Brandon and I are going to be hanging out up here if anybody uh you know, wants to get to know us or do Q&A, we're going to be hanging out for the next 15 or so. Thank you.
[25:03] Thank you.
Learn more about the Databricks Data and AI platform.
The information provided herein is for general informational purposes only and may not reflect the most current product capabilities or configurations.