Databricks and Palantir: Governed Operating Models for Enterprise Data
Summary
- Rich Products, a global food manufacturer with 13,000 employees and 4,000 product codes, unified Databricks and Palantir Foundry into a governed operating model built on four technical pillars: data federation from Unity Catalog into Foundry, workload identity governance, compute pushdown to Databricks, and AI and workflow integration.
- The Databricks and Palantir partnership has grown from approximately 30 joint customers a year ago to over 150 today, including more than 20 Fortune Global 500 companies collectively powering over 100 use cases.
- Rich Products replaced a legacy cost allocation system using a closed-loop architecture where governed data flows into Databricks, federates into Foundry for allocation logic and review, and validated outputs are published back to Databricks as the trusted source.
Databricks and Palantir: Governed Operating Models for Enterprise Data

Rich Products, a global food manufacturer with 13,000 employees and 4,000 product codes, shares how they unified Databricks and Palantir Foundry into a governed operating model. The partnership centers on four technical pillars: data federation from Unity Catalog into Foundry, workload identity governance, compute pushdown to Databricks, and AI and workflow integration.
Learn how Rich Products replaced a legacy cost allocation system using a closed-loop architecture: governed data ingestion into Databricks, federation into Foundry for allocation logic and review, and publication of validated outputs back to Databricks as the trusted source. Discover the six-point framework that prevents data duplication, over-engineering, and platform sprawl while enabling repeatable patterns across 100+ use cases.
🤝
Chapters
00:00Opening Remarks and Speaker Introductions01:43The Databricks and Palantir Partnership03:40Technical Integration: Four Core Pillars07:14From Partnership to Practice: Rich Products Operating Model08:47Databricks-First Playbook: Foundation and Philosophy11:07The Six-Point Framework and Reusable Patterns13:33Cost Allocation Use Case: Closed-Loop Architecture16:10Value Realized: Preventing Common Mistakes18:21Reusable Concepts and Takeaways20:33Four Key Takeaways for Your Organization23:10Next Steps: Resources and Getting Started
FAQs
What are the four technical pillars of the Databricks and Palantir Foundry integration?
The four pillars are: data federation from Unity Catalog into Foundry, workload identity governance, compute pushdown to Databricks, and AI and workflow integration. These pillars enable customers to leverage investments in one platform while benefiting from the capabilities of the other.
How did Rich Products use Databricks and Palantir Foundry to replace their cost allocation system?
Rich Products built a closed-loop architecture where data is governed and ingested into Databricks, federated into Palantir Foundry for allocation logic and review, and then validated outputs are published back to Databricks as the trusted source. This replaced a legacy system using a repeatable pattern designed to prevent data duplication and platform sprawl.
How many joint customers does the Databricks and Palantir partnership have?
The partnership grew from approximately 30 joint customers a year ago to over 150 today, including more than 20 Fortune Global 500 companies. These customers are collectively powering over 100 use cases using the integrated platforms.
What problems does the Databricks and Palantir integration solve for enterprises?
The integration streamlines interoperability between the two platforms so teams can leverage investments in one without duplicating work in the other. The six-point framework prevents data duplication, over-engineering, and platform sprawl while enabling repeatable patterns that Rich Products applies across its enterprise use cases.
Full transcript
[00:09] It's Thursday at Data and AI Summit. We've made it to Thursday. First and foremost, I want to thank everybody for being here with us this week and today. It is not lost on us how hard it is to take a week out of your busy schedules to come to Data and AI Summit, but all of you, the attendees, are who really
[00:25] make Summit special, so thank you. Before we can jump into any of the exciting stuff, we have an obligatory statement from our legal team. Any forward-looking statement that Brandon or I make today while on stage,
[00:41] you know, while we will try to hold to it, is not uh necessarily a promise. Anything unforeseen, of course, can change these statements. And finally, please do complete your surveys after today. For every session you attend, you will
[00:56] have a survey in your Databricks event app. Your surveys and your feedback are really how it is so what content will be at next year's Summit. So, please help us make next year's Summit great. Fill out your survey.
[01:12] Awesome. My name is Ben Abood. I'm the global tech lead for our Palantir partnership. I'm joined on stage by Brandon Lyle. Brandon. Hello, my name is Brandon Lyle. I lead our data and AI platforms teams at Rich Products. Today, I'm going to give an introduction
[01:27] and an overview of the Palantir partnership. And then I'm going to kick it over to Brandon, who's going to walk through a customer playbook for how Databricks and Palantir is used at Rich Products today.
[01:43] So, before we get into the what of the partnership, I want to talk a little bit about the why of the partnership. It should be no surprise that Palantir and Databricks have many joint customers. What we often saw was our platforms, the open data platform and Palantir Foundry
[01:58] and AIP, deployed side-by-side to really power our customers' businesses. We set out to streamline interoperability between our two platforms, so that as your teams make investments in one platform, you can easily leverage those investments in the other.
[02:16] What this means for our customers is they're realizing faster project delivery, they're able to consolidate their data architecture on the open data platform, and they're able to securely share data between Databricks and Palantir Foundry and AIP.
[02:37] A year ago at Data and AI Summit, I introduced our partnership. At that point in time, we had about 30 joint customers using the partnership integration. Today, that number is over 150. This includes over 20 Fortune Global 500
[02:52] companies. These customers are powering over 100 use cases in production. And they're doing so while minimizing data duplication, ensuring secure governance between our platforms,
[03:08] and without compromising on the strengths of either platform. The result of this partnership has been incredible. Our customers are very frequently realizing at least 20% overall efficiencies gained when they're
[03:23] deploying with our partnership architecture as opposed to a a traditional deployment model. So, I gave a little introduction to the partnership and how it came to be, and an update on how the partnership has been going over the past 12 months.
[03:40] Now, I want to talk about some of the technical details about what our partnership really is. Our partnership is a technology integration between Databricks and Palantir Foundry. This integration is centered around four pillars of product integration.
[03:57] The pillars are data federation, governance integration, compute federation or compute push down, and AI and workflow integration. Now, you can seamlessly federate data from Unity Catalog into Palantir Foundry
[04:14] and AIP to power the ontology and operational applications. Before our partnership, a very common deployment model that we saw was our customers would build their medallion architecture in Databricks. Then they would ingest some or all of
[04:29] this data into Foundry. Once that data got to Foundry, maybe additional transformation occurred. And then those data assets would need to be written back to Databricks. It was really challenging to keep these data assets in sync, and you were double paying for ingest.
[04:45] With our partnership today, Foundry can seamlessly federate data from Unity Catalog. There's no data movement. You can power the ontology directly from Unity Catalog. The second pillar of our partnership product integration is governance
[05:02] integration. Today, the governance integration covers how the platforms authenticate to each other. Primarily, we use workload identity federation, meaning that workloads on Foundry can run and hit Databricks APIs
[05:17] without the need for Databricks secrets. Something that we've started to scope is user identity federation. This would mean that user credentials could be passed from Foundry to Databricks so that Unity Catalog can truly be your single point of governance for integration.
[05:34] Additionally, Foundry can accept vendored credentials from Unity Catalog so that Foundry can directly read and write to Unity Catalog. Our third pillar, maybe our most exciting pillar, is compute pushdown or
[05:49] compute federation. With our partnership, you can now leverage the best in world Databricks compute engine, regardless of which platform you're using. If you're in Palantir Foundry, whether you're using pipeline builder or code repositories, you can now directly
[06:06] hit Databricks clusters to read and write to Unity Catalog. This has been incredible. This means that Foundry can federate data from Unity Catalog, so data can stay in place at read. And then at write time, you can actually invoke a Databricks cluster to perform the write directly to Unity
[06:23] Catalog. So from within Foundry, you can bring your own storage and bring your own compute in the form of Databricks. The fourth and final pillar of our partnership is AI and workflow integration. As your data science teams make
[06:39] investments in building, training, and deploying machine learning models in Databricks, you can now register those models directly in Palantir Foundry for use in your ontology and in your operational applications.
[06:55] So we've given a brief overview of the partnership, an update on where things are today, and a little technical deep dive on how that integration actually works. Now, I'm going to hand it over to Brandon Lyle's to talk through how our partnership is powering use cases at Rich Products. Thanks, Ben. And thank you all for joining us today.
[07:14] I want to transition from the partnership story into the customer operating model. What did it look like for us to actually make this real within an enterprise? The lens for today is not just about how the platforms work, but how we determine where they fit, how
[07:30] we govern them, and how do we create a model that was repeatable and useful for the business. A little bit about us. Rich Products is a global food manufacturing organization with a broad operating footprint. We're 100% family-owned, we have over
[07:46] 13,000 global associates, over 4,000 product codes, and we operate in over 110 countries, based and headquartered out of Buffalo, New York. We have a very wide range of product categories, as you can see here, which
[08:01] means that we deal with real complexity, not just in data, but in decisions, processes, and how work flows through the organization. So, at that scale, platform decisions cannot be case-by-case based on preference.
[08:16] They have to be governed, they have to be repeatable, and they have to be understandable enough for many teams to be able to operate within the same model. So, what is this talk about? This is not a feature tour,
[08:32] it's not a tool comparison, it's not a debate about what one platform can do versus another. This is about us sharing with you how we've decided to essentially converge on a Databricks-first governed operating model,
[08:47] and then only extend beyond that for a subset, narrow class of complex operational decision workflows where the business process required it.
[09:04] So, why does the model exist? We weren't solving for a tool or technology problem. We're actually solving for safety, scale, and speed and enterprise value realization.
[09:20] Safety. It's about finding a centralized governed platform, supported and governed integrated pathways, and full auditability. The business had to trust what came out on the other side. Scale. It's about reusable patterns,
[09:36] not solving for one use but for hundreds of use cases. Speed. Eliminating months long of architectural debates. Being able to have the team focus on creating and delivering value
[09:52] versus re-litigating the model every use case. Databricks is the foundation for our data, AI, and analytics at enterprise scale. And I'm going to talk about how we fought to preserve that while enabling composition with Foundry.
[10:13] So, when you think about Databricks, it is the substrate for enterprise data, AI development, applications for internal processes and operational workflows, AI, BI dashboards, you name it. But, it also can serve as the operational layer,
[10:28] agents, apps, orchestration, workflow execution. So, the question that we asked wasn't where should these capabilities sit from a tool standpoint? The question rather was what native capabilities can we
[10:45] leverage Databricks for? And is there a real justified governance reason to go beyond that? Every extension beyond native required governed justification.
[11:07] So, once we solidified and aligned on that foundation, we turned it into a playbook. And a playbook is intentionally simple. Because if an operating model requires a manual to follow, no one's going to follow it. So, the big ideas are six. The first one
[11:23] being Databricks first. The idea around Databricks first is go native and prove the exception. Secondly, patterns over projects. I think we all can appreciate that when we think about platforms, we're more interested in patterns for versus
[11:39] one-off project-based workloads. We want to be able to solve for 100 plus use cases, not one or two. Third, trusted publication. We wanted to be able to bring data back into our enterprise platform that was
[11:56] certified and trusted, so that it become a swamp and a dumping ground for intermediary data sets from Foundry. The next three ideas, federate, don't copy. Ben talked a little bit about this at the onset. We wanted to be able to virtualize and
[12:12] federate our data into Foundry versus driving duplication of data. We wanted to enable the lightest execution path, especially as we think about compute. Enable compute first push down as a first path.
[12:28] Only then, if that doesn't work, and there's cases where that doesn't work, we enable local compute in Foundry. Last but not least, observe everything. If it runs, we monitor it.
[12:43] Cost, compute, lineage, outcomes. We can't govern what we can't see. So, with this playbook, we needed a use case that will pressure test it, that will make it real. So, we converged and aligned on cost allocation, cost of service allocations.
[13:01] And the reason why we picked this use case is because it was high complexity. It required human judgment, collaboration across multiple teams. It was stateful. We have draft, validate, published outputs coming out of Foundry. And we had a hard deadline. We had to
[13:17] migrate off of a legacy system. So, we didn't have all the time in the world to iterate. We had to make it real. So, as you can see here, the design decision that emerged from
[13:33] this use case was this closed loop. Data comes from our ear piece system, comes into Databricks through ingestion. We then federate that data to Foundry where we build our allocation model. The output is then reviewed, validated, and
[13:50] approved. And then we publish that back into Databricks through a write-back catalog, which then is leveraged at consumption for downstream systems. And at month end, P&L data comes back into Databricks, becoming the certified trusted source of
[14:06] financial data. In this loop, Databricks appears three times. Ingestion, source data at the top, governed output coming back in the middle, and value, certified data
[14:24] at the bottom. The closed loop, the federation, the composition with Foundry made the lakehouse more valuable, not weaker.
[14:40] If we take a step back to architecture that we sort of put together was around, again, reusable patterns. So, not supporting an allocation use case only or the next use case, but the next 100 use cases beyond. So, you see here, we have many different source systems and and data sources that
[14:57] come into our enterprise data platform. We call it our EDP, which essentially is is Databricks lakehouse on Azure. We then, again, federate the data through virtual tables, leveraging VIN credentials,
[15:13] and JDBC as a fallback, to now enable teams to build data sets within Foundry. And only once those data sets have been certified and validated, do we then write them back into a dedicated catalog within UC.
[15:30] Not into any other uh curated or source catalogs, but a dedicated catalog. From there, many different consumer archetypes now have the ability to take advantage of that certified published data. Of AI, BI, Genie, I know that's not the
[15:46] name anymore. There's a lot of different changes this week, so I may be a bit behind. Databricks apps, even non-Databricks consumer interfaces, customer experiences to drive open reusability and consumption.
[16:10] So, when we think about this model and we think about this approach, our value was realized actually in the things that we prevented. As you see here, we prevented casual data duplication. No No casual copying of data, right? We federated and virtualized the data to
[16:26] Foundry. We prevented over-engineering. The playbook made it super clear and simple. We gave engineers, developers a decision framework on how to do work, how to operate work, ways of working, so that they didn't have to again re-litigate
[16:43] the approach and the model every use case. We had a repeated repeatable architecture framework. No month-long architectural debates. This was written before use case number two.
[17:01] And we prevented platform sprawl. There was intentionality in how we compose Databricks and Foundry. We didn't just add another tool to the stack, but we intentionally thought about the closed loop and ensuring that the lakehouse increased in value every time the loop closed.
[17:18] What do we enable? Again, we enable a centralized governance construct. Where composable workflows, whether they were inside of Foundry, outside of Foundry, inside of Databricks, outside of Databricks,
[17:33] still follow through a governed process. Unity Catalog became that governance spine and played the role that we wanted it to play. Approved outputs came back into Databricks so that we can enable broader enterprise
[17:49] consumption of data products. And last but not least, we enable speed to value. I mean, that's what the business wants. This is all great, but the business wants to go fast and drive value creation. This investment, this discipline up
[18:06] front enable that. So, you don't need necessarily our stack for beta. But, what you do need is a decision framework. As there's Here's a few things that anyone in this room
[18:21] can walk away with as reusable concepts and frames and mental models as you go into your organizations to think about your journey. First, again, it's Databricks native. Native until you can prove
[18:37] beyond the native pattern that you need to compose outside of it. Federate in, publish out. Again, we're going to federate and virtualize data into Foundry, and we're going to publish back out certified, validated data sets.
[18:55] This is important. Define the framework before use case number two. A lot of the times, what we see organizations do is they go off and do about 10 different use cases and then figure out, "Okay, how do we now get this under control?" Invest the time to think through the framework, your governance posture, before use case
[19:12] number two. Allow use case number one just like in our case to put pressure on your framework. So then again, you can move for speed, scale, and safety.
[19:27] Make governance executable. Shouldn't just be wikis, slides, documentation. It should be embedded in the operating model. It should be policies in your code. It should be just a way of working. And last but not least, close the loop. This is my favorite because this is
[19:44] where the lakehouse again gets more powerful and more valuable. And as we sit through a lot of sessions this week and all these great features that are being rolled out into the product, we can quickly adopt them because we're thinking about closing the loop and making the lakehouse more
[19:59] powerful by publishing certified validated data back into Databricks. And this is a huge mental model shift for us. We didn't want value to be created outside in a different system or tool and not ever come back inside of Databricks,
[20:17] which is our central platform. So four things you can take with you leaving this session today. Number one, Databricks is the governed foundation. Make that your central
[20:33] brain, your central engine of data, AI, analytics in your organization. Number two, when we compose, we compose on top of that foundation, not beside it. Compose on top, regardless of what sits on top,
[20:48] the foundation begins to continue to be hydrated with value and benefits. Number three, governance applies uniformly. Again, regardless if it's Foundry or Databricks, Ben again shared a lot of things that they're bringing to
[21:03] bear with governance. So we can have Unity Catalog remain that governance spine and apply governance uniformly across the entire ecosystem of platforms. And last but not least, the lakehouse gets stronger every time the business
[21:19] uses the model. Every time we close the loop, every time we introduce a new new use case, the lakehouse gets stronger. Let's make it simpler. One sentence.
[21:35] If there's one sentence that I want to leave you all with as you leave today's session, the goal is not to connect platforms merely, but the goal is to build a govern operating model that gets stronger every time the business uses it.
[21:52] That's it. That is the one thing if you can ground yourself on that, you're set up for success. I'm going to kick it back over to Ben, where Ben's going to share a little bit about how the partnership is scaling with partners
[22:07] and how you all can continue to stay engaged. Thank you. Thank you, Brandon. As you guys can see, Brandon and the
[22:22] rest of Rich Products have been incredibly thoughtful about how they want to deploy Palantir and Databricks next to each other. You know, I often think about things in terms of people, product, and processes. And no product uh a product is only as good as the people and the processes
[22:39] around it. At Rich Products, you have incredible people and they have defined a fantastic process for how to integrate these products together. So, Brandon, thank you for all the thoughtfulness you put into this. Now, I would be remiss if I didn't call out some of our thought leaders and
[22:54] system integrator partners who are helping us further the partnership by helping our customers integrate the two platforms together. This is just a small sample of some of my SI partners who are currently trained on the partnership.
[23:10] So, what's next? If you want to learn more about the partnership, there are a bunch of uh, you know, public-facing documentation and blog posts you can check out. Rory Patterson, Palantir's Chief of Staff, recently put out a blog post in October where he covers how over 100 customers are using the partnership to
[23:26] power their business. If you want to do a little bit of a technical deeper dive, there's a great YouTube video. It's Chad Walquist, an architect from Palantir, and myself. We cover an introduction to the partnership, and we give a demo of what the integration actually looks like in
[23:41] practice. You can find it on YouTube. And then finally, one of our partners, Blueprint, recently put out a blog post to describe what they're seeing in the field for how this partnership uh, is really helping to power their customers' businesses.
[23:56] If you want to get started today, there's a bunch of technical documentation online. I've got links to everything. On Palantir's website, you'll find the Databricks connector and all of the documentation under palantir.com/partnerships/databricks.
[24:14] Great first place to start. And finally, if you'd like to connect with me or my team, you can reach out to Taylor Melville, and he can help facilitate a connection. I'll give it over to Brandon to tell you to how to connect with Rich Products. Yes, if you would like to reach out to me, you can connect with me on LinkedIn. I will be uploading a supplementary post
[24:31] on today's session to probably go a little bit deeper. I know a lot of us are probably like, "I want more than that." So, I will leverage LinkedIn to share a little bit more about how to make this work. Awesome. Again, thank you, everybody, for spending the week with us at Data and AI
[24:46] Summit. It has been a pleasure talking to you guys about the Palantir and Databricks partnership. I think we've got about 15 minutes left. Brandon and I are going to be hanging out up here if anybody uh, you know wants to get to know us or do Q&A, we're going to be hanging out for the next 15 or so. Thank
[25:02] you. Thank you.
Learn more about the Databricks Data and AI platform.
The information provided herein is for general informational purposes only and may not reflect the most current product capabilities or configurations.