Polars and DuckDB with Unity Catalog: External Engine Integration
Summary
- Decathlon, a French sports retailer operating across 55 territories with petabytes of data, migrated from AWS Glue to Databricks and faced the challenge of maintaining compatibility with external compute engines like Polars and DuckDB while benefiting from Unity Catalog governance.
- The technical solution uses temporary credentials to enable read-write access from Polars and DuckDB to Unity Catalog external tables, while the Catalog Commit feature unlocks interoperability with managed tables that external engines cannot otherwise write to.
- Decathlon's strategy simplifies compute by consolidating most workloads to native Databricks compute while maintaining Polars on EKS for specific use cases, reducing the integration tax of managing multiple fragmented compute environments.
Polars and DuckDB with Unity Catalog: External Engine Integration

Unity Catalog provides unified governance, but many teams need to query external engines like Polars and DuckDB. Decathlon manages petabytes of data across 70+ countries and faced a critical challenge during migration from AWS Glue to Databricks: how to maintain external engine compatibility while achieving the governance and performance that Unity Catalog offers.
this video covers the technical implementation for running Polars and DuckDB against Unity Catalog tables. Learn how temporary credentials enable read-write access, the limitations of current external engines, and how the Catalog Commit feature unlocks managed table interoperability. Decathlon's engineers share their operational strategy for simplifying compute while maintaining Polars on EKS for specific workloads.
🤝
Chapters
00:00Introduction and Session Overview01:15Decathlon Company Overview and Data Platform Journey04:11From AWS Glue to Unity Catalog: The Well-Integrated Lakehouse06:06Foundation and Migration Strategy: External Tables First08:15Challenges: Integration Tax and Compute Fragmentation10:41Strategy Pivot: Simplification and Polars Exception12:22Unity Catalog Table Support: External vs Managed14:16Polars Live Demo: Reading and Writing External Tables18:50Polars Limitations and the Catalog Commit Solution22:37DuckDB Live Demo: Managed Table Integration and Lineage27:40Lessons Learned: Standardization Benefits and Future Roadmap
FAQs
Why does Decathlon need to run Polars and DuckDB against Unity Catalog?
Decathlon has specific workloads on EKS that run efficiently on Polars, and migrating all compute to native Databricks at once was not feasible. Maintaining external engine access to Unity Catalog tables allows the team to migrate incrementally while preserving compatibility with existing Polars-based pipelines.
How do external engines like Polars and DuckDB access Unity Catalog tables?
External engines can read Unity Catalog external tables using temporary credentials that Unity Catalog issues on request. These short-lived credentials grant scoped access to the underlying cloud storage, enabling read operations without exposing permanent credentials to the external compute environment.
What is the Catalog Commit feature and why does it matter for external engine interoperability?
Catalog Commit is a Unity Catalog capability that allows external engines to commit writes to managed Delta or Iceberg tables, which external engines otherwise cannot update directly. Without Catalog Commit, external engines are limited to reading external tables; with it, they gain full read-write interoperability with managed tables.
What was Decathlon's overall migration strategy from AWS Glue to Databricks?
Decathlon migrated external tables first to establish a Unity Catalog foundation without disrupting existing workloads, then evaluated which compute engines to consolidate. The team standardized most workloads on native Databricks compute while retaining Polars on EKS as a justified exception, reducing integration complexity across the platform.
Full transcript
[00:10] Hello everyone. So, my name is William. Uh I'm a staff for the plane engineer I had data bricks. And with Benjamin, we are going to do a lessons learned on a new eating catalog migration at Decathlon. And we are going to narrow our talk
[00:26] specifically into the interoperability of uh Unity catalog with external engines basically. Um so So, basically the how we have worked, it was a very long engagement together. So, basically I worked closely
[00:43] with Benjamin for pretty much a year to help the Decathlon migrating into Unity catalog. And basically the goal of this talk is, as I said, talking about specifically the interoperability of uh Unity catalog. Yeah. Hi everyone. I'm I'm Benjamin from
[01:00] Decathlon side. I'm data platform staff engineer. So, I I basically I lead the I lead the this topic for for Decathlon. And we were helped by principally by by William along the way this this year and
[01:15] the year before. So, today's agenda um I will I will present Deca- Decathlon. What's What's Decathlon? Just quick question in in the audience, who knows Decathlon? I see some some French Obviously, you know you
[01:31] know it, but okay, so some of you don't know, so I will I will present a bit uh what's Decathlon. Then I will present our strategy regarding the interoperability with Unity catalog and the migration. Then I I I will
[01:46] hand over to William. He will do some demos for you about this kind of inter- ability with the external engine with UC. And then I will talk about what's next.
[02:02] So, this is the Decathlon Decathlon store. So, Decathlon is a French company, big sports retailer. Um it's a it's a store in Shanghai, for for example. On the left-hand side, you you have all the opening the store opening
[02:17] um in in every country. So, it started 50 years ago. Actually, last this week it was the the anniversary. So, Decathlon is 50 years now. And almost 2,000 store across the across the the the globe.
[02:33] On 55 territories. The purpose of Decathlon is to bring people together through sports to make well-being accessible for all. So, basically, we want everyone in the planet to have access some
[02:50] to the best sport experience. To do that, Decathlon is designing its own product its own services is building the the the product across
[03:06] 43 production countries, shipping it, and then selling it in an omni-channel experience. So, this bring a lot of data. You can imagine value chain data, sales data, a lot of data.
[03:23] We are talking, yeah, petabytes of data. This kind of innovation the two that's the two I I I wanted to highlight because I liked I liked them very very much.
[03:39] The 2-second tents for its simplicity of deployment. And the Easy Breath mask for its integrate in some yeah, it's well integrated. So, I wanted to highlight that
[03:54] like best innovation or data platform wants wants to be the same, well integrated, easy to use, and high performance.
[04:11] So, let's talk a bit about the first about the journey of the of the data the data platform. It all started on AWS 12 years ago. And fun fact, the the first person that launched the the the data platform had to pay the bill, the monthly bill with his own credit card at the time. So, it was small, you know.
[04:27] Right now, I think he cannot afford the a monthly bill. Um so, along the way we we still we are we we still we are still on AWS, but we integrated Databricks uh 5 years ago. And
[04:44] a little by little, we we are more and more integrated in in Databricks. So, last year, we wanted to achieve the a big milestone for us. Uh as Ali said in in the keynote, if you remember, there is four chapter, and the
[05:00] chapter one is well integrated lakehouse. This is what we we wanted to to achieve. We went from AWS Glue uh uh um in Databricks as a metastore to Unity Catalog because like you like
[05:17] like you I think you know that Databricks is releasing more and more features. Everything is around Unity Catalog. So, we had to do this move fast. So, we we wanted to achieve this this big milestone for the having a uni- unified experience,
[05:33] unified governance, and uh counting on multiple new features to scale. So, last year, as Um, as William said, thank you by the way for the work that you did with us.
[05:51] We When when I say we for the two first part, it it's the data factory of of Decathlon. So, we are almost 70. Um, half of the team worked on on this on this migration for different from
[06:06] different teams. We had to to build a foundation um, before doing anything. So, when I said foundation, I mean a Terraform module, um, I mean, um, designing the the catalog, uh, shipping it with with Terraform as well,
[06:22] designing the permission the permissioning, uh, model. Um, moving from, uh, workspace level attribute assets like groups, like, um, users to account level. That's how Unity works. So, all
[06:37] of these are I I called it the foundation. Then, for our users to be able to to migrate easily, we synchronize er, each each table that we had in in, um, in Glue to Unity
[06:53] Catalog external table. Why external table and why we we did that at the time? Because we didn't want to move the data uh, and we because when it's external, it's just a pointer, you know, when you when you want to have a managed table,
[07:08] it's recreating the data in in another location. We didn't want to do that at that at this time. So, basically what what that did and what that allows for the our users, for our consumers, it's uh, it allow them to already
[07:25] migrating their, um, their pipeline in in their their consumption pipeline to Unity Catalog. And in parallel, the producers of the data could at their their pace migrate in right mode into Unity. So, this
[07:44] allows this parallel work. And the final step, we are we are almost there. Uh I expect to we finish uh by the end of the year the the whole migration and and get rid of the Glue ecosystem.
[07:59] We want to have managed table into Unity Catalog for every data products. So, as I said, we we have a long history now, more than 12
[08:15] 12 years. We we are we had a lot of compute uh engine. And the initial strategy with Unity Catalog was to allow flexibility for our user and keep their keep their beloved uh compute
[08:32] that they use. So, we wanted to make all these compute work with Unity Catalog at first. But uh I would lie if uh I say it's it was easy. Um
[08:48] it's Yeah, we had some friction. So, the first one I wanted to highlight is that um it has a cost uh for the admin um the admin teams, you know, to to manage infrastructure,
[09:03] um different infrastructure, um decouple compute layer, etc. etc. So, it it's it's a burning bandwidth for our engineering team. The second one is the I call it the integration tax because
[09:19] uh right now and at the time last year, the integration of external compute engine with Unity Catalog was not working perfectly, uh I I would say.
[09:34] I will I will take two example. One, Athena on AWS It's if you want if you want Athena to work with Unity Catalog with Delta table under the hood, you have to first use
[09:49] Delta uniform to generate Iceberg metadata because Athena is is with Unity is is just Iceberg compatible. Then you have to replicate your permissioning system from Unity Catalog to AWS Lake Formation. So it
[10:07] could be done for one to table if you you want but at scale with more than 12 or 2,000 table it's a nightmare. So that's was my my first example. The second one if you want to and William we are will demo that and will
[10:24] highlight the second point. If you want to to deal with manage table it's not as as easy as external table. So this is the two point that I want I wanted to highlight. So what we did
[10:41] I I am I said the initial strategy was full interoperability. We started like that and then we decided to simplify our our compute ecosystem
[10:59] drastically. Then we wanted to finish the the the migration to Unity Catalog as as fast as we can to unify the governance. These two points led to the the third one. Doing that we reduced a lot the
[11:16] friction. We reduced the integration tax. But we kept one exception. Not everything is on Databricks. We kept one. Why? We we kept polar on IKEA's
[11:32] because at this time because um yeah, at this time like I said, the serverless offer in Databricks for our side wasn't ready. Now it's another discussion. But at this time we wanted fast spin-up,
[11:49] cluster spin-up. We want We wanted cost effectiveness effectiveness for some use cases, so we used a mutualized EKS cluster. And uh Polar's on on top of that. So we kept for this the challenge of
[12:04] integrating uh an external engine to Unity Catalog. I will hand over to William to explain all of this uh how how how it works with uh two examples. Thank you, Benjamin. So um before diving into Polar specifically, I
[12:22] just want to uh step back and do a few reminders uh for everyone in the room about what is the support today in Databricks in terms of tables. So, as you know, in Databricks we have two types of tables in Unity Catalog. You have external tables and managed tables.
[12:39] Um I don't mention foreign tables and there are also some other table types. It's out of scope for the moment. I just want to focus on these two ones. And obviously, in tables you can read them, you can write, and you can create these types of tables, okay? So,
[12:55] um when we started with Decathlon, we had the full support for external tables. So, it means that by using an external engine, you can basically read them, write them, and create them, okay? However, just keep in mind something is
[13:13] that um depending on the tool, on the engine that you are using, the support may be a little bit different, okay? Um for example, in Polar specifically, you can read and write to external tables, but you cannot create them,
[13:28] basically. You have to deal with an API call. And depending on another tool also, it's you can have the create support. So, you have to be aware of that, obviously. Now, for managed tables, we have the read support even before
[13:45] catalog commit, it was possible. But, now you can actually write with an external engine to manage tables, but only if a specific feature is enabled on a managed table. That feature is called catalog commit. I will come back to that
[14:00] later. And you can also create managed table now with an API. Okay? So, if you want to implement the Unity Catalog and Polaris integration, I'm going to do a little bit of coding
[14:16] right now, and I'm going to highlight the actual implementation that we did. So, it's it comes with three steps. First step is you have to create a workspace client. So, if you don't know what the workspace client is, it's basically the Databricks SDK, something
[14:31] to interact with Databricks workspace APIs. And here we are just authenticating, okay? We are just saying, "Okay, this is the workspace URL, this is the token." So, I'm using, let's say, my own user. And now I have this instance. Once this is done, then I have to fetch
[14:48] the table ID in Unity Catalog. This is what the second line is doing, basically. I just want to fetch the information of the that particular table. And then the third line, it's what it's what's interesting. Basically, what we are doing is we are asking Unity Catalog to generate
[15:06] temporary table credential to read that particular table. So, what it means is that Unity Catalog, the service Unity Catalog, is going to give us um
[15:24] is going to give us a temporary AWS access key, a key ID, and also a token. And this allow us basically to interact with the AWS storage that is behind Unity Catalog. Okay? So, that way you don't need to have to manage a separate IAM role. You need to You don't need to
[15:41] manage separate permission. Everything is done by Unity Catalog. Okay? Unfortunately, you can't really see on the bottom, but basically we inject that storage option into the into the Polaris reader. Okay? And now
[15:57] Now, let me do a quick demo of all that. So, I'm going to my ID here and basically this is the code. Hopefully, you can see that well. Um Basically, what I'm doing here is just the exact same implementation. So, I'm
[16:13] just reading a from file my workspace URL. I'm also reading the token. I'm fetching that the table ID, creating a temporary credential, and all that. And I'm just going to call the Polaris reader and
[16:30] then I'm going to print the data frame. So, if I run these commands, normally and magically we should see All right. We should see uh we should see money. And if I go to the Databricks workspace, actually,
[16:46] if I can find it again up. Let me go back here. All right. So, as you can see this table and apologies for the one in the back. Uh okay. Yeah, we can do that.
[17:01] This is an external Delta table. Okay? This is the table that I was using in my in my code. And if I do the sample data, as you can see there's just these two rows. Okay? So, in my external engine I can I'm in my local computer by using Polaris I can read that.
[17:17] Now, what about if we do something else, which is actually writing? So, I have another piece of code here, which allow me to write to the same table here.
[17:32] Okay? And in here, the only difference is that the operation I allow a read-write operation. So, here I allow Unity Catalog to say, "Look, give me a credential so that I can write to that particular table." And then, you know, I'm injecting uh to
[17:48] uh to just just a single hole here. I run the code while I was speaking. And if I recall the the the the read code, then you can see uh the the new row that I just did.
[18:04] Okay? And if I go back to Databricks, I have obviously to reload the window to refresh that. And obviously, the compute stopped, so we have to wait. A bit. And normally,
[18:19] you should see something magical happening in just a few moments. All right, we can see it. So, this was just a very quick demo on how you implement read and write uh on how you implement read and write to
[18:34] um to Unity Catalog by using an external engine. Okay? We use by using Polars. So, going back to the slides here. Um what are the limitations of all that? Because this looks great on a demo, but there are a ton of limitation. And I'm
[18:50] here to highlight all of them. So, as I said, you cannot create tables manually. You have to use the SDK and API for this. Um if you are familiar with Polars, you may know that there is already a Unity Catalog extension that exists. But while
[19:07] testing, we realized that that extension is uh very impractical and experimental. And this is because you have uh a read-only implementation. There is no write support. So, if you need to write by using this
[19:22] uh UC extension, you still need to do the implementation that I just did. And so, it was actually much more convenient from our perspective to use the existing code and just add the boilerplate code that I've created so that we don't have another migration in the migration to
[19:38] do, okay? Um there are also some already known limitation is that if you create deletion vectors on Delta external table, Polars cannot read that. This is uh a known issue in the Rust implementation of Delta.
[19:55] Um also, one important thing is that we are losing the lineage in Unity Catalog. So, basically, in the demo that I did, you won't see the lineage information in Unity Catalog. So, you have to be aware of that. And lastly, obviously, the catalog commit support is not yet available in
[20:12] Polars, okay? Only some supported engine had that. So, um in addition to that, the implementation that I show you has a few downsides is that it relies on the Databricks SDK to generate table credentials. So, there is
[20:29] a dependency that you need to add to your project. You need also to generate a credential for each table. So, if you have like 50 tables, 15 tables, whatever, um it's becoming a little bit messy at some point. It only works on UC external tables.
[20:45] This is uh another limitation, but it's going to be fixed eventually. It's the If you apply a row filter and a column mask in that external table, you won't be able to interact with that table. It's in private preview, but this should be fixed at the at some point.
[21:03] However, not everything is black here. There is a little bit of hope. Because with the release of catalog commit, we expect that the Rust implementation of Delta will fix a lot of this issue. Like deletion vector, and also read and
[21:18] write to manage table without using some fancy implementation. Okay. So, um I've talked a lot about catalog commit, and just going to do just a one slider about what this is all about. So, basically,
[21:34] uh this feature it unlocks interoperability with Unity Catalog managed table. So, it means that the support for fine-grained access control, just like row filter and column mask, this will be available for external engines. So, we release the API for
[21:50] that, so that external engine can implement that. And lastly, you have the multi-statement and multi-table transactions available on managed tables in external engines as well. Obviously, the caveat with that is that the external engine needs to
[22:05] implement that feature. Okay? Um if you want to do that, you need to enable the You need to enable that table option on the bottom. So, delta feature catalog managed to true when you create the managed table. Okay? And just like that,
[22:20] it works. Okay? And you need to use a pretty recent DBR. So, you need to use DBR 16.4 plus to read that. And also, on the external engine you need to have a recent Delta or Spark version. Okay?
[22:37] So, now I'm going to do a quick demo on uh using DuckDB on my local machine and have a read and write uh read and write fashion uh on a managed table. So, I'm going to first to go back into the
[22:52] workspace here and go to my managed table. So, in here you can see that this table is a managed table on that catalog, that schema, and that table. If I click on sample data, you can see just the same row here.
[23:09] And if I go back to my terminal here, so if I do select star from manage table, I have the previous state, so I have a demo effect where it's everything is cash, so I need to reload that.
[23:25] It's all right. I should be able to do that. I need to reload all that. I have a demo effect, sorry. I have a demo effect. I'm sorry. So, this should work actually. Basically,
[23:40] what I did what I have to do normally if I detach the catalog, let me just to fix that in the demo quickly. Uh no. memory
[23:57] I will just do 5 minutes, and if I click attach again, hopefully that works. Bear with me, everyone. Okay. show all
[24:12] tables All right. We got that on manage table. And normally All right. Fixed that. So, as you can see,
[24:27] now we can have the row that I just show you in Databricks into DuckDB. So, as you can see, I had to go back because DuckDB has basically loaded the manage table inside of the memory, so that's why I had this error. So, I just had to
[24:44] uh quit the um the the the not the session, but basically um I I I had to quit the the the connection to Unity Catalog, reload the connection so that I can do a select star, okay? And now, if I do an inserting two
[25:00] Now it should it should actually work. I'm sorry if I don't see the bottom of the screen. Again, I'm just trying to do my best here. All right. So, the insert is done and if I do the select star now I should see something magical is that it actually
[25:17] works. Now let's go back to Databricks to that managed table. I reload that. And and voila. That works. So, just a quick demo on how the
[25:32] interoperability with managed table works. Maybe maybe you can show a bit more than than that for the the integration. If you go to the history for instance. Yeah. So, again I wanted to talk about that
[25:48] in by speaking, but there are a little bit of caveat in here. Is that if you look at the history here you can see that the user ID and the username has been unknown on that on that transaction. So, it means that
[26:03] in the table history you know that something has happened in the table, but you don't really know what's happening. So, the only thing that you know here is that you just know that there is a DuckDB somewhere that has done something on the table. So, this is one current limitation that you guys need to be
[26:19] aware of because this implementation here again is a little bit experimental at the moment. Um what else? And also one last thing about that implementation The lineage. Um You can show the lineage actually if you
[26:36] want. If you want I can show that. Uh and in the lineage if I click on the lineage graph there is not that much of of informa- information just because on the assets that writes the data. So, this is just the queries
[26:51] and the notebooks that I've created just to interact with that table. And just like also assets that reads the data, there is not that much because you don't have the lineage information and you don't really have a lot of information at the moment into the
[27:08] into the into the at the transaction level. That's what I call the integration tax. Yeah, exactly. This is kind of the limitations that you guys need to be aware of because before you push something to production, obviously. And lastly, also speaking about the DB
[27:24] just to finish. This is you can just do a read and a write at the moment. If you want to do an update and a delete, the integration doesn't work at the moment. So again, you need to be aware of that. Thank you everyone. Thank you William.
[27:40] I wanted to highlight the lesson learned here during the year. What what we learned the first thing is that the standardization actually didn't limit our team.
[27:57] It freed them. We were afraid of of that because everyone loved what what they used. It works. It worked. And they were afraid we were afraid to
[28:13] Yeah. Limit them and actually it's the opposite. So this is a win that we had. The second the second thing like William showed you, I think it's really clear now. Interoperability works.
[28:28] But at the cost. And we decided to take our chances our chances with Polars. Um and basically we maybe in the future we will
[28:45] need to to take other choices. Except if um the different community works together and make it perfect.
[29:03] So, yeah, simplification is the ultimate sophistication. If you simply if if you if you can afford to um to take time to integrate stuff, it's okay. But, if you cannot, as soon as something breaks, something is
[29:18] difficult, just simplify. And that's what we did. So, what's next? As I said, we'll evaluate of the the newest features from the different communities. We'll um unlock unlock stuff.
[29:35] Um we'll we'll see Unity Catalog features, Delta features, see how in our case Polar will integrate them or not. And depending on the results, we'll need as I said to to take some some some choices. Whether we
[29:52] keep um keep the interop interoperability work either go to Databricks full unification. We are constantly evolving our our platform.
[30:08] Um so, if you guys are interesting in in this kind of discussion, feel free in the discussion later on LinkedIn or anywhere. Uh we we are happy to to discuss with you how how it works on on
[30:24] your side. And uh yeah, evolve our our different platform. So, if you have some some question about that, I see that we have 9 minutes left. So, feel free, guys. And thank you very
[30:39] much for your attention.
Learn more about the Databricks Data and AI platform.
The information provided herein is for general informational purposes only and may not reflect the most current product capabilities or configurations.