Skip to main content

Open Lakehouse Architecture: GEODIS Modernizes Data Platform with Iceberg

Summary

  • GEODIS, one of the world's top logistics providers, migrated its data platform from a complex self-managed multi-vendor stack to managed services including Databricks, Privacera, and Astronomer, standardizing on Iceberg v3 as the cornerstone of interoperability.
  • The migration was executed use-case by use-case with no big-bang disruption, preserving existing investments while enabling Databricks, Starburst, and legacy Cloudera components to read, write, and modify the same data simultaneously.
  • Security policy is managed centrally through Privacera, and data products organized around data mesh principles enable controlled internal sharing and external data sharing with partners and customers via standardized protocols.

Open Lakehouse Architecture: GEODIS Modernizes Data Platform with Iceberg

Watch: Open Lakehouse Architecture: GEODIS Modernizes Data Platform with Iceberg
GEODIS, one of the world's largest logistics providers with 37 manufacturing sites, faced a critical challenge: data was siloed across operational systems in different regions, making it impossible to build global use cases or cross-sell initiatives. After six years running an internally-managed platform based on open source and multi-vendor components, operational complexity became unsustainable. The data layer management team was struggling with too many components to maintain, upgrade, and support while delivering no new business value.
GEODIS modernized its data platform by transitioning from self-managed infrastructure to managed services, adopting Databricks, Privacera, and Astronomer while standardizing on Iceberg v3 as the cornerstone of interoperability. The migration preserved existing investments through gradual, use-case-by-use-case transition, with zero big-bang disruption. By leveraging Iceberg's metadata portability, they enabled true multi-engine interoperability: Databricks, Starburst, and legacy Cloudera components can now read, write, and modify the same data, with security policy managed centrally through Privacera. Data products organized around data mesh principles enable controlled sharing internally and external data sharing via standardized protocols with partners and customers.
🤝

Chapters

FAQs

What is Apache Iceberg and why did GEODIS choose it as the foundation of its platform?

Apache Iceberg is an open table format that provides metadata portability across different compute engines. GEODIS chose Iceberg v3 because it enables true multi-engine interoperability — allowing Databricks, Starburst, and legacy Cloudera components to all read, write, and modify the same data without engine-specific lock-in.

How did GEODIS migrate its data platform without causing big-bang disruption?

GEODIS migrated use-case by use-case, transitioning workloads gradually rather than all at once. This approach preserved existing investments and allowed the team to validate interoperability at each step before decommissioning components of the legacy stack.

Why did GEODIS decide to move from a self-managed platform to managed services?

After six years operating a self-managed platform with open source and multi-vendor components, the operational complexity became unsustainable. The data layer management team was spending too much effort maintaining, upgrading, and supporting infrastructure instead of delivering new business value.

How does GEODIS manage security policy across multiple data engines?

Security policy is managed centrally through Privacera, which enforces access controls consistently regardless of which engine a user queries through. This allows GEODIS to maintain a single governance layer across Databricks, Starburst, and their legacy systems without duplicating policy configuration.

Full transcript

[00:09] Hello everyone. Thank you for joining this session today. I know that's the last day of this summit and probably you are quite exhausted by all the the event and conference we have seen. Uh so thanks for for being here. Um this morning I was reading a post on LinkedIn about interoperability because datab
[00:27] bricks and Microsoft announced interoperability between datab bricks and Microsoft fabric and what was interesting is to read the the comments and the the the debate in the comments of this post about what is really interoperability.
[00:43] uh many of you have different expectation about uh interoperability of data of data platform because we don't have all the the same goal the same objectives uh and it does not mean exactly the same the same thing it's a it's a big world interoperability uh today um I'm not
[01:01] pretending to present you what we can do in general I will tell you our story at Jodis about interoperability of our platform uh our choices uh the why we have built our platform in that
[01:18] way in particular around iceberg uh and explain you what does it mean interoperability for us and open lake layout architecture so first I can introduce myself I'm Dome I'm French uh I'm co and chief architect officer at jodis so managing both
[01:35] enterprise architecture data management practice and managing the data platform by itself First perhaps a few word on Jodis. You may not know this company. It's a French one. Uh we are a logistic provider, one of the biggest in the world. We are
[01:51] ranked around 10th place in the world. Uh doing around 11 billion turnover um 50k employees around the world. We are doing logistics activities in almost
[02:07] every country. And what does it mean logistic and what how does it link to to data is that picture that is interesting for you. We are doing basically all the the different logistic and transportation
[02:23] activity. Uh we are one of the the only logistic provider to do so. It means road transportation uh freight inter intercontinental freight forwarding uh supply chain management logistic so warehousing management and last mile
[02:40] delivery. Basically the entire supply chain can be delivered by our company. And in term of data ecosystem, what does it mean and why this is important in the way we have designed our data platform is you have to understand that all those
[02:56] activities are really siloed in operational systems in different operational systems and sometime in different region. It means that it's complex. It was complex. We are try we are aiming to simplify that but it was complex to share to cross data to cross
[03:12] information more than data across the company to build global use cases to have global ambition uh uh cross sales activities etc. It it was six years ago only a dream for us. That's why we have
[03:28] started this transformation about data platform through uh an open ecosystem. So we started that uh a little bit more than six six years ago with uh guiding principles that we have
[03:46] kept all over those years. The first one is to have an open platform. For us, it was it was key to have an open open platform with opensource project uh as a source of
[04:04] everything built to to to deliver this data platform. We wanted to start on the cloud. We we didn't have any experience in big data on premise before that. uh for this elasticity of of the cloud because we were starting from scratch
[04:21] basically without knowing where uh what will be the the target in term of sizing in term of volumes number of use cases flows etc. But at that time we wanted to keep the capability to run this stack everywhere.
[04:37] The cloud was quite new for us. It was not the only uh target we wanted to to have. So uh the the capability to deliver that even in private cloud was was still there. The another point is decoupled architecture something really key uh
[04:55] that can lead some complexity. We wanted to uh have a full life cycle of data from the extraction to consumption with clear border and clear step to uh address those those transformation and
[05:11] those operation on data. So uh uh we wanted to have interfaces standardized in between and interoperability between the modules that compose the data platform. An important point uh we wanted to have
[05:27] uh as well one extraction for one golden source and many usage and if we were able to do that in real time it was as well a really important target. So we can summarize those design
[05:43] principles with that. What does it mean for us open platform open lake layout house and interoperability? It means that we want to work in an open ecosystems of vendors covering all the
[05:59] data related services. uh you have uh on the left part of the of this slide the different services uh we want to deliver we want to include in our data platform from the persistent services. So where
[06:15] we store our data with the different type of storage for the different type of data and use cases to the bottom uh to the top sorry uh with the data engineering uh services within between everything needed to manage the life
[06:31] cycle end to end. So it can leads to more complexity because you can see that if we have several providers, we will have to be able to integrate those providers to make something uh to make this platform
[06:49] as a product use usable by everyone in the in the company. But the the the advantage for for us are really important to be able to manage our vendors without too much looking and the
[07:05] ability to swap those modules to replace it because the market is evolving. Data usage is evolving. Six years ago AI was not really there or at least AI was there for classical machine learning project but we were not talking about
[07:22] generative AI. So this this ecosystem the the the the composition of this platform was really a key pillar for for for designing things that was our legacy stack.
[07:38] Um you can see that many of those components in orange were delivered through cloud platform uh in the cloud but we were already using some satellite component like manage CFKA starburst
[07:54] platform data etc MongoDB as a service to compose this ecosystem of provider and it was the the initial state of that platform six years ago and we are able to uh uh onboard all our company on that
[08:12] platform. So all the the business line the different businesses you have seen on the on the one of the first chart were able to start working collectively on that platform.
[08:28] We we are uh we have started uh six years ago uh to organize ourself around data mesh principle. Even if data mesh principle were not really uh defined at that moment, we are not talking about data mesh. Uh but we we are organized like that because our company is really
[08:45] decentralized. We have several business lines. We have several region that need to be autonomous in the way they are managing their data. But we wanted to propose them a unique way of working on standardized technical platforms. Even if we are
[09:01] talking about several instance we were u standardizing the the tech stack. So here you can see the the different biggest business line ag uh with the the different uh way of of delivering logistics and we have put uh
[09:19] uh at the center of that the data product approach. You probably know uh everyone what's a data product. I think every company has his own definition of a data product. We have ours uh but it's a it's a really a key pillar in the way
[09:36] we are uh delivering our data assets across the company trying to have uh product approach uh for every sharable data. It does not mean that all the data is defined as a product but that that's
[09:52] a key pillar to to be able to share and to break those silos I mentioned earlier on the platform and the the the organization as of today working on that technical stack I have presented is
[10:07] organizing data mesh and data productization. Last year we wanted to accelerate uh the stack uh I I show you at the beginning was mainly internally managed
[10:24] and uh it was not anymore sustainable to continue to manage all those platforms internally because of different problems. Um DM is for our data layer management
[10:41] team. uh so the internal team managing that to the the the the technical stack were struggling to sorry I'm I'm making noise with my foot uh we're struggling to uh maintain the the oper the the platform operational in
[10:58] the in a good way because too much complexity of integrating those component too many things to do by ourself uh too many uh component to upgrade to maintain to to support and to to monitor Um so we launched that project to to
[11:16] move uh our data platform to a a managed service uh uh platform. It means that we we don't want to replace our component. We want to replace the uh the each modules part by the same product but in
[11:34] a a SAS or pass mode. Therefore, the low-level management of the platform, the infra VM stuff management uh will shift progressively from Jodis to an external provider like
[11:49] datab bricks. What we wanted to achieve is to really have a data focus uh to be able to really uh support our data engineering team, data analyst and data scientist uh and not spend time maintaining all the
[12:06] the stack for uh no value, no business value at the end. So here the you can see the what I I tried to to sell to my boss at that time. Uh what were the the the
[12:21] improvement limit the workload improve agility financial benefits uh of course uh focus internal team uh uh on on data management and of course the transition would be transparent. It was a promise
[12:40] uh at the end it's quite transparent. So we had several commitments. Um as mentioned we wanted to progressively move each modules without big bang. We
[12:56] didn't want to transform the way the data engineer and the datist work on the platform. We wanted to to have a a smooth progression. Uh we wanted to red not redevelop everything. uh that's why
[13:12] we try to keep uh almost the same middleware but in a manage mode. An important pillar is to keep the platform fully open and interoperable. That's why we bet it was early 2025
[13:29] on iceberg. We were already using iceberg on the previous platform but now iceberg became a really a cornerstone of our architecture. Keeping this capability to choose
[13:46] different hyperscaler that transformation uh uh brought some some limits. uh we were not able anymore to move on prem but I think
[14:01] designing a new data platform on prem in 2026 it's something quite unusual I would say um and the increase to our love hyperscaler or SAS provider will
[14:17] progress really increase so the the evolution of the stack started like that. That was the initial state of uh of the platform. End of 2025, we started to replace the orchestration
[14:35] modules. We were using airflow. We moved to a manager flow version with astronomer. The second piece of of this puzzle was the security modules. We wanted to be able to manage our security everywhere
[14:51] on the in the stack but as well in the manage mode. We move to previous era. And then the big change is changing the way we are delivering our data engineering capability to to our
[15:07] data engineers. For that we move to data bricks. the the overall interoperability of those modules was key and that's the the purpose of
[15:23] the migration strategy. How do we I I keep interoperability between all those modules. We started the project the build phase end of 2025.
[15:39] You have all seen probably a great picture of iceberg in the middle with all the the the engine around sparks flingo data snowflake. Everything is supposed to work. But at least for us it was the the first time we tested really
[15:55] this compatibility of bridging all those components around uh one lake one data lake sorry one lake is another product so what was the the migration strategy
[16:15] first step decoupling the orchestration uh we would the orchestration was inside the cloudera data platform We wanted to have orchestration independent of any engine processing data. That's why it was the the first move replacing airflows, migrating DAGs, adapting DAGs
[16:33] to to this new context of of open modules. Second step, security standardization. I mentioned the previous product. With that with that stuff we are able to manage security policy on all those
[16:48] component seamlessly uh and we are able to assure security everywhere in the platform. Third uh key pillar is iceberg as a first class citizen.
[17:05] So why keeping iceberg? Um data with iceberg the bet is that data is available on with any engine. That was the the theory. uh whatever the the engine, whatever the
[17:21] software, datab bricks, starburst that I could, every every tools uh is supposed to to support iceberg to be able to uh to continue to consume the the very same data data is still sharable inside the the
[17:38] organization but as well outside. I will come to that later on in this presentation. But iceberg for us is as well important to be able to share our data to the outside. So how do we did we do that? Uh first on
[17:56] our legacy stack we were already using iceberg not at 100% on all our data lake. Uh but uh it was a I would say a prerequisite to be able to to continue that migration. we were using still iceberg v2 because
[18:13] cloudera was not able to uh to support iceberg v3 at that time. I think it's always it's still the case. Um before moving to iceberg v3 in datab bricks you have to understand that for for
[18:30] our opinion uh iceberg v3 is really the version that bring interoperability between all those component. Before that it was really theoretical and and not really practical. We had to adapt uh some our uh well most
[18:48] of our spark uh jobs uh well we have a standard spark job framework for managing ingestion and transformation. We had to adapt that that uh standard spark job to the datab bricks context.
[19:04] But I would say that the relay effort was more to adapt to iceberg v3 than to adapt to database contact. It means that the portability uh of spark jobs between those two world is is quite good.
[19:24] So what are the steps when it comes to migrate something? Uh we were not doing that uh in a big burn big bag approach but more use case by use case. uh if you remember in the presentation the different domain each domain business domain I mean uh is able
[19:42] to uh u set the pace of his migration we have common deadlines for everyone but everyone is able to to migrate when he wants and what he wants. So first migrating orchestration. So
[19:59] decoupling orchestration from from our legacy stack then migrating the table the data by itself. Uh what is really interesting with iceberg is that you have access to
[20:15] the metadata. You don't even need uh access to the catalog or to share catalog between those two world to be able to migrate data. A job uh a spark job can read the the lake
[20:34] um understand the metadata the structure of the table and do the the proper migration job. So for that it's it's quite easy. We are working in two steps. first copying the the the data to manage
[20:49] uh datab bricks table and then upgrade to v3 we we even didn't use any uh accelerator from datab bricks it's a pure uh internal development uh I would say that uh one data engineer has been able to
[21:06] develop this kind of tool in a few weeks I would Oh, sorry. Oh, okay. Uh, then uh when the the data
[21:23] has been migrated uh to a managed table in Unity, we were able to move the Spark jobs from the former uh data engineering stack to data bricks. For that we have
[21:38] we have only to change the orchestration layer and to reconnect the job to the the new Spark engine with some adaptation in the D for sure but it's quite easy. We are not changing the logic and the migration is really
[21:55] smooth. And the last no not the last the next step sorry uh is to replicate the the security policy because the cataloges are changing. So you have to reimpport the policy but
[22:10] it's quite straightforward. And the last step is to update the data product. If you remember what I was explaining in term of data consumption because we are using data product we are not touching
[22:28] directly the the table uh the physical table I would say we are only touching the data product for us it's some kind of view so a view update to redirect to the new table and that's it there is no
[22:45] impact or big impact on the data consumption Okay. And that's uh u uh I don't have the good word sorry but we are redoing that for every use case every sources uh and as of today we have
[23:04] migrated almost 50 a little bit more than 50% of our current data stack uh and we are targeting to end September the the entire migration. So now it seems quite easy to do that.
[23:21] I've said that it's easy but it's not completely out of the box. There are things that are working almost out of the box. Security policy management, job orchestration, Unity catalog integration is quite a de facto standard uh in the
[23:37] market. So no big problem uh to manage that data storage sharing we are talking about uh uh basic uh uh storage in uh in in we are working in Azure but in S3 in
[23:52] Amazon it will would be exactly the same. Uh this part is is quite easy to integrate in between those component even for identity management it's it's really not a big deal. The code portability uh is easy because we are
[24:09] using standard on the market spark with it's mainly scala but we are using more and more pispark as well. DAX conversion for airflow and iceberg metadata access as I mentioned is really easy to to to manage.
[24:27] what was still immature. Um I mentioned that iceberg v3 was quite new end of 2025 uh when we started to work with data bricks I think uh iceberg v3 was private
[24:45] preview at that time it's not it's it's now g since June I think um but it it was really uh u important to move to iceberg v3 to have real interoperability and and and when
[25:02] we are talking about interoperability it means that we want to have all those component able to read to write and to modify the the data and the table inside the lake six months ago um
[25:18] I I would not have answered the question in the same way because we were able to read information sometimes modify inserting data But we were not even able to modify structure of the tables from all our component all
[25:35] our engine. As of today we are able to do so. Uh we are able to do DDL operation from data bricks from starboard from other components. It's it's it's really working. It's not anymore uh only a picture on some slide
[25:54] from the vendor. I can assure you that it's working. Um we had to modify uh some our data sets to to go to parquet. Uh it seems that for iceberg park is now a def facto
[26:09] standard and OC is is disappearing and sometime iceberg is not still by default in some tools you have to force a little bit the things to use iceberg as the default table format.
[26:29] There are still some incompatibility uh between cataloges uh uh and connector issue. Uh I can give you an example uh starburst using manage table inside inside datab bricks. Uh to
[26:47] do some detail operation uh datab bricks was expecting some parameters that were not sent by starburst. we had to adapt that and for that I really uh want to thanks datab
[27:05] bricks and the all the the the stakeholder data brick that have been able to adapt their uh the the way they are implementing iceberg probably not only for us but for our timing it was good uh they have been
[27:20] able to adapt those those really small parameters this small fine tuning to make this uh really able to to to work properly.
[27:36] I mentioned some adaptation in the code related to iceberg in particular to table partitioning behavior. it's working quite differently uh because we we have to move to manage tables something that I didn't mention that today if you want really uh to have this
[27:53] interoperability around icebear you have to use um manage table in um in unity catalog it's not it's still not available for for external tables uh it it was not our initial target but
[28:11] it's working and it bring quite a lot of value at the end for us because the the the iceberg table management with all the the operation of maintenance is automatized by uh by the manage table in unity. So really useful for that.
[28:28] uh and we the last point we had to do some adaptation uh on the the generic spark jobs uh to move between CDP and and data bricks. So everything today is working. There are
[28:45] still a few component in uh public preview uh in in those components. It's okay for us to to move uh and to to continue to use that. uh but I can assure you that this kind of architecture is working. We we can say
[29:02] that everything is interoperable around iceberg with with datab bricks and the other component. The next move uh we have not finished our transformation. We have still some component to move to
[29:18] a a manage u service uh with knifi with fast analytic storage. We are using Apache Kudu. Uh we have still other component that we want to move to SAS uh to simplify our our maintenance
[29:34] and iceberg is really becoming the cornerstone of interoperability. Uh I mentioned this ability to to have several components using iceberg but really important things that we are testing as of today with our colleague
[29:51] in the US in particular is external data sharing using standard protocol with iceber. Now everyone almost everyone is able to manage uh well to to to use iceber table. Uh so we are testing that
[30:08] direct sharing uh using uh standardized B protocol with with third parties with partners with customers outside to make those data platform able to discuss directly uh wi without any interface in
[30:25] between. That's the future uh picture perhaps the the future uh view of our data stack. Uh datab bricks announced quite a few things two
[30:42] days ago around fast analytic storage that is really interesting that we will test for sure. uh but at the end we have transformed our data platform keeping everything interoperable not breaking anything and by mid 2027
[31:03] we will have moved to a fully SAS or pass platform with external management of the infrastructure. So I mentioned that you have probably seen
[31:19] all those picture of iceberg as a interoperable uh uh item. uh and last year I think it was quite a lot of marketing lot of announce uh today mid 2026 I can assure you that
[31:37] it's working uh and uh it's not anymore a wish to have a real interoperability with open standard we are able to make those different vendors
[31:52] working together for for us is really important to to to keep the control on our data to keep the control on our data platform uh and to be able to try to to drive a
[32:08] strategy and not to follow only uh what the the market is doing. If you have any question uh I will be happy to answer. Thank you very much.

Learn more about the Databricks Data and AI platform.

The information provided herein is for general informational purposes only and may not reflect the most current product capabilities or configurations.