Hafnia Data Supermarket on Databricks with Lakebase, MCP and Marvis AI
Summary
- Hafnia replaced 10 fragmented data warehouses with a single governed Databricks platform organized around domain squads, creating a data supermarket where employees, applications, and AI agents all access the same governed data through different interfaces with the same governed answer.
- The DNA port application uses Databricks Lakebase as its Postgres backend, synchronizes operational writes with the lakehouse bidirectionally, and applies Unity Catalog governance across both systems, with over 160 AI-generated metric views grounding Genie spaces organized by business process.
- Hafnia hosts its own MCP server on Databricks Apps, and an internal AI agent called Marvis AI queries governed tools in parallel across multiple Genie spaces and is also accessible to external AI clients such as Claude.
Hafnia Data Supermarket on Databricks with Lakebase, MCP and Marvis AI

Hafnia's data supermarket on Databricks gives employees, applications and AI agents different ways to access the same governed data. this video explains how Hafnia replaced fragmented warehouses and disconnected applications with domain-owned data products spanning analytics, operational workflows and conversational AI.
See how the DNA port uses Databricks Lakebase as its Postgres application backend, synchronizes operational writes with the lakehouse and applies Unity Catalog governance across both systems. You will also learn how 160+ metric views ground Genie spaces, how Databricks Apps hosts the Hafnia MCP, and how Marvis AI queries governed tools from the DNA port and external clients such as Claude.
Learn more about Databricks Apps and Lakebase: https://www.databricks.com/blog/how-build-production-ready-data-and-ai-apps-databricks-apps-and-lakebase
Chapters
00:00Hafnia's Governed Data Supermarket02:34From 10 Data Warehouses to One Governed Platform04:10Organizational and Technical Transformation06:18Domain Squads Organized Around Business Processes07:53The Integration Tax Between Analytics and Applications09:46One Databricks Data + AI Platform for Data, Software and AI11:07DNA Port Demo Powered by Databricks Lakebase14:05Turning Analytics into Operational Applications15:41Lakebase and Lakehouse Bidirectional Synchronization18:41Unity Catalog Governance and Automated Deployment21:25Hafnia AI Assistant Powered by Genie23:55Embedded Genie Spaces Organized by Business Process25:40Why Governed AI Requires Metric Views28:05Generating 160+ Metric Views with AI29:43Hafnia MCP Architecture and Data Team Ownership30:56Marvis AI and Hafnia MCP Demo32:54Querying Multiple Genie Spaces in Parallel34:45Using Hafnia MCP from Claude and External AI Tools36:51Agent Memory, Identity and Unity Catalog Governance38:45Genie as Code, MLflow Evaluation and Deployment40:09How to Start with Metric Views, Genie and MCP
FAQs
What is Hafnia's data supermarket and why did they build it?
Hafnia's data supermarket is their analogy for a governed Databricks lakehouse designed to give employees, applications, and AI agents access to the same data through different interfaces while always returning a consistent governed answer. They built it after struggling with 10 siloed data warehouses, 10 different truths, and an inability to exchange data across departments or scale to AI use cases.
How does Hafnia use Databricks Lakebase in the DNA port application?
The DNA port application uses Lakebase as its Postgres application backend for operational writes, which are then synchronized bidirectionally with the Databricks lakehouse. Unity Catalog governance is applied across both Lakebase and the lakehouse so that data products remain consistent and discoverable across operational and analytical workloads.
What are metric views and why does Hafnia use more than 160 of them?
Metric views are governed SQL definitions that encode consistent business logic and serve as the grounding layer for Genie spaces and AI queries. Hafnia generated more than 160 metric views using AI, organized by business process, to ensure that governed AI always returns answers aligned with trusted business definitions rather than raw table data.
How does Hafnia's Marvis AI agent use the Databricks MCP server?
Marvis AI is Hafnia's internal AI agent that queries governed tools through the Hafnia MCP server, which is hosted on Databricks Apps. It can query multiple Genie spaces in parallel and is also accessible to external AI clients such as Claude, enabling governed data access from diverse AI tools without bypassing Unity Catalog controls.
Full transcript
[00:09] Self-service analytics taught our people how to shop. Now, we have a new kind of customer. Hi everyone. I'm Simon and thanks a lot for joining me today. So, I'm super excited to share our data journey. So, my talk is not a Databricks feature
[00:24] tour. It's a story about turning a governed lakehouse into a supermarket where human, AI, and application can shop. Three types of customer, one supermarket. But before I start, what do I mean by
[00:41] supermarket? We call it the data supermarket. This is our visual analogy used to transform the complex data lakehouse into an intuitive, organized, and easy to shop experience for all our employees.
[00:58] From 2021, uh our primary goal was to enable self-service analytics. Now, our domain expert and some of our employees are able to shop themself in the supermarket. So, hold this picture.
[01:13] Three customers, three way to shop the same supermarket. Different interfaces with the same govern answer.
[01:30] I'm Simon. I'm from France. I guess you got it and I live in Singapore since the past 11 years. So, I lead the data and AI team at Hafnia and today my talk is about building a gigantic business application on Databricks with Lake Base, MCP, and Marvis AI, our own internal
[01:47] agent. All right. Very fast, I won't spend long time here. Hafnia. So, Hafnia, we are a company. We are ship owner. So, we basically own very large vessels. We
[02:03] have 4,500 employees around the world. 300 of us are in the offices and 4,200 plus are crew member. We manage around 200 vessels and we own 120 tanker vessels. This is the biggest
[02:18] fleet in the world. So, we are a global shipping company. We're in Copenhagen, Singapore, Dubai, Monaco and Houston. And the data We are splitted among Singapore and Copenhagen.
[02:34] And we are of course daily powered by Databricks. All right. So, where we started? The challenge we faced few years ago. Every Hafnia department had its own warehouse, its own dashboard, its own
[02:51] Excel, its own database. It was 10 warehouses, 10 truths. Of course, we were facing the same issue with the siloed communications. We were not able to exchange data. With that setup, we were definitely not
[03:07] able to scale. And let's not even talk about AI. So, this is the boring failure mode of large enterprises. Data fragmentation and it's really not by design. It's because just nobody own the connection between the different
[03:23] departments. The trigger. In 2022, 3 years into the data team buildout, we were struggling with delivering value. And it wasn't a skill
[03:38] issue from the BI team. It was the way we were organized. Up until then, we measured ourselves on output, the number of report and dashboard we were able to ship and the number of requests we were of tickets and requests we were able to close.
[03:54] But, we went through a self-questioning phase. We were feeling something was wrong. We were not able to deliver at scale. So, after this self-questioning phase, we started measuring measuring ourselves uh on outcomes. What changed in the business because of our what?
[04:10] Work. That single reframe is what unlocked everything that came next. The supermarket, the squad, the application development, the MCP that you will see later. None of it would have happened without that trigger.
[04:26] How to scale? Our answer to that trigger. We had two shifts in parallel. The organizational shift and the technical shift. And trust me, the tech shift was the easy part. Um so, the org shift, as I said earlier,
[04:41] we introduced the concept of data supermarket into the company as a reusable and governed data product platform. Again, our users are from the shipping industry. When you speak about data warehouses, it's really not understandable. When you shift the discussion through data supermarket,
[04:58] everybody is able to uh understand what we mean. We move from a centralized BI team with 10 BI developer and one manager to three autonomous domain squad. We have uh three squad and we'll drill down later. We shifted from reactive delivery to
[05:14] tactical planning. What does it mean? Reactive delivery is basically the first person uh from the business that come to us will be the first one that will be served. And again, this is not scalable. The tech shift now. We move from SQL Server
[05:29] to uh and Synapse to the lake lakehouse paradigm, of course, on Databricks. So, we choose to go towards one platform to manage it all across BI, data engineering, AI, and application development. And that's really the topic
[05:45] of this talk today. Our engineering team loved it. Uh it was not a a downgrade for us. It was clearly an upgrade. We had to learn many new skills. We had to uh really reshuffle the way we were working, but it was amazing because now we're working on the latest technology
[06:02] on the Databricks platform. The real The real challenge here was redefining the roles. Moving from ticket taker ticket taker to domain owners. That's what unlocked leverage. All right. So, let's drill down to the
[06:18] organizational shift. So, we've organized ourselves into three domain squads. We stopped organizing around technologies and department. We started organizing around business processes. Every squad owns its shapes end to end.
[06:34] From data extraction to dashboards to operational application and AI experiences. The commercial squad is owning the whole process and the tech stack. This is very important. And our conversation with the business changed from can can IT build this to
[06:51] what deliver the most value. I will give you a quick example on the commercial squad. I'm not sure if you're familiar with the shipping industry, but the first process global process we're working on is the pool management. So, basically, it's like a real estate
[07:06] agency. If you have a house that you don't want to manage, you just give it to the real estate estate agency, they will manage it for you and collect the fees. We do the exact same thing with our vessels. So, if one of you own a vessel and don't want to manage it, I'm open for
[07:22] discussion after the talk. So, then the next step it's about the position planning. Uh it's the strategy part. Where do we send our fleet? Do we send the fleet to the east? Do the send the fleet to the west? All right. So, the next uh process in
[07:37] the commercial squad is to take care of the fixture and compliance. A contract in the shipping industry is called a fixture. So, before signing a contract, we need to ensure that everything is in order. And here we are not involving only one department. And that's what that's what
[07:53] is interesting with this approach. It's a a process is owned by different department with main uh process owner. All right, and so on and so on uh in this squad until the claims management process. This is the same uh approach for the technical uh squad and the finance and
[08:10] super squads. Ownership moved moved from system to business outcomes, and this is very important for the next part. All right, so here it will look familiar. Okay, the uh Ali uh slides were much more interesting
[08:26] with the map, but we find the same concept here. On the tech shift part, we have two worlds with a neat integration tax in between. The analytics and AI world, the applicative world. The analytics world, our data super markets, runs on the lake house. We have
[08:43] a classic architecture, a medallion architecture. It runs with DBT. And uh we manage all the dashboard and the genie space in this analytics and AI world. But we have also developed some internal application
[08:59] uh since the past 10 years. So, each time we develop a new application, we need to to pay this integration tax. Same for the scalability that was not great. The lake house became the source of truth, but applications remain disconnected, internal applications.
[09:15] Integration tax on every new project to pay. Uh dupli- we had to duplicate the data from the applications to our bronze layer. The scaling use the scaling uh use cases increased the complexity, and clearly it was the architecture that became a
[09:31] bottleneck. We were starting to ingest a lot of new applications. Every new use case delivered value, but another tax to pay. So, and I'm French, and our national game after soccer is how to pay less tax.
[09:46] So, I tackled it. All right. So, our answer to it, it's one platform. So, we don't ship dashboards anymore. We ship a triptych. As I said, first for the data, we have the data supermarket. For the software part, we have the DNA
[10:02] port, and we'll have a demo right now. And then on the AI part, the Hafnia MCP. The DNA port run on a Nuxt and Vue front end. We have a fast API back end for the API part. And it runs on Lakebase.
[10:17] The Hafnia MCP, it's deployed through a Databricks app. And uh it's directly uh hosted in Databricks, so we don't leave the platform. So, we start treating them as three views of the same product. Uh we have hired software and AI engineer uh to
[10:34] extend the data team capability. So, one platform, not three stacks. Uh the data for decision, the software for actions, and the AI for conversation and automation. What's extremely interesting here, it's
[10:50] the Unity Catalogs help us to have the same governance, same identity model, the same ownership per squad for the same business outcome value. All right. So, the DNA port powered by Lakebase.
[11:07] This is the first customer. The first customer was to enable self-service analytics to shop the data supermarket and break silos. And this is how we answer to this uh question and problematic. The DNA port, it's our one-stop shop for
[11:24] data product product across all domains. So, I will do a quick demo on that. All the data you will see in the demo are not in the lakehouse. All the data you will see from those dashboards are coming from Lakebase.
[11:42] Lakebase, if you if you don't know it, I prefer to to this. It's uh the the Postgres database service from Databricks. All right, so let's play this demo. So, this is the DNA part. This is how it looks. You will easily recognize the three domain, technical, commercial,
[11:58] financial, and our AI agent in the middle. Very simple way to navigate. Every user have access to every layers. Here, let's drill down to the financial doc. We are able to display financial data through a view uh dashboard. So, we
[12:14] developed it ourselves. Of course, with uh AI agent. And this dashboard is quite interesting. It's the peer comparison. How do we do, Hafnia, compared to our peers on the market? Okay, it's a lot of data. That's pretty interesting. The CFO loves it, and the finance manager as well.
[12:31] But, sometimes it's a lot to process, and you don't want to think so much and read all the slides. So, here we have directly integrated AI within the platform. So, we have a small uh chatbot that will tell us It's not even a chatbot. It's just a window that will tell us
[12:46] uh what do we see on the screen. If we don't want to have to read uh uh in details. All right. So, this is how we simply integrate AI with application and uh analytics. This is extremely simple use case. We will go only through simple use cases today.
[13:03] So, this is our fleet, the war fleet. So, the map of uh where each vessel is. And now, I would like just to show that in this uh platform, we have also integrated some Power BI dashboards. Uh in the past years, we've been developing Power BI dashboards, and we
[13:18] don't want to lose them. We don't want to be able to still use them, of course. So, they are available in that platform. All right. So, now we will go through a dashboard that we have also developed ourselves. So we have dashboard on data data bricks AI BI, on Power BI, and also
[13:34] built-in dashboard directly from the platform. So here this is the clean product clean petroleum product and the dirty petroleum product dashboard. Uh so we are able to navigate the data. It's quite interesting. Uh
[13:49] Yeah. So now it's the next step. We will go towards the world of applications. Here we are still at the end of the food chain. Basically, we are developing a dashboard, but we can't interact much beside just interpreting the data. So here we will uh
[14:05] I will pause here. So we have the war fleet. Like this it seems just a table, quite simple. But behind the scene, we've been working hard on uh different data sources to collect a very rich data set from all the vessels, all the maritime vessels in
[14:21] the world. So here we have a data set of 8,000 vessels, and it's coming from five different sources. So it's a very rich data set. Classic data engineering stuff. Usually, in the previous world, we were only able to interact with those data.
[14:37] But what if on top of it, I want to enrich this data set without having to leave this platform, add uh or update the data in another app. I want to be able to interact with my data directly from the DNA port.
[14:52] So here we will enrich the flag of the vessels uh here Andrea and set it to Singapore. We did not just read in the DNA port, we just write also. So we did a write operation, and this is extremely simple example. We have much
[15:07] more complicated applications, but I want to keep it simple today uh for everybody to be able to uh get some idea and some inspiration on what's possible. We basically developed a full CRM uh tailor-made for our commercial department, for the pool managers.
[15:26] All right. So, what did we see just now? This is the round trip loop. So, application and analytics are on the same governance plane. So, the step number one, the DBT runs. Our gold layer is refreshed. We have fresh data.
[15:41] Great. This is the classic data engineering part. But, the second step, what's great? We have a an automatic synchronization with Lakehouse tables. The data we have in the gold layer in the lakehouse are now available as a referential in Lakehouse.
[15:58] We will use the referential directly in the DNA port. So, here we can interact fast with the data. It is an OLTP system, which is what we want. The user can interact with the data through the API, and we can write back to Lakehouse. That's what we did when we've updated the flag
[16:15] of the vessel. And then, again, this data from Lakehouse is now available for the SQL warehouse uh for the silver layer. And this is quite great. Basically, with the SQL warehouse serverless compute, we can get
[16:30] access to the data that are in Lakehouse. And this is how we have uh we stop paying the integration tax. So, this is the unlock. No data ever leaves Databricks. The operational right becomes the analytical input without ever leaving
[16:47] the platform. Again, no integration tax to pay. From analytics to application. Again, Lakehouse sync Lakehouse sync, sorry.
[17:02] It's the local no-code uh feature. So, we don't have to update the uh lake uh the the the DNA port. It is automatic. So, the data, when they are refreshed in the DNA port, they are also few seconds later refreshed in the Lakehouse application in the DNA port.
[17:18] This is what the business sees. Numbers from lakebase and served on a next application. The second trip, user actions records are synced back to the lakehouse. So now, this flag is available for data
[17:34] enrichment in our lakehouse. On top of it, as I said, we can have access uh to the lakebase database with SQL server warehouse serverless, which is great. And we have a good example on the left, a very simple one. We can select from the lakebase database, the
[17:49] port the world fleet data set, and join it with the gold uh layer on the core vessel and core customer. On top of it, we can even use uh AI functions. Here, in this example, it's an AI query running on Cloud Sonnet, and you can submit the prompt that you want.
[18:06] So this is a a great way to enrich your data. And then, once the the gold layer is ready, we can uh serve it on the AI BI dashboard, Genie Spaces, or River City L.
[18:24] This is the second lakebase unlock. Unity Catalog doesn't stop at Delta. It reaches into Postgres as well. So OAuth is integrated with Unity Catalog. This is the same row policy that governs a Delta read, that governs a Postgres read. Same identity, same audit trail. And this is
[18:41] big for us. Um we are, since last year, on the New York Stock Exchange, and we need to really be careful and be very uh precious with our logs and who access the platform. And the Unity Catalog is is helping us a lot.
[19:01] Databricks asset bundle. So that's how we deploy our artifacts into the different environments. So uh the Databricks asset bundle is configured with the databricks.yaml. The CI/CD run on GitHub. And then we deploy the schema migration with Alembic.
[19:17] The schema migration on Lake Base. Lake Flow Sync configuration is also embedded in the Databricks asset bundle. And also the Databricks app that hosts the MCP is also in there. So we have a big monolith application.
[19:32] Few years ago, I will not be in favor of doing that, but now with AI, you will see why it's maybe a good thing. Three things we love and three things we love to see. So Lake Base, what works? And the
[19:49] integration tax we don't have to pay anymore. Production right latency. So we started a year ago developing this platform and at the beginning we were using Lake Base a Delta table as a data source for interaction with
[20:04] our applications and it was a nightmare. We had concurrency issue. It was slow. Well, it's anyway not a good practice to build an app on top of an OLAP data load. We don't have any custom code to
[20:21] maintain for the Lake Base Sync and this is a great thing for us as well. And the Databricks asset bundles deploys everything as an atomic operation. Our wish list for Databricks, I don't know if we have some Databricks product team from Lake Base here, but
[20:38] anyway, since yesterday this slide is a bit expired, but because some of my wishes has been have been granted. So the data API on the standard tier. So we use the standard tier. We don't use auto scale. But now the data API the auto scale sorry will be mandatory from
[20:54] next month. So we will have to move there. Native Lake Base migration CLI. So we are data engineers and to maintain the code the deployment code the migrations through Alembic is not something native for us.
[21:10] The observability it's now in preview, so I guess this one is also is also granted. But overall, Lake Base dramatically reduced the operational integration complexity for our internal applications.
[21:25] So everything together, Lake Base, the Lake House, one DBT code base with the MCP, plus the coding AI agent on top of it, we can scale extremely fast when we want to deploy a new feature in the data engineering part in the Lake
[21:40] House and to build a dashboard on top of it. It's extremely fast to develop from A to Z a use case. All right, our second customer. Our second customer in the supermarket.
[21:56] Hafnia AI assistant is powered by Genie. So same, let's go through a quick demo. But before just to remind you, the first customer was a self-service analytics customer. Now let's tackle the AI assistant.
[22:15] Same supermarket, three customer, and the same govern answer at the end. All right. So we are back to the supermarket. So now you know the interface and we will go to the commercial uh domain.
[22:30] We will see on the left here and that's very important. If you remember correctly, we've defined the commercial squad in three in few different domains, three different process. The first process is the pool management, second process, position planning, fixture and
[22:46] compliance, and till claims management. And this is how we've organized the taxonomy of the DNA part. This is extremely important for us to think on how we will build the taxonomy of the app. And this is how what we will follow also for the Genies.
[23:08] Okay, so let's open again the CPP DPP transition dashboard and we will add some filters. This kind of dashboard are very important for the chartering team, the people who are doing the strategy of the company. Before this kind of dashboard, they were trying to get the data, but it was more
[23:23] good feeling than data data-based uh decision. Here you can see something quite interesting. Since the Strait of Hormuz blockade, we have reverse on the two shafts. And this is this kind of information
[23:39] that are very important for our businesses. Without this kind of data, they will not be able to notice this kind of change. All right, so we'll add some filters. So, what we want to do with the Genies
[23:55] here, it's for each dashboard, we want to have the corresponding Genie. So, here on the screen, you will see the different Genies that we have integrated directly in the DNA port. Our users, they don't have to leave the platform uh to get access to Genie. So, that's we
[24:11] use the embedding feature of Genie here. Okay. So, let's go to the position planning Genie space. So, this is a Genie space. You know Genie now. Uh I think it's quite popular for this summit. So, what's I want to show it's this
[24:28] thing. We have the commercial position planning that that has the same taxonomy than the menu on the left. This is how we've organized the Genies. Not by department. We didn't do the same mistake, but we organized the Genies per business process.
[24:49] All right, so after we will see how the Genie is configured. So, we follow the best practice. We develop those Genies with the Databricks Dev Kit. Um so, yeah. Here you can see the CPP DPP transition metric views that correspond to the dashboard we were uh, analyzing earlier. So, we have we try to
[25:06] have one metric view per dashboard. The business will trust Genie only if they can ground the data, only if they can verify. Once the trust is there, then it's much easier to make make them use the AI tools. But, without the trust, you will do Genies, but
[25:22] it will be difficult to have our user trusting it. All right. So, here is the classic interface. It's just the embedding Genie, and we ask the question about the CPP and DPP transition uh, since 2026. All right. So, here we have a classic operation on the Genie
[25:40] space. And I will pose to a very simple details that has its importance. All right. So, let's look at the query. The query we see a function that we are not really used to see.
[25:55] Uh, here we use the measure function name. So, what is this? Before I start, uh, plausible versus correct. This is the demo effect.
[26:11] When we started to show Genies, it was amazing. We had the wow effect. It's magic. AI can answer to my data. For many years, we shaped data warehouses for human. We now have customer that can read barcode. The pace of data access will be
[26:29] completely different, and the number of combination possible also different with the AI. So, the lakehouse was great for dashboards and and Excel analysis and few other use case. Our scheme schemas were tuned for SQL query and private table.
[26:45] Uh, but now we have Genie. So, we've started to plug Genie on Delta table. The demo looked great until you read the numbers. It was rubbish. Possible shape and made up content.
[27:01] AI need labels, semantics, and govern measure. And the govern measure part is extremely important. So, what was our answer to it? And the fix was not a better point. It was a massive data model redesign towards metric views.
[27:19] So, we now have 160 plus metric view in prod, and one metric view is part of one domain. It is owned by one squad. It's part of one Genie space in the DNA port menu. So, once a metric view exist, the Genie
[27:35] tool on top is trivial. Trust me, the whole sweating part is on building those data model. That's how a metric view looks in DBT. I don't know among you if us some of you are using
[27:50] DBT, but this is basically the same query syntax than creating a a metric view on Databricks. All right. So, here the metric view are very enriched with comments. But, what I would like to show to you
[28:05] it's on the bottom part. The measure are very well defined. And once Genie queries a metric view, it cannot invent new measures. And this is very important. That's how we want to have govern measures. All right. So, stocking the shelves.
[28:21] So, of course, we are a team of 10 people and we did not handwrite 160 metric views. Opus and Codex generated them, and us, the human, we decided the taxonomy. So, as an input, we have our DBT project
[28:37] with all our models. So, it has all the context of the supermarket, and this is extremely important to give access to this context to AI to generate very rich metric view without having to start from scratch the explanation about the the purpose of this metric view.
[28:53] We use Databricks Dev Kit and Databricks MCP. The Dev Kit for the best practices on building metric view engine spaces and the MCP to get a sample of data which really helps to enhance the quality uh when Opus or Codex will
[29:08] generate our our metric views. As I said already, we follow the DNA port navigation menu. So, it makes it easy. We know what metric view to build. One metric view equals one dashboard. Then for the engine, we use the AI to
[29:23] generate the code. We use Databricks Asset Bundler to deploy everything and GitHub as a repo and CI/CD tool. So, in output, we have now 160 metric views splitted among the different domain squads.
[29:43] All right. Our third customer, the Hafnia MCP. The Hafnia MCP is powered by Databricks apps and calling the Marvis uh sorry, calling the Genie spaces. But before I go to the Hafnia MCP, I would like to share a small story about our AI identity crisis.
[30:01] Most important architectural decision we made. What we choose not to build as a data engineering team. We started with an agent on agent bricks. It worked. But as a data engineering team, who should we be? A data team is supposed to deliver data.
[30:19] Data engineers were maintaining a thin agent wrapper on top of our own data and we were not feeling like it was our job to do that. So, we drew a line. Data engineers ship tools and app and AI engineers ship supermarkets AI shoppers.
[30:40] So, the technical vision of it, the data team deliver the MCP, not the agents. Within the MCP, we have tools and the Hafnia MCP is the durable data contract. It is owned by the data engineers in the team. And then, the Marvis agent that lives in the DNA part that we will see right after is managed
[30:56] by our AI and application engineer. All right. So, let's do a quick uh demo also on the Marvis AI agent within the DNA part and our Hafnia MCP within VS code, within GitHub CLI.
[31:17] Okay. So, let's go back to the commercial A. And open Let's open the same dashboard. So, I will not open a genie directly, but here I will open our agents, Marvis AI. So, we will do the same kind of filtering to find the same data than before.
[31:33] So, we'll filter on 2026. And what's is interesting is for instance, as a user, you maybe don't know which genie to have access to when you have a question. You know it's commercial related, but you don't know exactly what part of the commercial domain it is. So, this is where it's extremely
[31:49] interesting. Here, the Hafnia MCP, we again follow the same taxonomy than the menu in the DNA part. Then it it makes sense for our user to navigate in the tools. We will focus for this question on the commercial genies and it will select preselect the list of tools related to
[32:06] commercial domain. All right. So, I will ask the exact same question than before and we're supposed to find the same answer. While I I run the answer, I will filter I run the question, sorry. I will filter uh it on uh
[32:23] including the sanction vessels. So, remember this number here. We have 278 272 and 218 CPP and DPP vessels. So, we want to be able to validate this data with our uh with our MCP.
[32:38] Again, we need to build trust with our user. It was very difficult at the beginning to deploy AI when they were not able to verify the data. So, here we find the same numbers than what we saw in the dashboard. So, we can go step by step, slowly but
[32:54] surely. So, we've deployed the genies based on the metric view dashboard by dashboard. All right. So, another use case that is much more complicated. Here you saw that in the previous graph, the Strait of Hormuz blockade had an
[33:10] impact on our business. We saw in this graph because it was right under our eyes, but we won't be able to understand the whole impact of it. So, we can ask a question and among all the commercial genies, try to have a summary on what was the
[33:25] impact in the claims manner in the claims process, for the pool process, etc. etc. So, let's run this uh question to our agents. Our agent is connected to the Hafnia MCP. With the same list of tool.
[33:43] Now, the question is quite compre- quite comprehensive and you can see that we will spin off different genies in parallel, different MCP tools. And under each MCP, we have a genie space. Of course, I have accelerated the video.
[33:59] It's not that fast. It takes 3 to 4 minutes to generate. But, since we have grounded each genie, we are sure that those data are correct. And this is something this is something extremely important. No rubbish data.
[34:14] All right. So, now our agents running on GPT 5.5 collect back all the answers from our different genies and is able to write a very comprehensive summary on the global impact of the Hormuz blockade
[34:30] for our business. This kind of dashboard will take hours or maybe even days without so many informations. And that was uh really a game-changer for our business. And this is something that is well uh well appreciated.
[34:45] All right. So, the second use case. We stayed in our own ecosystem, the DNA port, the lakehouse, Databricks, basically. Now, what if we want to have access to the Hafnia MCP outside uh our own
[35:01] ecosystem? From Claude code, from Claude desktop, from ChatGPT, or from here uh GitHub CLI. So, I'm logged in with my own account. And I've also added the MCP, the Hafnia MCP, as a as a new MCP in my session.
[35:17] Initially, I was working on Codex, and I will ask the MCP to switch to Claude just to show you that we are completely agnostic to the model we are using to query our data. All right. So, let's switch to Claude Opus 4.7.
[35:33] It's anyway a bit better to do this kind of analysis with a data analyst role. Okay. So, let's just send the same question than the question we first sent in the DNA port. So, of course, we want
[35:48] to find the same answer. That's the whole purpose. So, you will see that it will similarly, same way, trigger uh call to the tool, the query position planning MCP. So, that's great. Uh this is how the MCP makes the data available from anywhere.
[36:05] Then we can add the follow-up question, among those vessels, which one are sanctioned? All right. So, now we will have a a detailed answer. Very easy to use.
[36:20] Again, it can be used from Claude desktop, from Copilot, uh Microsoft. So, it works. Now, I can ask, "Okay, summarize this data uh in an email format." It will not use an MCP, of course. It will use the global knowledge of Claude, and it will uh write a
[36:36] comprehensive email. And this is where we'll have all the power of MCPs. So let's say I have also connected Work IQ MCP, which is the Microsoft uh MCP. It's It has access to Hotmail, to Teams, to all the Microsoft suite.
[36:51] And we could ask uh here uh to Claude to generate an email in my draft based on the data coming from Hafnia MCP. So we can have those MCP working together to do like beautiful automation, and this is with this
[37:06] solid foundation that we'll be able to build agents at scale. All right. Now you might wonder how AI shops in the supermarket. This is the picture. So first on the left, we have
[37:22] the DNA port agents, and it's a most amazing use case for Lake Base. Lake Base, it's a fast system. It's an OLTP system. It will not be possible with the Lakehouse Delta table to have uh this kind of uh frequency. It will store the conversation memory, and also the agent
[37:39] state management. We also have access to the MCP with the external AI tool tool as I just said. And the MCP will use the on behalf of of identity model. Uh an MCP will be plugged with tools. We
[37:55] have 18 Genie space in uh our environment. We have AI knowledge with for the policies, the SOPs, and all the reports. It's a lot. And we have also the ontology. It's not the ontology that has been introduced yesterday, the Genie ontology. It's the Neo4j ontology. So it's how to navigate
[38:13] the connection bit uh bit among our objects. Genie is good for BI kind of question. Select, sum, uh group by, blah blah blah. Ontology is great to navigate the connections.
[38:28] Uh each genie is plugged on metric view, this is the rule. No genie plugged on Delta table. And a metric view is of course plugged on a Delta uh table for the data retrieval. This is the authorization chain here um in the bottom with uh Unity Catalog.
[38:45] Unity Catalog enforces the policy at every hop. Uh a finance user and a commercial user can ask the same question, but they won't have access to the same tool. So, the answer might be different. But this is done by purpose. This is our governance model.
[39:02] So, from prototype to software life cycle, so again, if I want to give some feedback to the genie uh product team, genie space and metric view can be managed as code, but it remains partially partially brittle. Uh if you want to maintain a genie space today as a code, it's a big JSON file,
[39:18] basically. So, we would like to have something a bit more maintainable and uh uh engineering soft uh software asset kind of thing. The environment consistency still requires manual governance. It takes some time to ensure we don't drift.
[39:34] And the benchmark and LLM judge evaluation, it it's it's a great enablement with MLflow. We don't have to manage this. And as data guy, we don't want to have to manage the MLflow things. Uh what we need, it's a declarative declarative genie as code. Uh this is really something important.
[39:50] And to be able to deploy genie uh in the dub. Today, we deploy them through a custom Python script. All right. Then, Monday morning, what to take back?
[40:09] You don't need a 10-person team to start. Uh one aisle, one shelf, and one self-service lane. Start with one domain, one process, own it, build the corresponding metric views. Build the genie on top of it, connect it to a genie to a MCP, and then it's going
[40:26] to be already a great achievement. And then you can scale from here. Two lessons. The AI isn't the hard part, the data model is. And Lake Base is what made operational and analytical the same store on
[40:43] Databricks. For the data team, start by delivering tools. Do not focus on agents yet. Harness your tool first. Ship an MCP that anyone, human, AI, partners company can plug into.
[40:58] Once you control the tools and the underlying data are managed, you can go for agentic applications. Thank you very much. Thanks for your time. Thanks for coming.
[41:15] And if you remember one sentence from this whole talk, remember this one. The AI isn't the hard part, the data model is. Thank you very much.
Learn more about the Databricks Data and AI platform.
The information provided herein is for general informational purposes only and may not reflect the most current product capabilities or configurations.