Skip to main content

One Platform to Replace Them All: Fonterra's Databricks Migration and Genie Strategy

Summary

  • Fonterra, the world's largest dairy exporter, migrated from legacy SSIS and SQL Server infrastructure—where 50% of engineering capacity was consumed by maintenance—to a unified Databricks platform, cutting BAU overhead to 20% and enabling self-sufficient teams.
  • The team deployed two production Genie spaces in just 3 weeks by starting with targeted use cases, co-designing with business stakeholders, and iterating rapidly, with Genie surfacing insights not initially obvious even to the engineers who built the underlying data models.
  • Fonterra's roadmap combines Genie with supervisor agents, AI/BI dashboards, Databricks Apps, Delta Sharing, and MLflow to create a single unified platform for all data discovery, eliminating the report sprawl that characterized the legacy environment.

One Platform to Replace Them All: Fonterra's Databricks Migration and Genie Strategy

Watch: One Platform to Replace Them All: Fonterra's Databricks Migration and Genie Strategy
Fonterra, the world's largest dairy exporter, operated Market Analytics System (MAS) on legacy on-premises infrastructure: SSIS, SQL Server, and 50% of engineering capacity spent on maintenance. Complex data pipelines, manual integrations, and unmanaged ML models on personal laptops blocked every new initiative.
Learn how Fonterra rebuilt its entire data platform on Databricks using a purposeful, iterative strategy. Discover how Genie spaces turned data engineers into empowered curators, enabling business users to ask questions in natural language and get instant insights without writing SQL. See how the team deployed two production Genie spaces in just 3 weeks by starting small, co-designing with business stakeholders, and leveraging learnings from peers across the organization. Explore the roadmap for combining Genie with supervisor agents, AI/BI dashboards, Databricks apps, and Delta Sharing to create one unified platform for all data access.
🤝

Chapters

FAQs

Why did Fonterra choose to rebuild rather than lift-and-shift to Databricks?

Fonterra chose a purposeful rebuild strategy because a lift-and-shift of their legacy SSIS and SQL Server pipelines would have replicated existing technical debt rather than solving the underlying maintenance burden. By rebuilding on Databricks, the team reduced BAU overhead from 50% to 20% and established a modern, extensible foundation capable of supporting AI workloads.

How did Fonterra deploy two Genie spaces in 3 weeks?

Fonterra succeeded by starting with clearly scoped use cases—milk supply forecasting and price forecast analysis—co-designing each space directly with the business stakeholders who would use it, and iterating quickly based on feedback. Learnings from the first Genie space were applied directly into the second, creating a repeatable deployment pattern the team could apply across further use cases.

What insights did Genie surface that data engineers had missed?

In this video, Fonterra describes cases where Genie identified patterns and correlations that were not initially obvious even to the engineers who designed the underlying data models. This reinforced the value of enabling business users to query data directly in natural language rather than filtering all analytical questions through data analyst intermediaries.

What is Fonterra's roadmap for its unified Databricks platform?

Fonterra plans to combine Genie with supervisor agents for more complex analytical workflows and build Databricks Apps for user-facing interfaces. The roadmap also includes AI/BI dashboards, Lakebase for operational data workloads, Delta Sharing for external data access, and MLflow to scale ML model demand as data science initiatives grow across the ingredients and trading business.

Full transcript

[00:07] Good afternoon all. Good afternoon. Um so for those you're wondering about the one platform to replace them all, I do come all the way from New Zealand um and I was told in order to keep my residency, I have to make one Lord of the Rings reference. I thought let's sneak it right into the title, you know? One platform to rule, I mean, replace them all, right?
[00:24] Jokes aside though, um what I'm really going to be talking about today is how uh we essentially went from a bunch of systems including a lot of on-premise systems um and took that all into Databricks. I'm going to be talking about how we're getting value out of that right now, the things we're doing right now, the wins we're getting. And
[00:40] then lastly, I'm going to be talking about um essentially what are we doing next? What are we looking at, right? What What's coming up? So I'm going to try and take us on that whole journey in these next 40 minutes. Um hopefully try and get that all in. Before I get to that, let me introduce myself. Um my name's Kurt Richter. I'm
[00:55] the data engineering manager at Fonterra. Definitely different hairstyle back then. Um but I've been in the role for about a year and a half now. Um I specifically work in ingredients and trading data and analytics. Um I've been fortunate fortunate enough to be in the data industry for nearly 20
[01:11] years. I've been a BI developer, I've been a data analyst, I've been a architect and pretty much everything in between including recently engineer, right? Now a little bit about Fonterra and specifically what this talk is, ingredients trading or what this talk
[01:27] is. A little bit about who we are. So Fonterra is a dairy cooperative. It's owned and supplied by thousands of farming families across New Zealand and it's really such an inspirational story. These are farmers that had a can-do attitude and a will to work together to become a global leader. They knew that
[01:43] they could do a lot better if they worked together to become a global leader. When I talk about a global leader, this is what I'm talking about. As of FY25, we had $26 billion revenue and we had nearly 16,000 employees
[01:59] all across the world. We sold to markets from everywhere from Europe all the way to America and everywhere in between. We really are today a global leader in the in the in the dairy space. Honing down a little bit, like I said, I'm in ingredients and trading.
[02:15] Um going to really try not oversimplify this, but essentially we're the value engine of Fonterra. Um our job is really how do we get maximum value out of every single drop of milk, right? Um and we do this in a couple of ways. You know, we look at uh market signals,
[02:30] customer demand, pricing, risk insights. We look at what what products we can make, what streams we can make, what channels there are. Um we look at where we're going to be selling to. So, we get a whole bunch of information to try and decide what we can do with every single drop to get the maximum value out of it,
[02:46] right? So, that's the the ingredients and trading. Um again, like I said, I report into the data analytics arm of that that department. So, we really look at all the data for all these teams. Cool. So, like I said, I'm going to take you on a little bit of a journey. Um try and take you where we came from.
[03:03] Um and that was 10 years ago, just over 10 years ago, 2015, there was a system called market analytics system, or everybody knows it as MAS. It was a platform that was built to meet all those functions. We'd get all the data in there, but it was made up of
[03:18] four kind of key components. Snaffler, custom bespoke software that they built. It had AWS Lambda, a whole bunch a whole bunch of state functions. And it was really designed to go and fetch data, whether was web scraping, APIs, databases, file drops. This the
[03:34] its whole role was to go and fetch the data, right? It would then hand the data over to SSIS on prem. Um SSIS, you know, kind of predates your your medallion architecture in some ways. So, this is just a big massive engine of hugely uh huge amounts of
[03:49] transformations, right? Um then dropped it off again uh in the basement somewhere. There was a a a server sitting with SQL Server 2019. And we had 300 tables plus minus um sitting on that server. Now, I know you 300's not a lot, right? But again, that
[04:06] goes back to SSIS. SSIS was this beast of a a thing. Some of those data flows you'd open, you'd go grab coffee, and come back, and it still wouldn't be open, right? So, um a lot of that transformation logic was happening there. So, you didn't have multiple tables and things like that. Um and then last at the end was just
[04:21] Power BI Fabric, which is obviously the reporting the end layer. We built some apps on that. Adjacent to this So, that So, that was the MAS system. That That That's what we called MAS. Those Those things together made MAS. But adjacent to this, we had a data science team who were building ML models
[04:37] largely on their laptops, on their own environments, trying to manage them through that. Um we had an application team building lightweight Again, it was taking this data, going to um
[04:52] uh to Postgres, and then building these applications on top of that. And thirdly, we had operational systems actually using this data. So, this was a critical system, right? It was really important, not only from a reporting So, it wasn't just analytical. We actually were feeding
[05:07] um a data to actual operational systems. So, it was really important, really, you know, uh key to the business. It being an on-prem system system and using a bunch of legacy tech, it had a few challenges.
[05:25] Disjointed daisy chain architecture, multiple points of failure, right? Like I said, you had all these different tools all trying to work together, and some of them didn't work natively together. So, you're trying to force it to work natively together. So, when something broke, you're literally going back down that whole daisy chain to try and figure out where it broke, what's wrong with it, right?
[05:41] BAU overhead was roughly 50%. I was shocked yesterday if if any of you caught the keynote where they said most companies are still battling with this. Um I think they said a survey was done where it was something like 53% is is the average uh BIU overhead. I shock will sing it at over 50%.
[05:58] Scaling limitations. My first week in the company, I remember getting on a call with Microsoft. I didn't know what was going on. I didn't know anything about the systems. I didn't even know what SSIS was. Um but I got a call on uh with Microsoft because one of our key jobs was failing. Um we couldn't really throw any more
[06:14] resources at the server. It was already a beefy server. Um we're going to Microsoft saying, "Can you help us? We don't know what else to optimize here." Um and eventually we actually just had to kind of move other jobs out of the way and say, "We can't run you now. We'll run you at a different time, right?" So, we definitely were starting to hit those massive scaling limitations.
[06:30] Oh, this is one that's haunted my whole career. Server maintenance, backups, patching. You know, every time you've got on-prem servers, you've got to go through this. Once a month, you've got to have a engineer on call cuz you know something's going to break. So, um definitely not something you want in today's world.
[06:47] Little governance or data quality. And that's just native to everything being inside of um uh SSIS, right? How How do you govern that? There's so much logic. There's these massive packages. How do you actually govern that, right? So, it was just a native thing and across all that stack, how do you govern govern
[07:02] across all of it? So, there's little data quality. There's little There's no checking. There's no rules or anything like that. Um hard-coded configs, right? This is another thing that as we went through migration, how many times we had to look back and go, "Oh, man. Where did they get this number from? Like, did they just imagine it one
[07:17] day? Did somebody in the business, you know?" Difficult to onboard and train new engineers. Uh before I started, there was two new engineers that started before me. And by the time I got there, they'd been there for a couple of months. They were still struggling to understand these systems, right? And they're really good. They're really smart engineers. Just because of how
[07:33] much complexity was in this whole system. Very low documentation. This one was very important, right? It also was why we had difficulty to onboard, but this was built by an external company. They came in and built the system.
[07:49] Little documentation, so it created critical dependency. Whenever something was wrong, we had to go to them, right? We had to go, "Oh, please, can you help us? Please, can you help us?" right? Again, definitely not something I subscribe to. This one was big. Using tools that were outside of support, right? Um the SSIS version we were on was actually busy
[08:05] being sunset. So, whether we had moved to another platform or not, we would have had to do something, right? We had to actually migrate somehow. And for those of you that have worked in migration projects, they're never as easy as they sound, right? So, "Oh, just change your version. No, it's easy." Doesn't work like that.
[08:21] Like I said, code and ML models living on people's laptops, completely unmanaged, just every data scientist building his own model in isolation on his laptop, um trying to see what he can scrape out of the resources on there, you know? Um And there were many more challenges, but these were the main ones. These were the main ones that
[08:37] gave us pause and went, "Actually, you know what? We need to do something," right? So, I'm going to talk about where we are and kind of what we did, right? So, so we acknowledged those challenges.
[08:53] Early 2024, this was before I was even there. They got a company to come in and go, "Please, can you do an assessment of the system? Please, can you give us a third-party view? Tell us what you think about it," right? No surprise, they looked at all those challenges and the many more, and they said, "You guys need to get off the system," right? So, so their recommendation was get off the system.
[09:10] So, late 2024, they started a project. They said, "Cool, let's start scoping this out. Let's see what we can do. How can we get off this?" Um and they gave us essentially options. They gave us couple of recommendations of platforms to go to. One of them was Databricks, um
[09:25] at the time, our central IT had already started moving to Databricks, and they were already starting to get a lot of benefit, right? So, from us, it was very easy for us to go, "Well, they're already starting to set guidelines. They're already starting to set guardrails we can leverage, we can use, we can adopt, right? So, it was a very easy decision. Let's go with Databricks.
[09:46] Early 2025, that's when I came on board. The project was officially started. We said, "Yep, we're going to get everything off of Mars. We're going to get rid of all that all those ETLs, all the SSIS, everything that's on premise needs to go. Needs to be done." Middle of this year, we had everything running inside of
[10:02] Databricks. Mars was still around and still being used. We just had a period where we were making sure we were literally parallel running them. We're making sure that whatever we're giving to the users is right. We did a bunch of UAT, but everything was already running in Databricks. This is a cool cool brag.
[10:19] 2 weeks ago, I was on the call with the infrastructure team asking them to go send somebody downstairs to turn off the server. Done. Kill it. It's done, right? Um that that for me is a really impressive thing, right? Like this is a project that came under budget
[10:35] and under time. And I know you say, "Well, that's quite a long time. A year and a year and a bit to get it off. That That's a reasonably long time, right?" We made the choice right at the beginning that this was not going to be a lift and shift. We decided we're going to rebuild. We're going to challenge assumptions. We're going to ask
[10:51] questions of this data. We're going to look at the way we source the data. Where we were getting a web scrape, could we get an API? Where we were getting an API, could we get delta sharing? You know, we we challenged all the citations. We went through all those sources. Said, "Do we even need this data? Is the business even using it,
[11:07] right?" So, we built it from the ground up. And our core goal was to have everything that Mars was doing inside of Databricks, right? And I can say from a at least the top three, we've completely successfully done that, right? All of that stuff is now running inside Databricks. So, sourcing the
[11:22] data, moving the data through our medallion architecture, getting data to serving layer, it's all done in Databricks, right? That's what it looks like now. And there's other tools that are always on the hinges of these things that you don't talk about, right? Like servers and stuff like that. But everything has
[11:39] just been replaced by Databricks. But what does that mean? Like, yeah, cool. We got the data on Databricks. What does that mean? What does What does it mean to the business? What What What did they get from that? Well, one, our BAU and and bug fixes is down to less than 20%. That was me being
[11:55] I'd rather say 20 cuz there's that one sprint that we get a little bit more, but it's actually less than 10%. And we have a lot less resources managing it. We no longer require external parties. Like, how great is that? We're self-sufficient. We don't need to be going, "Oh, it's broken. Please, can you come help us, right?" Everything is
[12:10] managed by my team. We don't need external parties to be helping us manage our stuff, right? Obviously, that also leads to reduced cost. Data is easier to find. And more importantly, you can go all the way and track its lineage and see where it comes from, right?
[12:27] Like, previously, we wouldn't be able to I wouldn't be able to tell you where the silo came from. My source table, wouldn't have a clue which source it came from. Unless I go through everything, try and read the XML, try and try and do a bunch of work just to find out where it came from. Now we've got that. Comes natively. We don't have to do anything special. It's just out of the box.
[12:42] Like I mentioned, we had all these integration points. We don't have to two cuz Databricks does everything, right? It literally does everything. We're doing everything inside Databricks. Another big one, easy deployments and releases. We use data that should be automation bundles, but data automation bundles, right? DABs. Makes it so easy
[12:58] to deploy. We We came from deploying maybe once every 2 weeks, maybe once a month because you had to had so many interdependencies. We do deployments all the time now. It's It's just second nature. We're We're just sort of having full CI/CD pipelines where where they just run automatically, right? One of the the goals we're getting to.
[13:18] Project delivery time reduced. This was a big one. We actually had one of the consumers come up to us um in our town hall and actually talk about how a couple of months ago, whenever they put in a request, it would take weeks, months. This is a a non-technical he he doesn't know about
[13:33] data bricks, doesn't know he's just talking about how he gets his enhancements, his features. He says he was so impressed that now we're down to days and sometimes hours to have an enhancement or feature, right? Because of simplicity.
[13:50] Delta sharing. Whenever I talk to one of our partners or one of our producers of data, I'm straight away asking, do you have a Delta share? We we deal with a lot of data. Like I said, we had over 130 source profiles where we're sourcing data from. Um as soon as I have a conversation, my first thing is, do you have Delta share? Or I guess open share. Um
[14:05] but yeah, so it's it's it's really one of the things that we can do now easily to get data. We don't even have to build half the time. It's just Oh, yeah, just share it with us and then we'll get it through the rest of the layers, right? Most important, and this is the one that I I put a lot of value in. We had
[14:20] engineers who are excited and engaged, right? They weren't working on legacy tools. They weren't working on legacy platforms. They're building stuff. They were having fun. They're seeing how far they can push it, right? Um they were completely engaged again from, you know, them coming on and trying to learn these old systems to
[14:36] come and play on Databricks rather, right? Um and that's the one that I'm most I put the most value in. It's just seeing that joy. So, that's cool. We're winning. Databricks is great. We got all the data there. Doing an awesome job. Very happy. What now? I mean, we're used to doing all BAU
[14:53] and bug fix. What do we spend our time on now? At this point we said, okay, let's stop. Let's go look at what the rest of Fonterra is doing, right? So, let's go have a quick peek at what all the other departments are doing, right? What what are they doing in in central? What are they doing in finance and and all these other departments?
[15:09] I'm going to play you a quick video just to give you an example of some of the stuff they're doing. Fonterra has invested in strengthening how SharePoint, Teams, and Viva Engage sites are managed to ensure we design fit-for-purpose sites, reduce
[15:26] duplication, improve findability, and better manage risk. This also supports how we continue to embed AI into the way we work. But finding the information you need to manage your SharePoint sites hasn't always been easy.
[15:42] You often need to search across reports, guidance, and tools just to understand what action to take and where to start. The SharePoint Insights agent in Co-op GPT changes that. It gives you one central place to ask questions and get fast, clear answers,
[15:59] along with recommended actions to help you manage your sites effectively. The agent lives inside Co-op GPT, Fonterra's secure AI platform. It brings together SharePoint Insights data so you can quickly access the information you need when you need it.
[16:15] Getting started is easy. Open Co-op GPT in your browser, select the @SharePointInsights agent, and ask your question. Simple. Now, let's take a look at a few examples. Here is how it might work for a SharePoint owner wanting to find out
[16:30] what sites they own and the retention period assigned to each of their sites. Knowing their sites' retention matters to ensure their content is kept and disposed of correctly. This prompt will provide them with a list of site names, the site URL for accessing the site directly, and the sites' retention
[16:46] period. They could also take it a step further by asking which of these sites require a second owner to be added to the sites' permissions. Having at least two owners is a minimum requirement across Fonterra's SharePoint sites.
[17:02] Some other questions a user might like to ask are, "Which of my sites are inactive or rarely used?" Or, "Do I own any sites that could be deleted?" The agent doesn't just give you answers. It also suggests actions like adding
[17:17] owners or deleting sites. The SharePoint Insights agent is a fast and reliable starting point. But like any AI tool, it's important to check responses for accuracy, especially before taking action. You can also provide feedback using the
[17:33] thumbs up or down feature to help improve the agent over time. This means less time searching, clearer guidance on what to do next, and a simpler way to manage your SharePoint sites, all in one place.
[17:50] The SharePoint Insights agent is here to help. Open Co-op GPT, select the agent, and start asking your questions today. Cool. So, there's two things I really want to draw your attention in that video. One,
[18:05] Fonterra as a whole is embracing the the AI and giving AI agents to the users, right? You would have seen they they talked about Co-op GPT, and then essentially that's just a kind of marketplace for you to go and grab whatever agents you need, right? So, we're really embracing that. But the second thing is, did you see when all of those things returned, all those prompts
[18:20] returned, it returned Genie space used, right? So, at the back of this there was a Genie space, right? That that's what all that information gets to. So, we're like, cool, man. We want to do that, right? Tell me more about this. So, so we went to ask them, "Well, what are you doing, right?" And this is what they told me that the
[18:35] problem that they're trying to solve is they had a sprawl of um SharePoint sites, and it was growing every single day, and it was growing rapidly, right? It just became a really under-managed swamp of documents, right? Like you had SharePoint sites for everything, right?
[18:51] With this, they now had a way to manage it. Now to take that under-managed swamp and start putting some management layers on it, right? Start getting to what are the sites we own, you know, what can we do? Um you could see trends in site creation, so you could start targeting those things. Um you could manage the health. You could see
[19:07] when last has somebody been to the site, right? Um it also gave a whole bunch of unexpected benefits, which are really good, right? It changed behaviors. Now, instead of people just automatically going "I don't know if I have a site for this. I'm just going to go create a new one." People were going to this and going,
[19:23] "Actually, let let me check. Is there a site for this, right?" You know? Or "Is my site actually being used? Can I just delete this? I don't know who uses documentation." Or "It's actually out of date, right?" It gave people that that ability to manage it, right? Now, from my side, I'm going, "That's
[19:39] one of the cases." We spoke to tons like that, right? So, it wasn't only one. It was just one of them, but we spoke to a lot of guys doing some things. I'm going, "Oh, this Genie space is cool. What can we do with this? I want in." Like, "I want to play with it." So, I went back to ingredients and trading. Um, and we said, "Let's set ourselves
[19:56] some goals of things we want to Genie." So, so we said, "Let's try and solve some problems, right?" So, so we just put some generic ones on a table that we knew we know comes often. Said, "How do we uh locate the right data set, right?" It's always difficult to locate the right data set. It's difficult to write or adapt SQL, Excel, Power BI queries. We've got a
[20:12] bunch of analysts constantly trying to do this, constantly coming to us, constantly asking for help. Exporting results, right? Always got to be exporting results. You you always end up in a Excel or in Power BI or in something, and then you're doing additional analysis on top of it, right? To try and get the data or get the data
[20:29] right. And then you're building charts, right? So, ultimately, this has got to go somewhere. You put it into a chart. You've done all this work. All that great data is ready in Databricks. We did all the work for you, but it's still going out, right? So, we said, "We want to try and solve that problem." Another one, interpret results. How many times are analysts literally looking at
[20:45] a graph and then trying to explain what happened or or or look at tables trying to explain what happened. We want an easy way to be able to do that. That That's a problem we need to solve. And then write commentary, right? Same thing. You want to put it into a board paper. You want to put it into a presentation. You're narrating. You're sitting there
[21:00] going, "This is what I see, right?" Or somebody's going saying it. So, we're like, "Surely, these are easy challenges, right?" We also said some of the things that we want to and in in combination with that, some things we're hoping it'll get to or or it'll be able to achieve, right? Ask questions in natural language. Of
[21:16] course. Naturally. Instantly explore visuals and data, right? This is about getting the people closer to the data, right? So So these were the challenges that we said we we want to see if we can do these things. See key insights highlighted. Review draft commentary. Reuse generated SQL.
[21:32] Fast, thorough, repeatable analysis. Those things we said we want, right? And and based on our our uh talking to other teams, it seemed very doable, right? We said we could do it. So how do we do that? Well, we went to pick the use case. We went to business and we said, "Anybody having these problems? Anybody
[21:48] having any of these? Or anybody want any of this?" Surprisingly, they all did. All of our business units had problems like that. So we went and we picked one. And came out our first use case, milk supply forecast. Obviously, as part of the supply chain, it's very important
[22:04] to know where your product how much product you're getting, where it's coming from. Um we had the data scientists, like I said, that built great models. We had now moved all those models inside of Databricks. They're all sitting there, they're all living there. Great. Awesome. But they still weren't getting to the people that needed them, right? That information still wasn't quite making it
[22:20] the last way there. Um so what we did, we got one business analyst, very technical business analyst. We coached him up a little bit, told him about Genie spaces. And he sat alongside the business and within 2 weeks, one sprint, we had our
[22:36] first Genie space up and running, right? Being used by the business. And we were just we were happy to just play around and just use it as a complete POC. But within 2 weeks, we had our first Genie space, right? What did we do with that? We said,
[22:52] "That's pretty cool, but but again, beginner's luck, right? No way we can repeat that. I was just we just got lucky as easy use case." So we chose another one, price forecast analysis, right? And this is about seeing what the markets are doing, trying to see how we can get the best price for the milk, right? At what regions, right? So, there's a lot of
[23:08] data. Again, a lot of modeling had to be done. Same problems, though. The people that need to use the data were struggling to use it, right? There's heaps of data, they were struggling. Week later, we had our second use case. Genie space. Another Genie space, right?
[23:24] We gave it to the business users. I'm not going to focus too much on those Genie spaces. They're pretty cool. What's more meaningful to me is when I went back to the business, I said, "Cool, you guys have had these Genie spaces now for a little while. Tell me about it, right? Give me some comments." This is what came out.
[23:44] Genie is my new best friend at work. I love that, right? What What more could you ask for than than a comment like that, right? Like, and somebody who resonates with that, I using Genie all the time now. Like, that was great. I was like, "Ah, you you just made my day. You you're just trying to, you know, flatter me." Um It removes friction between curiosity
[23:59] and insight. This one is important. This one is great, right? I went back to the user that said this, right? He emailed me, sent this, he said this, he said "Can you talk to me? What do you mean by that? Like, I I think I kind of get it, but can you maybe just elaborate?" And he said to me,
[24:15] "What often happens, right? Is they'll have a hunch about something, right? They'll go, 'Does this correlate? Does this Does this have any impact? Is there any insight in this?' But because of the amount of time and effort that could have zero value, they wouldn't even go down that path, right? They wouldn't even start to
[24:30] investigate. They would just ignore their curiosity. Or vice versa, they would they would go down that path, and they would go and investigate, and they'll do they'll spend plenty of time and effort trying to validate this, and then it would be worth no value, right? It would have no correlation at all.
[24:46] With Genie, they don't have that problem. It gives them insights. They can get to data really quickly, right? They can be curious. They can go and see, "Should I invest more into this? Do I need to go deep in this, right? At least it gives them the beginning step, right? So so they have better informed whether
[25:01] they should actually spend the time. Next one, with Genie within a few minutes had the data needed, tested visuals, had some generic commentary to review before release. It picked up insights that weren't initially obvious on review of the tables. I want to focus on that last part, right? I'm a firm believer of man in the
[25:17] middle. I'm I I don't think that this is going to replace people. I think this works alongside people, right? Most of the time, a Genie space won't know that market behaviors are happening because something halfway across the world is happening or somebody said something.
[25:33] Genie spaces might not have that insight. Your analysts know that. They'll be able to add to that. But if you treat this as like a co-worker and it's going and finding things that you may not have thought of because it goes so deep into the data, right? It's It's now picking up things that they didn't even think to correlate, right? And change the way
[25:49] they're actually working and thinking about things. And they're starting to pick up more of these correlations, right? Supports ad hoc analysis. Again, like I said, we we trying to give the power to the people, right? We we we want the people using the data, not not always have to come to an engineering team or an visualization team to build.
[26:07] Let them do some analysis, right? Let them play with the data. Prepare sound-based commentary to communicate insights clearly. Like I said, they were previously typing out everything they saw. Now they get it done for them, right? So for them, they can spend a lot more time doing that deep analysis, right? It's freed up a
[26:22] lot of time, right? So now some key learnings and principles that we took through this, right? Target specific use cases. We went for something small. I know me and I know a lot of data engineers who always try and solve world hunger. We
[26:38] try to go for the biggest thing with the most bells and whistles as big as we can get, right? Target the small use cases. Let them become your advocates, right? This is spreading like wildfire now throughout the business because these guys going, "Oh, look what we have. Look what we have, right?
[26:55] Co-design alongside the business. This is really important. As soon as something was wrong, they were like, "Whoa, whoa, whoa, whoa, whoa. That number's wrong, right?" I would have never seen it. I don't know the data as intimately as they do, right? Work with them. Like I said, we had our BA sit with them for those 3 weeks, and he was literally every single time he made a change, he would show it to them
[27:10] and go, "What do you think of this? Is this the type of question you would ask, right? Would you call this something else?" So, it's working with them, right? Not not just running off and building and trying to build the perfect perfect thing, right? This is such a simple concept, but it gets forgotten so often.
[27:26] Get it to user as soon as possible. Allow it to iterate and enhance. I I can't tell you how many projects I've been where they just build and build and build and build and build, and they get to the end, and the user's like, "Well, I don't not quite what I wanted, right?" It's such a simple concept, but it always gets forgotten.
[27:42] One thing that we did is we wanted an user. Even in the beginning, we said to them, "It's going to be terrible. It's going to be horrible. Just try it, and tell us how we can make it better, right?" Get it to them quickly. Don't get them the finished product. Build it with them. Acknowledge that there's potential for
[27:57] for for mistakes, right? Work to minimize them as much as possible. But, with that being said, I think we've created a little bit of a false dichotomy here, right? Um One of the big things that we learned through this migration project is we'd have to go to people and go, "Yeah, that
[28:14] report you've been using for the last 5 years, the numbers have actually been wrong. You know, you've been using the wrong data." But, people have this real image of what I've been using for 5 years is right, even though it's built by humans, we don't understand what happened in the background, it's been right, and AI can be wrong.
[28:29] It's a false dichotomy. They can both be equally wrong. And I had to do that multiple times through this migration project, go, "By the way, we're going to be rolling over next week. Your report will show different numbers. Here's why. Here's the calculation." And actually explaining to them that it has just been plain wrong, right?
[28:47] Leverage some of the great work already by other teams within Frontera. This is one that applies these conferences. This is why these these conferences so powerful, right? You get to see what other people are doing. You get to leverage that. It's very seldom you have to start on a journey in the beginning. For us, it was the internal team. We could go to other departments. We could
[29:03] take their learnings. Um we could apply that. So that's why we had two gene spaces in 3 weeks, right? Um never mind the big journey before to get there, but from that point, we could build really quickly because we had learnings from other people, right? And we wanted to use those learnings. We didn't want to start from scratch and try and make the
[29:19] same mistakes, right? Very important for me, monitor usage to make sure that it's meeting needs. If not, go back to drawing board. Um if any of you know Warhammer, there's a concept of a gray pile of shame. Um and essentially, you get a bunch of mini models, you you you want to paint them,
[29:35] you get them ready, and then you put them to a corner. And then you get the new models, and you put them to corner, and you put them to corner until you have a big gray pile of shame. I feel like a lot of organizations did this with reports. A lot of organizations were just building thousands and thousands of
[29:50] reports, and nobody ever checked whether actually being used. So then you just have this library of reports, and and yay, we have 2,000 reports, what whatever it is, you know? Um I don't want to go down that same route. I think we should learn from that. I think I'm checking these things all the time. And these two workspaces that we got set up
[30:07] in 3 weeks are being used nearly daily by multiple users, right? And I want them to be. And if ever they're not, I'm okay to go and say, let's get rid of it. I'm more than okay to say, let's just scrap it. Or better, why are you not using? Did we miss a brief in the beginning? Is it too slow? Do you not
[30:23] trust the data? But monitor it. Don't let it become that gray pile of shame. So now, where we heading? I have to throw the disclaimer. This is obviously before all the keynote speeches. Some of these things I'm definitely going to be playing with the new toys, so some of them supersede that. Um but
[30:42] kind of broken into three pillars. So, the data engine engineering visualization. More genius spaces. Why wouldn't we? We we getting such great wins with the business as is. Why would we not carry on with that, right? We definitely going to do that. But again, I don't want just a massive genius spaces, right? So, we
[30:58] starting to look at things where we can do things smarter like, you know, putting on supervisor agents on top, right? Attaching knowledge assistance, right? Doing all these other things around it so you don't just have a sprawl of genius spaces just because we can, right? I want them to be purposeful focus. I want them to be driven by the business, right?
[31:13] Um of course, the more wins you get, the more data people want, right? So, we ever loading data. So, we looking at how can we optimize that, right? Um one thing we recently been playing with is SDP, Spark declarative pipeline. Seeing how can that help us move through those layers. How can it help us transition?
[31:29] Um like I said earlier, changing to open shares, right? It's just so easy. It's so quick. And more and more of the our producers we're having to connect APIs or or other sources. Our customers going, "Hey, by the way, we've got a Delta share now, right? Why don't you use that?"
[31:44] Definitely something we'd like to move more towards. Um I talked about those three blocks. The last one that we haven't quite got right is getting rid of Power BI. That's still there. Still being used quite throughout the organization. Um but we are already starting to move all the logic out of
[31:59] the fabric warehouses into Databricks. We're already saying, like, let's let's let make Databricks do all the heavy lifting. I really really want to get the AI BI dashboards. Um but it's not quite there. We're playing with metric views now. Um they're not they're just not there, but it's definitely something that I'm
[32:16] pushing towards, right? Again, I want one platform that looks after everything, right? All the way through. No need to go out. And of course, we're playing with all the agents. What data team in today's world isn't, right? So, literally all of them. The next pillar, and this is the one I'm
[32:31] super excited about, right? Um is the trading applications. So, like I said in the beginning, we've got guys building applications, they're building in lightweight the light uh uh lightweight Python frameworks. Um this again sitting on Postgres. We've got all this capability in Databricks now, right?
[32:48] So, first thing we said is these apps have a lot of heavy lifting. They're doing a lot of data stuff, right? We need to move all of that to Databricks. Doesn't make sense. We've got infinite scale, infinite compute. Why are we trying to do it in this little app that has no scaling or very low scaling or very difficult or costly
[33:04] scaling, right? So, the first thing we're doing is moving all of that logic into Databricks for the applications. We're also starting to migrate some of the apps that actually run in Databricks, right? You've seen how much stuff in the keynotes is about Databricks apps. I'm a big fan of this. Um me and my team are building a lot more
[33:20] apps and and we we're starting to see how we can link them all together and things like that. But, some of those applications have already been built. We recognize that we've got all the data there. Let's tackle the next problem. Those applications are not necessarily up to scratch.
[33:38] At the very least, if I achieve nothing else, I want to get those applications pointing to lake base. I would much rather have them all inside of Databricks using lake base, but the very least if we're going to have them in Azure and they need to sit there, why not use lake base? You have one Rback model. You don't need to maintain many Rback models. You don't need to need to maintain infrastructure as code.
[33:54] You've got it all built in, right? Um so, why not use lake base, right? So, that's one of the things that I'm very keen on doing. And then the last one is our data science and machine learning. Like I said, came from guys literally working on their laptops locally.
[34:09] We now have all models in Databricks. We're migrating them all to MLflow. Um there's one or two we still have to do, but all our new models are being built in Databricks. They're all being built in the MLflow root framework. We had to get that stood up. And this all happened natively while we were doing the Mez migration. We were building all these
[34:25] capability in the background, right? It wasn't a deliverable, um but it was happening in the background, right? Like I said, we may not be there to replace Power BI, but any of our models and that that we're building now, we're checking a Genie space or AIBI dashboard on top of
[34:42] it straight there and then, right? So, you've got that visualizations, you can see those models, and you can see the results quick and easy getting it to the people that matter. You would have seen in the talk this morning if you watched the the keynotes, um talked about how ML the the the need for
[34:58] it has gone up. I think it was 50x or something like that. We're experiencing We're experiencing that. We have a lot of people wanting more and more models, right? We We don't have the largest team. So, we're resorting down to how can we start vibe coding these models? How can we start using these frontier models,
[35:14] start building our own ML models, right? Um and super excited to take back to the ML team when I go back all the stuff that's come out of the keynotes this morning, right? What is our end goal, though? What What What what is my end goal, right? What
[35:29] What do I want to achieve? What do I want to achieve? What do I want to achieve? That. Doesn't matter if a user wants an application, a dashboard, a Genie space, an agent, shouldn't matter. You should have one place that you go
[35:45] to. What would you like to know? Go there, ask it, it gives you the answer, right? And it's not an unachievable goal. It's really achievable today, right? We We've got most of the groundwork for that. But that's my end goal. Doesn't matter who you are in the business, you can go
[36:00] there, and you can get your answers. Just want to give a bit of shout out big shout out to the team. I've tried to condense what we've done into 40 minutes. Um the real value or the real exciting part is the how, and I'm more than happy to talk to anybody about it. But if you meet meet any of these people
[36:17] or you reach out to any of them, these are the people that made it happen, right? So, they would be awesome, more than happy to talk the how and get into the nitty-gritty. Um very smart people. Um there is two other people that aren't on here that we definitely could not have done it without. Um, and that's our Databricks account executive and our
[36:34] architect, um, Will and Dean. Um, without them I've been on so many calls going, "Hey guys, um, we're stuck on this. How do we do this?" Most of the time it was a case of either Will the architect went off and built it or Dean was saying, "Oh, just wait. There will There will be a feature
[36:49] next week that that solves that problem, right?" So, um, yeah, but big shoutout to them. Definitely couldn't have been done without them. And that's all for me and I think we're pretty close in time.

Learn more about the Databricks Data and AI platform.

The information provided herein is for general informational purposes only and may not reflect the most current product capabilities or configurations.