Skip to main content

Syngenta's data mesh strategy: Databricks for global agricultural intelligence

Summary

  • Syngenta is building a multi-domain data mesh on Databricks to unify agricultural data from R&D, production, and supply chain into actionable intelligence that supports crop protection research and better farming outcomes.
  • The company uses conversational AI to democratize agricultural expertise for farmers worldwide, enabling non-technical users to access complex research insights without requiring deep data skills.
  • Syngenta's transformation lessons emphasize starting with high-value R&D use cases, preventing data debt during scaling, and fully retiring legacy systems as a strategic imperative for achieving real transformation.

Syngenta's data mesh strategy: Databricks for global agricultural intelligence

Watch: Syngenta's data mesh strategy: Databricks for global agricultural intelligence
Global agriculture faces a defining challenge: producing enough food for growing population while soil degradation threatens arable land. Syngenta, a leader in agricultural technology, is tackling this through data mesh and Databricks, unifying complex data from labs, fields, supply chains, and farms into actionable intelligence that drives better crop outcomes.
this video covers Syngenta's journey building a multi-domain data mesh across R&D, production, and supply chain, starting with high-value research use cases and expanding through organizational momentum. Learn how conversational AI democratizes agricultural expertise for farmers worldwide, why preventing data debt while scaling matters more than rapid integration, and the strategic imperative to fully retire legacy systems for real transformation.
🤝

Chapters

FAQs

How is Syngenta using Databricks to tackle global food security challenges?

Syngenta built a multi-domain data mesh on Databricks to unify complex agricultural data spanning labs, fields, supply chains, and farms into actionable intelligence. This platform supports research into new crop protection products and improves understanding of field conditions, helping address soil degradation threatening food production for a growing global population.

What is a data mesh and how did Syngenta implement one?

A data mesh distributes data ownership across domains rather than centralizing it in a single team. Syngenta built their mesh starting with high-value R&D use cases and expanded through organizational momentum, covering domains including research, production, and supply chain on the Databricks Data and AI platform.

How does Syngenta use conversational AI to help farmers?

Syngenta uses conversational AI to democratize agricultural expertise, allowing farmers worldwide to access research insights without needing deep technical knowledge. This approach bridges the gap between complex scientific data and practical on-farm decision-making.

What lessons did Syngenta learn from implementing a data mesh at scale?

Syngenta identified several key lessons, including starting with high-value research use cases to build organizational momentum before expanding. They also emphasized that preventing data debt during scaling matters more than rapid integration, and that fully retiring legacy systems is a strategic imperative for achieving real transformation.

Full transcript

[00:09] Hello. I have a big ask, particularly after lunch. Can you all just stand up, please? It won't be long, I promise. I'm not going to talk about what you're probably most interested in, which is the World Cup,
[00:25] but actually the other side of this picture. What you see there is totally degraded soil. That soil isn't usable anymore to grow food.
[00:42] And we keep losing soil to degradation. And I'd like you to guess how long does it take to lose one more soccer pitch of actually arable soil to drought or to other factors that make it unusable. From now, when you think it's lost,
[00:59] please sit down. How long does it take? Glad about the optimism in the room. So,
[01:14] and are you just like standing? Could also be. Uh you can stand up again if you want, but really, we're losing four soccer fields per second. It's incredible. And 1/3 of the soil on this planet is
[01:30] already gone. It's unusable to grow produce food. And that, while population still growing, by 2050, we expect about 10 billion people on this planet. So, it's no longer a question of just
[01:46] nice and little farms, but how do we actually supply population with enough food with these very limited resources? We're lucky, though, because we're not starting from nothing.
[02:05] Agriculture is super data rich. Not like 100 years ago, right? There was almost no data. Today, just one single run over a field can easily produce 1 GB of data per
[02:22] acre. So, that's what a modern tractor would collect. But, there's also a lot of data, very small and big, both from satellites to sensors in the soil that will capture information about the state of the soil, the microbiome, uh
[02:40] how healthy it actually is, and what the soil needs to be as fertile as we need it to be to grow our food. So, we don't really have a challenge of having enough data, and the whole industry doesn't. But, we do have a challenge of
[02:56] connecting that data. At Senta, we do a lot of research into new products to protect crops, but also research into new crops themselves. And for both of this, it's critical that we understand what goes on in those
[03:12] fields. How do we do this? Let me start with where everything starts. When a farmer puts their seed into the soil, it goes underground.
[03:28] It's Actually, there's dirt on top of it. You don't see it. You know it's there, kind of, but it's gone. And ideally, after a couple of days or weeks, depending on the crop, you'll see a little seedling.
[03:43] And at some point, you will have a healthy crop emerging. Our data must not do this. Our data must not go underground. Because we need it from the moment that it's produced.
[04:00] And this can be very diverse. It could be a lab scientist, right? Discovering a molecular breakthrough. So, that data must be available. We cannot hide it. There may be a field scientist in a field struggling with some pest outbreak.
[04:15] We need to know this pest is there. But also understand what's being done about it. We may have a logistics manager who's dealing with a bottleneck in supply. All of this actually is critical to produce enough food for all of us.
[04:32] It's nice to connect this with three lines, right? It's probably not that simple. Which is why we need to talk about Databricks and all the work that we are doing all together. Um because it's not just three lines. Let's maybe have a look into what really is in that box and what we are doing and
[04:51] what the industry is actually doing to drive agricultural intelligence beyond just artificial intelligence. And I'll start in the lab. In those labs, historically, innovation was very sequential.
[05:06] Looking for new products means we're looking for something that works. When there is something that works that can protect a crop, it's great. But next, we need to understand if it's safe, right? We also need to understand if it's actually realistically producible.
[05:24] And then and then and then. And this takes a lot of time that we don't have anymore. Our climate is changing very quickly. Growth conditions are changing quickly. We don't have 10, 12 years to figure out a new solution to a new problem. So, at Syngenta, we gave a special
[05:41] mission to some of our great scientists, uh just a handful, and that was a data first research project. We said, "We know you struggle with data whatsoever, right? Go for it. Invent a
[05:57] new product without doing experiments. Just use data that's readily available, as we hoped. And see how far you can get, right? You're not allowed to do experiments unless you need to produce some more data. But don't trial and error your way
[06:14] into a new product. See what's possible." And what we learned was it took them up to 9 months to make some of our old historical data available, usable for their science. First to find it, second to actually
[06:31] overlay metadata. So, 10 20 years ago, people didn't really care about capturing metadata along with some maybe chemical experiments. And making this useful took a long time. But we succeeded. We understood what data is needed. But it was a couple of
[06:48] people just hacking their way through in a way. And that inspired us to understand what it really takes to use data at scale. So, we really started at a very practical case. Uh there was no view of let's just
[07:04] digitalize, verify all the data we have. But go all in. Find a new product and see what it takes. And from this data, actually we moved into our first lake house, which was for R&D.
[07:19] And that is spread across five domains, where we capture data now from the field. Like experiments in the field, but also real data from actual uh crop growing. Um data from the labs,
[07:35] experiments that we will be running in our labs. Of course, we need to understand our products themselves and our product portfolio and very importantly safety and regulatory. So, across these five domains we have now organized our R&D data
[07:52] so that we and our scientists in particular are able to access them instantly. So, that's not a gradual improvement. I can't express it in percentages. We're just skipping the whole pregnancy. So, it's from 9 months to zero.
[08:08] And and that's now the availability of actually the wealth of our historical data and importantly the data that we are currently creating in all the new work. And in R&D this has helped us not just to be faster
[08:26] but it is helping to actually create new products that were inconceivable before. So, these products actually are now all touched by AI. Every single crop protection product that we are working on
[08:43] is touched by data, is touched by AI uh across the research cycle. And with this we've widened the search space. We're able to understand much more opportunities that humans could not convey. And we are now on track to actually drive
[08:59] our portfolio through that power of what we've collected now in in our R&D hub. Now, if we start stop there we wouldn't be talking data mesh, right? One hub doesn't make a mesh. Uh this one has been extremely valuable
[09:15] for us. But here's our next problem. This is a very awful picture in many ways. It may look nice, but I don't like it because this warehouse is full. Right? Which means we have a lot of capital
[09:31] locked into this warehouse. And it also means all our products are not where they need to be, which is in the field. And of course, we can use AI to forecast demand much smarter. But no AI can forecast demand right
[09:48] without enough data. So for this particular problem, not very hard to imagine, there's another big data issue. How do we integrate data from our SAP systems, from Salesforce, integrating this with weather data
[10:03] actually, understanding what's needed where, logistics data, where our products are stuck or where they currently are. Um So not surprisingly, we figured we have a template. So R&D is doing this. Let's go for it.
[10:21] And we managed to transition our logistics now into very data-driven, data-product driven, um more science than guessing in many ways. With the second hub into our mesh, um
[10:38] reflecting that data. Uh Now achieving predictive supply where we have reduced some of our overstocks in the warehouses, um but also moving into very differentiated supply chain. So this is all going on as we speak. Um we
[10:56] have a very mature technology ecosystem now spanning R&D and PNS. Uh and very importantly, with this first follower in production and supply, the magic happened.
[11:11] And very frankly, this was not strategy. This is something that genuinely has happened to us. Yes, we kind of wanted it. We have been driving it. But having R&D
[11:28] and a long time nothing was good. Was a very useful data hub in itself. Production supply coming along. Great. And this is the first follower. And I think that's a good learning for
[11:44] any build-up of a data mesh across the company because the rest comes quickly. As soon as we had two major parts of the company on that mesh, there's almost a bit of FOMO and everybody else is joining in very quickly.
[12:02] So, this allowed us to move much faster after having these two big parts of the company already in. Um where the rest, including sales, customers, understanding the fields in detail, but also corporate was joining in to complete our mesh across the whole
[12:19] company. So, now our final, and nothing is final in our IT world, but let's call it preliminary final uh view of our ecosystem here. This is actually where we have R&D with
[12:37] multiple sources, um very, very uh sometimes immediate and IoT data sources from sensors to legacy data from old experiments or lab scientists capturing data.
[12:52] Production supply field, customers, and corporate, obviously, uh very importantly as well, bring our finance teams or HR teams into the picture. Uh so, we run a reasonably simple, I might say, architecture across the mesh
[13:09] with those four big hubs. Um and with that, we are in a good position to actually connect everything that matters for us for driving agricultural intelligence, so that we're able to make a real impact on the field.
[13:28] Now, where do we take all this? Does any of this matter? Probably doesn't until we land it here. And this is where reality kicks in, the fields.
[13:45] On this planet, there's about a hundred million large farms. It's farms over 2 hectares. And many of those growers are very, very tech savvy. These farms are a data business of their own in many ways.
[14:00] Um not all of them, but many are. And that's where actually I'm super proud and excited about our community in farming. Uh where there's so much wealth of wisdom that it doesn't take much, I think, to
[14:17] bring it together much more uh to make a much bigger impact on how fast and at what quality and what productivity our food is grown. Um so, we are working with those large farms. And we are working through software, Cropwise, um helping farmers
[14:34] manage their farms very actively. Um we're also sharing some of our raw data, open-sourcing some of our code, and I find it critically important to be in that ecosystem, which by the way is something that we share with the Databricks team as an as
[14:51] an ethos, um that these open architectures and the community around this is critical. So, those big farms are central to driving the intelligence across the fields.
[15:06] And we cannot stop just there, of course. There's not just a hundred big farms, but there's more than 500 million smallholder farmers. And they are incredibly critical
[15:23] for global food security, supplying local communities, but also for their own livelihoods. Right? These are families whose whole existence essentially depends on the cropping that they do. And that's massively fragile.
[15:41] There may just be a little pest, just some insect pressure, and you will lose the whole crop of the season. Um which puts the whole family into big strain. But also the community around them um may not be supplied with quality food that year.
[15:56] So, these smallholders, in many ways, are as important or more important locally than the bigger farms. And they have way less means to technically optimize what is going on.
[16:14] And they have no access to our science or any of the big science on the planet in many ways, or at least they did not have. Because that's where conversational AI comes in. That for us is a massive way of democratizing access to data.
[16:30] So, LLMs are not helping us do better science in the fields necessarily. But what they do do is actually allowing people to access that wisdom, to understand what's going on in their plants, to have a much better view of
[16:48] the practices they could apply to protect their plants much better. And to actually tap a little bit into that global knowledge um in the very, very local context. And that's the power of bringing AI on
[17:04] top of now very structured data, uh bringing this all together. So, it's not just AI, of course. There's a lot of data science in what we do as a company, but exposing this has benefited massively from the past couple of years
[17:19] and from what's now developing of course even even more fast. talk through couple of those transformation lessons.
[17:34] And this to me there's five. I mean you can do any number in a way, right? But we have those big big drivers for for value at the very beginning. You remember I mentioned in R&D our
[17:50] challenge to scientists do something good with data as a new product. How's that possible? Just go through and find your way. Uh start at where there is the value. And I'm very happy we did that because it's
[18:06] a massive rabbit hole to just digitalize all our R&D information for example. I don't think we would have succeeded frankly um to just go and try to clean it all up and bring it all together. Um so this way looking after value first
[18:22] we were able to focus on what data matters and bring that in first and quickly and prove value of course which then accelerates the whole rest of the journey. Um same for production supply. Again, where's the value? How can we
[18:38] move this quickly? And then we start there, right? Building function specific agents, problem specific agents underpinned by data that is required. Um we're not integrating the whole world yet. So I I'll come to this in a moment.
[18:54] So start at value. Secondly, find the first follower. Because the journey was slow in the first years. You know, R&D was fast and good but isolated. Uh and we only really accelerated after
[19:09] number two was on board. As I said, there is no mesh with just one app, but two apps almost make a mesh, and two apps certainly get the whole community going. It's very critical though, of course, that the community gets prepared for
[19:25] this. And this is where I'll get you now. Because what I frankly underestimated a bit is the search that will happen. Databricks makes it incredibly easy to
[19:41] scale technically. It don't help us with our people, right? So, how do we actually manage the scale of technology at the speed that our people can cope with? It's so tempting, right? Once you have a
[19:56] pattern, you have a template, and you can integrate more and more and scale and grow. Um but this surge technically needs to be managed first. Uh we are putting a lot of focus on actually not bringing in any of our data debt, um which will be hard to fix in
[20:14] hindsight, as we are all experiencing time and time again. Uh so, this this time first time right. Um but also, as we scale the technology, we are scaling governance, right? And governance not just in rules, but the
[20:30] whole operating model. And in this operating model, we're putting a critical focus on actually changing how we operate as a company. And what I've heard is in many talks, I guess, as well. Like, don't put AI on a broken process.
[20:47] It won't fix it. Uh and that's again tempting because it's so easy to scale, but what does it really mean to change the way we work through the power of data and AI? And bringing in that perspective very early has been a critical success factor,
[21:04] combined with the focus on value, to make sure that we scale in ways that actually make sense and are sustainable. Last and it always comes last, but very often it comes 10 years later.
[21:20] Uh cleaning up the legacy. Um I'm taking some risk in talking to you about it because we've just started. But I want to mention that we made that decision and I'm very happy to also be in exchange with you about this.
[21:37] We said, "We will never do this if we don't do this decisively." So we have just kicked off work to exit our legacy data infrastructure by end of the year. Which is daunting to me. But I feel there's no other way. This
[21:52] long tail will haunt us for forever, right? If we're not decisive about moving completely into the mesh that we've built. So we're in the middle of this journey. Uh I said, "Happy to take your learnings as well or questions. Um
[22:08] but it is critical I think to bring this conviction. Um we cannot just let this trickle out. I don't think it will ever fully happen. And so the strategic value of the mesh that we've created will be fully realized once we have all
[22:24] critical data in. Um and for now of course we already have all the high value high value cases supported through data in the mesh. I'll leave us with this. Um
[22:40] maybe just a call to anybody here here who's in the industry or touching the industry. Uh mentioned earlier, please let's share more, work more together with data. And data isn't there to be buried underground. Uh and as an industry we can only make
[22:57] progress with more open source, more shared data. Uh there's a lot of pre-competitive data that makes a lot of sense for all of us. Uh and building that layer together so our agricultural trends can be much smarter, uh has more insights from
[23:14] everywhere. Um we're committed at Syngenta where we keep sharing. Uh and I really wish for a community that is even more active. We see this in many other industries. In agriculture, there's so much potential. Uh and I think we're on the
[23:30] cusp of finally really getting there.

Learn more about the Databricks Data and AI platform.

The information provided herein is for general informational purposes only and may not reflect the most current product capabilities or configurations.