Skip to main content

SAP to Databricks: Building Enterprise AI with SAP Business Data Cloud

Summary

  • SAP Business Data Cloud automates the transformation of SAP data into lakehouse-ready data products, managing bronze and silver layer ingestion and eliminating manual extraction pipelines that previously required months of curation.
  • Adidas manages 12 petabytes across 250 distributed data products and uses the SAP-Databricks integration to reduce Integrated Business Planning cycle times from hours to minutes through zero-copy Delta Sharing.
  • Hana Cloud Data Lake Files provides object store-backed persistence, and Lakehouse Engine supports Spark-based integration across 40-plus source formats, allowing organizations to scale AI workloads directly on SAP's most important enterprise data.

SAP to Databricks: Building Enterprise AI with SAP Business Data Cloud

Watch: SAP to Databricks: Building Enterprise AI with SAP Business Data Cloud
As enterprise data grows from gigabytes to petabytes, integrating SAP systems with analytics platforms has required complex pipelines, redundant copies, and months of curation. SAP Business Data Cloud simplifies this by automatically transforming SAP data into lakehouse-ready data products. Niclas Schlautkoetter, SVP at SAP, and Dmitry Luchnik, Enterprise Architect at Adidas, demonstrate how zero-copy sharing between SAP and Databricks accelerates AI transformation.
See how Adidas manages 12 petabytes across 250 distributed data products, integrating Integrated Business Planning data to reduce cycle times from hours to minutes. Learn how SAP Business Data Cloud manages bronze and silver layer data products automatically, eliminating manual extraction and transformation. Explore Hana Cloud Data Lake Files for object store-backed persistence, Delta Sharing for governed data sharing via Unity Catalog, and Lakehouse Engine for simplified Spark-based data integration across 40 plus source formats.
🤝

Chapters

FAQs

What is SAP Business Data Cloud?

SAP Business Data Cloud is SAP's flagship product that automatically transforms SAP data into lakehouse-ready data products, managing bronze and silver layer data automatically to eliminate manual extraction and transformation. It integrates with Databricks through Delta Sharing and Unity Catalog, enabling governed zero-copy data sharing at enterprise scale.

How does Adidas use Databricks with SAP data?

Adidas manages 12 petabytes across 250 distributed data products and uses SAP Business Data Cloud with Databricks to integrate Integrated Business Planning data, reducing supply chain cycle times from hours to minutes. The architecture leverages Delta Sharing for governed data sharing via Unity Catalog and Lakehouse Engine for Spark-based integration across more than 40 source formats.

What is Delta Sharing and how does it work with SAP?

Delta Sharing is an open protocol for secure, real-time data sharing that eliminates the need to copy data between systems. In the SAP-Databricks integration, Delta Sharing connects SAP Business Data Cloud to Databricks Unity Catalog, enabling governed data access without data duplication or complex extraction pipelines.

What is Lakehouse Engine and why did Adidas adopt it?

Lakehouse Engine is a simplified Spark-based integration framework that supports data ingestion across 40-plus source formats, enabling organizations to bring diverse SAP and non-SAP data into the lakehouse. Adidas adopted it as part of their architecture to simplify data integration across their 250 distributed data products while maintaining governance and consistency.

Full transcript

[00:08] Yeah, okay, let's start. Let's start the session from SAP to the lakehouse Adidas blueprint for AI powered enterprise data. And I'm I'm very happy that Dimitri is joining me or I'm joining Dimitri here on stage for for that session because the journey for of SAP at Adidas together with Databricks
[00:25] is also something that is going on now for quite a while. And yeah, we have been discussing lots of architectures in the past, also prior to us launching our flagship product, which is SAP Business Data Cloud, which you will see also with some architecture examples. And now this
[00:41] is finally coming to life and also bringing lots of value to to Adidas. And on how this is implemented at Adidas, what the basic concepts are from an architecture perspective, we will show you also in this session here now. And yeah, let's let's dive directly into the
[00:57] contact into the content. Um yeah, so I'm Niklas. I'm globally responsible for the adoption of data and AI at at SAP. So helping our customers make best value and get best insights out of SAP data.
[01:12] And yeah, mainly covering solutions and platforms like SAP Business AI platform, Business Data Cloud and related solutions. Yeah, Dimitri. Thanks, Niklas. I'm really happy to be here. I'm with Adidas already 14
[01:29] This year will be 14 years. And since over this time I have participated in establishing several different data platforms. The latest one being the lakehouse. My colleague is also here. And now I'm with enterprise architecture looking into the data and AI integration landscape.
[01:44] Yeah, thanks, Dimitri. So let's jump into the context. So yeah, just a couple of words for those of you who don't know SAP, which I believe cannot be too many too many people here in the room, yeah, but SAP, of course, largest provider of
[01:59] enterprise software in the world, leading enterprise application software provider with SAP Business Suite. You can see it down here, yeah. Adidas is one of those 437 customers globally. We are
[02:15] present in all industries, including retail with all the complexity that's behind retail and also manufacturing, of course. And yeah, our systems operate the majority of of the world's global commerce, yeah, and that's why, of course, our systems
[02:31] are of quite some importance and typically generate the most important data that our customers are dealing with, yeah. So, if we are discussing data architectures with customers, we often hear, "Okay, nowadays, maybe not the majority of data is generated by SAP
[02:48] systems, yeah, especially in retail when you think about social media generating data or maybe in oil and gas in other industries." So, the vast majority of data, often we hear, is not generated in SAP systems, but typically the data that is steering your company, that is
[03:04] generating the most valuable KPIs in terms of revenue, of margin, of stock levels, of goods movements, these kind of things are typically derived from SAP systems, and especially the combination of SAP data with non-SAP data is really
[03:19] bringing value to the company, yeah, and then in the end also allows you to optimize things like optimizing a margin and revenue or reducing cost and these kind of things, so that's where we where we are coming from and of course, SAP is a
[03:35] multi-decade SAP customer, yeah, having embarked on the very early ERP journey, is one of the first HANA customers, yeah, meaning SAP's in-memory technology that we have for many years been focusing on, which we are also now shifting into a more lakehouse-centric approach,
[03:52] as you will see that. Um but in the end this is what you can see here is what I what I just explained. There are the mission critical SAP data that our systems generate are steering enterprises globally and they are spanning across all domains and all line of businesses
[04:09] if you will. Yeah, so from finance over spend, supply chain and that's the topic of today also Dimitri right the supply chain and also HCM meaning HR is something that is of of utmost importance and these kind
[04:25] of data sets, of course, you do not only want to process in SAP technology, yeah, but we are partnering with companies like Databricks to derive maximum insights and make use of the of the state-of-the-art tooling that you have all seen here and that yeah, most of you are probably
[04:41] also using already here. Yeah, and how to make best use of SAP data, how to get easy access to SAP data in a way as easy as never before. That's something we will bring you a little bit closer in that session and especially on how SAP is interacting with
[04:58] the Databricks platform for use cases like supply chain with sub IBP and also success factors is something that is that is now bringing Adidas architecture also to the next level. Yeah, over to you Dimitri.
[05:14] Cool. Thank you. I will not introduce Adidas. Actually, I will. It's Adidas not Adidas. But you all know our great products some of them I'm representing here. But I will introduce our analytics landscape. So Adidas is a global company, right? We work on all continents and also our
[05:32] analytics landscape is distributed because every market is responsible for their profit, every market responsible for their supply chain for to some extent to the on the product range. So the system which has to support it It to be scalable, It has to be able to take
[05:49] care of billions of data records, right? And and the petabytes of information. Also, from usage perspective, yeah, you know, it's like imagine all this summit which is using using our platform on daily level. Um next one.
[06:05] Mhm. Core of the our analytics landscape is the lakehouse. Lakehouse is 5-year-old baby. It's not a baby anymore. It's pretty big. It contains yeah, 12 petabytes of information. So, uh thousands of people is
[06:22] thousand people are using this and it's connected to almost every Adidas system in our landscape. Um so, now lakehouse replaced a lot of different platforms which we're building over over
[06:38] years. So, some of them are now already decommissioned. Some of them are on decommissioned landscape. And now lakehouse like truly became the our combined like from this vision, the warehouse, the home for ML use cases, the home for big data, and now of course
[06:53] the host of agentic applications like Genie. Next one. Mhm. And in the course of you've seen on previous slide, 250 data products. Data products is term from from the data mesh paradigm. Just a few words about this one. So, in
[07:11] the past like luckily a long time ago, so the the paradigm for us was that there is a central team who can take the data from multiple domains, transform, and bring it to central systems. Now it's not like this. Uh click.
[07:26] Yeah? And now we moved into this distributed data landscape where business domains who know their data, who know how this data where is the data originating from, where the data is moving to, how the data is used. So, they now own their own data product. So,
[07:42] data engineering idea, data engineering know-how is distributed across multiple multiple teams. This is very important for when you talk about scaling and when you talk about bringing new technology, so how fast is the and how steep or not steep is the ramp-up phase for new
[07:58] technologies and for knowledge, right? And how we can then share. Because like having multiple teams, we want to simplify a lot in the landscape. So, Niklas, we here talk about this partnership, BDC. Please tell us about. Yeah, let's have a quick intro into what
[08:15] is Business Data Cloud as as one of the key components of of what SAP announced also at SAP Sapphire, the next evolution is Business AI platform and Business Data Cloud is of course a key component to it. I cannot believe it's only 5 years, Dimitri, that you that you just
[08:31] mentioned it. I thought it was was much longer. But that shows also this the strong growth that of course Databricks is also having in our customer base, right? So, we see a huge adoption of Databricks at SAP SAP's customers and that's why that was also one of the
[08:47] first partnerships in the data management and data platform space that we went into besides others that we have also now embarked on with all big hyperscalers and also Snowflake, right? But of course Databricks is a special one that was the initial partnership of Business Data
[09:03] Cloud and has of course a special place also in the overall architecture because we also have SAP Databricks as an OEM component as part of our platform. But of course we are treating that the same way as we treat customers who have a big enterprise external Databricks platform
[09:20] already in place and you will also see an architecture example on how that looks like for example on on Azure, yeah? But let's have a few words on Business Data Cloud, what the main concepts are. In the past when customers wanted to put Databricks workloads on SAP data, that
[09:37] went along with complex data extraction mechanism. So, somehow you needed to get access to the SAP data, to the APIs that we had, you needed to create data pipelines to extract data physically out on various levels, table level, view level, application level,
[09:54] Hana level, or database level, and so on. So, quite quite some complexity. That was the first big challenge that we that we saw in the market. The second big challenge was getting the data and the business semantics of the SAP data sets back together in the target
[10:09] until you were really able to make use of that and put sufficient compute workloads from Databricks onto these data sets. Yeah. So, for some of you who might know the SAP data structures, that's quite complex. Yeah, we are talking about header header tables and item tables of a sales order, for
[10:25] example, that only makes sense if they are joined together. Yeah, together with maybe a table that is containing customer data, that is maybe containing also plant data or store data, and so on. So, it's quite quite complex. Right? So, customers had to put lots of time, money, effort in A, data integration,
[10:44] and B, into data curation and bringing the data back together until they were really able to use that. So, that's the first thing we were overcoming with Business Data Cloud, as you can see here, by us integrating SAP data automatically as part of the platform
[10:59] into not in-memory technology, and that's the second big innovation, but also into an SAP centric lakehouse architecture, which is founded on object store technology. So, the same way also Databricks lakehouse is built upon object store, Delta table, and compute,
[11:15] we are following the same concepts, and our friends from Databricks have said last year, "Finally, SAP also goes lakehouse." Yeah, and that's what we are doing. So, taking away the need from the customer to integrate data manually, and by creating SAP managed data products,
[11:32] taking also away the need to curate and model the data the right way so that it's available for consumption. Yeah, and that's something we are providing out of the box and that's one of the key value propositions of business data cloud. The second one is
[11:47] is it really is Hana in-memory technology? So, Hana is a database as SAP's flagship database technology. Is that really the right place to store each and every data set? How many petabytes were you talking about? 12. 12 petabytes, yeah? Into
[12:03] always-on in-memory database like SAP Hana. And of course, the majority of you of you would say, "No, it is not." From a TCO perspective, from a scalability perspective, the lakehouse has proven to be the best possible architecture that we can have nowadays, right? And I mean,
[12:19] our friends from Databricks they announced things like also lakehouse RT, real-time. So, all these things are evolving and that's something we are also participating and this is why we also have the primary persistency for all SAP data, no matter if it's coming from S/4HANA, if it's coming from
[12:36] SuccessFactors, Ariba, Concur, Fieldglass, the primary persistency of the platform will be object store technology, yeah? And in in in our case, we are partnering with all big hyperscalers because the technology in the engine room underneath of that is
[12:51] either AWS S3, Azure Data Lake Gen2, or Google Cloud Storage, yeah? So, that's a big innovation which is more of a foundational pattern here within the platform. Then we introduce something like the knowledge core. Somehow you need to process the data. You need to put your
[13:07] the compute of your choice onto that data. And of course, also we are not a pure data layer that is intended to to to provide data to external platforms, but we are also offering state-of-the-art compute engines to process the data efficiently for very
[13:22] SAP-centric use cases, of course, out of the box utilizing, for example Hana as a data as a in memory database technology to to process the data very efficiently as as we can see it here, which is based upon Hana Cloud, yeah.
[13:38] Besides the workloads that of course are executed in SAP and also enterprise data bricks. But we have our own modeling tools as well with SAP Data Sphere, which is also in use at Adidas for various SAP centric use cases. And also
[13:53] of course the majority I would assume in this room here is operating data bricks in combination with most likely Power BI I would say, yeah. The same way of course SAP has a very strong BI tool in place, which is very much focused on SAP data and very much also
[14:10] focused on planning workloads, which is a super strong business case of course for most customers when it comes to financial planning, workforce planning and and others, yeah. So those engines are consuming the data products that we have in the platform. And in the past you had to create multiple copies as I
[14:26] said via data replication, physical data movement into the lakehouse of data bricks in order to also apply the data bricks workload and compute onto the same data sets. And that's not necessary any longer. Now with business data cloud you are working on one set of the data
[14:42] without duplicating it and without replicating and keeping it redundant. Yeah. Why is it so important to also have the data in the right format at the right place in real time ideally because of course also SAP goes agentic and we
[14:58] are preparing our SAP data out of the box for all agentic and AI use cases because that's of course something that is creating instant and immediate value to our customers besides let's say classical BI workloads, classical data warehousing workloads, classical let's
[15:14] say proco data engineering workloads because the purpose is nowadays to put AI AI workloads and AI capabilities on on top of that data, yeah. So, that's those are the basic concept. Of course, we get into much more we can go into much more detail and you see a much more
[15:31] detailed architecture, of course, in a hybrid environment with SAP and Databricks. But first of all, Dmitri, tell us how is that applied at at Adidas? So, you remember Niklas was talking about complexity of integration SAP data into Lakehouse. So, this is our spider
[15:47] web uh which shows like how it is and this is very very much simplified. So, if we talk about SAP, it's you've seen I've showed the slide with 200 systems which are Lakehouse connected to and SAP you can say like just one of them, but it's pretty big in terms of the impact and
[16:03] the so then the value of this data. So, you see like SAP data is accountable for about 30 30% of the volume, but if you look into the usage, so like two of three reports or two of three applications and queries which are coming to the Lakehouse, they somehow
[16:18] consume one of the other SAP SAP data set. And to get data there, it's tricky. So, sometimes we have to go direct path, sometimes we need to go through the BW, sometimes we need to route the data through Data Sphere. So, this creates this spider web of interfaces. And of
[16:34] course, if BDC would come up a bit later, so our architecture would look a bit simpler like this. And make it a bit more uniform. So, with introduction of SAP
[16:50] uh HGLF with with the Delta Lake files, everything could look simpler, right? So, we can go we can route information into the HGLF and then use it use it inside the Lakehouse, right? And then the same data is stays available also inside SAP technological stack with
[17:05] Data Sphere and SAP Analytics Cloud. Uh with this we would get of course the not only simplicity, but the speed, time to market, which is very important for for fast fast-moving uh pace of our business.
[17:21] Uh all right. But, I forgot something. I was heading getting a bit ahead of myself. HDLF, Nicholas, what's HDLF? Yeah, HDLF, and this is the name SAP's name and branding for our object store technology, which we have
[17:37] branded under the Hana portfolio. Yeah, HDLF stands for Hana Cloud Data Lake Files, and it's the MC object store that's based upon S3, Azure Data Lake, or Google Cloud Storage, as I said, with a wrapper around that, which makes it highly integrated with the rest of the
[17:52] in-memory database ecosystem. That's HDLF. That's the basis of Business Data Cloud, and also the basis of SAP's lakehouse approach. Yeah, that's where all the data products that we generate in that lakehouse with means of Spark
[18:07] reside at. And there are, of course, various advantages to that. Yeah, if we are using, of course, open formats like parquet, we can apply Delta table on top of that, and then apply Delta Sharing Protocol on top of that to share the data out,
[18:23] rather than replicating and physically moving it into, for example, the Databricks object store. Yeah, but how about the idea to make that data not only available to Datasphere, to SAC for processing within our own ecosystem, but
[18:38] with with a with one mouse click, sharing that data directly to Unity Catalog of of Adidas, for example, and make the data available there for all the workloads that you can put on the data assets that you have under governance in Unity Catalog. Yeah,
[18:54] so that's a strong value proposition of Hana Hana Cloud Data Lake Files, of HDLF, and also, of course, the Delta Share Protocol that we have put on top of that. And that not only holds true for the data that we are integrating proactively on behalf of the customer by
[19:09] us managing the integration of S/4HANA data, of SuccessFactors data, of Ariba data, and so on, but also of course custom-specific data products that you can use in order to maybe realize very custom-specific use cases. And the next
[19:24] slide will show you an architecture as we are discussing it for example with customers who have a separate Databricks instance uh running on Azure on the on the very left side here, and Business Data Cloud in the center of it, which
[19:39] holds all the SAP data. And to to guide you through through this one a little bit, yeah, you can see that we are having our source systems both in the cloud, on premise down below, and we are integrating data into HDFS, what is
[19:55] which is the SAP object store on that picture here, where we claim to already have a very efficient mechanism to create bronze and silver layer data products on behalf of the customer, which are then ready for consumption out of the box with the means and the
[20:11] toolings that you can see up here, yeah. Um that's not only limited to the source systems that we are uh uh offering to our customers from a transactional perspective, meaning the Business Suite systems like S/4, like ERP, and so on, yeah, but also um and I
[20:28] assume someone or many of you can also relate to that, at least SAP customers can. We have a very long data warehousing history with SAP Business Warehouse in place, yeah. So, the good classical BW system, and I think you guys are also still operating a BW
[20:45] system, yeah. Um maybe not as big as with other customers, but we roughly have still, I would say, 25,000 BW customers out there who are relying critical workloads in the area of of enterprise BI as well as planning
[21:00] uh in on on that classical on-premise code stack, yeah. So, that's something customers can also not get away so easily from because they are operating those critical maybe also planning workloads, consolidation workloads, so very much enterprise business processes in in that data warehousing stack,
[21:17] right? But of course, in the area of AI, of cloud-based data warehouses, also like Snowflake, like Databricks here, but also Google, AWS, and Azure, yeah, all the customers of course want to get away from the classical BW, yeah, but
[21:32] the the question is how how disruptive can that be, yeah? And with this path of bringing BW also into the cloud environment of business data cloud, we are providing a minimum non-disruptive way to modernize your BW by first in the first step bringing the
[21:48] BW system in there, bringing the data of the BW system onto onto the object store layer, applying delta table delta sharing protocol on top of that to super easily also share BW data with Databricks and applying the workloads of BW also onto onto that data assets,
[22:07] yeah? With a clear goal to of course decommission the BW system over time as fast as possible by at the same time safeguarding the workloads as well as the the data assets and the models that you have created maybe over decades in that BW system, yeah? So also that data will
[22:23] land in the object store technology, it can be consumed with Datasphere SAC, intelligent applications, and most importantly with AI agents and Joule, which is our co-pilot on the one side, but without any copying of the data, you share the data out via mechanisms that
[22:41] we call BDC connect, and that one is just a reverse proxy that is allowing you to share exactly the metadata information with Unity Catalog in in in Databricks, and all of a sudden that data is available in your in your Databricks environment, which sits in this case on Azure, but
[22:56] you could easily of course replace Azure with AWS and also Google, so uh cloud-agnostics of of um of Databricks, yeah? So, that's a very strong value proposition for for many customers because it's like I said, overcoming the need to physically integrate data, to
[23:13] prepare the data with much time and effort because one important aspect, the source system data at SAP, also changes in a in a very fast pass pace, yeah? So, with every new release, we are updating the data models. Maybe for certain tables, we are adding new columns and so
[23:30] on. All of these changes in the source system data models break existing physical data pipelines. Nothing to worry about because we are taking care about the life cycle of SAP data in form of data products in the object store, and you will always have stable and
[23:45] governed SAP data, no matter what kind of changes are lying underneath of that. And what we can of course also see here, one important source system is for many customers who are of course also in the manufacturing and supply chain business active, is sub IBP, integrated business
[24:01] planning, yeah? And these are the cases that that we will also show you now, um which are very relevant for Adidas. It's the SAP IBP integration and also SuccessFactors integration via the platform of Business Data Cloud and share that data in a zero copy way also
[24:18] with the Adidas Databricks instance, yeah? So, the question would be, how can that technology then help your specific use case, Demetri? What are the benefits that you can use? Can you use it with your uh with your designed medallion architecture that you have designed at
[24:34] Adidas, yeah? So, you most of you, I think, who are using Databricks, you're using medallion architecture, right? And for us, uh BDC jumps into jumps into this picture, and we need to find a place actually how to integrate it. So, do we use zero copy? How we how we use it as which layer in
[24:51] this medallion, right? Is it the help or not help or not? So, but here we also needed to consider so, what is our usage of of the data across different layers. So, you see that's what what we measured across the across the landscape that uh for every terabyte ingested into the
[25:07] bronze, we have 100 times over reads from that bronze, and then 1,000 times over reads from silver, and so on. And it might look a bit counterintuitive, but if you think about all the reconciliations, all the checks, and multiple reloads, and data synchronizations which are happening,
[25:23] and then the usage of the silver layer across multiple use cases, this starts makes starts make sense. So, when we looked into uh how BDC can integrate with the lakehouse, we thought about this as a trusted source for SAP uh for SAP data,
[25:40] right? And considering also the cross cross-regional latencies, we prefer either integrations through bronze, integrations through silver uh silver layer depending on the use case and the size of the use case, and uh uh amount of amount of those data loads in between
[25:56] and between the layers. So, that's about medallion architecture. Um Lucas? Yeah, I mean, when also thinking about what is SAP bringing to the table, and what is the key differentiator also for
[26:12] SAP when we think about the business content that we are bringing to the platform. So, again, SAP is not competing on the level, of course, of a technology platform with Databricks, with other platform providers, like also Snowflake and the big hyperscalers, right? That's not our our core business,
[26:29] and not our core domain. Yeah, our core domain is business data, business processes, and business systems that generate data. So, why not offering our customers also out-of-the-box content that they can use, extend, enhance also with AI capabilities to ideally
[26:46] um make use of SAP data, and benefit as fast as possible um from the insights that these data can can generate. And this is why we have introduced the concept of intelligent packages with domain content that you can apply on top
[27:02] of the data products that we have proactively ingested into the platform and curated to a high degree so that you would not need to do every from everything from the scratch neither in SAP data sphere for example, but also maybe for very SAP centric use cases,
[27:17] you would also not need to model these from the scratch in data bricks, but rather maybe consume immediately what's out there in the box for very specific LOB use cases and that's what we call, as you can see here of course, analytic models, stories which are then
[27:34] deriving where you derive insights from. The whole domain of planning is of course in an enterprise something that's of utmost importance. We are providing out of the box context for the critical business planning capabilities here, yeah. So that's that's another value
[27:49] prop- position that we are right now also further enhancing and infusing with everything AI, yeah. So the next releases that we will be bringing into the platform will consist of AI agents which allow you to create maybe SAC dashboards, widgets, analytic models,
[28:07] planning models out of the box which much must much without much user interference. And the thing is um is it really nowadays required that you would model and dashboard by scratch by hand, yeah? Maybe maybe through a consultant. Those times are over, yeah?
[28:23] In the agentic world, these are tasks that are highly specific, highly individual maybe for each one of you in the audience can be performed and generate exactly the visualization and the BI dashboard of your needs highly individual with with almost no
[28:39] effort, yeah? That's the area where SAP is evolving in and yeah, what is the what is the application of intelligent applications or the applicants of intelligent applications at Adidas, Dimitri. So, for intelligent applications is
[28:56] for us again as I said BDC is a trusted source and intelligent application it can be a trusted source for some specific data sets, right? We just tested by SCP and saves a lot of effort on our side. So, our usage for intelligent app is we deploy, we connect and we can use.
[29:14] So, this is this is the value proposition what we see and we apply it also for some of the some of the use cases. Um looking to IBP. An example for IBP. Yeah, I mean this is like what I said, right? So, we are offering these kind of things for a
[29:31] variety of source systems and IBP is is for sure a super important one especially for these high volume scenarios that our customers in the area of IBP have and yeah, Dimitri explain it a little bit more at IBP at Adidas.
[29:46] So, IBP stands for integrated business planning. It supports our supply chain and to understand what actually it does, I need to quickly zoom in and to like what is the complexity and scale of the of Adidas supply chain. So, like normal year Adidas produces and ships about 1
[30:02] billion products, right? So, number so it's shoes, it's t-shirts, right? It's balls and socks and so on. Whatever is shipped and we don't produce in one factory in one country, right? It's distributed network about 300 suppliers globally. Then it has to be shipped into all our
[30:17] markets and all the into all continents. So, it's like a supply chain is is pretty pretty complicated. Uh also when we talk about the end distribution so the outbound outbound supply, right? So, it goes to our wholesale partners, it goes to our own retail stores and
[30:35] also don't forget about adidas.com. Maybe you can use it today uh to buy some of our awesome products. It's uh, millions and tens of millions of packages which are delivered globally. So, if you think about it, like since the moment we showed this slide about what 4,000 items were
[30:51] already produced, right? And 32400 parcels delivered. All right. So, what IBP is doing, right? And what is our planning inside supply chain? So, um planning needs to marry what our
[31:09] customers would like to get, what we and our supply chain can deliver, can handle, right? And create this realistic realistic demand plan. We don't run it once a year. We don't run it once a month. We run it once a week. And there are two um
[31:25] uh, two important deadlines and timelines, right? So, like middle day of Friday, when all the plans are fixed, right? So, there is just 60 hours, 66 hours till Monday morning when the uh, planning next planning cycle
[31:41] supposed to be finished. All the magic in IBP supposed to happen. All the data extracted to the lake house. All the reports prepared. So, that the planners and people responsible for supply chain can look into this data, right? And figure out uh, if plan makes sense, make
[31:58] some adjustment, and so on, and transform those uh, suggested plans into purchase commitment, financial commitment. All right. Um So, we run it for the first time. Results were a bit disappointing. So, we over
[32:15] like the planning plus reporting preparation went well over into Tuesday, which doesn't leave enough time for the planners actually to do to do their job, right? And then all discussions started, like how complicated this IBP supposed to be, what kind of sacrifices we need to make to fit the planning run into
[32:34] this window. And uh yeah, so we can look like preparation for reporting many hours. Unfortunately, that by that moment that was already so heavily optimized, it wasn't possible to squeeze squeeze a second out of the out of the reporting preparation. So, it's a lot of complexity and the sacrifice sacrifices
[32:50] for business logic supposed to get into the planning run. But then suddenly IBP comes in, right? And uh Next one, so we tested. Ooh. So much better, right? So, uh open source sharing hours became
[33:06] minutes. Every preparation can be done much much faster, right? So, I mean, only BDC would not be able to solve all our IBP problems. There still was some optimization for the for the business logic which we needed to do, but this buffer of multiple hours, right? It
[33:21] helped to actually find some reasonable business solutions for uh for supply chain complexity and supply chain planning. So, that's the IBP story. IBP story, but there's more to that, right? And I mean, that's a very good example on how you are saving so much
[33:36] time also on integrating IBP data into an analytical platform until you can really create value and insights out of that, yeah? That's one example. The next example is, of course, we have a huge customer base operating our HR system, which is called SAP
[33:51] SuccessFactors. And SAP SuccessFactors was also acquired by SAP. So, no native SAP technology, different database technology, different um APIs and integration mechanisms. And also that is something we are taking care of by SAP
[34:07] ingesting HR data from SuccessFactors in the platform ready for consumption. And that's also uh was an important use case also at at Adidas. That uh and it is, by the way, one of of many um use cases that also many customers are looking forward to.
[34:24] So, yeah, correct. Uh with SuccessFactors, the challenge for us was slightly different, right? So, the team uh so, we measure a lot of information about about people. Niklas said also right also correct success factors some
[34:39] technology some data models. Those data models they good for for maintaining the business of of HR maintaining the relationship with people and so on but it's not very good for analytics, right? And then if the team is not big enough then taking this
[34:55] like learning what is the data model of success factors, right? So figuring out what is best data model for analysis. It's an effort, right? And it would require so this learning curve would be would be very steep. Intelligent app in this case they helping, right? So they take the
[35:11] heavy logic of transforming some data structure we don't even know need to care what it is into something more suitable for the for the report creation report building. Something in the star schema or snowflake schema something which we can put the reports on tell
[35:26] Genie how to how to use this data and so on. So as said, right? So heavy lifting is happening inside the intelligent app. The team needs to just press couple of buttons deploy and and then use. But it's not the full story, right? So the how smart how
[35:41] intelligent the app is it has no clue about Adidas specifics, right? So it knows about the success factors data models and success factors KPIs. I'll give you an example for instance so Adidas measures KPI called women in leadership, right? So success factors
[35:57] knows gender split success factors knows the our grades, right? So which positions occupied by what people but has no idea like which of those positions are actually making any kind of contribution to the strategy, right? Which can be classified as leadership. So on this mapping is like one mapping
[36:13] file which we can easily bring into into the context and into the reporting landscape. Okay. Yeah, interesting Dimitri. That's exactly what is relevant for so many customers, yeah? And if you think about that was just IBP and SuccessFactors as
[36:31] a data sources in the in the platform that we can share then with Databricks and process the data further, yeah? But if you think also about the vast ecosystem that's still out there with Ariba, with Fieldglass, with other solutions that are out there, yeah? It's as easy as never before to consume SAP
[36:48] data in enterprise Databricks, yeah? And Dimitri, tell us about a little bit about the minimum setups that you said you have realized that uh Yeah, correct. So, when we talk about the uh minimal setup, so that's that's important important for us, right? So,
[37:04] it's on the platform of our size, any kind of inefficiency, any kind of scalability problems that would create a lot of frictions, right? And the platform will, yeah, over overload, overheat. We don't want that. Uh and so, that's why we created from the beginning
[37:21] we created the thing which is called uh Lakehouse Engine. It can take data from known formats. Uh it's about 40 formats now supported, maybe even more, so you can have a look. There is a QR code, scan it. Uh and it's it's open source by Adidas,
[37:37] right? So, what it is, it's um framework, Spark-based framework for data engineers. You can easily specify from this data source, get the data into the lakehouse. It takes over all all non-business value-added uh adding
[37:53] boilerplate code, how to handle deltas, how to produce change data feeds, how to transform information, how to how to handle Kafka keys, and so on. So, it like it helps to to organize this stuff. And uh now, so just imagine we want to load data from BDC. BDC offers data in the
[38:10] Delta Lake format in cloud storage. Then what happens for us, so integration of this data Nikos, please. Mhm. Click the next one. Integration of our data into the lakehouse takes that amount of line of code and that amount of complexity. So, this is like
[38:26] you import lake house engine, you specify, okay, take from this delta share, take this table in the delta format, store it, please, in this location, in this file, and use a merge merge command. So, it will take care of all the complexity, right? And that's
[38:42] where we seen the huge improvement. That's the that was the magic of those IBP, right? From complicated extraction and transformation logic to very performance and cost and speed optimized uh
[38:57] uh processing. So, I mean, if you want to take one note out of this out of this session, it would be this one, right? BDC makes SAP data natively supported by lake house. Mhm.
[39:13] Yeah, good point, Dmitri, and uh probably no single session at Databricks Summit without a picture of of some code. Yeah, yeah. But uh yeah, Dmitri, last question. I think we are almost almost at the end of the presentation, yeah? One last question about scalability, yeah? That's
[39:28] also something. So, scalability is like in everyone's mouth here and on on top of everyone's mind at at the Databricks Summit, yeah? How does that uh relate to the Adidas blueprint? How can you leverage AI and scale AI across the different business units that you have? So, uh to answer this, maybe we we
[39:45] quickly look into the picture of the data products, right? So, we know our data products, we know their scope, and what I like the most with the BDC ID and the data products in BDC and how BDC manages the data is that basically with BDC, you are not connecting to a huge
[40:01] box of SAP systems. There, through one service principle, you have access to thousands of tables, and then some data engineer would bring something which he don't need he or she do not need for that, right? So, BDC allows to make those atomic packages of the uh BDC data products, which are tailored
[40:18] for the specific purpose, right? So, then you ingest, you connect like one by one. You tailor, you have a clear lineage, right? And then you can add and add them uh, you know, kind of endlessly. So, from that perspective, this also fits well
[40:34] into our Adidas landscape and into our Adidas lakehouse. So, yeah, that would be All right, sorry. Then, thanks a lot, Dimitri, for sharing the Adidas journey, yeah, which is about to continue, of course, um, in our joint partnership between
[40:50] Databricks, SAP, and Adidas, yeah, and we invite you, of course, to also follow the same path, and, uh, yeah, please reach out, visit us at the booth. Thanks, Dimitri. Thank you, Nicholas, and thank you, everybody.

Learn more about the Databricks Data and AI platform.

The information provided herein is for general informational purposes only and may not reflect the most current product capabilities or configurations.