SAP Enterprise Cloud Services Cuts SIEM Costs 70% With Databricks
Summary
- SAP Enterprise Cloud Services runs a private cloud hosting highly customized ERP systems for major enterprises including Mercedes, BMW, Walmart, and Coca-Cola, and built a multi-cloud security intelligence platform on Databricks that reduced SIEM storage and query costs by 70%.
- The Kraken architecture uses Cribl and Kafka to bridge protected customer VPC networks into Databricks while maintaining regional data residency, with Spark Declarative Pipelines and Delta Live Tables normalizing terabytes of streaming security logs per day through bronze, silver, and gold tables across AWS, GCP, Azure, and Alibaba Cloud.
- SAP is developing AI-agent-assisted SIEM source onboarding that automatically generates parsers and OCSF mappings, reducing the previous two-week manual process to approximately one hour, alongside LakeWatch, Genie, and Agent Bricks capabilities on the roadmap.
SAP Enterprise Cloud Services Cuts SIEM Costs 70% With Databricks

SAP Enterprise Cloud Services uses Databricks to support multi-cloud security intelligence while reducing SIEM storage and query costs by 70%. This case study explains how SAP restores visibility into private cloud environments, processes terabytes of streaming security logs and maintains regional data residency across AWS, GCP, Azure and Alibaba Cloud.
See how Cribl and Kafka bridge protected networks into Databricks, where Spark Declarative Pipelines and Delta Live Tables normalize logs into bronze, silver and gold tables. The session compares Databricks costs with Microsoft Sentinel, covers detections through Anomali Logic and explores faster source onboarding with agents, OCSF mappings, Lakewatch, Genie and Agent Bricks.
Delta Sharing FAQ: https://www.databricks.com/blog/delta-sharing-top-10-questions-answered-part-1
Databricks operational excellence best practices: https://docs.databricks.com/aws/en/lakehouse-architecture/operational-excellence/best-practices
Chapters
00:00SAP Enterprise Cloud Services and Private Cloud02:21Restoring SIEM Visibility With Log Surf and Cribl03:08Real-Time Security Analytics Requirements04:44Why SAP Selected Databricks05:48Multi-Region Deployment and Agent Bricks06:20Legacy SIEM Cost Analysis09:48Anomali Logic and Databricks SIEM Migration10:58Multi-Cloud Security Data Ingestion Challenge11:46Kraken Architecture With Kafka and VPCs13:08Databricks Jobs and Log Normalization14:12Bronze Envelope and Delta Live Tables17:06Bronze, Silver and Gold Detection Architecture17:5570% SIEM Cost Reduction and Faster Detection18:57Cyber Risk Intelligence Hub and Genie20:40The Security Data Onboarding Bottleneck22:14Automating SIEM Source Onboarding23:03Agent-Generated Parsers and OCSF Mappings23:54Cutting Onboarding From 2 Weeks to 1 Hour24:44Lakewatch Normalization Proof of Concept
FAQs
Why did SAP Enterprise Cloud Services need a new security analytics platform?
SAP's private cloud manages highly customized ERP systems for major enterprises, but customers lost visibility into the runtime and OS layers when SAP took over management. SAP saw an opportunity to productize security insights using Databricks so customers receive governed dashboards and detections out of the box rather than rebuilding the same capabilities themselves.
How does SAP's Kraken architecture ingest security data into Databricks?
Cribl and Kafka bridge logs from protected customer VPC networks into Databricks while respecting data residency requirements, keeping tenant data within its designated cloud infrastructure. Spark Declarative Pipelines and Delta Live Tables normalize the streaming logs through bronze, silver, and gold tables to support real-time detection and dashboard serving.
What cost savings did SAP achieve by migrating from its legacy SIEM to Databricks?
SAP Enterprise Cloud Services achieved a 70% reduction in SIEM storage and query costs by moving to Databricks for its security analytics platform. This cost reduction was a key factor alongside the ability to handle terabyte-per-day streaming scale and support AI and machine learning features on the same platform.
How is SAP using AI agents to speed up security log source onboarding?
SAP is developing an AI-assisted onboarding workflow where agents automatically generate parsers and OCSF mappings for new log sources, targeting a reduction from approximately two weeks of manual effort to approximately one hour. This capability is paired with a LakeWatch normalization proof of concept to further streamline security data standardization.
Full transcript
[00:08] Does anyone not know the brand SAP? Everybody knows SAP. SAP actually runs the most critical business processes in the economy. All the big brands use SAP: Mercedes, BMW, Walmart, Coca-Cola, and they use them in the most critical processes like ERP systems.
[00:25] An ERP system is the system of financial records. If it doesn't work, you don't know what you owe or are owed, and can't do your financial statements. So it's very important these systems are extremely reliable and safe. Within SAP there's an organization called Enterprise Cloud Services, which runs the private cloud of SAP.
[00:59] It is very strategic because SAP has a priority of migrating customers from on premise into the cloud, a big multi-year project for a company like BMW or Coca-Cola. With this private cloud, SAP takes the highly customized on-premise ERP system and deploys a tenant in AWS, GCP and Azure for the customer, running it on their behalf.
[01:31] This makes migration much easier, and SAP manages a big part of the stack, so it's less effort on the customer. However, customers lose visibility into areas like the runtime, middleware and OS because SAP manages them. They had this access previously but now rely on SAP to tell them the system is safe, and many customers didn't want that.
[02:21] So they implemented a solution called Log Surf using Cribl to route these logs into an object store, so customers can ingest them into their existing SIEM or data platform and regain visibility and trust that SAP keeps the system safe. SAP realized many customers were rebuilding the same insights, like a dashboard showing where people are connecting from.
[03:08] They could productize that so you get these insights out of the box, and they were looking for a solution to run as the engine behind all of that, with a few requirements. One was real-time ETL, in the nature of streaming logs; when somebody connects from North Korea you want to know right away, not the next day.
[03:56] Second, it needs to serve dashboards efficiently. Third, terabyte-per-day scale, so dashboard serving and streaming must be cost-effective at scale. AI: they wanted to build AI or machine-learning-based features for anomaly detection, learning from the data rather than just flagging North Korea. Multi-cloud and multi-region: SAP runs in the big hyperscalers, so build once and deploy anywhere.
[04:44] Last, data residency. With the private cloud there's a promise that data doesn't leave the tenant, so the solution reuses the cloud infrastructure within the tenant instead of ingesting into a third-party solution. They found Databricks as a solution that achieves these: Spark declarative pipelines for real-time ETL, Databricks SQL to serve dashboards, Delta sharing to create global KPIs from regional aggregates.
[05:48] They use declarative asset bundles to develop once and deploy to multiple regions keeping data processing within regions, Terraform to manage infrastructure, and are evaluating Agent Bricks to go beyond dashboards and let end users ask questions in natural language. From this experience they learned Databricks' capabilities and thought about their internal SIEM within ECS.
[06:20] Their internal SIEM has many capabilities: detections, triage, investigation. They wanted those properties plus open formats and reusing the lake. A big factor was cost: their data was growing 40% year over year, creating huge additional cost in the existing SIEM. I created an analysis to give an idea of what a legacy SIEM costs compared to building in Databricks.
[07:25] Take Microsoft Sentinel as an example, since it has a public pricing calculator. In Azure Sentinel the most important part is the analytics tier, where all the SIEM magic happens, and ingesting 1 TB costs $2,480 at list price, expensive if you ingest 30 terabytes a day. Vendors like Splunk and Azure Sentinel introduced a cheaper data lake tier at $150 per terabyte, but you lose the SIEM features.
[08:12] In Databricks you pay for the uptime of compute, so ingesting terabytes over a day with a cluster running 24/7 is $32 or $64, and storing on the lake is about $26 per terabyte per month, huge differences in cost. Same for querying: in Sentinel's analytics tier it's included, but querying the data lake costs about $5 per data volume at list price.
[09:02] In Databricks you pay for uptime, so a one-minute query is about $0.46, or with a 5-minute auto termination, about $2.30. These numbers don't directly translate to a project, but give a sense of the cost difference. The big thing is the analytics tier in Sentinel and Splunk is valuable with features like out-of-the-box detections that Databricks hasn't had for years.
[09:48] They used a third-party solution, Anomali Logic, which provides these features out of the box and connects to data platforms like Splunk or Sentinel. They moved their security workflows from their SIEM into Anomali Logic, and as the next step could migrate their existing SIEM into Databricks as the compute engine, realizing cost benefits plus a broader data platform. Ian developed the entire platform.
[10:58] When we started, the organization is very big with tens of thousands of customers, and to get to that data we have to bridge complex networks and security guardrails for all the customers' data. Luckily we have the Cribl infrastructure already in place, so we did our work on the integration and found this isn't a tooling problem, it's about the data.
[11:46] We have data in all regions, multi-cloud in AWS, GCP, Azure, Alibaba and some on-prem, and need a way to bridge that data and get it into the lake. Because of these complex networks we built a system called Kraken, a VPC with subnets and a Kafka cluster, a regional appliance we plug into a network next to the data source, just hooking up the cables to start ingesting.
[13:08] At the bottom we've done extensions to Kafka to support different protocols, webhooks, HTTP, even OpenTelemetry, because we're bridging a lot of secure gateways. Then the security intelligence fragment, Databricks in the center, pulling data in using a Databricks job running 24/7. As data lands in the regional clusters, it pulls it centrally and does the normalization.
[14:12] For example, a raw syslog gets parsed and normalized into our raw format, our data contract for going into the lake, a JSON object with a bronze envelope around it, and that becomes our contract. Once transformed, we put it back onto Kafka and ingest it through the Delta Live Table pipeline, so all logs regardless of source are shaped into the bronze envelope and go into a single bronze table.
[15:14] We can query the data in less than 20 seconds with this pipeline. Everything on this DAG is code hand-cranked by an engineer. The bronze is the first entry point, and we've not changed that code in over 12 months; it's very stable. Any new logs in JSON format that went through the parsers are supported straight away without changing the entry point code.
[15:47] It's only when we bring in the domain tables to the left of the DAG, in this case supporting Anomali Logic for detections, that from the bronze it fans out to security domain tables, cataloging where the data needs to go: a cloud trail log to the cloud table, a WAF log to the web table. You can query the bronze data quite easily before any logic on the right side of the DAG.
[17:06] This is our infrastructure on AWS with bronze, silver and gold tables where we do all our detections. We're doing POCs, one with Lakewatch onboarding and one with the Anomali Logic blueprints, analyzing the best approach to speed up onboarding. Once all the detections are running, these become our measurable outcomes, the security intelligence we produce from this data domain.
[17:55] The impact is a 70% cost reduction in storing and executing queries. Because data comes in fast, we decreased the time to detect, plus reduced detection engineering by leveraging Anomali Logic. This moves us to the CISO's vision, ultimately to make better decisions. Our current state is the data domains, a cookie-cutter solution we stamp in each cloud, doing processing and detections at the source.
[18:57] We want to bring all that intelligence to our cyber risk intelligence hub, where we ask questions on the data. The ontology in Genie we saw this morning fits this model. The idea is to build a cyber risk knowledge graph, so once the intelligence is in the database we understand the relationships and ask those questions. The threat landscape continuously changes, so it's a 24/7 process recalculating risk.
[19:45] This gets us to the CISO's vision, the ability to make better decisions with cyber risk intelligence, where the CISO can ask the ultimate question: given our current threat landscape, assets, vulnerabilities, controls and business priorities, what are the top actions we should take this week? A traditional SIEM can't answer that, but a cyber risk intelligence hub can.
[20:40] Right now our biggest challenge has been getting the data into the system fast enough. A lot of the code so far has been done by engineers by hand, and onboarding isn't just the code, you need to bring pipelines together with validation. The technology side is the easy bit, but the organization and governance side, the people and process, act like a multiplier, creating inertia.
[21:27] Data onboarding is 2 weeks. In the grand scheme 2 weeks doesn't seem long, but when you scale to 100 sources it takes approximately 200 weeks at this 2-week cycle, so the math stops mathing. We recognized this early and asked for resource, but were told to do more with the resource we've got, which is common.
[22:14] So we asked how to do more with the same resource, and landed on onboarding being a set of steps: start with a data source, analyze it, infer what the data needs to map to, convert that into code and test it, and map the bronze to a gold structure with classification inference. All these add up to a repeatable process plus knowledge work, an opportunity for automation.
[23:03] We proved this: my engineers spent a week creating the agent for the parser generation, used an existing source we'd coded by hand, and validated against what the agent was doing, getting about 85% accuracy in code generation, generated in 5 minutes. We did the same for the OCSF inference, mapping to a gold table, getting it down to 5 minutes with 90% accuracy.
[23:54] Putting these figures together, it's going to take about an hour to onboard a data source, so scaling to 100 sources is about 100 hours, a significant time and cost saving. We're not the only ones; a slide from Anomali Logic quotes 15 minutes to onboard new feeds, in line with what we're seeing.
[24:44] Finally, we're doing proof of concepts with Lakewatch, and we're very interested in the normalization out of the box, where as soon as your data is in bronze you can do all the normalization through point-and-click. We're on the design partnership, feeding into the project and giving feedback, and it complements what we're doing in terms of speeding up how long it takes to get data in. That's ultimately my message: trying to speed up that onboarding.
Learn more about the Databricks Data and AI platform.
The information provided herein is for general informational purposes only and may not reflect the most current product capabilities or configurations.