Scaling Streaming Analytics: SEGA's Lakeflow Journey with 2 Billion Events
Summary
- SEGA Europe migrated from a custom Scala Spark pipeline to Spark Declarative Pipelines on Databricks, eliminating manual event handling and reducing deployment labor from weeks to hours while maintaining sub-minute latency for 40,000 events per second.
- The new bronze-silver-gold medallion architecture handles automatic schema evolution across 140 games and 40 million players, and processed 2 billion events during the Football Manager 2026 launch without operator intervention.
- SEGA's gold layer uses Materialized Views and Metric Views to provide consistent KPIs for player segmentation, personalization, and game launch analytics, giving the data services team a reliable foundation for production use cases.
Scaling Streaming Analytics: SEGA's Lakeflow Journey with 2 Billion Events

SEGA Europe processes billions of game telemetry events daily across 140 games and 40 million players for real-time analytics and player experience optimization. The data services team ingests 40,000 events per second with sub-minute latency. When a custom Scala Spark pipeline proved difficult to maintain, SEGA migrated to Spark Declarative Pipelines, adopting a resilient medallion architecture that automatically handles schema evolution.
Learn how to build production streaming pipelines for game-scale workloads using Databricks. SEGA's migration eliminated manual event handling, reduced deployment labor from weeks to hours, and enabled automatic schema evolution. During the Football Manager 2026 launch, the pipeline ingested 2 billion events in one week without operator intervention. this video covers Kinesis stream ingestion, bronze-silver-gold medallion layers, fan-out architectures, metric views for consistent KPIs, monitoring, and optimization strategies.
🤝
Chapters
00:00Introduction: Inside SEGA's Lakeflow Journey00:55SEGA Scale: 140 Games, 40M Players, Petabytes04:14Data Services Team and Data Ingestion08:06Use Cases: Segmentation and Personalization09:56Migration Journey: From Legacy Scala to SDP12:25New Architecture: Bronze, Silver, Gold Layers15:23Gold Layer: Materialized and Metric Views16:43Game Launch Lifecycle and Planning18:52Football Manager 2026: Launch Deep Dive20:46Results: Scale, Labor Savings, Monitoring22:21Lessons Learned: Partnership and Best Practices
FAQs
Why did SEGA migrate from their custom Scala Spark pipeline to Spark Declarative Pipelines?
SEGA's custom Scala Spark pipeline was difficult to maintain and required significant manual effort to handle schema changes and new event types from each game title. Spark Declarative Pipelines on Databricks provided automatic schema evolution and a more maintainable architecture, reducing deployment labor from weeks to hours.
What scale of data does SEGA Europe process through its Databricks streaming pipeline?
SEGA Europe's data services team ingests 40,000 game telemetry events per second with sub-minute latency, supporting analytics across 140 games and 40 million players at petabyte scale. During the Football Manager 2026 launch, the pipeline processed 2 billion events in a single week without operator intervention.
What is the medallion architecture SEGA uses in its streaming pipeline?
SEGA's pipeline uses a three-layer medallion architecture: a bronze layer for raw Kinesis stream ingestion, a silver layer for cleansed and structured data, and a gold layer with Materialized Views and Metric Views that provide consistent KPIs for use cases such as player segmentation and personalization.
How did SEGA use Databricks during the Football Manager 2026 launch?
During the Football Manager 2026 launch, SEGA's Spark Declarative Pipeline ingested 2 billion events over one week without requiring operator intervention, demonstrating the resilience of the new architecture under launch-scale load. The real-time data enabled the data services team to monitor player adoption and game experience throughout the launch period.
Full transcript
[00:07] Welcome everyone. Sorry that took a bit of time. We had a bit of a technical glitch and now it's all sorted. Um, good morning. So, uh, thanks for joining us today. And joining our session. This is inside our Lake Flow journey. We're going to talk about how Sega scaled their
[00:23] streaming analytics across different game titles. Um, if you were here last year, you would have probably, uh, seen the talk from Sega where they would have discussed about how they did the migration, you know, what motivated them. And and today, we're
[00:40] going to see what's happening, you know, what is the, uh, today's state of after the migration, how it's going. And we're also going to look into, uh, the game launch aspect, how this pipeline is actually helping them to do
[00:55] game launches at scale. So, uh, for today's agenda, yeah, we're going to look into who Sega are, um, and what is the scale of data that they actually dealing with, and who are the data services team within Sega. We have
[01:12] Anna from the data services team. So, um, we're going to see what they do, what kind of data do they collect, and what are the use cases that they actually power from the data that they collect. Uh, we're going to take a sneak peek into the migration journey again because I would like to revisit the
[01:28] topic and tell you how, um, important that migration was for Sega. And, um, of course, looking into the game launch. So, we're going to take the example of Football Manager 2026, uh, and, uh, walk you through how the whole game launch looked like, um, for within Sega.
[01:45] And of course, uh, we are going to share what the advantages were. And of course, we will be going to close it with some of the lessons learned, uh, from this migration. A bit about me. I am Sanjay Ashok. I am a solution architect at Databricks. I've
[02:01] had the pleasure to work with Sega for over a year now and I've been in the front seat helping the team deliver some of the most interesting use cases on Databricks. Um but make no mistake, everything today that you're going to see is like, you know, fully developed by the data
[02:17] services team and they're the ones who are actually maintaining it. With that, over to you, Anna. Hi everyone. Um my name is Anna Leventchuk. I'm a data developer at the data services team over in Sega Europe. Um Sega's uh
[02:34] internationally recognized, whether it's for the 80s consoles or Sonic the Hedgehog. Um but today we also look after some amazing IP um from a growing uh European studio base. Um the franchises
[02:49] who look after the franchises that hopefully you are already familiar with. Um Thank you. You can sit. Yeah. Cheers.
[03:06] Just to introduce you to a few of them, we have uh Creative Assembly's amazing Total War series, the latest uh iteration being the Warhammer 3. We Sports Interactive look after our Football Manager and Two Point Studios uh who are located a bit south of London uh with their amazing Two Point Museum
[03:24] um game. And Rovio uh is actually the latest studio to join our cohort famous for their Angry Birds. We also look after some Japanese IP over in Sega Europe, so we're responsible for things like localization and marketing
[03:41] to our European player base. Um you'll recognize hopefully some franchises like Yakuza and Persona over there. Um some of our very popular titles. What um is important to note is with such a wide range of studios and game IP, um,
[03:57] all of these studios operate differently at a different cadence. They put They have different requirements and put out their games, um, in a very, yeah, individual way. Um, so we're never serving only one product, we're serving many.
[04:14] Sorry. Yeah. So, who are our data services team? Um, we are looking after the central data platform over in Sega Europe. We've been, um, a Databricks uh, client for about 6 years now and are heavily invested in the platform. Uh, we ingest anything from sales data to, um,
[04:31] marketing preferences to, uh, game telemetry and we serve that to our end customers, which might be business planning, looking to see how the latest launch went, uh, game designers wanting to understand how our players are behaving, um, or marketing wanting to
[04:47] understand how to segment our users and, uh, market better. Thanks. Just to put some numbers to that and kind of give you the scale of what we're dealing with. Um, we're ingesting telemetry for over 140 games. Um, that's
[05:02] 40 million plus, uh, players, um, which is petabytes of data. At peak times, we are ingesting I should really get a hang of this guy. Um, at peak times we're ingesting 40K, um, events per second. Um,
[05:19] that's across 451 types. Uh, we're looking to serve that at less than 5-minute latency. Um, everything that we show you today has been built to be able to serve these numbers reliably. Um, and yeah.
[05:37] So, what Let's have a little quick look at our data. What What Why does it matter to Sega? Um, like I mentioned before, we're ingesting anything from sales to your um, marketing preferences to your in-game habits.
[05:56] We want to understand how far you get into that level. Do you get stuck anywhere? It really helps us instruct like game balancing in future iterations, progression, and just try to make your experience as good as possible. We also have a look at your UI interactions. So, do you make use of that little hint
[06:12] button at the top of the screen? Or do you just brute force your way through the level? Do you engage with the DLC announcement in the menu? How can we make sure that you're aware of it? Or maybe you're not. Maybe we just carry on. Ownership, we also have a look are you a
[06:31] franchise VIP? Do you engage with every iteration that we put out? Or do you engage more during the free weekends and try our games that way? Um No, all good.
[06:52] Yeah. Sorry about that. Thank you, Sanjay. Some of the data challenges that we face. So, obviously ingesting such a big so much data from such a variety of studios comes with its own set of challenges. One thing that we love to look at is our player reviews.
[07:09] We do read them. So, here we have a look at a Warhammer 3 review which the inbuilt Data Bricks AI sentiment analysis function classified as positive correctly. The two-hour museum very addictive to our eyes, very good glowing review.
[07:28] The function the inbuilt function thought it was a bit negative. You might come across a lot of Warhammer 3 reviews might mention things like death and war and kill. We need to provide our models with context. A lot of the games might appear violent to the
[07:45] AI functionality, but we want to make sure that it does not count those towards Yeah, the thing that we serve to our end clients. So, we would like it to We want We want it to know that that is a positive review for us at Sega. Yeah.
[08:06] Some of the use cases that we provide to our end users, so user segmentation, which I've already mentioned, super useful to both our marketing and um our game designers. So, we want to know how to best serve each cohort of our or subset of players.
[08:22] Are you a quick completionist? Are you a coaster? Are you a flirt? Um This really helps us understand who is playing the game and what may maybe they're missing from this iteration of the game. How can we make sure that the next one is better?
[08:37] Um yeah. Another one is player stats. Everybody loves a bit of community. Player stats, it really engage helps boost the community and adds a little competitive spin to the games as well, which is always fun
[08:53] for our players. Um the player stats also make it into our marketing. So, marketing love to kind of boost the emails or any kind of communication that we sent out with some
[09:09] personalization. So, I'm sure a lot of you are familiar with Spotify Wrapped. We try to do a little bit of that in our marketing side. So, yeah, look out in your emails. Hopefully, you'll get one, too.
[09:24] Um Sanjeev, do you want to cover AI BI? For sure. Yeah. Thanks for sharing that, Anna. So, this is where it all comes together, right? So, you have the AI BI dashboards, you know, that's the end consumption. You've seen we have a single govern layer that that's the data that's actually feeding into the
[09:40] segmentation, feeding into all the dashboards that's being consumed at Sega. So, um for instance, a live ops director comes and asks a question. So, uh they can get an answer directly on the data instead of them going and filing a ticket, getting an analyst to review, and then the you know, so that's
[09:56] going to take some time. So, um with that we're going to look into how the lake flow migration look like, right? So, this was a recap um what happened, but the last year talk I was mentioning, you know, we shared a bit more about how
[10:12] the migration itself happened. Uh but here I just wanted to do like a high-level recap so that you know, you're aware of um you could appreciate better where uh you know, how uh Sega's done this. What motivated Sega, right? So, if you look at the actual challenge that they're facing, they had um legacy
[10:29] custom Spark um streaming pipeline, and it was all written in Scala, of course, and there were a lot of bespoke code that was involved with the team, and that was actually made it really difficult for them to understand, right? It was hard to understand. Um it was
[10:45] quite complex. Um the team has to spend a lot of time in maintaining it. And it wasn't fully automated. For instance, a new event comes in the pipeline, the team has to spend a lot of time to actually rewrite some of this, redeploy. So, there was actually
[11:01] time and effort involved in in the old legacy architecture. So, there wasn't a good monitoring in place. So, uh there were no the team wasn't able to capture the signals, understand how the uh data has been growing. So, uh which actually made the whole system a bit rigid as
[11:18] well as expensive. At some point, they weren't able to audit what exactly was happening, right? And that's the point when actually say I Sega had the question. So, how do we future-proof our data ingest? And the answer to that was moving away
[11:35] from the old legacy architecture to Lake Flow Spark declarative pipelines. So, um the team has been able to, you know, they they did the migration 2024. Uh since then they've been able to deploy pipelines at scale. They're able to
[11:52] easily monitor. And um all the Scala code that the team really didn't want to maintain had a hard time in maintaining is gone now. Uh and they are every new events that's coming from the studios, it's all automatic. So, there is no manual intervention involved.
[12:09] Um And this is exactly a future proof architecture looks like, right? So, fast forward to 2026. Now, if you take a look at what how the architecture looks like, um this is a classic Medallion architecture. So, you have three layers.
[12:25] I mean, bronze, uh silver, and gold. Uh the um yep, and and all the dependencies and transformation are maintained like, you know, managed declaratively. So, let's actually double click into the bronze layer itself, right? So, how does
[12:41] the pipeline look? So, we've taken some screenshot of the actual pipeline, um the the prod pipelines. So, um every new event that comes into the bronze layer gets processed and it's routed onward automatically, right? And
[12:59] um Sega reads this directly from Kinesis uh with a streaming live view. And there's also legacy data that um Sega Sega ingests. And all of these are preprocessed onto the same format. And um the key point here is adding a new
[13:15] event doesn't require rewriting the whole pipeline. So, um and in the bronze bronze layer, they also do event filtering, um on the fly encryption, stream grouping. So, all of these are happening in this layer,
[13:30] right? So, once this is done, they we move to the uh silver layer, which is a classic fan out architecture. So, what happens here is all the events that we receive from bronze are converted to um
[13:47] a table per event type style architecture. So, essentially every event that actually comes in is being fanned out to uh a table in itself. So, um this is absolutely critical um because this is what makes it consumable
[14:03] at downstream, right? So, all these tables makes it easy for the team to build gold layers um at downstream. So, and and if you look at generally how the data grows um at Sega, uh it's purely
[14:20] through the studios, of course, right? So, we have new event that's happening, we have a new um let's say game launch that that's happening. That actually increases the total number of events that the pipeline uh is going to ingest, and thereby uh you know, so the
[14:35] architecture actually treats the schema as additive. So, uh what do you mean by that? It's you know, you find a new column in the payload, add it to the mega schema, and the table grows uh to accommodate all the new event that the studios send to Sega.
[14:51] And uh in the schema evolution step, um what Sega does is like, you know, they sample the latest record per event and inspect it for uh any new column that's been actually sent from the studios. So, thereby, you know, you are not losing anything from, you know, what you get
[15:07] from the studios. At the same time, you're also ensuring that you're able to manage uh those volumes of data uh that's coming in, right, at scale. And uh of course, touching on the gold layer, so Sega uses both materialized
[15:23] views and metric views. So, they use materialized views for performance. I mean, the pre-computed cached data, so that's directly fed into the dashboards. And they also use a heavy heavy adopters of metric views, which is the
[15:38] in in the semantic layer, they build all the KPIs to standardize the KPIs, which are consumed consumed downstream by the dashboards and Genie. So, um governed it's it's like, you know, consistent definitions. For instance, if you ask anyone in the company ask what
[15:54] and who who is an active player, you would get the same answer. So, you know, from from the end consumption. So, um Sega hosts all its BI capabil- capabilities on Databricks, and they're also a big users of Genie. They have democratized
[16:10] the self-serve analytics within the company. So, anyone from any department can just come and ask a question directly on top of the data in plain natural language and, you know, get the right answer.
[16:25] Cool. So, we have seen how, you know, how the migration went. What's the new architecture now in place? Um now, let's come to the point where we need to ensure this architecture is working for them, right? How do you stress test this? So, there is no better way to stress
[16:43] test a streaming platform uh rather than have, you know, putting it through the game launch itself. So, if you look at the game launch life cycle, I would like to show you how it looks like. So, the team typically sits with the studios who are
[17:01] um gathering, you know, developing the game itself. And if you ask me when the whole involvement happens, it's actually right from the start when they're designing the game. So, the team is involved early on. They sit with the studios to gather the business requirements. This could be
[17:17] something like a new button, a new game mode. It could be a new in-game helper. And then they gather all of this and Sega internally decides where or you know, which KPIs to capture, like which metrics to capture.
[17:32] And then once once they do that, they move to a sandbox testing environment, right? So, they take all those business requirements. They for instance, something like even category okay. Yeah. So, even category they look into the data formatting and
[17:50] they also look into if if there's anything that's missing from the the the from the requirements and ensure that what kind of volume are we going to expect for the new release or you know, new game update. Um
[18:05] with that being said, I'm really interested treat this like a launch command center, right? So, that's how Sega operates, you know, working with the studios. With that actually, I'm really interested to know how the launch command center look like for Football Manager 2026.
[18:21] Sure. Um yeah, let's let's have a look. Um So, for every launch that we do, we make sure to as mentioned, test everything in sandbox. For Sports Interactive's launch of Football Manager
[18:36] 2026, we it was the first launch that we did post migration to STP. Um what we did is hook up our development build to the events. We find somebody in Sports Interactive, preferably more than one person to go
[18:52] and play test the game for about 2 to 3 hours. Um and we see what comes through. Are any events missing? Are any fields missing? Are any fields malformed? You'd be surprised. There's There's a lot to review.
[19:07] Then we want to know what would what would that look like for 500,000 players. So, we scale it and we kind of look for any bottlenecks, we look for any issues. Once that's all done and we're happy with the with the development built, we do the same with production. So, we do the exact same
[19:22] cycle with production and once that is ready, we're ready for a launch. Great. Thank you. So, the game launch Yeah.
[19:38] Once the game launch is a go ahead, we're ready to change our what we're doing once again. We're now doing live monitoring, live debugging, working together with the studios to kind of find any bugs in the
[19:53] first week or so just to make sure that the launch runs as smoothly as possible. So, we're we're yeah, looking after fluctuating scale, make sure that the streams stream groups are balanced. And that is that is what our game looks like.
[20:14] Just to again give you some numbers in the first week of FM 26, we ingested over 2 billion events. That's That was 79 distinct event types and a fan out to over 100 to 166 silver tables. Day one
[20:30] alone saw 250 250 million events coming through and that peaked at about day six to about 325 billion. Wonderful. Thanks for sharing that. So, concretely, if you look at what did Sega
[20:46] gain by doing this? It's it's actually four things, right? So, the first one in the old system, this spike would have actually crashed the drivers since the throughput is quite high and it would have required the team to
[21:01] actually look into what's happening in the pipeline, a lot of manual intervention that we saw earlier. So, if you look at it now, there's a rebalancing process in place and no single pipeline actually gets too overwhelmed. So, no two big events are sent to the
[21:17] same pipeline at this at once. So, there's auto scaling. So, we automatically scale once we see there's a spike that's happening. One of the crucial advantages that we saw was in labor hours. So, previously we would have two plus devs working around the clock with the
[21:34] studios at least a week before the launch and then a week after. With the launch of Football 2020 Football Manager 2026, we saw one dev working for a couple of hours a day and that's a really really big change for us. Um Sorry.
[21:49] Yeah, which is amazing, you know, like from several engine to just a few hours. So, that's that's a great improvement. For sure. Also, the monitoring which comes pretty much out of the box with Databricks allowed us to have a much smoother deployment. We able to
[22:06] monitor from the get-go. Sports Interactive was the and we were able to yeah, find bugs as they came. All right. So, let's look at the key to success. So, Anna, would you like to share
[22:21] what are the lessons learned and if someone's going to do this migration now, what are the key takeaways that they should you know, look for? For sure. So, these are some of the very important things that we learned while doing the migration. One of them was make sure to partner up. We were really lucky to work with Databricks partners
[22:37] called Advancing Analytics who sent us two great guys to help us do the migration. They were already SDP experts and um were able to kind of hold our hand through the migration and it meant that we did it in much quicker and smoother than we would have
[22:54] done it on our own. And I think it took us about yeah, just over a month to do the whole thing. Um second is keep an open dialogue. So, we're working very closely with the studios uh to plan for these launches. Um, we want to know um what events are going to be coming
[23:09] through, but as often happens about a week or two before the launch, um there'll be like so many more requests coming through. Can we change this? Can we do that instead? Um, do we have capacity for another event? Which with SDP, we do. Um, third is um
[23:26] make sure to use um DABS. Um, so the auto automation bundles from Databricks were really, really useful for us to make sure that we're um adhering to best practice with CICD. Um, it is very tempting to go into your notebook and make that one quick change and then kind of like see how it goes, but um we don't
[23:43] want anything uh going into production that we um don't want or hasn't been tested. So, uh using DABS was a really big plus for us. Um, also invest up front. Um, I know it's the case for many companies or many um teams kind of considering any big
[24:00] migrations with within Databricks uh to kind of like potentially wait until the last moment possible just to make sure that um it is really, really required. Um, we were really happy that we did this up front. We didn't uh wait. We kind of took the risk and it really, really paid off. Um,
[24:17] we're really happy that we did the migration when we did and had gone through the iteration of testing it on Football Manager 2026. And with a really big game coming out um Warhammer 40K kind of on the horizon, we are sure that we're ready and we're really excited to see the numbers coming through.
[24:34] Um, also make sure to optimize. We made sure to monitor the newly migrated pipeline for at least a month or two afterwards. Um, we still do, but the first two months were really, really crucial in trying to find out um any ways that we could save costs. We actually found a bug which meant that
[24:49] one of the events or sorry, a few of the events were assigned to only a single event stream. Um, so finding that, fixing that meant that we saved a lot of costs. Finally, to document and all, I feel like with the junior ontology and Genie just being
[25:05] able to kind of explain everything to you in simple terms has been such a big game-changer, but making sure that you document, it will not only make Genie better, it will you will also get thanked by your junior devs cuz you'll get to upskill them much
[25:20] easier. Yeah, thank you. So, from a brittle legacy pipeline that we had to one that can handle a 2 billion event launch week. We're really happy with SDP and thank
[25:37] you guys for listening. Yep. Thank you. Yeah, and don't forget to use the survey fill the survey in the
[25:53] app. Yeah, thanks again.
Learn more about the Databricks Data and AI platform.
The information provided herein is for general informational purposes only and may not reflect the most current product capabilities or configurations.