Skip to main content

Databricks at Massive Scale: How Supercell Powers 300 Million Players

Summary

  • Supercell, a mobile game developer with 300 million monthly active players and approximately 5 petabytes of data, migrated from self-hosted Spark on Kubernetes to a managed Databricks Data and AI platform with serverless compute, Delta Lake, Unity Catalog, and Workflows.
  • The platform team's mission is to make data accessible to every employee, enabling game designers, data scientists, and engineers to perform analytics, data engineering, and machine learning without managing infrastructure.
  • Supercell's data science team uses Databricks to build and operate production-grade machine learning models for player safety, including a layered detection pipeline that identifies child safety signals and grooming behavior at global scale.

Databricks at Massive Scale: How Supercell Powers 300 Million Players

Watch: Databricks at Massive Scale: How Supercell Powers 300 Million Players
Scaling data infrastructure to support 300 million monthly active players requires more than just bigger storage and compute. Supercell uses Databricks to democratize data access, manage massive ingestion pipelines, ensure consistent governance, and enable everyone from game designers to data scientists to make hypothesis-driven decisions.
Hear from Boris Nechaev, Head of Data Platform at Supercell, about their infrastructure journey from self-hosted Spark on Kubernetes to serverless Databricks with Delta Lake, Unity Catalog, and Workflows. Ilari Vaha-Pietila, Data Scientist, shares how they build production-grade machine learning models for player safety, detect child grooming and high-risk signals through layered detection pipelines, and balance rapid innovation with the operational rigor required at massive scale.
🤝

Chapters

FAQs

Why did Supercell migrate from self-hosted Spark to Databricks?

Supercell was self-hosting Spark on Kubernetes, which required significant infrastructure management that distracted the platform team from delivering value to the business. The Databricks proof-of-concept focused on managed Spark and notebook capabilities that would reduce operational overhead and accelerate data access across all teams.

How does Supercell use machine learning for player safety?

Supercell's data science team built layered detection pipelines on Databricks that identify child safety signals including grooming behavior among their 300 million monthly players. The models are orchestrated through Databricks Workflows with MLflow handling model lifecycle management and production monitoring.

What does Supercell's data platform architecture look like today?

Supercell's current platform runs on serverless Databricks compute with Delta Lake for storage, Unity Catalog for governance, and Workflows for production job orchestration. Their approximately 5 petabytes of data supports everything from game analytics and data engineering to production machine learning models for trust and safety.

Why don't simple detection approaches work for player safety at Supercell's scale?

At 300 million monthly players, simple rule-based or single-model approaches generate too many false positives or miss the edge cases that only emerge at massive scale. Supercell's layered detection pipeline combines multiple specialized models that progressively filter signals, making the system both accurate and operationally sustainable at global scale.

Full transcript

[00:07] Hello everyone. Thank you for joining us today. Um my name is Boris. Uh I'm from Supercell. I'll uh say a few words about our company on the first slide. And I will share with you our journey with uh Databricks. And then also I have Ilari here with me on stage
[00:23] who will talk about uh the ways that we create safe playing environment for our our players. But yes, promised uh Supercell is a mobile uh game developer located in um in uh Helsinki, in Finland. Uh we are about 16 years old um
[00:42] and we now have uh six live games. You can see some of the logos. Hopefully some of you had a chance to play our games at any point. I also have stickers by the way. Find me later if you wish to have a sticker that you could put on your laptop, let's say. Um so yeah, we are around a thousand
[00:58] people now. And as also to the title of the presentation said, we have around 300 million monthly players worldwide. So we we get our um our games played from across the world. Uh I was just at the um at the live stream with Ilari
[01:15] Kaplan and he asked what was sort of like, I don't know, the um the the the biggest surprise, you know, on on the geo side of of our demographic. And the the funniest thing in the world that I remember is that once we found a data point which corresponded, you know, with a geo code that corresponded to
[01:30] International Space Station. I am not sure if it's true or not. So I don't know, maybe you know how data goes. Sometimes it could get buggy, but I don't know, maybe someone didn't want to miss their battle up there. Um So I'm with the data platform team.
[01:46] Uh this is our mission. We aspire to make it easy for everyone at the company to work with data. And when I say to work with data, it's uh doing various things. It's of course asking uh data for, you know, something getting some insights to help them make decisions.
[02:02] It's of course to transform the data to get data in from various sources to transform the data into a useful format. Good old data engineering. It's to consume data one way or another. So all pretty much all of it that that we facilitate from within our team. So
[02:17] we have maybe around 5 petabytes give or take, you know, sometimes we do a spring cleaning and suddenly it's less, sometimes it accumulates a little bit more. But give or take and yeah, we we run a bunch of ETL and have a bunch of production tables and and many many dev
[02:32] tables. So at Supercell we are mobile games company and of course primarily what we do is games and vast majority of our data comes from our games. When a player plays the game we we can collect sort of what they are doing especially from the server side. There is a client side on the device and
[02:49] there is server side. So we prefer to collect from the server side because it's much more reliable. But then besides games we have many other domains which sort of feed the data into our data platform. It's player support and finance and trust and safety
[03:04] and Hillary for example is from the trust and safety team and you will learn more about that later today. And marketing and our Supercell ID which is the account service and you know, community we work with influencers, YouTubers and what not. And of course security. So some of the announcements today that we heard on the
[03:20] security side are quite exciting for us and we will look into that later. So yeah, we try to optimize for convenience, speed and openness. Openness for us means basically that everyone in the company has access at least read access. So we do not give right access.
[03:37] To the left and to the right, but at least read access to data bricks or whatever dashboards we have so everyone can can work with the with the data. And we want of course to make it fast and and and and very convenient. And of course we try to activate the
[03:52] data in some form. So be it to facilitate better decisions or to drive growth and whatnot. So this is us in the data platform team. Then a bit of a backstory. Oh my actually my fonts have broken. I should have exported a PDF maybe. But
[04:09] backstory so how how we got started and what was before Databricks for us. So we ran a we used Hadoop you know in the good old times MapReduce was a thing and and Hadoop with I think the the language was Hive and
[04:25] Pig was competing one. So like long ago we were doing that more than more than a decade ago we were doing that. But then Spark came across and we quite quickly realized just how advantageous it is to to Hadoop and how much more and how much
[04:40] better things that we can do on Spark. So we became so believers in Spark. And then we thought okay how how do we do Spark and then there were a few engineers and Kubernetes was a new and sexy thing and Spark was a new and sexy thing and they were like let's put Spark
[04:56] on Kubernetes. So we did that those few engineers did that and in addition then data scientists and analysts were like can we have notebooks everyone you know was you know with notebooks so we self hosted Jupyter notebooks. Zero out of 10 I cannot recommend
[05:11] self-hosting pain a lot of pain and a lot of suffering. We were stuck like again another anecdote is that we were stuck on a very old version of Spark for a long time just because those two engineers of course you know how it always happens they moved on to other projects and teams and whatnot. And
[05:28] there was a ungodly amount of Pearl scripts that were somehow binding all these together. So like all the best things that you can imagine that go with self-hosting so it it was fun in the beginning but then not fun at all. And then we started thinking right we
[05:44] survived with this for a year or two and then when it became impossible, people were asking, "I want new version of Spark. Can you give me a new version of Spark?" And you're like, "No." Then that was was not not a good answer, of course. So, then we started thinking, "What do we do?" And then we started looking at the at
[06:00] the sort of market broader market. Uh and we were believers in Spark in Spark, right? As I said, like at least in my opinion, it still remains one of the most scalable uh technologies for data processing. So, we started thinking, "Okay, who knows Spark well?" You know,
[06:16] and there is a company founded by creators of Spark out there. Um so, we we started talking to Databricks, uh ran a POC. Um number one goal for us, we outlined I have some email somewhere where we outlined uh what we were after with a
[06:31] with a pre-sales engineer who was working with us. So, number one was managed Spark. We were just like, "Please solve the pain of self-hosting for us. We will never want to self-host anything like this again." Uh and of course, you know, in the process, unlock the scale. We knew that
[06:47] uh that there will be more data, you know, data coming in from from all sides, so we wanted to make sure that we are future-proof in our ability to to work with data. Then notebooks. Uh yeah, Jupyter was as great as Jupyter was, uh still there were quite a bit of
[07:03] things, you know, there's like rough edges and someone always needed to duck duck tape and you know, bubble gum things together. So, also we didn't want to uh self-host Jupyter notebooks. And Databricks came from the get-go with a native notebook experience. Uh so,
[07:18] this is our second the second thing that we were after, and this made our analysts and data scientists very happy. And this was a big step in data democratization, uh making it possible for everyone in the company to work with data. Um so, it helped us a lot. And then third uh was we knew that there was a
[07:35] lot of things already at that time in 2021 uh coming on Databricks, and we knew that many more things will be coming uh in the future. So we said, "Okay, let's figure out the first two things and then we will, you know, start looking into the rest of the ecosystem and start building
[07:51] on top." So yeah, so then the POC was successful and we proceeded, as you might have guessed. But at the time at the same time at that time we still had an external orchestrator. Yeah, the files were just, you know, in S3 as external tables, parquet files, not Delta. That still was in the future
[08:08] for us. Meta store external meta store, it was Glue for us. And yes, we only used all-purpose classic compute and jobs. So no serverless. I don't think serverless was a thing back then in 2021. I think it came later. Then what we did,
[08:23] so after the POC was successful, we started laying down the foundation. So if like a year or so ahead, we converted all our parquet to Delta with all the nice things that Delta offered. We quite liked them. We ran a few tests and we quite liked what we
[08:39] what we saw. So we definitely decided that it would make sense for us to convert to Delta. And at the same time of converting all parquet files to Delta, we migrated to UC basically. So we deprecated the old meta store and we just registered
[08:55] everything in in UC. So these two things went hand in hand for us and I'm actually quite happy that we did it that way. At least it it did not cause us any sort of disruptions or or troubles. And doing one not doing another would have been
[09:12] in some sense, you know, not maybe a good choice for the future because we would have anyway needed to do this in the future. So I think these two migrations hand in hand made a lot of sense for us. And yes, as as I as I write here in slides, you can see what this unlocked for us, those those two
[09:28] things. Then yes, so this this setup really really good foundations for us and then we started looking at other things. Ecosystem kept evolving, new things were coming out. Workflows, so that that is where we got rid of the external orchestrator and
[09:44] just moved everything to workflows and and and jobs. And and still very happy with that. Serverless came around and we were like, "Ooh, like not only we don't need to self-host anything of course at this moment at this point, but also we wouldn't need to specify, you know, the
[10:01] size of the computer that we need and whatnot." With serverless, you just sort of throw in the workloads and it just on its own magically and correctly and hopefully optimally figures out what what what needs to be done. And then yeah, this year we were converting many of our tables from
[10:16] external to managed because that also unlocks that kind of like a important building block for many of the other things. And yes, in in combination all those things helped us save money and made things faster. I have some some numbers here on
[10:31] the slides. And massive complexity. Again, like already during self-hosting there was a lot of complexity that went away of course quite early, but even after that we had, you know, like Docker images and you needed to think about the lineage and which which image inherits
[10:47] from which image and and and what is being installed, all the dependencies and whatnot, and how we expressed pipelines, how we did CI for all these pipelines, and you know, startup types was was were not where we wanted them to be without serverless. So yeah, this this
[11:04] unlocked quite a lot of good things for us. Um then maybe this is more or less where we are right now. So we can do a lot of cool things. MLflow and I think you will hear from from Ilari a little bit on that.
[11:20] Then yes, we can use GPUs for, you know, training training things. And like what what what for example we do with this is Uh, have sort of like a headless uh, game client and we can use it to train
[11:36] artificial bots basically to play our games. So, of course that that requires a lot of compute so this is reinforcement learning and requires GPUs so we could use that. Yeah, then we have quite a lot of AI BI dashboards not exclusively we also use Tableau
[11:52] but quite many AI BI dashboards now exist for us in in Databricks. Yeah, we take we make use of pipelines and streams to get data in to monitor quality to do real-time calculations and
[12:07] real-time calculations so we can you know track the stats on on most important KPIs such as new registrations, new players and and and revenue. Yeah, app adoption is growing quite explosively for us. Lots of people are excited especially with the advent of
[12:24] white coding suddenly everyone can just oh, I always dreamt of having something like this and then they can back Gordon. We have guidelines and and sort of ways of deploying that and then Dabs of course lots of things are are are are done orchestrated by and deployed by
[12:40] Dabs. And yes, liquid clustering and we are subscribed to a lot of preview features so very excited about many of the announcements also made today so we are trying to stay ahead of the curve and experiment. And of course governance model of course UC enables very solid gives foundations
[12:56] for the governance model but the company keeps growing complexity keeps growing and we always need to sort of like iterate and and try to future proof how do we do governance on the platform. And yes, this is maybe a little bit forward looking into the future how what what we
[13:13] could could still do in the future and what we sort of where our thoughts are right now forward looking sort of we we now have the capability to do pretty much whatever you want with the data but now comes the question of how do we actually activate it? For example, how do we personalize our games
[13:30] for our players to have a sort of more bespoke and personal experience and cater to the players better and in the process hopefully make our games more fun. We are doubling down on experimentation, AB testing. We want to be hypothesis driven. We want to learn. We want to
[13:46] create this virtuous learning loop where we set out a hypothesis that we believe that this this feature would would make it more fun for players to play the game and improve let's say engagement metrics or how how long the session length let's say or or number of logins
[14:02] and whatnot. So then we experiment and then we learn from that and then spread this learnings across all games. Yes, player harm prevention again you will hear that from from from Ilari. Then we are we are trying to sort of now now somehow the complexity a little bit exploded and we are trying to get our
[14:19] heads around observability. Also also was was was happy to hear about some of the of the things that are upcoming from Databricks side on this and we are building certain things is like everything from cost observability to resources, how do people use the platform, what do they deploy and and
[14:36] whatnot whatnot. And of course, spring cleaning is something is a new so we can have a good grasp of of what can we sort of deprovision. Then yes, AI is of course you know an elephant in the room and we are thinking hard on on what is the right way for us to to embrace that.
[14:53] Knowledge base and decision companion. So the high level vision for us is that we can with with the help of AI and and the data that we have in Databricks, we can give everyone in the company sort of a decision companion and they can
[15:09] sort of talk to the companion, of course AI companion and it will both take the data that is needed for a certain decision. It will tap into sort of company brain like what is the the knowledge in the company, and maybe
[15:24] hopefully account for biases because when when humans make decisions, we are all prone to multiple biases. So, hopefully this uh companion would account for that, and then again, help everyone make as best decisions as possible. Um then, 360 view of player is a little bit on the CDP side. Also,
[15:40] there was uh something about that today in the in the keynote, and we're excited about that. So, we are looking on how do we collect that data sort of in the uh structured in the right way to enable marketing use cases. And yes, we're very happy to be with Data Bricks on this journey, and very happy to see new
[15:57] things upcoming and experiment, and then yes, see where this takes us. All right. Thanks, everyone. This was my part. All right. Uh thanks, Boris. Yeah, that was the the platform. Uh
[16:13] Then, a bit more on how we actually utilize this platform to protect our players. So, I'm going to go into a one specific use case, and how we can scale it out to the full full player base. And just for scale, 300 million monthly users,
[16:30] uh population of US is somewhere close to 350 million. So, we're dealing with that amount of scale on on on this platform. So, uh that also drives a bit on few of the decisions we've made with the kind of structure, how we do things.
[16:47] But before we go there, uh bit of what is actually our team's job. What do we try to do? What does Trust and Safety try to do? We try to keep our players safe. We uh try to protect in integrity of our
[17:02] games, and we try to encourage and foster positive player behavior. And in the slides, there's the part of reducing risk of real-world harm, and that is the one case I am about to dive a bit deeper deeper in
[17:19] this how we actually do this kind of protecting from real life harm. But before we go into that, uh there are few key design principles how we actually uh
[17:35] try to do these kind of things. So, first of all, we use different types of signals and systems uh layered systems to identify high-risk safety signals, which means not a single model, not a single signal decides what happens. It's always a combination of different types of
[17:51] things. We also minimize what is escalated for review, so we don't make our human workforce, for example, read unnecessary content because the content is usually very heavy heavy to read on. And of course, like the final call is always on trained subject subject matter
[18:08] experts, so they have the proper training and they know what they're doing on this. And all of this comes from the kind of from the first design phases, we always design these with the privacy, security, and legal uh obligations in mind.
[18:24] And build in the kind of how we would then improve the whole system. How would we scale this out? What kind of pilots we do to get this rolling? But on to the case. So, uh one concrete example is how we use uh
[18:42] data to detect high-risk child safety signals or how we use data to detect uh child grooming. Uh it is one of the hardest problems that our team actually deals with. Uh if you just think about the problem itself, as well as uh in a mathematical or
[18:59] machine learning type of sense, you're dealing with basically anomaly detection from the whole player base, and you have to do it in a multilingual setting. You cannot deal it just with English. You have to support all of the languages we do. So, uh it is very difficult
[19:16] problem to solve. And I mean, sure, we've done this before. One of the uh one of the things uh we did before we were relying a lot on on user reports. And don't get me wrong, user reports do a lot of the heavy lifting on our trust
[19:31] and safety operations. But especially on the child safety signals, there are two big reasons why this doesn't work out as a only signal. So, of course, bad actors, there's some in the game. There's always
[19:47] some in any kind of community. And well, would bad actors report themselves? Obviously not. That's uh reason number one. And of course, the second reason is if you talk about the grooming, the interaction is built to
[20:02] feel like a friendship, like a trust. So, the person in risk doesn't usually know what is happening. So, therefore, we cannot just wait for reports. We need to do them proactively and detect those proactively so we can take action.
[20:18] So, how you would then uh start to do with the such a high scale? Well, you could try keyword searches. Uh that was the kind of first naive idea how we start to do it, but we very soon found out it is not a good idea. Uh
[20:34] because, yeah, keyword search, while it is a valuable signal, it is a bit too noisy signal. So, it is very you get a lot of a lot of kind of noise within your signal. You need to be very
[20:49] precise. But that isn't enough for trying to kind of reliably detect something. But yeah, it's year 2026. So, why are we using keywords? What about we go with the other obvious? Why Why we just use, you know, LLMs?
[21:08] Go there and uh let them flag the risk. That's pretty obvious, right? Yeah? Well, uh we were talking about scale. And we were talking about anomaly detection. So, when you talk about scale and anomaly detection, you're like, all
[21:24] right, how much waste will it be to run everything through big LLMs? So, you cannot go because majority of the content is non-risky. And if you would go in the this route, you would hit the problems with operations, so latency, maintenance,
[21:39] scaling. And of course cost. It would cost a lot to run this kind of full full volume. And then for auditability and privacy uh considerations, it is very they are mostly black box. It's very difficult to know what exactly
[21:56] caused the decision within LLM. So, you will lose a lot of auditability if you just go with just this as your sole signal. So, what we do, we use layered approach using different types of signals. So,
[22:12] you can think this as a kind of funnel going from the top where we use a lot of uh fast models, fast uh heuristics, and going deeper and deeper into more precise, more compute heavy problems. So,
[22:27] right at the top of the funnel, we use techniques such as keyword searches, uh player features, basically something that could be described as a SQL kind of uh operations. And then, after we get the signals from
[22:43] those, we aim for high recall on the first part. So, which means we want to retain as much as the violating parts while discarding some of these. So, we optimize for recall in the top. Then, when we go into further in the middle, we go into the more specialized
[23:00] models using more compute. We use embeddings, GPUs, we use uh gradient boost models in there. We use cosine similarities, all that kind of cool stuff. But those will be the more specialized one that will detect specific risk. They
[23:15] will detect certain kinds of features from the uh from the risk that the previous step detected. And then finally, we send the most the riskiest ones to the human review.
[23:30] That will make the kind of final call. Was this actually violating or not? But the key point is that we try to aggressively drive down the volume that will go on the further and further down into our into our pipeline. So that we're responsible but as well not wasting too
[23:46] many resources and have the scale that we can do to compute on. And this is kind of where Databricks comes in in in a way. So the recall optimized small models, they are if they're models they're saved in Unity Catalog with the
[24:02] uh with the versioning, those kind of things. And we call them through there using mostly serverless architecture, the one that Boris described earlier. For the ones we go a bit deeper with heavy compute, we use GPU clusters uh for those ones. And still we use the
[24:18] models that we have saved in MLflow using MLflow in Unity Catalog. And then the larger models, the the big multi-hundred billion parameter models are then the high precision ones that only touch a small subset of all the
[24:35] content we go through. So this is mostly like a game of economics, how much you can compute, how much budget do you have for the time scale, latency, those kind of things.
[24:51] And then with the orchestration part, we're not using Jupyter notebooks per se. We're using Daps to control our flows, our retraining pipelines, inference pipelines. And uh
[25:07] every time we process something with human labeling, we get more training data. We have scheduled retraining uh flows that will then improve the model straight away. Uh we use GitHub actions to control the CICD so that we get the split between
[25:24] environments and then we make the production uh deployments there. And then, of course, the nice thing about Daps in this one is that we get the the dashboards and alerts from the pipeline are built in. So, when we deploy the model, we deploy the
[25:40] dashboards at the same time so that we will get the kind of full monitoring at the same time. And sure, you kind of need to go into something like this or this kind of route because now you're dealing with the production grade
[25:55] uh safety system instead of just, you know, random notebooks. So, to make this maintainable and so that you can have a team working on this, you need some kind of structure in there. And how we actually went from those first trial notebooks
[26:12] to the full grade system was pretty deliberate. We had a lot of gates. We had a uh pilots trying out different things if they are actually working. So, the first ones, they were actually Jupiter Jupiter notebooks and we were trying out does this even does this work? Uh
[26:29] And yeah, first pilots were done only in English in a very uh limited subgroup of of players. After we validated, all right, this is working. This is working better than our previous version of of way of doing, then we started to look at, all right, should we actually try to scale this
[26:46] thing forward? And yeah, it's uh at that point, we also realized, oh, we actually need to pay a lot more attention to how we train our our agents who will actually
[27:02] deal with the final result of the pipeline. So, we partnered with the experts on delivering the full training, developing that. We optimized a lot of the tooling we use for reviewing these kind of reports. And sure, when dealing with scaling from single language to
[27:18] multiple different types of languages or markets, we needed a lot more hardening on the scaling parts of the pipeline. And yeah, we are we expanded gradually from just first from Euro languages, few games, to actually
[27:35] operating on a global scale. So, for example, the market in Southeast Asia is very different than from US or from from Europe. And after those kind of we got it scaled out, then we start to optimize it very
[27:52] heavily. All right, what kind of different prompts can we use? What kind of model retrains should we use? Should we add more features to the models we use? And upgrading the infrastructure. So, this kind of in my opinion, you cannot go straight
[28:07] away, oh, I just turned this on and now I'm globally uh detecting safety signals. You need to kind of earn it through testing and making sure you're also the organization and the operations are ready for this, instead of just putting
[28:22] out a a model out. So, to kind of summarize uh all of this is that if you build a uh high-risk child safety signal
[28:39] detection system, it is a lot more than just the technical parts of the things. Sure, Databricks and uh all the tooling itself made it a lot easier uh compared to what it was in 20 uh
[28:54] 2017. I'm happy that I was not doing it back then. I was doing it now. But uh it also involves a lot more pieces that are within within the organization. You need to have the training in place.
[29:10] Uh across across different kind of regions. You need to be integrating the tools and processes. You need to uh be in touch with legal as well as security so you know what you're actually getting into.
[29:27] And big part is a big part is actually building the knowledge and resilience within the company to deal this kind of stuff. And one thing I also want to highlight that because we have this kind of layered approach
[29:43] from different types of algorithm diff- different types of pipelines the uh the quality assurance and auditability between each uh kind of step gets more and more important in
[29:59] in this kind of system.

Learn more about the Databricks Data and AI platform.

The information provided herein is for general informational purposes only and may not reflect the most current product capabilities or configurations.