Skip to main content

Databricks serverless SQL and Unity Catalog: Riot Games' analytics strategy

Summary

  • Riot Games achieved $45,000 in monthly cost savings by migrating from a legacy Databricks deployment to a serverless-first architecture using Databricks Serverless SQL Warehouse and Unity Catalog.
  • Unity Catalog enabled Riot Games to establish unified governance, data lineage, and auditability, shifting the team's philosophy from restricting access to actively enabling users with visibility and managed resources.
  • Transparent cost tracking through Unity Catalog and standardized DBT workflows empowered individual teams to own and optimize their own usage, reducing the operational burden on the central data platform team.

Databricks serverless SQL and Unity Catalog: Riot Games' analytics strategy

Watch: Databricks serverless SQL and Unity Catalog: Riot Games' analytics strategy
At Riot Games, scaling analytics required a fundamental shift in platform strategy: from restricting user access and managing infrastructure to enabling teams with visibility and managed resources. By adopting Databricks Serverless SQL Warehouse paired with Unity Catalog governance, Riot Games achieved $45,000 in monthly cost savings while improving user experience through faster query startup times and reducing operational burden on their small data platform team.
this video covers Riot's technical journey from legacy Databricks deployments through Enterprise Edition to a unified serverless-first architecture. Learn how Unity Catalog enabled governance, lineage, and auditability; how Databricks job clusters, SQL Warehouse serverless, and standardized DBT workflows improved performance; and how transparent cost tracking empowered individual teams to optimize their own usage rather than placing that burden on the platform team.
🤝

Chapters

FAQs

What cost savings did Riot Games achieve with Databricks Serverless SQL Warehouse?

Riot Games achieved $45,000 in monthly cost savings after adopting Databricks Serverless SQL Warehouse paired with Unity Catalog governance. The shift also delivered faster query startup times and reduced the operational burden on their small central data platform team.

How did Unity Catalog help Riot Games improve data governance?

Unity Catalog provided Riot Games with unified governance, data lineage, and auditability across their platform. This foundation allowed the team to shift from restricting user access to enabling teams with visibility and managed resources.

Why did Riot Games choose a serverless-first architecture for analytics?

Riot Games adopted serverless SQL warehouse to address challenges from their legacy Databricks environment, including slow query startup times and high operational overhead. Serverless compute eliminated the need for manual cluster management, freeing the small data platform team to focus on enabling business users.

How did Riot Games standardize data engineering after migrating to Databricks?

After migrating, Riot Games standardized their data engineering workflows using DBT alongside Databricks Serverless SQL Warehouse and job clusters. Transparent cost tracking through Unity Catalog empowered individual teams to own and optimize their own usage rather than relying on the platform team.

Full transcript

[00:08] Well, uh welcome. Uh my name is Ken. Um to kind of start it off, I just want to see a raise of hands. Uh how many of you are familiar with what Riot Games does and what we do? Cool. Uh I wasn't expecting that much cuz I
[00:24] thought we were not that familiar uh known. Um but to some of you who did not raise your hands, we actually uh owns the games of League of Legends, TFT, um Valorant, and recently announced 2XKO. Um so,
[00:40] uh if you have not played those games, I would assume you all have phones. Uh those of you laughing know what I'm talking about. Um no, but on a serious note, um if uh you haven't tried our games and you actually have a phone too that it plays
[00:55] on, and feel free to look up uh Wild Rift and TFT. Um for PC, we have League of Legends and Valorant and 2XKO. Uh both Valorant and 2XKO is are also on consoles. So, uh with another round of hands, how many of you actually played or still playing
[01:11] the games? Cool. Awesome. So, thank you for uh allowing me to be here cuz you guys are you know, paying the bills, but no. Um no, but seriously notes, uh you guys are like the players are what we focus on, so we really appreciate you playing the
[01:27] game. Um and as part of these sections, obviously everyone in here is my audience. Um what we are trying to do is uh share our journey um of going from like a very legacy uh data bricks environment to uh adopting
[01:44] serverless uh SQL warehouse and how we are going about that journey uh in the hope that, hey, what we share with you allows you to make uh you or your and your organization to make a decision on like how you wanted to adopt it. Um so with that,
[02:01] we'll get started. Uh one last thing notes, uh there will be a survey towards the end. Please do fill it. Um it will help me help Databricks so that we can do better next time. Um So yes, so today's topic is embracing serverless equals at Riot Games. Uh
[02:19] and we'll walk you through the journeys of how we get there. Um and to kind of formally start, my name is Ken Lane. I'm the engineering lead uh in a central data platform team at Riot. Um and I've been at Riot for um 7 years. All the time has been on the same team.
[02:35] And I have my uh partner in crime. Hi, my name is Demetri. I'm the staff software engineer at uh on Riot working on data platform. I've been at Riot for 13 years. Yep. So we're the OGs. So for those who knows. Um but yeah, so what are we uh on the
[02:52] data platform team? Um so we are part of a data foundation that kind of serve almost all Riot's data related uh side of things. So we own pretty much two primary scope. Um the first one is a data ingestion pipeline. So we do
[03:08] provide a uh centralized uh data collection services for all games and non-games um and get the data into a data lake. That is not the focus of today's. Um but if you interested in that and wanted to have conversation and chat after section, uh welcome to to hit me
[03:25] up. Um the other side of the teams kind of the ownership of on the team is actually administrating um Databricks uh as a platform. So that goes from um permission grants, uh infrastructures like provisioning infrastructures, um
[03:42] managing the lake, the catalog, um and all the way down to like how uh people are actually using Databricks and the platform side. So. Uh this is kind of the quote from Dimitri. Um we already have a lake house with Databricks,
[03:58] um and we're trying to make a lake home. And just to be clear, lake home is not uh official term from Databricks. It's not part of features. Uh for those with kids, probably will understand by thinking about it as, "Hey, now we have a house.
[04:14] Uh how can I turn start getting people into the house? Uh with kids, how can you govern the kids not to do weird things uh as a parent, right?" So, um cool. So, to start off, we'll do a little bit kind of a history of like
[04:30] where we how we got started it. Uh we start have adopted Databricks since 2016. Um at the time, we are one of the early adopters for Databricks. Uh so, we're on like single tenants, custom deployments, custom images,
[04:47] custom uh runtime. Uh so, it is to some degree it's convenient, um but at the same time it comes with a lot of like pain points, right? And at the time, we don't anticipate it a lot of users. Uh
[05:02] and and the focus is primarily on the data engineering and uh data scientist not as much as the lot for uh insights. Um some users of insight, but not much. So, because of that, we kind of do a
[05:17] um what we call the share compute model. So, we spin up uh two or three or multiple of like highly available over-provisioned clusters for them to use. Uh as you can imagine how wasteful that would be, but at the same time, Riot is a global uh company, so we have teams
[05:34] who are on uh East Co- uh West Coast in Europe as well as in Asia, like Singapore, um and Korea. So, um it basically the servers are has to be always up to serve the various demands
[05:51] um at any given time. Now, because it's a shared model, like shared compute, we have kind of two common issues. One is the noisy neighbor. Um I don't know how many of you familiar with Hive back in the days. It's a custom uh as a catalog. Um if you do a
[06:07] select stores limit 10, what would happen? Does anyone knows? Yeah, well. Uh in case you know, it would actually scan the whole uh all files in that table. So, if you have a large table, it would cause the whole cluster to be basically
[06:25] can't be used. And uh often time you have to actually have restarted. So, that created a lot of problems for users in the sense that uh when insights potentially not aware, not understanding how the underlying works, uh they're probably just with SQL background, um
[06:40] and just running a query that basically affect a lot of people. So, and we also often let people schedule production jobs on those clusters, and it will cause a lot of uh uh support uh on the team. Uh one thing I forgot to mention too is our team is
[06:55] very small. Um our team right now, even now, is like only five. Technically, five engineers to support um to manage Databricks as a platform as a whole. Um and we're serving right now about 30 to 50
[07:13] data engineers and data scientists, probably more than that now, uh and over 200 insights, um probably even more now. So, um So, for those who are just know Hive, uh there's really no governance. Uh and at the time we don't have a way of doing a
[07:28] environment separation. So, it's the just a single production, or everything is just doing dev, production, dev, production in the same bucket. Uh and then at some point we make a transition to Glue um first, so before E2. Um with Glue, it's a single catalog.
[07:47] It helps a lot, um, for us to manage from a catalog management point of view. But, what the pivot point is really on the E2. For those who are not familiar with what it is, uh, it's Enterprise Edition. It is a multi-tenant environments that you all now should be using now.
[08:04] Um, and when we make that transition, we purposely make specific, uh, changes to help kind of improve the overall areas of certain certain areas, but it also still comes with limitations. And at the time, Unity Catalog was not,
[08:19] I believe at the time Unity Catalog was in very early previews. So, it's not officially GA. There's lack of features that we just could use it at the time. Um, so, with the experience that we had before with the limitation, like if I go back one screen, like the noisy neighbor
[08:36] problems. Uh, one thing I did not mention is on the resource contention. Um, for those who know how familiar you are with the, um, the resources, every single worker takes up two IPs. So, it's, um, as people like getting, because we don't
[08:52] have governance, a lot of people get, uh, over permissions, like too much permissions. So, they actually spin up clusters and it will cause us run out of IPs. Um, so, with that kind of, with those things that we, um, keep in mind, so, we
[09:08] make some improvements around the access control specifically, along with like creating a shared compute model with, uh, team-based resources. What that means is, um, we, instead of a once or few single super large, uh, computes,
[09:24] uh, now it's assigned to team. So, with the use of taggings, um, to kind of assign like the actual team itself, uh, that started getting that visibility into like, hey, how much people are actually using at team level. Previously we had zero. We
[09:41] we don't know who's using it. We You can technically, but very difficult to kind of trace it down. Um so, as part of the improvement, we may um we improved the ACL uh using the IAM pass-through. I don't know how many of you are familiar with it, but it's basically about by AWS IAM.
[09:59] Uh and IAM, we try to implement it all back with it cuz we don't have UC. And doing so requires us to use the uh IAM roles and policies, and it has a very very tight uh restrictions. And we literally just run out of possibility of
[10:16] like running individuals enough uh permissions with the limitation on number of policies and number of roles attached to. So, um and then with the team-based resources management, we uh still lack the visibility on usage at
[10:33] the individual level. Uh even like we don't have visibility on who created what uh because we're not on UC yet. And also, it is possible to audit um like who logged in or who did what through Cloud Shell, but that get really really expensive uh when you turn them on.
[10:49] So, uh and as we migrate to E2, one of the thing we did is the separation environments. So, now we actually have different workspaces. Um but that's I mean, we have a dedicated workspaces for production jobs. Even
[11:04] then, we still lack the lineage of like, hey, how does one table goes towards a a or like what what tables is used by what job. So, that kind of a lineage uh still not exist. Um and then obviously on the catalog
[11:19] management, one of the benefit about using the catalog is the allows you to create multiple catalogs that help us with the data isolations, which we don't have yet at the time. So, yeah. So, and we also not on medallion
[11:35] architectures cuz it was relatively new and we have to carry over all the legacy stuff, so it makes it very difficult at the time. So, that is the kind of the history of, you know, how what we had before. With that, I'll have Dimitri talk about what
[11:52] now. Cool. Thank you for that, Ken. So, we're going to transition to act one of building our foundation, and our foundation is going to be built on Unity Catalog. We're going to use it as the cornerstone to build a house, and we're going to take our house from, you know,
[12:08] like how like house to like home. Uh why are we choosing Unity Catalog? Uh there's a couple of different reasons. So, I'll go through them all, right? The first one is the very big one is we actually can have ACLs on our data. Uh we've been doing SQL style grants since
[12:23] 1970s, uh and it just made sense to bring it in and use it uh on the data that we have in the cloud now, right? Uh we knew that our backs and a backs were a feature that's coming on uh UUC, and it's something we are we're planning on using it. We wanted to use it at the time, and now that's actually
[12:40] it's here now. Uh and then we also now have a lineage with our with our with all of our data. We can see from our reports, from our dashboards all the way to the source, to the raw data, how they connect. We can we can build that DAG and see how our data how our data's relate to each
[12:56] other. We have auditability tools with system tables. So, we have we track who's doing what queries, who's accessing what data. That's kept in a system tables that is easy for us to pull data out of. We no longer have to uh go through uh users' uh history in cloud
[13:12] trail or in your cloud providers' uh audit logs. We have that in tables in UC now. Uh we're able to build medallion architecture because we have catalog isolation level, catalog abstraction. Uh in the past, we had to use
[13:29] different naming conventions in order for us to separate our environments in Glue or set up, you know, multiple Hive Metastore for us to have different environments. Now, we have one UC catalog one UC Metastore that has catalog architecture underneath it and then schema stable so that makes it easier for us to actually structure our
[13:45] data using medallion architecture. Furthermore, we can actually isolate our data by using catalog and workspace binding. We can protect our data and make sure that if there's a workspace that's for specific group, that workspace only has access to that group's data and then
[14:00] off we go. We we we're ensuring that we're not exposing our data to somebody who's not supposed to see it. And then the last one is we can actually do data sharing now for for different regions. As we mentioned Riot is a global company. We have offices throughout the world. Most of our data
[14:16] resided in America, but with Delta sharing we're able to easily easily share it to other regions and we are able to choose what we're sharing out. So, as we were building our foundation, next thing we we focused on enforcement
[14:32] and being cost conscious. As we mentioned, we are small team. We're supporting a very large organization. So, we're trying to figure out how do we structure our Databricks setup to to to to ensure that we can manage it, but also give our users the ability to track their usage, track and
[14:48] and do auditability and do financial tracking. So, what we've ended up doing is we provisioned team-based compute. Each team gets you know, a set of clusters, set of folders, set of team spaces in Databricks that are catalogs and schemas. We are enforcing tagging on all of our
[15:04] clusters and then what we've also done, we've limited the ability for our users to create clusters themselves. Instead, we've given them cluster policies that are t-shirt size. We have a small cluster, medium cluster, large cluster that they can spin up and then you use
[15:19] this for their purposes. Um we additionally try to align a lot of our usage and workload on uh based on where it fits best within Databricks ecosystem. So, it's not really rocket surgery there. Uh jobs should run on job
[15:35] job clusters either ephemeral or uh schedule jobs. We still had some some use cases for all-purpose clusters. Those are mostly used for development and experimentation. And then we had also ended up moving a lot of our workloads onto SQL Warehouse.
[15:52] These are majority SQL-based workloads as well as BI tools like Tableau uh things like that. And what we ended up finding out as we started analyzing our cost and usage is that majority of our cost was being driven by SQL Warehouse.
[16:11] So, we wanted to figure out what was driving that cost for SQL Warehouse. Why is it so expensive? We have this, you know, tool from Databricks where we like its performance. We like the our users are enjoying it, but it's costing us an arm and a leg. Um I want to I also remind you where we are at the state. We had just adopted
[16:26] UC. We just started getting the lineage data. We just started getting uh usage, but we have don't have a full history of our uh performance. So, we don't know exactly we can go back on the historical data to to analyze our uh
[16:41] um how our data's been been been used by SQL Warehouses. So, a lot of a lot of our optimization was based on perception, feels, and kind of like our our base knowledge.
[16:56] So, as we trying to optimize UC, we try to think a few different things that we could do, right? We could have a we want to have a cost-first mindset. So, what does that mean? We can have uh auto-scaling clusters with auto-termination, and that will save us a lot of money, but unfortunately, that's not going to lead to a good user experience because uh the clusters will
[17:13] be terminating all the time. The user needs to spin it back up. The spinning up will take a couple of minutes uh during which time the user is just sitting in front of their computer and just looking at the screen and you know, trying to figure out what's going on there. That's one option. Another option that we could do is
[17:29] uh have always-on clusters, optimized for low latency use cases, optimized for not having any user interaction interruptions. We'll have faster results. We'll have always-on clusters, but unfortunately, that's going to result in us basically lighting a bunch of money on fire.
[17:50] So, what we think about solving this, right? We um We know Databricks has this product called serverless. They've been pitching it to us. They've been asking us to use it. Uh we want to use it, but at this at this point, we don't necessarily know much about it. We know serverless is generally costs more. Uh we don't have any experience with its
[18:07] stability or reliability, and we don't really know how to set up policies for serverless to make sure we can do proper financial attribution. The next as we're thinking about that, we basically had a moment when we
[18:22] realized that we just had to kind of step back to reality and reassess our thinking at the time. Um as mentioned, we are small team, three to five engineers supporting a global ecosystem, supporting multiple different roles. We're supporting data engineers. We're supporting analysts. We're supporting executives.
[18:38] Um a gamut of highly technical users to non-highly technical users. Um what we've basically realized is that if we continue managing Databricks for our users, it's going to be slow, and it's not going to unlock any of the value for them.
[18:54] So, instead of managing usage, uh managing cost, and using uh managing our infrastructure, we wanted to switch to instead enabling our users by surfacing their cost to them, surfacing their usage, and then going back on using managed infrastructure,
[19:10] and having Databricks manage that for us. And what that does is transfer the count accountability from our team onto our customer teams and letting them figure out by surfacing the usage to them, by surfacing the cost, letting them figure out what kind of value they're deriving
[19:26] from their workloads and letting them them make the decisions whether what they're doing is valuable or not. So, that was act one. We built the foundation. Uh next step would be expansion. However, before we want to
[19:41] expand, we want to figure out we want to learn a little bit about our ecosystem and figure out what we have going on there. And then with Unity Catalog with governance, now we have observability. We have we actually have the tools in our toolbox that allows us to analyze
[19:56] our usage. So, you looking at the SQL warehouse, we can see query patterns, we can see cluster utilization, we can figure out how our clusters were being used by time of day, by by day of week. We can see all of that. And then we also have access to system tables that allows
[20:13] us to extract the data about cost and who's actually driving the majority of this of this cost and how that's how that's looking. So, what do we realize by looking at all that? We realize we have an opportunity. Uh with SQL warehouse pro, uh we have
[20:30] over provisioned, oversized clusters that are always on because they're they were optimized for a lower lower latency use case. But that means that they're uh mostly wasted during off-peak times because they're just sitting there idle. But uh
[20:46] during on-peak times, because those clusters still need to scale, uh whatever user starts tries to tries to use them, those clusters need to scale and that scale it takes takes some minutes, so it's still not a great user experience. Um we had we also have a 5-minute startup time on the SQL warehouse pro,
[21:01] which during that time we are incurring AWS cost, but we're not really getting any value. And then we've uh set up a 60 to 120 minutes termination on our SQL warehouses because we did not initially want to impact user experience during the day. So like the typical pattern
[21:16] that we saw is, you know, an analyst would submit a query to the cluster, they would be running, they would either get a result or or not, the query still be running, they would, you know, get up, get a cup of coffee, they would come back. If we had a short termination time on the cluster, the cluster would be terminated. Now they have to start it and it's just like
[21:33] all of that interruption was just not getting us to a place where we need to be. So instead, what we realized is that with serverless, we we can have elastic scaling, we can have basically have no infra management, we can have cluster spinning up within a few seconds.
[21:49] And uh the the beauty of that is that we're only going to be paying on a per use basis. Because these clusters are so fast to start up, we're only paying for them when they're used, we can actually terminate them whenever we want. And then when our users we need to use them again, they'll come back, they'll
[22:05] they'll they'll hit the button to start it, it will take a couple of seconds, it will be off and running. So serverless for Riot, it was not a blind leap of faith, it was a data-driven decision enabled by governance and the analysis that we did.
[22:24] So after we've done the analysis, we actually want to build our house, we want to surround it with roof and walls. And that roof and walls is going to be serverless SQL warehouse. What we ended up doing is transition our 30 plus SQL warehouse Pro into SQL warehouse serverless.
[22:40] We have no startup wait times. We've reduced our auto termination from 120 minutes to 5 minutes. By doing all of that, we realized about $45,000 in monthly savings. What you can see on the graph below is
[22:56] orange bar is yellow bar is SQL warehouse Pro, blue bar is SQL warehouse serverless. You can see, you know, orange bar shrinking. Blue bar going up. Overall, the trend is going down as well, but we don't see don't see on this graph is the complete elimination of our cloud provider costs.
[23:13] So, now we're just paying Databricks. We're not paying uh AWS for the compute. We're instead we're relying on Databricks to manage that for us. So, how's it living in our house now? We're living under one roof, right? We have a better user experience. We have
[23:29] 10 seconds on average startup time. We have availability on auto scaling. We have reduced query contention because now we're able to stamp out more of these SQL warehouses for specific use cases and not worry about keeping them on and running all the time because we
[23:45] we're only paying on a per use case. We have minimal idle time and we're also closely aligning with the demand patterns of our users throughout the day. And what is all real realized for us is the simplified platform operations, lower management. We don't need need to, you know, worry
[24:01] about managing EC2 costs. Uh managing, you know, all of the uh AWS setup. And it's for the the other big big shift for us is uh was a mostly a mindset uh change for our team where we went from restricting
[24:16] everybody uh asking them to come to us with any use case that we will evaluate, we'll provision the resources for you, we'll figure out how to use it. Instead, we had a mindset change where we're going to try to enable our team to drive the value from Databricks and from their data. We'll let you to we'll surface the cost
[24:33] for you, we'll figure it we'll let you manage your own stuff, and we will track track the cost for you. So, so you know that what you what what you what you're paying for and what kind of value that you're getting out of it.
[24:49] So, now that we've uh you know, surrounded the we built the walls, we built the roof, next step actually comes expansion. We want to we want to have a bigger house, we have more people living on our block. We are actually a right of standardizing over common a couple common common patterns for data engineering. Uh we are using DBT for our
[25:06] ETL SDLC. Uh uh DBT is hooked up to dedicated uh SQL Warehouse serverless per project. Uh if you're not familiar with DBT, it's a stream-wide SQL-based modeling tool that all of our analysts and data engineers are using to kind of build a common
[25:22] vernacular and language around creating ETLs, building models. Um additionally, we're able to uh we are standardizing all of our internal and external JDBC and ODBC connections for things like Tableau, Atlan, uh other
[25:38] other products that we have in our ecosystem. All going to be going through OG JDBC and ODBC through uh SQL Warehouse serverless. Uh we're going to have better obser- observability into the connections that we're getting into our ecosystems that our users are setting up, but still have a lower
[25:54] management because we don't need to provision, you know, a dedicated resources for each one of those connection types. And it's going to be easier integration and attribution for us because we are able to track them in system tables so we could see where where all these connections are coming from, how how many connections are being made.
[26:10] Um we do still have a couple opportunities. Um mostly, we've been using SQL Warehouse serverless. Uh we do recognize we there's opportunities to use uh serverless for jobs, uh serverless for personal compute, serverless for Spark Clever Pipelines.
[26:25] Uh Right is very excited about all of the serverless products that Databricks is putting out, and we're excited to put even more of our workloads on them. And then with that, I will transition it to Ken to talk about a few more challenges, and then close it up.
[26:41] Thank you. So, one thing I wanted to call out um is part of that uh internal external integrations with JDBC. Right, we got a lot of observability onto exactly like when
[26:57] people connected to it to know what the sources is. I don't know how many are familiar with the UI, but it will literally tell you whether it's like it's a Tableau connections or Allen connections or whether it's like directly using JDBC drivers. That information helps us a lot to understand
[27:14] the usage patterns and and allow us also detect potentially like unknown connections. Where previously we were using it to using a all purpose before SQL Warehouse, right? We were using all purpose as the connection because you
[27:30] can set up JDBC connections over it. And we're basically blinded by it. So, I I believe there's a I believe you know you can actually turn off JDBC and all purpose now, but I could be wrong. So, we have to double check.
[27:45] Um So, uh even we're now at using Databricks um SQL serverless. We still have a lot of challenges and as Dimitri mentioned, right? Like we transferred that accountability to our
[28:02] external teams, right? At Riot. Um and we need to make sure is that hey, um they have the right uh processes and understanding how to drive their own governance um because it would not be
[28:18] effective uh if they don't, right? Like if they don't go in and look at the uh dashboards understanding the usage or understanding like questioning like why I'm spending so much money on, you know, doing certain things or why this individual is doing uh uh uh
[28:35] like a running ML servers uh servers to to spend like $1,000, right? Like if they don't use like offer operating in the same process with the process, uh it's not going to be effective. So, one of the challenges we have to continuously working with those team
[28:52] leads to making sure that that is probably implemented and properly dedicated. So, that is still like in in the active working working with them. From a technical point of view though, like
[29:08] serverless, it's great. One thing that we won't be able to do is actually connected to Riot network, the internal network. So, we're still working with Databricks on figuring out how to easily manage private network. Not that it cannot, right? It can. It just that I
[29:25] think managing it right now is quite a bit of challenge or a lot of manual process. So, we're still like walking through that. And for certain features in Databricks that uses serverless, they might be
[29:41] the tap up because we rely on tagging to do attribution to understanding usage at the team level. Without tagging publication to all the features that being used, it makes it very difficult for us to even know who's doing it. So, that is kind of the key
[29:57] to like how we can know about new usage and and new cost to to to usage. And then for Spark, for those who are familiar with the
[30:13] other logs and metrics, you currently we still have no no visibility to like the actual Spark UI when it comes to the serverless. So, you're not going to be able to see a lot of Spark metrics. So, that means like it's limited to certain usage
[30:31] of our preventing of doing a full serverless adoptions as well. Um As I mentioned serverless when it was available, it was like free for all. And if you notice, everyone can just go in there and click it and get their own servers. That also became a problem for
[30:46] us because with the lack of tagging propagations or enforcement of policies to create those tags so that we know the usage, we got burned by like, "Hey, we just got a spike in uh usage in serverless
[31:01] without knowing exactly what or who." So, we're still working through those challenges, um but hopefully in this uh this year's summit, there's some new information about, you know, management of that will be coming out, so we'll see.
[31:17] Um So, comes to the finale, um excuse me. So, uh we are um Data Bricks as we used to use Data Bricks heavily. It's also where all of our like on AWS, so we actually work closely together um
[31:34] to make the serverless possible and manageable at Riot. Um now serverless is kind of like at a gaming domain strategies, and we've been collecting a lot of feedbacks from our end users, we call Rioters, um that, you know, they're they're very happy about the the the experience,
[31:50] right? Like they faster to inside, do not worry about clusters, uh like how to size a cluster it properly, they get immediate you know, the the computer available for them. Um and cost optimizations as a is is more of a result
[32:07] coming as during that adoptions, but from our point of view, we are trying to just push for serverless as much as possible when it makes sense, right? Like certain usage may not be possible, may not make sense. Um so, but the the the goal here is,
[32:22] "Hey, we'll look at user experience first, what does that mean?" And ensuring that we have proper visibility to those usage. So, that's kind of how we're driving it. Um and then with Unity Catalog as kind of the foundation, it also basically allow
[32:37] us to centralize the uh data governance side. And that actually go beyond Databricks within the Databricks. For those of you who are familiar with the administration of Databricks, there's a cloud credentials like I don't know if you're familiar with cloud credentials, that allows you to kind of managed
[32:56] access to AWS uh services that is not just S3, right? Like you get storage credential to manage S3, you have volumes to manage S3, but if I need access to like RDS or
[33:12] um DynamoDB, Elasticsearch, you typically would just go through using the Python library, right? Um and that already lose track of your lineage and governance. So, with cloud credentials, it'll actually centralize your management as well. So, that has
[33:28] actually been really good. Um so, within time, I mean, as I mentioned, we try to bring visibility to our users and you know, hope that it will help them make better decisions on themselves so that we don't have to manage it at cuz we just couldn't at a scale.
[33:44] Um Yeah. Hopefully you all know, right? Riot value creativities as well as hyper focus on players. We know that we need to move fast when it comes to like innovations and we as a data platform, even though we don't have
[34:00] a direct interaction with players, we treat our writers the same way. Um and we know that putting restrictions on the platform could only leads to like slowing things down and in return, we're not bringing better values for our players.
[34:16] So, and serverless has been great user experience for writers and and surprisingly unsurprisingly, how it has helped drive a lot of like the adoptions recently particularly with AI. So,
[34:31] and and I think like serverless has helped tremendously of not having that burden on our team with that. So, with that, this is the end. Um and I really thank you all for coming
[34:49] in and hearing us today. Now, I have about 4 minutes. So, if anyone have any questions, I'm happy to answers. Yes, I don't have a mic, sorry. You have to yell, though. I'll try and talk louder. Uh so, you you you were talking about accountability and visibility. How are y'all organizing
[35:07] what groups you show that back to for that that visibility? So, the question is how like we have this governance process and how we kind of like we know what group to kind of talk to and stuff like that. So, um we target it at much as a higher level
[35:25] as possible because so like you you can imagine we have League of Legends Studios, we have Valorant Studios, right? Like we try to serve that all the way up to that data leader within that pillar. Right? So, basically working with them
[35:42] giving them the definition of what governance is, giving them the process to drive it themselves, giving them understanding the value of doing so. Because they also always asking the question is like, "Hey, why am I burning that much money on data, right?" So,
[35:59] it's like, "Well, here's all what it is. You have to tell us what is important cuz you need to know." And so that they can take that information back and you know, have the conversation of whoever, you know, spending all this
[36:14] money and making sure that the usage is proper. Hopefully that answers your question. Yes. Uh when you standardize data engineering, what guys do you think do you see benefits there as well, or is it just with your analytics
[36:30] So, the question is when we standardize on DBT, what What was the performance and cost benefits there? Uh what's the cost performance of that? So, uh
[36:46] the primary is the performance um because we at one point started using DBT in equation with all purpose, right? If I recall correctly. Um as you can imagine, the constant of evaluation of models, creating a model, is really slow because
[37:03] um the the auto scales and the optimization, particularly like benefit of SQL warehouse, is that it has a separate SQL query engine that is optimized for running SQL uh queries. So, like there is only out of the box difference.
[37:20] Um and and so, like that is kind of the performance gain. Um as we get more adoptions, we haven't really go back to say like evaluate what would could have been, right? We just believe in that, hey, it is driving value because now the insights who pretty much prior to was
[37:37] under the hood, right? All they do is build the models, test it, works, publish it, productionizing. So, that I think is valuable, right? Like I that yeah. Yeah, hopefully I answered your questions.
[37:52] Uh in the back. Uh 1 minute. Yeah. Uh you mentioned that uh you know, you have systems in place to attribute costs to different teams asking that they forecast their pipelines and their queries. Are there ways that you have to surface
[38:08] estimates before they incur those costs? Is that way they can make a decision before they spend money? Or is it all retroactive? It's right now it's retroactive. We we don't do any forecast. Um Um, if that's what the question was. We would love something like that.
[38:24] Yeah, we would I mean, I'm staring at them. Uh So, we One thing I did not probably did not mention, right? Like, part of the governance process is to have a regular cadence with those data leaders in each of the pillars
[38:39] to actually have the, you know, the conversation about like your usage. Right? So, we do set up like a monthly kind of cadence um, with them to review the whole thing right now until to a point where like they are comfortable to do it monthly on
[38:55] themselves. So, at that point like they're just asking themselves of like, "Hey, why am I doing that much?" Or they could come to us regularly about like, "Hey, how we can help them optimize on some of those." So. La last questions, I
[39:12] I guess I can't say at that point like I'd love for a way to have some guidelines on or like rules on what you can run on real estate versus not. Cuz I feel like you could at least cap the maybe this is like cluster sizing or something, but you know, I feel like
[39:27] it's pretty simple or easy for a user an inside user to rack up a big bill on actual It would be. And And this is one of the things that were There was a There were We're We're working on Databricks with that We were We asked the same questions like, "How can I
[39:43] limit the usage with serverless?" Um And And I was The one of the thing is uh There Is it available? Do we know like the power top the budget policies at one point? I don't know if you heard about that. Um,
[39:58] basically it's uh outside of your IAM like the Sorry, the user policies or the ABAC policies is specifically for a uh serverless budget for management policy. You can actually set the number of the cost uh with that policy and assign it to either
[40:14] individual users or group and that will cap them on their use. Yeah, but at one point though is because we don't know exactly how much they need, it's very difficult to also give that allocation in the shortly,
[40:30] too. So, but at the end of the day, I think UC has the uh as a central governance platform, it give us a lot of those cost data, lineage data, and then, you know, audit trails to help us build a governance
[40:47] ecosystem, too, help us understand all those usage. So, uh with that, that's all. Thank you all for coming. I'll be still around if you have any questions, feel free to stop. Thank you all.

Learn more about the Databricks Data and AI platform.

The information provided herein is for general informational purposes only and may not reflect the most current product capabilities or configurations.