Skip to main content

From Eight Hours to Eight Minutes: Automating A/B Test Analysis

Summary

  • Sega HARDlight, a UK gaming studio developing live mobile games with about one million daily players, reduced A/B test analysis time from 8 hours to 8 minutes per experiment by building a unified automated system on the Databricks Data and AI platform.
  • The solution combines Lakeflow pipelines, MLflow for model versioning, Unity Catalog for governance, and AI/BI dashboards with Genie to automate statistical analysis across four families of outcome metrics using both Frequentist and Bayesian approaches.
  • Standardization and automation enabled HARDlight to scale from analyzing experiments one at a time to running dozens in parallel while maintaining full statistical rigor and enabling non-data-scientist stakeholders to interpret results directly.

From Eight Hours to Eight Minutes: Automating A/B Test Analysis

Watch: From Eight Hours to Eight Minutes: Automating A/B Test Analysis
Mobile game studios run dozens of experiments simultaneously, but analysts spend 8+ hours per test on manual analysis, statistical validation, and interpretation. Sega HARDlight faced a critical challenge: they had built sophisticated statistical models for A/B testing, but the volume of experiments was overwhelming a small analytical team. Their solution required not just faster tools, but a complete rethinking of the workflow, from data ingestion through statistical analysis to decision-making.
Joel Dias (Senior Data Scientist) and Sanjay Ashok (Solutions Architect at Databricks) walk through HARDlight's six-year experimentation journey, from early ad-hoc queries to a production-grade automated system. Discover how they reduced analysis time from 8 hours to 8 minutes per experiment using a unified architecture combining Lakeflow pipelines, MLflow for model versioning, Unity Catalog for governance, and AI/BI dashboards with Genie. Learn the statistical methods they implemented across four metric families using both Frequentist and Bayesian approaches, and how standardization and automation enabled them to scale from single experiments to running dozens in parallel while maintaining analytical rigor.

Chapters

FAQs

What bottleneck led Sega HARDlight to automate A/B test analysis?

HARDlight's small analytics team was running dozens of simultaneous experiments across its live games, but each test required 8 hours of manual statistical validation, metric computation, and interpretation. The volume of experiments made this unsustainable, creating a bottleneck that delayed game design decisions and made it impossible to scale the experimentation program.

What statistical methods does HARDlight use in its automated experimentation system?

HARDlight implements both Frequentist and Bayesian statistical approaches across four families of outcome metrics. This video explains that supporting multiple methodologies was important because different experiment types and business questions call for different statistical frameworks, and the automated system applies the appropriate method based on the metric family being evaluated.

What Databricks components make up HARDlight's unified experimentation architecture?

The system uses Lakeflow pipelines for data ingestion and transformation, MLflow for versioning the statistical models that power analysis, Unity Catalog for governance and access control, and AI/BI dashboards with Genie for visualization and natural language querying of results. This video describes these four components as the complete transformation of HARDlight's experimentation workflow.

How did automation change who can interpret A/B test results at HARDlight?

Before automation, interpreting results required hands-on data science expertise because the statistical validation steps were manual and complex. After implementing the unified system, this video describes a democratization of experimentation where game designers, producers, and other non-data-scientist stakeholders can read and act on results directly from the AI/BI dashboard.

Full transcript

[00:08] Very good a very good morning everyone. Uh it's nice to have you all here uh today. Uh it's an absolute pleasure to present the work uh that Sega Hardlight and data bricks have been doing over the last uh few years uh in front of you and I'm glad that you're all here today to listen to our story. U I'm Joel Das. Uh
[00:25] I'm the senior data scientist uh with Sega Hardlight. Yeah, I'm Sanjay Ashoke. I'm a solutions architect at data bricks. So for today, uh you have quick look at the agenda.
[00:45] So we have uh the first half is going to cover the Sega Hardlight's journey, right? Who Sega Hardlight are, what are the games that they work with, and how their experimentation journey has been uh so far. And uh the second half we're going to look at while they were
[01:00] building out this experimentation platform what are the challenges that they faced especially around the scaling side of things and how did they end up building a unified system to solve all of their challenges. Uh we'll also have a dashboard demo of what they've built
[01:16] and of course we're going to conclude the talk with what was the impact and what was the learning for hardlight from this journey. Uh we so we start with a little introduction of uh who we are. Uh who is
[01:32] Sega Hardlight? Uh we are a small gaming studio. We are based in Lamington Spa uh in UK. Uh we are a part of uh Sega and we report to Sega of Japan. Uh we are a bunch of um artists uh designers, engineers, uh
[01:49] quality assurers, analyst, operators, marketeers uh and community engagers. Um uh we build and develop live gaming um live games. These are on different platforms uh such as uh Google play uh
[02:06] Google play store uh the app store uh there's uh Netflix there's prime and also uh bought three odd games on Apple Arcade. Um one of our most recent titles was Sonic Dream Team. Uh Sonic Dream Team was released in 2024.
[02:23] Uh it was uh recognized as one of the best mobile games uh for the year by develop star awards in UK and it was also uh a finalist uh in the Apple Arcade uh uh titles for the year uh for one of the best games released. Uh we run live events in all of our titles. Um
[02:41] these are live games that about a million players globally log into daily. So we have to keep these games engaging uh the content fresh. Uh we create and release a stunning new characters. Uh that's one of our core compet competencies. Uh you'll see some of
[02:58] these characters through the slide deck. Uh these are all made inhouse uh by a very dedicated and talented team of artists and designers. Um the player experience uh is at the heart of everything we do. Uh it's at the heart of all the decisions we make and we work
[03:15] uh continuously on improving this player experience for our players. Um, we run experiments on two of our titles. Uh, that's, uh, Sonic Dash and Sonic Forces. They're both available on
[03:31] Google uh, Play and uh, the App Store. Uh, Sonic Dash. Uh, it's an endless runner where you run on a track, you dodge obstacles, uh, jump over bad mix, destroy bad nicks. The main uh objective here is to collect uh character cards
[03:46] and unlock new characters in the game uh to free uh your animal friends that are captured by Dr. Eggman and his bad nicks and you build houses for them. Uh then there's Sonic Forces. Uh that's a multiplayer racing game where you compete with three other players uh from
[04:03] across the world in a very uh quick battle. You use powerups uh for attacks. You collect power-ups for attacks, defenses, for speed boosts. Um you um uh you use them strategically and the main objective in this game is to
[04:20] again uh level up uh collect trophies, unlock characters by engaging in event missions. Um yeah, and um yeah, I think the characters are the main part of this game. Um we run different types of experiments on both these games. We
[04:36] release new features. So one of our objectives is to make sure that the features we release uh are tuned uh uh when they when they go out uh to the players. Uh we try and improve uh the first-time user experience as much as we can. So that's a continuously iterating
[04:52] process. And then there's the game economy which is quite challenging to um keep configuring because every time we release a new feature the economy changes. Uh econ the game economies mimic real world economies. There's supply demand. Uh so balance getting that balance right is quite uh critical
[05:09] to having that uh ideal game experience in the game. Um so there's uh the experimentation journey. So as a studio we've taken
[05:26] quite a journey to get where we are here today. It's not the final destination but it's been a journey for about six years now. And we thought this is a good checkpoint uh for us to share where we've reached um with you. Um
[05:42] this is uh a slide uh in a very spoiler format because it basically tells you what we're going to run through uh in this presentation. So there's about uh it starts in 2020. Uh this is when the product boss uh Gabbor Armati he uh
[05:59] briefed us about this new mission that we need to take on to run experiments uh to make the games better to run more of more and more of them in a given time frame and very soon uh so around late 2020 we hit the first set of
[06:16] bottlenecks. um this visual is not uh as optimized as I would like to uh because 2000 in the the first 3 years were really slow. So those two milestones should be much closer together. In 2023
[06:32] uh that's when uh we uh released the first um dashboard or a live monitoring dashboard. Uh this is the this was built on Tableau. So we'll go a little bit into that. And then we hit the first set of bottlenecks after we had this live monitoring dashboard. And then we came
[06:48] back stronger uh with uh going deeper into statistical analysis and findings. And then uh comes in the point where we uh uh adopt data bricks and tie in this framework into uh tie in all these different solutions into a single
[07:05] framework that we now use today. Um right so so this is the the first bit which is the the rapid experimentation mandate.
[07:20] So this is the boss telling us that we need to run more experiments. Um as a studio we I think we we take pride in being quite uh talented in building and shipping new features um that players uh to to keep improving the player
[07:36] experience. But what we struggled with uh is fine-tuning these experiences. Uh we wanted uh the ability to make an adjustment uh to uh the player experience. Very much like tuning a m like a musical instrument where you make a fine adjustment. You listen to the
[07:52] results of it and then you want to tune it again. Listen to the results and do this until there's a balance and there's a harmony and that's the ideal player experience. It's a continuous process. uh but we uh need to keep doing that to keep on uh to keep ahead of the game.
[08:08] The mobile game market is also quite competitive uh overall uh and this happened around 2020 2021. Uh there were quite a few big games that went out there. Um so making these incremental improvements to the game was necessary.
[08:24] uh we believe that these small changes compound over time and uh eventually lead uh to quite a big factor in terms of both the game experience as well as the business uh meeting business objectives.
[08:42] Um this is the early experiments flow. So as soon as the mandate went out we started uh running experiments. Um the challenge here was uh getting from 0 to one. uh the issues were faced are nothing like what we uh would would be uh
[08:58] it's nothing like we can see around us today. We live in this very advanced space uh surrounded by great quality data and AI. Uh so the the kind of problems we dealt with was making sure that we're getting the right experiment ID uh with the account ID of a player
[09:15] making sure that the player is enrolled into the right experiment group. making sure that they stay in that experiment group uh through the uh duration of the experiment and when that experiment ends making sure that they um are assigned back into the mainstream experience of
[09:33] the game. Um we had an analyst called Rebecca Hansen who worked tirelessly just making sure that these little pieces of data coming through correctly and making sure that the data quality was something that we could work with. Um we did things like measuring
[09:49] retention, their engagement levels, um whether they interacted with the features that we hypothesized that they would um and whether this in this increased engagement led to um better monetization for the business. Uh we
[10:04] were writing ad hoc queries um per metric uh outputting excel workbooks many of them uh having two and fro discussions with the with the wider team um producing reports uh having meetings
[10:20] there was a lot of two and fro and then eventually coming to a decision. Uh as you can see this was quite a long process but this was also uh probably our most proudest milestone because it was that getting from zero to one uh which is probably the the most difficult part of every uh journey.
[10:41] Um the first set of bottlenecks. So when we started doing these experiments um we were uh like I mentioned we were running quite a few uh SQL queries to monitor what what was going on. uh we had no visibility uh over these experiments. So
[10:56] when the experiment was live uh we didn't know whether the players were in the group, whether they were getting the right treatment, how were they responding, uh how were how is it impacting the business metrics. Uh it was the only way we could do that was for an analyst to go in uh to Athena uh
[11:13] write a SQL query and figure things out. These queries were sitting in our notebooks uh in a local uh desktops and then we would pick them up, run them, figure things out on on the go. Um when the experiment was over, um it had to be
[11:28] tasked up. Uh so we had to go through our sprint planning. There was a task put in for an analyst. They would pick up uh these few SQL queries from the previous time they did it, run uh these queries, produce reports, etc., etc. Very soon uh these overheads grew uh
[11:45] beyond our control. Uh and the greatest overhead of all was the overhead of time because we are a small team and managing our time and resources effectively is key. Um to address this um need for
[12:01] monitoring and visibility. This is when we used uh this is when we uh uh built a live monitoring dashboard on Tableau. Um uh we uh u uh we wrote one SQL query. The SQL query was like a parent query
[12:17] that unified these different maybe about 20 or 30 odd queries into a single output where a we um all the game engagement metrics that we could think of. How is a player? How many runs a player is completing? How many
[12:33] characters are they unlocking? How many missions uh were they completing? uh how much currency were they spending, receiving etc. How many purchases were they making ads they were watching? Uh and this has all fed into this one output table that fed into Tableau. Uh
[12:50] the key design u that continues into its existing form today was to treat an experiment uh the same way that we would treat a clinical trial. So there was a treatment group and there was a control group. Um something that uh I haven't
[13:06] come across but maybe uh I haven't looked far enough is this concept of time. So we treated all the players on the day that they were enrolled into the experiment as day zero and that factor of time. So how much time have they been in the experiment versus the impacts or
[13:23] the engagement and the changes that they were seeing in the treatment groups versus the control groups was key uh to this dashboard. Um there was we've used this uh dashboard for about two years. Uh my colleague Yon Lee has spent um I have he he works uh
[13:43] quite a lot. So he's spent tireless amount of time on just iterating through this for the two years and we reached a state where it was uh our max uh Tableau skills and we're very proud of where we reached with this one. Um but uh soon uh we had a second set of
[14:02] bottlenecks um and that was uh statistical validation of results. Uh we could see players responding uh to the treatment groups some of them more visibly than the others. Uh but we were not a we are not convinced that this directional changes that we were seeing
[14:19] were actually uh not just noise. Uh we also felt the need for having some sort of a probability of improvement. uh with this mandate to run experiments in quick succession of each other. Uh there was also this pressure to start and stop experiments uh at an optimum uh cadence
[14:37] so that we are able to maximize uh this iterative process. Uh this meant that we could not run an experiment for too long and something that we are very conscious about is what happens to uh the treatment group beyond 14, 21, 30, 60,
[14:54] 90 days. uh because things can change quite rapidly. A player might like an experience that we are testing out but that it might not be the best experience for a life for a lifetime value of that player. Uh so this uh almost this lack of confidence uh statistical confidence
[15:12] came across in the results that we shared with the wider team and the wider team pushed back uh against the experiment for very valid reasons. So then we dived uh deep to seek statistical confidence as a team. We had to level up uh do our research uh figure
[15:30] out what we need to figure out. Um there were quite a few principles that we agreed after doing uh our individual homeworks and researches and that's uh to agree that um not to pick a
[15:46] statistical religion. So there were quite a few debates around should we go with a frequentist approach or a basian approach but we decided quite early on uh that we need both of them to have the strongest uh results. So frequentist
[16:01] method it gives us P values and FDR correction. These are essentials when you're when you're testing about 100 metric segment combinations but P values they are also opaque uh to game designers when they're making decisions. uh Beijing gives us something that they
[16:16] can work with. For example, there's a so and so percentage uh up improvement in the day seven retention of the variant group as compared to the control group. Um both of these metrics, the frequent and
[16:32] the basian metrics, they show up in the dashboards and and we look at both of them to make decisions. Uh player churn rates in free-to-play mobile games is quite high. uh that could be across different industries as well. So it might be a challenge quite a few of you must might be facing today. So most of
[16:49] the curves that we get have this strong density right in the beginning and then a very long tail. Uh to deal with this we uh adopted mincerization. Uh and this is at a segment uh and an experiment level. So we cut off the long tails
[17:04] based on percentiles that are tailored for every uh cohort. Um we have minimum sample size thresholds. Uh these are at metric and segment levels. Uh when the sample size is not large enough, we uh publish in the dashboard that we don't
[17:20] have enough of sample size to make any sort of a conclusive uh decision. Um there are about four families of outcomes that we focus on. There's conversion, there's retention, there's numerical metrics, and then there's lifetime value. Conversion is just
[17:37] simply how many players uh converted to spenders during the in the duration of the experiment both in the control group and the variance group. Uh since it's a boolean uh you convert it or not. Uh we use kai square and beijian beta. Uh kai
[17:55] square tells us that this difference is unlikely to be chance. uh Beijian beta tells us that there's uh 94% probability of improvement um of the variant being actually better. We run both. Uh retention uh did the user come back? Uh
[18:11] this was uh also quite a tough one to figure out. Uh because we did not just want to look at on day seven was the retention improved or not. Uh we wanted to understand the full survival curve. um when do players drop off and does a change uh keep them engaged for longer.
[18:28] Uh this is quite a universal metric as well. So used across industries. Uh since this this is a curve over time. Uh we adopted capital mere survival analysis with a log rank test. Uh this compares the full shape of how players leave and not just uh the output of a
[18:44] single day. Uh and then there's numerical metrics. Uh this is everything that's on a continuous uh that's a continuous number. So you've got ad uh revenues, ad impressions, inapp purchases, inapp purchase revenue. Um
[19:00] engagement metrics, the number of runs, missions, characters unlocked, etc. Um here we use Welch D test. Uh this compares averages while also accounting for the the the the variance between the groups which we found quite effective as
[19:16] compared to just a simple uh t test. And the most uh important metric of them all um almost like the northstar and everything in uh and like everything in statistics is debatable. Uh but this does present probably the strongest
[19:33] argument and that's the lifetime value. How much revenue a player will generate over their lifetime which is heavily influenced by how long they will stay in the game. I think of this as a single indicator because it is that balanced uh number between not just uh being good
[19:48] for the business, monetizing better but also an outcome of players actually liking uh the treatment uh in uh uh the the test group. I keep using these terms interchangeably so please don't get confused. There's treatment group,
[20:04] variant group and the test group. Uh we project this forward using an exponential uh curve fit. Uh essentially we model how spending decays over time uh to predict a few uh lifetime value. The confidence comes from the curve with
[20:19] covariance uh propagated into a wall style 95% confidence in interval on the difference. If that interval doesn't cross zero the lifetime value difference is declared statistically significant. In simpler words, u we fit a revenue
[20:34] curve per group um measure the uncertainty on each and check uh whether the gap between the groups is large enough that it's unlikely to be noise. All this being said uh as a studio uh we also firmly believe that the statistical
[20:49] evidence that we get um has to inform judgment and not replace it. Uh so when a decision is made using statistics, we still value and encourage thoughts, opinions that are different to what the numbers say. Uh and we take that into consideration to have a a holistic
[21:06] decision made from any sort of um experience that we roll out to the players. Thanks Joel. Yeah. Uh thank you Joel. So Joel just showed us how hardlight grew their
[21:22] experimentation team like you know the journey right so we did do a bit of deep dive into the analysis the list of all the models that they used so uh the live dashboards the Beijian and the frequentist confidence on top I think by 2025 hardlight were genuinely good at
[21:39] running experiments uh but the problem wasn't the quality in itself as you would see they experimented and explored so many different statistical methods to achieve the quality that they would like. But the biggest challenge was actually the volume. So um here was the
[21:56] situation. It was a small analytical team, a handful of people supporting multiple live Sonic games and the number of experiments were running at one at a time but it it was increasing steadily and all of these involved manual let's
[22:13] say manual processes where someone has to run all these analysis by hand ensure that these are running fine and validate at the end. So the first thing that breaks in this kind of a setup is not just the statistic itself but you know
[22:28] it's the people right that that was the ultimate bottleneck. So hardlight had one question. So how do they how are we going to run these carefully handbuilt analysis to dozens of you know dozens of
[22:44] experiments running in parallel without actually losing the rigor that Joel just pointed out earlier. So we're going to look at the workflow. This is how it looked like before hardlight made the transformation right. So they had the the first step starts
[23:02] with a manual experimentation step where an analyst would gather all the test IDs, update the queries and publishes the dashboard and then the daily KPI tracking on the legacy dashboard was done uh which you also saw earlier which
[23:18] was followed by a data extraction step. So this is essentially just pulling the player and the gaming metadata and and also the experiment metadata using the custom SQL. And then there was a statistical analysis piece. This was a
[23:34] Python script which was usually run by the uh analyst and this is usually it it was running in sometimes in the local machine and um yeah essentially to identify the outputs and then finally we
[23:49] they also had a human step where uh they would interpret these outputs and decides whether these are significant or not. So maybe a quick show of hands. How many of you have spent so much time in you know running these AB tests and realized you
[24:06] have wasted almost like close to more than 4 to 8 hours. Okay. Yeah. I I see I see quite a few quite a few hands up and this is exactly the challenge the hardlights were facing because they were spending more than close to eight hours. This these are
[24:22] skilled analyst time where um you know on one single experiment and imagine how that is going to be a big blocker for them as they scale uh the the experiment and also the other big challenge you know if you see there is a manual inter interpretation step. So if an analyst changes the whole interpretation also
[24:39] changes because two analysts can look at the same result and arrive at uh reach different thresholds and their interpretation could be completely different making the whole process slow and inconsistent.
[24:57] So what did hardlight do like they completely uh they were rethinking the workflow. So u when something's slow the initial instinct is for you to speed that process up. So they could have ensured they have written a better Python script. They could have written something let's say invested more time
[25:13] in a particular tool but heartlight decided to completely re revamp the whole workflow instead of spending more time on um improving single tools. So um fast forward one unified system.
[25:30] So three tools, five hands off handoff that you saw earlier. Everything replaced with one system. Uh you could actually see this is uh the final architecture. The it let me walk you through all of these steps, right? So it starts with a raw player telemetry. So
[25:47] the from the gaming servers the they were ingesting the data. This is every login, every purchase, every session coming off from the from the gaming data itself. And then that lands through um to the platform you know through lakeflow into a medallion
[26:03] architecture and then which is bronze, silver and gold. Um yeah so we had the bronze layer was raw events of course and then the silver layer is where the feature pre-processing cleaning and preparing the data itself was happening and then
[26:20] ultimately it was sent to the gold layer where yeah ultimately it was sent to the goal layer where they had premputed AB test metrics and model outcomes like lifetime value and engagement were ready so hard by using materialized views so
[26:36] that you know these again premputed views so that uh the downstream dashboard could be much more performant and on top of the goal sets the uh the the statistical layer. So we did see the Python script and the initial modeling earlier. So which used to be that which
[26:53] used to run in a local laptop now it runs within the platform. So still has of course Beijian inference bootstrapping winsization to tame the revenue outliers lifetime value with exponential DK. So ultimately the same models you know
[27:11] run per test or automatically without any manual intervention and all of these steps were orchestrated by the lakeflow jobs. Um these were running uh in in schedule refreshes
[27:26] automatically every single day. So the results flow into the aibbi dashboard ultimately which kepts refreshed every daily right. So Genie also sits on top of these dashboards in case anyone would like to doubleclick into particular analysis and underneath all of it uh is
[27:45] the most important layer that is unity catalog. So which underpins the governance and every single piece that we see here is fully governed by unity catalog. So which gives you that one lineage from the whole raw data that's coming in uh to all the way till the
[28:01] dashboards. So what changed you know we we have seen like okay that's a new architecture but there were four key pieces four pillars that made easy for hardlight right. So the first thing was the standardization we did see that now we have a single
[28:18] pipeline we have like to repeat against like the Beijian bootstrapping all of these methods are running on every single test automatically with the same assumptions every time right so the analyst doesn't have to rewrite anything
[28:35] from the scratch so for every single experiment they simply have to just review the result which significantly improved the um time for the analyst And then we have the automation layer. So again they're using lakeflow pipelines to ingest the data into these
[28:51] metallion architecture. So they had um ultimately every single time the the statistical tests uh don't run on noise. So every single time the data lands uh they would um the test once it reaches the sample size target the workflow is
[29:08] automatically kicked on right. So it it will keep running uh every single day. And the most important part is the dashboard where refreshes nightly and while the test is still alive. So you were you were watching it as the tests were unfolding it you were getting those
[29:24] live results as well. And then the accessibility part which made this is super important because as a stakeholder in in the gaming business as Joel mentioned earlier you know you've run so so many experiments and someone like a live ops director or a PM wants to
[29:40] understand what exactly is happening they would just come to the dashboard. This is my personal favorite as well. Um this is the having an LLM summary that actually summarizes the whole uh experiment for you on top of the dashboard. So if a PM wants to
[29:55] understand just the highlights of what the experiment is, they can simply read through it. And if an analyst wants to double click it, they could always go in and um treat the dashboard in treated like multiple layer and do a deeper analysis if they're interested. And of
[30:12] course, Genie is also going to be there tied to the dashboard which will help you answer any questions in natural language. And ultimately we did see this governance. So we saw about Unity catalog. I would also like to mention MLflow. So uh hotlight used MLflow to
[30:28] version assets um and all the models everything end to end from uh for this whole life cycle. So this is a quick snapshot of how the dashboard looks like. Um you could see we did mention about the LLM summaries
[30:44] at the top and then um you could you could see a quick on on the left hand side on the red you can see there is a quick highlights showing there is variant B shows a positive lifetime value impact um with no negative signals
[31:00] etc right and then you can actually if you look at look down it's a progressive layer that keeps um explaining what the experiment is it goes a bit deeper deeper into the statistical outcomes talking about the p values, the lifetime value projection, uh confidence
[31:16] intervals, etc. So the real dashboard if you look at it on the left you you could actually see the experiment didn't have a clear winner. So the system actually says so it doesn't really make up any any anything it actually gives the
[31:31] optionality for you to choose uh how the experiment need you you need to run the experiment. So essentially looking at a result like this uh you could actually roll out run the experiment for a while collect more data and then come back once you
[31:49] achieve that statist statistical significant number. Cool. Uh quickly I want to show you about the um um how it looks like. It's the same timeline what was before but now you could see the manual step initially which was
[32:05] there before now it's completely changed by uh the automated pipelines and then we have um the next step that is having a statistical modeling piece and then which is again fully governed analytical layer with materialized result and uh
[32:21] single source of truth of course and then we have our AIBI dashboards on top of it we have a genie as well and finally ly you know those output gives you that decisions as well. So um I want to quickly give it back to Joel so he
[32:37] could show you a demo. Right. So this brings us to today um the dashboard that uh the team at Hardlight is actively using. Uh this is pretty fresh. This came out in about uh last
[32:54] month. We are still iterating through the last final bugs. Uh but it's there. We've used it for about three experiments already. It's changing things. Uh I'm just uh being conscious of time. Uh I've take I'm going to use a recording instead of doing uh showing
[33:09] you the dashboard just because I'm going to have to switch windows and it might just break everything. So I hope this works. Okay.
[33:32] Okay. So this is playing. Um so this is the dashboard. There are three tabs in it. Uh the two uh key design principles here. There's one is the layering. So it's like a layered cake. Uh there are many u so you can dive into as many layers as you feel the need for. But right at the top is the
[33:48] LLM summary. So you can get a complete uh context of what is in the dashboard just from uh reading the LLM summary itself. Uh there are three tabs. The first one tells you about the health of the experiment. Are players enrolled into the groups? Uh are both the groups
[34:04] similar to each other? Um are there any anomalies that need to be flagged? The second one is just about player metrics. Uh this one tells you uh when the when the test is running uh how are the players engaging uh in terms of how many ads they're watching uh how much ad
[34:20] revenue is being generated the inapp purchases that they're making in app purchases uh revenue that's being generated uh and which elements in the game uh are they engaging with uh compared uh to each other. So the control versus the variant groups and
[34:35] then there's the final um layer uh which is the final analysis. Uh this is the main bit here. Uh so when the test is completed, we publish an executive summary. Uh this tells us
[34:50] about the four uh family of um metrics that we analyze and output for. This tells us what happened whether there was a winner based on uh both the frequentist and the basian approaches. If any of them flagged a metric as significantly changed better or worse,
[35:07] this is this is mentioned here. And then as you go deeper uh there's uh also an LLM summary within the four uh categories. So one for retention, one for lifetime value, for conversion and for the numerical metrics.
[35:30] Yeah, I'm just going to zoom past it. Uh, this seems um much more easier now than what it was when we first envisioned it. There are small elements like this drop down over here uh for the experiment which was uh almost an impossible uh feat uh to uh to
[35:47] achieve about a year back. We never imagined that we would be able to switch between experiments just by having a drop down and switching between them. Um I think the main objective behind this dashboard is to democratize uh running experiments. So before it was an analyst
[36:04] a datalled um effort but now the objective is to step back and to let the live operations team to let the designers to let the product managers take ownership of the experiments uh start off a configuration using the live operations team monitor it see the
[36:21] results and make a decision and then the analyst can step in only when they are needed or to make an iteration or an improvement to the process um and I think we're there uh which is a nice uh checkpoint to be at uh and we're happy to share that checkpoint with you today.
[36:45] Yeah. Um quick quickly touching on the impact. So uh yeah we save close to more than eight hours on analyst time to 8 minutes. U all the experiments running daily to 2x experiment scale because we are saving that much time. Hardlight is able to do more experiments today. And of course having a single unified system
[37:01] brings all the team to a single place uh so that the it's easier for everyone in the team to run these experiments. So it's definitely not just the uh I would like to also highlight the whole hardlight team here um who have put a
[37:17] lot of time and effort in building this uh over the years and yeah yeah and just a a big shout out to uh Richard Kh uh who's been u almost the brain behind this entire process uh he's adopted data bricks to an extent that he
[37:34] knows the ins and out of it can and can make anything happen at just a request in a very simple phrase. Yeah. Okay. So, thank you all for listening today. So, we just put our uh LinkedIn uh connect. So, if you're interested to
[37:49] connect and learn more, feel free to reach out to us. Thank Thanks again. Thank you.

Learn more about the Databricks Data and AI platform.

The information provided herein is for general informational purposes only and may not reflect the most current product capabilities or configurations.