Skip to main content

Run Databricks More Efficiently on AMD CPUs: Performance Benchmarks and Real-World Telemetry

Summary

  • AMD's fifth-generation Turin CPUs show 43–62% better performance on Databricks TPCDS benchmarks compared to Intel, enabling organizations to consolidate workloads, reduce per-core licensing costs, and meet SLAs with lower CPU utilization.
  • AMD's internal GPU fleet telemetry team rebuilt a pipeline that previously suffered 35% compute underutilization and 48-hour failure recovery times into one that recovers in under 1 hour, processing billions of daily data samples on Databricks running on AMD Turin CPUs.
  • The telemetry pipeline follows a Databricks medallion architecture collecting via Prometheus and Grafana Mimir through bronze, silver, and gold layers, with AMD offering a pilot program for proof-of-concept deployments on Azure.

Run Databricks More Efficiently on AMD CPUs: Performance Benchmarks and Real-World Telemetry

Watch: Run Databricks More Efficiently on AMD CPUs: Performance Benchmarks and Real-World Telemetry
CPU choice directly impacts Databricks performance, cost, and infrastructure utilization. In this video, AMD engineers compare the latest Turin CPU generation against Intel across industry benchmarks, then demonstrate how their internal GPU fleet telemetry team processes billions of daily samples on Databricks running on AMD CPUs. With Databricks TPCDS benchmarks showing 43-62% better performance and the telemetry pipeline reduced to under 1 hour recovery time from 48 hours, learn how to consolidate workloads, lower per-core licensing costs, and meet SLAs with significantly lower CPU utilization.
Discover the architecture behind collection via Prometheus and Grafana Mimir, ingestion through Databricks lakehouse, and transformation using medallion patterns across bronze, silver, and gold layers. Explore the AMD-Databricks pilot program for validating performance improvements through proof-of-concept deployments on Azure.

Chapters

FAQs

How much better do AMD Turin CPUs perform compared to Intel on Databricks workloads?

Databricks TPCDS benchmarks show AMD Turin CPUs achieving 43–62% better performance compared to Intel across relevant workloads. This video explains that better CPU performance translates to the ability to consolidate workloads into fewer virtual machines, reduce per-core licensing costs, and meet SLAs with lower overall CPU utilization.

What is AMD's GPU fleet telemetry pipeline and how does it use Databricks?

AMD's GPU fleet telemetry team processes billions of daily data samples from GPU infrastructure using a Databricks pipeline built on AMD Turin CPUs, following a medallion architecture with collection via Prometheus and Grafana Mimir and transformations across bronze, silver, and gold layers. Before the rebuild, the pipeline suffered 35% compute underutilization and required 48 hours to recover from failures; it now recovers in under 1 hour.

What is the AMD-Databricks pilot program?

The AMD-Databricks pilot program allows organizations to validate performance improvements from AMD Turin CPUs through proof-of-concept deployments on Azure. This video describes it as a path for customers to test whether AMD-based instances deliver the expected performance gains for their specific Databricks workloads before committing to a migration.

Why does CPU choice matter for Databricks workload performance and cost?

CPU performance directly determines how quickly Databricks jobs complete and how efficiently compute resources are utilized, affecting both SLA adherence and infrastructure costs. This video explains that a more performant CPU allows teams to consolidate into fewer VMs or servers, reduce per-core software licensing costs, and handle growing AI inference and data workloads without proportional cost increases.

Full transcript

[00:09] Hey uh welcome all. Uh I think this is the last session before lunch. Uh so thank you all for making your time here. Uh I appreciate it. I hope to make it as huh productive for you as possible. Um so I'm Shiva Gumorti. I'm senior business development manager at AMD. Uh I I focus on partnerships with databases
[00:26] and analytics partners. uh and I cover cloud on-prem uh everything around it and data bricks is one of my major one of my major ISP partners that I uh work with uh some housekeeping um please do complete the surveys for this uh you
[00:43] should be having the links in the app okay so we are here to talk about um uh database uh I'm sorry the data bricks uh implementation on AMDbased CPUs And we
[00:58] can we will be telling you how good we are on data bricks workloads. And I have here Nilan from our AMD GPU team who's going to tell us hi everyone. How yeah go ahead. Yeah I'm a data architect with the data center GPU team and I'll be talking about a use case where we used uh the
[01:15] new generation of Turan CPUs for processing a bunch of data and our whole processing is on data bricks. So we'll get to that later. Okay. Okay. So why are CPUs relevant or important to
[01:31] consider for databases or in general data management, right? One of the things is uh you know if you have a a big uh requirement in terms of performance uh CPUs are the real power host that are running all of your data workloads today. In the past we have not
[01:48] had uh a choice of CPUs before and uh you know the other big other company that was doing CPUs pretty much that was the only choice around but now we have a choice and in fact it's a better choice that I'm going to talk more about and now it becomes more and more important with all the emerging workloads uh like
[02:05] AI inference and all that stuff that the choice of CPU has to be the choice has to be made really really uh you know is the choice becomes really important that's what I can say so performance is one of the key criterias. It allows you a performance CPU allows you to expand your capabilities uh much faster and
[02:22] much efficiently. You can potentially with your deployment you can consolidate into fewer servers or fewer VMs as well if you have a very performant CPUs and cores that are available to you which means you can do the same amount of work in lesser CPUs or maybe you want to do with more work
[02:38] with the same amount of CPUs as well and also you'll save a lot on licensing costs. uh licensing cost uh you know as you may know pretty much all the database companies out there uh license per core uh you know it's true with a
[02:53] lot of major uh other major ISVS or software as well so licensing costs are humongous and it's per core and if you reduce if you maximize the performance that you get out of one single core you're actually effectively uh getting better ROI on your investments lastly if
[03:11] you have SLAs to uh track like basically you know you don't want go more than 75% of CPU usage and all that stuff. You know, having a very powerful CPU, you will see those CPU usages numbers actually come really down and that will help you meet your SLAs even though your workloads are exploding in fact on every
[03:28] core. So that's why you you need to consider CPUs as a very important choice when you go for your deployment and AMD epic is will help you get there uh with all the performance that I'm going to talk about. Uh so just a little bit of a
[03:45] history here uh relentless execution since 2017. Um we have been we introduced the chiplet architecture way back in 2017 and since then it's been an upward curve. Uh and it's it's it's been a phenomenal journey across we have not
[04:01] missed a single milestone or a or a release schedule uh since then. uh we started with the first gen epic in 2017 and today we are in the fifth gen epic uh which is we have code named Turin and uh that we're going to see more of how good it is compared to the competition in the next few slides
[04:19] the and with that the market has noticed uh today we are at about 46% uh market share uh overall uh and a lot of it has come not just because our competition has tumbled it has actually come because of our own merits and the chiplet
[04:34] architecture that we introduced way back then. It has been a pretty solid climb over the period of uh 7 to 8 years that I would say and today there's quite a bit of deployments all around the world. Uh every OEM platform now carries an AMD
[04:50] based server. Uh on the enterprise side some of the biggest customers uh have actually shifted to AMD epic CPUs for their uh workloads. uh initial obviously there was inertia and but they've been pleasantly surprised with the performance and the cost savings that
[05:05] you're seeing with AMD platforms some of the biggest uh you know what you use today probably use Uber today you saw Netflix show the back end is all AMD Apex CPUs majority of them is actually running on AMD Apex CPUs today and lastly the cloud every cloud provider
[05:22] now has AMD based instances I'm going to talk a little bit about it in the next slides but it is very and I'll also tell you how to what whether it's named platform or not. Uh just jumping into Azure because data bricks is pretty uh big on Azure. Uh
[05:38] here's how the road map our road map mapped to the deployment of AMDbased instances on uh on Azure. Uh so right from the first gen epic we have been part of uh Microsoft Azure and we have multiple instances available today as
[05:54] well. Uh the last the latest generation fifth gen epic is available on uh Azure today. Uh it's the V7 series if you're very familiar with the with Azure's instances. And how do you find if it's an AMD instance or not, right? So you just look
[06:11] for the letter A, right? It's as simple as that. Anything any instance that is got a letter A as the second letter in its uh in in Azure Microsoft Azure is based on an AMD CPU, right? And you can also search it up a little if you want to find more information on what exact
[06:26] CPU uh is used. It's the same is the case on a uh AWS as well. If you look for a letter A in AWS instances that also is is uh AMDbased instance. GCP is slightly different. They use the letter D but I'll talk about it in another
[06:42] session on Google Cloud. So these are all the instances available today. We have computer general purpose instances. We have memory optimized instances which give you twice the number of memory basically doubles the memory to core ratio. We have
[06:57] confidential computing instances on Azure. Uh high performance computing compute optimized ones. I'll talk about a few of them in the next slide but this is the o overall map of all the instances that we have on AMD Azure on Azure with AMD based instances.
[07:14] So the most recent one the fifth gen epic which we uh which uh we released recently uh we have Azure instances on that as I mentioned it is the V7 series uh we have quite a bit of uh exciting instances there uh I'll talk one the D
[07:30] instances and the E instances are the regular ones the upgrade to their original um original series that they had but the computer optimized ones with the letter F at the beginning which is which stands actually for full physical course is the one that's the most interesting and the most performant core
[07:46] out there in the market today if you want to use an am uh if you want to use an Azure and what it is is that if you're using an E instance um when you ask for two vCPUs you're actually getting two hyperthreaded cores but if you request an F instance and you
[08:02] request two vCPUs on an F instance you're actually getting two full physical cores non-hyperthreaded right so those are more performant than the regular the hyperthreaded ones And that actually increases the performance a lot. And if you're looking
[08:17] at license cost savings and also that's the best instance to go to because you're going to get the maximum performance out of it which means you can actually consolidate and get less uh cost on those instances. Now I talked a lot about performance. Uh we have done a wide range of performance uh workloads
[08:34] on these instances. uh at AMD we have more than 150 to 200 different uh benchmarks that they run but we just took a small subset of it uh and which we think is kind of represents what the overall overall uh you know requirement in the market on the cloud side on Azure
[08:51] is uh so some of them you will see familiar names like MySQL uh and you know SQL server and all those things but look at all the performance deltas compared to the latest that is available on with Intel on Azure everywhere we see a big big tower there
[09:08] and overall like you know we can say that the D instances compared to a equivalent D instance of Intel from V6 it's about 46% better performance just this is performance based on those benchmarks and the E instance is about 47% better performance on on uh when
[09:24] compared to the equivalent uh Intel instances but the last set of bars is is the real is the one that I was talking about the F instances that's about man 80 93% right better performance sorry I
[09:40] didn't memorize those numbers but 93% on average better performance than the Intel part right and they are the next slide will also you may ask me like okay you know maybe they are more expensive right there's a small delta in perform numbers in terms of dollars per hour that you will see on
[09:56] that but let's start here if you look at these two charts I go I'm going to go back and forth the performance per dollar is this chart you'll see a little bit bump in uh the performance per dollar that's because D instances are about 10 to 12% cheaper than the Intel instance I just talked about and with
[10:12] that also we are seeing obviously a big bigger performance per dollar number here the E instances are also about typically about 10 to 10 to 12% cheaper the F instances are about like 5 to 8% more expensive than the Intel instance but the delta you get is like more than
[10:28] 25% on top of that so we have multiple choices we can let our phops guys have the fun and get you the best instances for your deployment and and I'm pretty sure any of these instances will get you uh a big performance per dollar or
[10:44] benefit uh over your existing deployments that are there today. Uh coming to data bricks, I'm going to show you some data bricks numbers as well. Those were all industry numbers. Um data bricks we have a lot of instances enabled today. Uh the we we
[10:59] the these are all still based on the four fourth generation epic. They are not the fifth generation epic but we are told that the fifth generation epic also will be enabled very soon. The v7s will be enabled very soon. Uh so do look out for that but you will get the full benefit
[11:15] when that comes out. We are working uh actively with data bricks to do that. Now what do you get with data bricks benchmarks right? We ran TPCDS uh a benchmark based on TPCDS. Uh it's a very well-known analytics performance thing. We have seen bench we've seen the
[11:30] benchmark wars wars between data bricks and snowflake. Uh TPCDS is a very well is a popular and well-known uh benchmark in this market. And what we see here is the our fourth gen our fifth generation AMD Apex CPUs uh obviously privately
[11:46] enabled at this point were able to beat the Intel uh processor existing one that's by 43%. And the F instances actually give you 62% better performance as well. On the that side we have the performance per dollar as well considering the cost that is involved in
[12:02] this. So that gives you a bigger bump. uh it's more than 61% higher performance per dollar with the E instances and also a substantial 54% better performance uh with the F instances as well. So uh you know when when you go for your new
[12:18] deployment obviously you're looking to save money uh just look up AMD based instances specifically and if you are interested in any pilots or something like that we'd be happy to help you out as well. Uh I'm going to hand off the mic to Nilen for talking about GP telemetry. It's a very very interesting
[12:35] use case that we had internally at AMD using these exact met uh instances and you can see how how much we benefited from moving to AMD based instances. Thank you. Thanks. So I don't have a close
[12:54] one for you. So I'll be talking about the GPU, data center, telemetry, data analytics and we have used data bricks extensively for data processing on uh the CPUs which Shiva talked about. So we were using DA
[13:10] as V6 primarily. Uh we started our journey with that and then transitioned to the Turin CPU and V uh saw significant boost in performance drop in cost per performance and also uh the availability is much better. But to get
[13:26] to that I have to first talk about the use case and describe the use case what it's about. Uh we are focusing only on the data pipeline portion of the GPU fleet telemetry but the fleet telemetry as a use case is much more nuanced like
[13:42] uh all the training inferencing happening into this mod as Ali mentioned token maxing is going on. So as a user I need to be notified or as a data center owner I need to be notified on the GPU allocation. If the counters
[14:00] that show the GPU degradation for a node in the data center doesn't get alerted in well in advance or in proper time. It's much difficult to take action. So we have to deise a pipeline. It's much
[14:15] useful to have a pipeline which in which the data flows from the GPU rack via some processing pipeline gets uh processed and stored well within the SLA so that we can take action for allocation of the GPUs within the rack.
[14:31] In the agenda today we'll be talking about uh why GPU telemetry is important. The you can see there are multiple sections but we won't not go over all of those. The most important ones are the architecture and the transformation stage.
[14:48] The pipeline basically involves using Rafana Mimir uh as the backend data store which is a short-term data store when the data gets collected or the metrics and the counters get collected
[15:07] when we scrape them maybe at a second interval or a subsecond interval but the data processing architecture we have on data bricks which I'll talk to uh which I'll get to next basically process it at a subminute interval stores it on data bricks iceberg v3 and then makes it available for downstream dashboarding
[15:23] reporting and decision making uh for analytics in the gold layer that's in the transform stage we have described that in detail in the data model section where we have a complete medallion bronze silver gold uh architecture where we use the telemetry metrics to take
[15:38] some decisioning on whether to swap out a GPU make some changes or uh any other GPU addition sorry metric addition or removal in that particular node.
[16:00] So the goal for us was to build a single pane of glass or basically a real-time data intelligence layer which can give us complete vision for the fleet telemetry. So if you see the numbers that we're talking about there's 35% underutilization tax which basically
[16:16] talks about any idle GPU GPU sitting idle on floor work not being done there's no inferencing happening on it wasted cycle for those GPUs for big corporations purchasing billions of dollars of GPUs or racks
[16:31] uh 35% under utilization it's a pretty high multi-million dollar tax which we are trying to fix over Also meantime between failures the recovery window is about 48 hours uh where we complete the data collection we
[16:47] do some analytics and then we take action whether the GPU needs to be swapped out and we're trying to fix that as well the power wall based on the current estimate from Goldman this year about 13.6 6 gawatt of data center power wall
[17:05] has been created or will be enabled by end of 2026. So that's a pretty significant number and at this scale power is just a binding constraint. We have to ensure that whatever power is being leveraged we're getting the right bang for the buck. So
[17:22] each of those number uh you see on the top maps to a single pillar which we have tried or taken an effort to solve. like for predictive health we are trying to forecast the node degradation even before they happen. Uh we are uh so as the pipeline has been
[17:39] created and we have validated the numbers we were able to bring down the recovery window from 48 hours down to under an hour. That's a significant improvement using internal benchmarking and also we have understood the type of metrics we need to enable uh or the data
[17:56] points we need to capture so that we can also address the power shifting and the power concerns. I've listed out the main problems we identified in this whole journey. So the
[18:11] problem can be summarized into top four and primarily the GPU access uh basically gates uh the time to market for any such enablement. Underutilization as we saw before is a direct problem. It has multi-million dollar attacks which for sure needs to
[18:28] be solved because idle capacity means GPU is sitting idle on the floor. Nobody's using it and yeah that's a big loss. So the main challenge is because we don't have a global consolidated view of what all jobs are running on a rack or a GPU or
[18:45] if they're actually being utilized what the utilization level it's very difficult to take decision of uh running inferencing on a particular rack on what GPU and how the job should be spread out. So fragmentation is basically the root cause and
[19:01] like we are running uh jobs across 30 uh this particular pipeline collected a scrape data across 30 plus sites and uh running a little over 5,000 different nodes which had exporters uh different label schemes uh different
[19:17] scrape intervals some of the exporters generating data at 15 seconds some generating at 10 seconds some generating minute level uh metrics it's very difficult to create a uniform uh configuration where we can collect data from different exporters uniformly and then analyze it
[19:33] together. So the consequence is the GPU utilization which we measured at a different interval and not uniform. It's difficult to for us to take action in that regard. In a situation like that also uh there are basic operational
[19:49] questions which remains unanswered if the data is not collected in the right format like we cannot allocate capacity across the racks uniformly and if you if we don't have uh vision or if we don't have a complete picture of what all jobs
[20:06] or workloads are running on those racks it's very difficult to reclaim the idle capacity which we can't see. So those are the problem areas we have dealt in this particular pipeline and we have tried to solve and objectives basically uh comes down to getting those uh data
[20:24] points as early as possible not in some BI dashboards which would be reviewed by leadership later to take action. It has to be super fast has to be just in time and the datadriven placement of jobs or allocation of jobs for the rack is very
[20:39] critical in situation like this like GPU assignment has to be by business priority and not a contained capacity which gets in Q and it's assigned in round robin way.
[20:55] So over here I will talk about a basic example data pipeline uh which we have enabled. So starting from left you can see we have exporters. We have a bunch of different exporters. There's node exporter GPU a and broadcom and all our inband exporters which is scraped using
[21:11] an agent uh graphana alloy. So alloy agent scrapes uh different data points uh the metrics basically from the nodes uh AMD uh data center rack something like a Helios can have multiple node uh
[21:26] multiple GPUs. So some of these exporters uh scrape data at infra level so at infrastructure which is at a node level and some of these are GPU level exporters which scrape GPU uh level information. So we can enable any of
[21:43] those exporters as we want. It's up to the user and scrape data at whatever granularity they want. And in this pipeline telemetry pipeline the data moves through four different hops. We collect data. Collection happens at the exporter level. It gets stored for short-term in the mimir which
[22:00] is again a time scale data store from graphana uh horizontal scalable and that's a short-term storage uh data directly goes from different exporters via alloy agent to mimir uh data bricks is where the majority of the processing
[22:15] happens which we are more interested in and then there's a final MLBI layer but the actionable inside or the loop back happen directly from data bricks uh to take some corrective action against the exporters if needed. So yeah as I said we have collection of
[22:31] different exporters uh GPU A and egg broad I will give a drill down of these exporters in a later slide and each of them run inband so there's no out of band collection there is no memory sorry no network latency uh delay for that so everything is happening inband uh we are
[22:48] exposing the metrics locally graph analoy is collecting a single agent is collecting everything inband so dropping rules whitelisting everything happens here. Scale-wise, we have a little over 5,000 nodes uh scraped uh just for this
[23:04] pipeline and a little over 30 different sites like 5,000 nodes spread across 30 different locations. We have the script configuration is managed by githops which also handles uh cardinality and that's how the fragmentation of data is managed as well because we remove
[23:20] fragmentation at source at the collection level and then the m2 database pipeline is via a federated query approach. We can also do a API based approach if needed. Now the decision of this particular approach was to enable a lakehouse based path in
[23:37] parallel to the memory based path. uh not as a replacement. So the operation plane control is operation plane is controlled by mimir DB which has shorter data retention and the actual uh and basically shorter data retention
[23:52] basically means uh visualizing the data in graphana for some live queries and some adoc analysis but for long-term data warehouse or the lakehouse is on data bricks which uh and both of them have very different consumers retention policies and SLAs uh for our use case at
[24:10] least In this example end to end I go even a bit deeper explaining the different collection storage and transform stages. So as you can see there's node exporter
[24:26] uh which basically collects everything inband and it runs system on test and then exporters export the data out to Prometheus memory. Uh there's CPU, memory, disk, network bunch of different host level metrics.
[24:43] The GPU exporter collects uh all the metrics at a GPU level and since there's no separate out of band path, we only rely on the independent exporters. Although we have a capability to enable out of band expos exporters as well for whoever is interested. If you want to
[24:58] also calculate the network latency delay and you have multiple different racks working in sync. So you can capture data transfer rate between uh inter rack. This is everything inter in interact. Uh we already described the storage part
[25:16] where we use graphana mim for short-term storage and visualization happens in graphana. uh with thousands of nodes and several GPUs and NIC's each. The cardinality is uh important problem. Uh so we try to solve cardinality as well using GitHub
[25:33] provisioning because the data is per node each node having multiple GPUs and each GPU can have multiple lane and the GPU can even be partitioned. So it's per G uh per node per GPU per uh partition
[25:50] per lane uh per metric per minute and it can even go granular. So the data volume is significant and for us to take corrective action or some preemptive action we have to go to that granularity and understand what exactly is going wrong.
[26:06] The transformation stage is where uh we pull the data from mimir because it's a short-term storage that is live data lives there for only a couple of hours. We pull everything to data bricks and we start building our bronze layer.
[26:27] So in the bronze layer uh that's a step one. We first ingest, we pull the high resolution time series data uh into data bricks and data is long lived in bronze. We don't touch anything. We don't do measure of processing. Uh it is just a
[26:43] landing zone within data bricks and we do the heavy processing or aggregation in the silver layer which comes next. If I have to talk about the data volume, we scrape close to a billion raw samples a day. And this is just a test pipeline which has been developed and it's not yet in production. So the scale is not
[26:59] that big. But even with that, we have about a billion raw samples per day which gets reduced to about 50 million 50 million of aggregated rows in the first level of aggregation down to 4 millions in the gold layer. So the raw
[27:16] transfer from mimir to bronze uh it's 1 billion to 50 million uh aggregation 50 to 4 in gold and then from gold 4 million to 50k records reduction because the dashboards whenever we visualize it
[27:32] they do further aggregate the data uh as and when in the time period by weekly or by month. So the volume across the full pipeline the stage like uh all this aggregation uh this is where we use data bricks. So
[27:49] we have used DAS v6 as a primary uh compute node. We have uh efficient uh job cluster firstly validated on allpurpose and then we have set up a job cluster to do that. We use uh D8 as V6
[28:05] as a driver 32 as V6 as driver sorry worker and we migrated off to V7 and we found significant boost in the performance for this pipeline and like I will talk about the performance
[28:21] improvement later but I want to point out the two major pieces which carry most of the engineering cost and the DB cost in this overall pipeline. First the counter to rate conversion most of the raw pipeline
[28:42] most of the raw pipeline uh or monotonic rate conversion so the counters for the exporters they reset at the top of the R and they restart or they overflow so we have to keep a tab of that so there's a lot of backend calculation uh additional status symbols on top of the actual
[28:58] metrics which gets propagated So that has to be applied before the aggregation. So bronze layer doesn't just store. We have to have some additional logic along with autoloader so that we aggregate those data points
[29:13] uh into the bronze layer as well. Aggregating the raw values first and taking the rate uh afterwards because so that we can reset the counter was is very critical and will be very critical for whoever develops a fleet telemetry
[29:28] pipeline because a lot of the data will be out of order and that is a normal situation. We have tried fixing it. Yeah, there is no traditional fix for it. So you will arrive you will get out of order data and we need to have such counters to balance that uh act. And
[29:46] then there's time alignment and dduplication which uh happens because of the out of order data. So the exporters do not scrape uh on the same phase. Different exporters since we have about 5,000 different nodes across 30 locations each
[30:01] scrape at different frequency. Some might generate a sample one sample per minute others might generate uh 60 sample per minute. So understanding the sample frequency, understanding the script frequency so that you have different out of order uh deal cubes if
[30:17] needed uh so that we can handle those data points and aggregate in a uniform single logic. Uh doing that is very critical. Coming to the medallion piece, that's where a majority of the processing
[30:34] happens and which takes the longest time for us. bronze as I said is a raw landing zone. So one to one mirror of whatever we see in mime uh no major transformation and data remains here for a long time. We have enabled schema on read but uh we do not delete anything so
[30:51] it's append only so that in longterm if audit is needed refill or backfill is needed that's where brunt is for we can repopulate silver or gold as and when required uh nothing changes. So we have schema on read because a lot of the times
[31:07] exporters uh do the schema change. So a field might be added, a label might be changed and for that we need to uh ingest those captures and since we have everything on iceberg v3 that enables us to do schema evolution uh do time travel
[31:24] if we have to train a model against a certain version of a table and also it's uh adaptive to dynamic schema changes uh since we have enabled and we moved from delta to iceberg v3 for that matter. So yeah that's one more reason silver is where the majority of the data
[31:40] volume reduction happens. So from bronze to silver we see about 20x volume uh getting reduced uh from billion to 50 1 billion plus to 50 million because we do aggregate a lot of the data and a lot of the data gets uh either rejected or
[31:56] cleaned up in the stage. So the typing dduplication alignment from different transform stages happen over here. Gold is where we do further reduction about 12x from silver because we aggregate by node by source where it is coming from
[32:14] and also a couple of other business metrics. The reason of the three tires is pretty straightforward for anybody who has used medallion because uh the landing transformation and the serving is uh very critical to each of their
[32:31] audiences. So we have to keep them separate. I would also like to talk about the different metrics. So the type of metrics we uh capture which is useful like for compute load uh we have
[32:46] infrastructure telemetry data which is per node uh like system level host CPU load uh memory pressure storage throughput latency all that is uh captured at a node level. So there's one single data point which flows every
[33:01] minute or maybe per second depending on the in uh metric frequency. GPU telemetry is calculated per accelerator. So a node might have multiple or tens or yeah hundreds of GPUs within the rack
[33:17] depending on the type and there will be that many metric rows per GPU per partition or per lane and those are GPU metric points and there we also and on top of that we also collect some error count another performance signal which
[33:32] gets appended to the data to take some enhanced uh decision. The chart over here is just illustrative just to showcase uh the percentage of let's say uh for power 78 references to we are calculating we are capturing 78% of all
[33:48] the power metrics which loyent can scrape or the exporter exports we are not calculating everything because the remaining ones are not useful for our business use case so persisting them in the long-term storage doesn't make sense so that's what the number signifies
[34:08] And coming to the actual improvement in the pipeline where we are running the data where we are processing the uh telemetry data. So we started off our journey with standard V6 uh DS 8 uh I think V6 and moved to the similar V7 because of that movement we have se
[34:26] seen significant improvement in the runtime. So about 35% CPU performance and that's the reduction in runtime reduction in the cost of running those allpurpose jobs. Uh we have seen 20% or more higher IOPS 11% higher uh remote
[34:45] storage throughput. We are seeing a bunch of other improvements as well. But more important the ones which I more care about is the reduction in cost and much faster job runtime on data bricks.
[35:04] There's a bunch of links in the deck over here. You can go through the documentation portal like if you want to go deeper, please go through the references. There are documentation portal from AMD ROCOM that's a software stack which runs on the data center GPU. There's in instinct MI accelerators and
[35:21] also graphana mimir which talks about how the data gets captured what's an alloy agent how the data is stored in meir and yeah things like that and obviously everybody knows about the data bricks lakehouse platform so that's about it thanks we'll open for any
[35:36] questions uh thanks ninja uh that was good uh hopefully that gives you a good idea of overall like what the performance capabilities of the platform is AMD based Azure instances Also thanks for one of those use cases that we internally use as well.

Learn more about the Databricks Data and AI platform.

The information provided herein is for general informational purposes only and may not reflect the most current product capabilities or configurations.