Skip to main content

4,500 Jobs, No Blind Spots: AccuWeather's Serverless Databricks Migration with Datadog

Summary

  • AccuWeather operates 4,500 weekly LakeFlow jobs delivering forecasts to 1.5 billion people, and migrating to Serverless Databricks Workflows cut compute costs by 50% while eliminating the cluster startup delays behind 90-minute incident response times.
  • Datadog Data Observability applies domain-specific logic across five pillars—data quality, pipeline health, infrastructure, cost, and usage—cutting unactionable alerts by 50% and reducing incident response time by 80%.
  • A pre-flight pattern that validates input data before expensive jobs spin up, combined with ML-based anomaly detection, eliminates manual threshold management and gives AccuWeather comprehensive pipeline coverage without constant human intervention.

4,500 Jobs, No Blind Spots: AccuWeather's Serverless Databricks Migration with Datadog

Watch: 4,500 Jobs, No Blind Spots: AccuWeather's Serverless Databricks Migration with Datadog
AccuWeather operates 4,500 weekly LakeFlow jobs delivering forecasts to 1.5 billion people daily. When pipelines fail, severe weather warnings don't reach those who need them. The team faced 90-minute incident response delays caused by noisy alerts and no structured observability across their data stack.
Learn how AccuWeather migrated to Serverless Databricks Workflows, cutting compute costs in half while maintaining sub-second response times. Discover how Datadog Data Observability applies domain-specific logic to reduce unactionable alerts by 50% and detect failures across five key pillars: data quality, pipeline health, infrastructure, cost, and usage. Explore the pre-flight pattern that catches missing data before expensive jobs spin up, and how anomaly detection with machine learning eliminates manual threshold management.
🤝

Chapters

FAQs

Why did AccuWeather migrate to Serverless Databricks Workflows?

AccuWeather's cluster startup times were creating delays that hurt time-sensitive weather forecast pipelines, and noisy alerts were causing incident response delays of over 90 minutes, which had critical impact on customers making operational decisions based on forecast data. Migrating to Serverless Databricks Workflows eliminated startup latency and cut compute costs by 50% while maintaining the sub-second response times the business requires.

What is the pre-flight pattern for data pipeline monitoring?

The pre-flight pattern checks whether required input data is present and valid before allowing expensive downstream jobs to start processing. This approach prevents jobs from consuming compute and failing mid-run due to missing upstream data, which is especially important for AccuWeather where a missing weather feed can cascade failures across all 4,500 weekly jobs.

How does Datadog Data Observability reduce alert fatigue at AccuWeather?

Datadog applies domain-specific logic that correlates signals across five pillars—data quality, pipeline health, infrastructure, cost, and usage—before firing an alert, cutting unactionable notifications by 50%. ML-based anomaly detection identifies thresholds automatically, removing the need to manually tune alert rules as data volumes and seasonal patterns change across AccuWeather's weather data pipelines.

What are the five pillars of data observability in AccuWeather's implementation?

The five pillars monitored by Datadog Data Observability at AccuWeather are data quality, pipeline health, infrastructure health, cost, and usage. Together these pillars give the team comprehensive coverage of their data stack, enabling 80% faster incident response and ensuring that forecast data reaches AccuWeather's 1.5 billion end users without interruption.

Full transcript

[00:08] Hello everyone. Welcome to our talk 4500 jobs no blind spots AccuWeather serverless Databricks migration to with DataDog. So, I recently became a dad. And as other parents in the audience might know when this happens you end up
[00:24] buying a lot of stuff a lot of different products gadgets constant Amazon orders. And there's a product that I actually love called Nanit. Does anyone actually know this product have they ever used it before? Can you raise your hand?
[00:40] Wow, no one. Okay. Well, it is so Nanit is a baby baby monitoring app and camera that allows you to you know track how your baby's sleeping in their crib and it's a great product. It has a lot of good features such as like being able to know how the baby's
[00:56] breathing tracks like temperature humidity in the room. Really good quality easy to use. But it does have one problem. And that problem is that the notifications are extremely noisy. It is extremely sensitive to like the
[01:11] slightest movement sounds it gets a lot of redundant notifications right when like at the same time and this is actually what this looks like. This is my notification screen like two days ago. You can see sound detected motion detected a lot of them have the same timestamp.
[01:28] And I say that um It is even worse because often times I'm just standing like this already with my daughter and I keep getting those notifications. Now um The impact of this is I will say I probably don't it leads to me
[01:45] sometimes not responding to these as fast as I maybe should. And overall you're probably asking like why am I talking about this? The reason is this is a similar problem AccuWeather it actually facing um operating their data pipelines, where they were really getting these noisy
[02:00] alerts that made it difficult to know what to action on. Um and this actually led to like delays or incident response delays of over 90 minutes, which can have critical impact to their customers. Um their customers rely on this forecast data to make critical operational decisions for their business.
[02:16] Um such as, you know, should we uh close the manufacturing site tomorrow? Um or is an airline going to need to uh cancel a flight back from the Data and AI Summit? Or maybe even worse, um are we going to have to postpone a World Cup match? Um and so, that's what we're going to
[02:31] cover today. We're going to talk a little bit about how AccuWeather solved this problem, um as well as how they implemented data observability. Um before that, we'll talk a little more about AccuWeather's uh data architecture. They are operating 4,500 uh Lake Flow jobs at scale. And then, also as part of
[02:48] that, they've been adopting serverless and there's some learnings we'll share there. Um and we'll also talk about Datadog's data observability solution, um and how we can help uh Databricks customers deliver data um more reliably at at better cost, too. So, um
[03:04] Hi, nice to meet you. I'm Ryan Warrior. I am a senior product manager at Datadog uh for data observability. And I'm Travis Teague, the data operations manager at AccuWeather. And just some quick intros to our companies, uh quick background. Uh
[03:19] Datadog is a um so, I work at Datadog. Datadog is a monitoring security platform. Uh basically, we uh monitor every piece of infrastructure software that you're running, um networks in your company, making sure they're healthy, um reliable. And when things break, we give
[03:34] you the tools to actually fix it and solve it. And the same thing goes from the security side, where we help you detect any kind of threats or attacks, fix them, and ideally like prevent them before they even happen. Um people love uh we have these really good visualization dashboards. That's one thing uh you see here that customers
[03:50] love, as well as in the bottom right hand corner, we have a really adorable logo, Bits, which is a dog. And And for dog lovers in the audience, uh it's based on a Jack Russell Terrier. Um and so, in terms of the breadth of our offerings for monitoring security, we have You can see that there's a lot
[04:06] of products. I'm not going to go through all of them here. Uh but today we'll be talking about that third column, data observability, which is our newer offering for observability for data teams. And I'm also happy to announce that uh as of a couple days ago, we are the Databricks observability partner of the
[04:21] year for 2026. And now, Travis. So, a little bit about AccuWeather. Who are we Who are we? Well, um we're a weather company, but we're more than that. AccuWeather was founded 64 years ago by our founder and now uh executive uh chairman, Joel Dr. Joel
[04:38] Nyers, with the sole purpose of providing a superior forecast for to empower businesses to better manage the weather impact on their operations and personnel. We're a privately held company whose sole mission is to save lives and protect property.
[04:53] Now, alongside our um our website, mobile app, and APIs, we also have data offerings. So, we offer 35 plus years of historical data. This includes historical normals, current conditions data, forecast data, and most recently historical forecast data. And all of
[05:10] this is available in the Databricks Marketplace. So, let's get into it. Weather data is not your traditional data set. Uh it's not a normal database. Uh living in, you know, nice pretty tables. It's a multi-dimensional grid that never stops
[05:25] moving. Have you ever gone outside one day and, you know, it was nice and hot and then the cold front hit and then all of a sudden you need your big puffy jacket? That's weather. It's living. It's ever evolving. So, to understand why we chose Databricks, we first need to understand the problem we were
[05:42] facing. Weather data comes in every format imaginable. It's structured, semi-structured, unstructured. Um it's gridded model data. It's uh station observations, radar data, satellite data. Um it comes in binary
[05:57] formats. It's a mess. And no single traditional tool is able to handle all of this cleanly in one place. The third challenge is you can't separate data engineering from machine learning in these cases. For at AccuWeather to produce the forecast that
[06:14] you trust, we have to ingest observational data to validate, train, and alongside our forecast data to produce and blend the forecast into a reliable forecast that over 2 billion people who depend on AccuWeather every
[06:30] day trust. So, what did we build? Here's a nice little picture. Um Just kidding. Uh so, we built a single source of truth on the data lake governed by Databricks and Unity Catalog and Lake Flow jobs to
[06:46] manage and run over 4,500 jobs a week. Now, increasingly, we've been able to migrate a lot of this to serverless, which I'll get into in a minute. Now, this enabled us to publish curated products like our historical, current conditions, and forecast products in the
[07:01] data lake. So, now one platform owns ingestion, transformation, and distribution. So, new consumers are a grant, not a project. And so, in this screenshot, you can see the lineage uh of just one job that we run on Databricks. Uh and this is just one of
[07:18] many similarly architected jobs. Um and a nice little callout here, uh this lineage is native in Databricks. I didn't have to piece all of this together. Um I just found the job that I wanted to display, screenshotted it, and voila.
[07:33] So, as we continued to build our our pipelines, we hit another problem. We had friction with job performance and cost. So, what did we do? The first issue was cluster startup time. Um for traditional job compute,
[07:49] everyone knows that spinning up a cluster can take anywhere from 5 to 8 minutes, sometimes even 15 if you're running a really big cluster. Um now, this is just unacceptable when you have a job that needs to run every 5 minutes. The cluster takes 8 minutes to spin up and you need to run it every 5
[08:05] minutes, math doesn't math. That pro And so, as you can see, that problem compounds really quickly. Now, imagine that job fails and you have to repair it. Well, now you're waiting even more time, spending more money to spin the job cluster up and it's it's just a
[08:20] huge problem. The second thing um we looked at Oh, yeah, sorry. Uh the second thing we looked at was pools, but these were brittle um and they didn't really fix the issue for our use case. It wasn't the right
[08:35] tool for the problem. Now, for those low latency jobs where we couldn't afford any startup, we ended up having to default to an always-on classic compute, which as most people probably know is even more expensive
[08:50] than job compute. So, now we're paying for idle compute when the job isn't even running just because we can't afford the latency of a traditional job cluster startup. So, we needed a better path forward. Introduce serverless.
[09:07] Now, we go through some criteria before we move a job to serverless, and these are the three things that we look at. One, time KPIs. Does the job require low latency startup or low latency performance? If that's a yes, we immediately look at serverless. Now, the
[09:24] second thing you have to consider um is can the job even run on serverless? I talked about how much of a mess weather data is to work with, um and they require it requires custom tooling to be installed. Unfortunately, as of right now, you cannot install these tools on serverless, which means
[09:40] we have to go back to traditional job clusters. But, just because a job can't run on serverless doesn't mean you can't use serverless. I mentioned earlier that, you know, if a job fails and you have to repair it, now you're waiting for the job to spin up. It takes time, it costs
[09:56] money if you have a big cluster. Um but let's say your job failed for missing data. Well, if you know the data is missing, why are you spinning up the job in the first place? You eliminate the failure, you eliminate the latency of the startup, and you eliminate the cost of the
[10:12] cluster even starting up in the first place. So, the next thing we look at is can we run pre or and or post data checks? Can we check to see if the data exists using serverless? And if it doesn't exist, don't run the job. That saves us so much
[10:27] time and so much money because now we're not waiting for these clusters to spin up and eventually fail, costing us money from the infrastructure and the DBU usage just to have the job cluster come up, um because we know the data is missing. So,
[10:43] the results of all of this was we were able to migrate uh most of our jobs to either fully serverless or a serverless hybrid. And in some cases, we saw our compute reduced by 50%. We kept or even improved our time KPIs and reduced the number of failed jobs to missing data to
[11:00] nearly zero. Something to note uh that I want to call out about the reduced compute costs is that while our DBU usage may have increased slightly, the significant savings comes from we're no longer having to pay our cloud provider for the infrastructure that gets that has to be
[11:15] spun up for the traditional uh jobs. So, now, with everything running on Databricks, uh keeping tabs on the health and performance of these complex pipelines at scale became another challenge. We needed something that offered end-to-end visibility, something purpose-built for
[11:31] the complexity of what we had built. So, Ryan, can you tell us a little bit about how Datadog has approached this problem? Yeah, certainly. So, uh I'm first going to maybe start with this background on data observability and
[11:46] what is data observability? And uh to answer this question, what do we all do today? So, uh someone just shout out. When we have a question, what do we do? Who do we ask?
[12:01] There we go. ChatGPT. We ask AI, right? Um and uh since we are at the Databricks Data and AI Summit, uh which AI should I ask? Genie, yes. So, that's what I did. I asked Genie, what is data uh data observability? And it gave a good answer. It said, "Data observability
[12:17] refers to the ability to understand, monitor, manage the health and reliability of data across your entire data stack. It extends principles from application observability uh to data pipelines and data sets." And it talks about uh several key dimensions of uh data quality measures you want to do.
[12:32] So, one is freshness. Is data arriving on time? Um is data actually the the expected volume, like number of rows? Um does it have the correct schema? So, the right amount of right columns, right structure? Um for specific columns, is the distribution of the data um certain
[12:49] statistics around average, mean, max correct? And then it talks about lineage, right? Where does the data come from? Where is the data going? Um what's what is impacted? Who is impacted? So, this answer is correct. But I would also argue that it is a bit incomplete,
[13:04] right? Because if you do look at these things and measure it, um it does tell you something's wrong, but it might it's limited in being able to tell you why did it fail and actually how do you go resolve it, right? Um you also want to be able to uh monitor um
[13:19] upstream processes, ingestion, the jobs, other uh other parts of the maybe the infrastructure, um the clusters that are being running on. There's a lot of ways that things could fail and data might end up and not arrive. Uh, but if you just measure the data itself, you're kind of getting a limited view on how to actually attack that problem.
[13:35] Um, the other aspect is all the data might be healthy, it might be uh, flowing fine, but maybe your Databricks bill doubles in the next month because your data volumes did increase. Um, and or maybe you made it uh, configuration change to your cluster and now it's actually uh, you're using more
[13:50] compute. And uh, you might actually be breaking your budget. And if you're just measuring these things, that's also not giving you the complete view. Um, but don't just take uh, my word for it. Um, this is actually what Gartner says about data observability. So, they say that uh, data observability
[14:06] tools enable organizations to understand the state of the health of their data, um, their pipelines, um, the data infrastructure, as well as associated financial costs. And they have this kind of uh, five-pillar um, like template around how they think about the different things you want to measure, right? Your financial
[14:22] allocation, the data itself, the actual pipeline and processing infrastructure, and then how data is used. And this really brings a more comprehensive view to data observability. And so, um, this is how also how we at Datadog think about it. So, first of all, like as Genie said, data quality is
[14:39] important. That is the foundation. You do want to measure what's actually happening with the data. You do need to have the lineage. Um, but you also then want to measure the pipeline health, right? So, is the data processing happening as expected? Are the code and the queries, um, like succeeding? Are they not uh, you
[14:55] know, running with longer duration that'll break your SLAs? Um, the next layer is actually cost and usage. So, you also want to make sure uh, if there is a a pipeline issue or a data issue, but no one's using the data, does it matter?
[15:10] Um, I would argue no. I'd say like if it if no one's using it, it's probably not worth the time to go investigate it and attack it. And similarly, um, you know, are you is your data pipeline running within budget? Is it not doubling in cost due to you know, changes from data sources or the infrastructure.
[15:27] And then finally, um you also do want to measure what's happening with the underlying data infrastructure. Um you want to be able to know um if clusters are basically crashing, that's also going to make the data not arrive on time. Um and similarly, if if warehouses are queuing up queries and the queries aren't running, that's also going to
[15:43] hurt your data delivery. But I'll now turn it over to Travis again to talk a little bit how about how they approach data observability. Yeah, thanks, Ryan. So, now we have a tool to our problem. So, to how we at AccuWeather do data
[15:59] observability, it really starts with why we chose Datadog in the first place. Um when we first started looking for a monitoring platform, um we, like many other teams, did not really have Databricks and data observability uh on our requirements list. We wanted something that could ingest logs uh and
[16:15] monitor my systems, VMs, and processes. That's really That was all that I wanted at the time. Um but it was during our conversations with Datadog that we learned about their data observability product, which does all of the amazing things that Ryan just said. Um
[16:31] and that ended up being one of the main selling points, knowing that we could bring all of our monitoring across every system into one platform, a single pane of glass, if you will. Now, once uh now inside this platform, we can
[16:47] monitor ingest systems bringing data into our lakehouse. Uh we can monitor Databricks jobs and pipelines. We can monitor data quality, job cost, and performance. Now, I come from a pretty heavy DevOps background, so my first question with anything that I build is,
[17:03] "Well, can I automate this?" Um so, naturally, that was my question when uh when we were talking to Datadog about this. Um the desire And this desire compounds because our data operations team at AccuWeather is pretty small, and we're tasked with building
[17:18] the monitoring around our Databricks ecosystem. So, across 4,500 weekly jobs, uh that's a pretty big ask. So, we rely on this kind of automation to just to keep up. The answer
[17:33] was yes. Datadog has APIs, SDKs, and MCP servers to help automate everything. So, in the world of increasing AI tools, our our team continues to ask or or continues to explore how to use AI to automate these thing these setups and
[17:50] deployments. And speaking of AI's impact on data operations, I'm going to give it back to Ryan uh to talk more about how Datadog has implemented this. Yeah, so going back to another question. So, what changes with AI, right? With respect to data operations and running
[18:06] data platforms. Um obviously a lot of things, but I think I'll highlight just a few uh right now. So, um first off, I think with with the uh ability to actually use coding agents and and actually author new pipelines, the surface area is going to go pretty
[18:22] quick. You're going to be able to build more pipelines, uh tools like LakeFlow Connect get more data in. There's just going to be more and more data pipelines that actually produce more data to more consumers, and that's going to grow pretty fast, right? So, you have more real estate, more surface area to cover. Um secondly, the AI agents are if you're
[18:39] going to be using AI agents on top of this data, going to be automating more of your business, it's actually also going to increase the importance of these data pipelines to your business, which also then increases the risk if they break. So, there's going to be uh your data pipelines are going to become more important and a bigger risk point for um
[18:55] you know, for for your operations. And then, related to the third point around AI agents, there's also going to be uh a cost element. So, um right, Ali in the keynote was talking about how with token costs right now are unsustainable, right? For all businesses, and that's a problem that
[19:10] needs to be solved. Um but token costs are not the only thing that goes up with using AI, right? If If your AI is actually making lots of inefficient queries to your SQL warehouse or um, making extra usage to Lake Base, right? It's going to also drive up costs of other parts of your data platform as
[19:26] well, and that's also something that you are going to have to pay attention to. And this is actually not just theoretical, right? It is happening now. Um, so we at Datadog did a survey recently of over 100 data engineers, and basically nine out of 10 of them said
[19:42] that they are already operating having like AI running on top of their data warehouse. And so the right now today, it's already happening. So, with that, I'm going to do a little bit of an overview of um, how Datadog does data observability and talk a little bit
[19:57] about some of our capabilities to help you run reliable and cost-efficient data pipelines. Uh, I'll cover very quickly in three categories. So, one around detection, like how do we actually help you detect many different types of issues early on. Um, then next around how do you actually
[20:12] investigate and solve those issues? And third around how we help you understand and optimize the costs of your pipelines. So, starting with detection. Um, I talked about how that foundational layer, right, was data quality. And so, we have the ability to detect data quality anomalies across many different
[20:28] metrics. So, right now what we're looking at here is um, a data freshness monitor. So, did the data arrive on time, right? And you can see we have kind of an expected bounds here highlighted green with that red line and saying when the data is basically not being updated within its expected 2-hour window.
[20:44] Um, one thing I'll call out about how we do this is all of it is anomaly detection based, right? So, it's all using machine learning statistical models behind it. Um, that takes into account the different patterns of your data, seasonality. Um, and why this is actually better is it allows you to not
[21:00] have to write lots of custom rules around actually like checking your data, right? It It us to actually profile the data and actually detect if something is um you know uh, deviates from the norm. And so it saves you time on kind of management configuration of all these different checks. So this is this is one thing. Um then
[21:17] what then the next layer is around the pipeline. Like what happens upstream of the you know your gold layer of the data you're delivering. We can also do checks on all the jobs in processing, right? And Travis will talk a little bit more about that as well. Um and how that was valuable to AccuWeather, but you want to detect issues early, right? Early on. Uh
[21:34] you don't want to wait hours if you if you have large pipelines that run for hours. So we can detect the upstream kind of failures of the jobs, jobs that are running too long. Um and those are things we can also detect as well. Um and then infrastructure issues. Right? So here we're looking at all the
[21:49] different uh Databricks clusters. We can actually track kind of the health of all these clusters, what their utilization, and give you insight into if they're kind of running past capacity and or failing. Um and similarly for warehouses, right? So your SQL warehouses um if you have a lot of queries queuing up, um if you
[22:06] need to make changes to your warehouse around configuration of you know the the auto scaling or the sizing, um we can also help you indicate that as well. So a lot of different Basically, we give this detection across all layers of of the data stack. Um and now Travis, I'll turn it over to you to talk about how you've used
[22:21] Datadog data observability detection. Yeah, so one of the things that we heavily rely upon is the alert functionality, the monitors that Ryan went over. Um So one of the So one of the great things that Datadog does, it allows you to apply business and domain logic uh to
[22:38] your jobs, which is not something you can really do natively on Databricks without building a custom engineered workflow. Um so for example, we have a job that runs every 5 minutes. Um and because of the frequency of this job, if it fails once or twice or even multiple times a day, it's really not a
[22:54] big issue. The next run's going to pick it up, refresh the data. So we're really only concerned after several consecutive failures. And in the screenshot, this is an out-of-the-box monitor that Datadog offers. So, in this one, we want to know if the job fails five times in a row.
[23:15] Uh here's a little graph showing you kind of what this job does. As you can see, it's pretty noisy. Imagine getting a PagerDuty notification for every single one of these failures. Trust me, I have it at one point. It was not fun. Um but this monitor now ensures that we're only alerted when there's actually
[23:31] a problem. So, across these several days, yes, we had several fai- failures um throughout, you know, throughout that week, but none of those constituted an issue. Another thing that you can do with these monitors um is correlated monitors. As you can see, this one is
[23:47] just slightly more complex. Um but it adds so much benefit to our teams. This is applying domain forecasting domain logic to our to our monitoring. So, if I talked about blen- blending models together, this
[24:03] often involves multiple models that all have the same forecast horizon. If we If one forecast model fails, it's not an issue. We have two, three, maybe even four to back it up and cover it. It only becomes an issue when we start missing
[24:19] multiple forecast models for the same time period. We also have the ability to always alert for what we call major model hours. Not all models are equal. Um so, at our 0 6 12 and 18 hours, we can ensure that the most uh the most major models are always
[24:37] covered. So, with both of these monitors, the alert always means that something is wrong. And the results speak for themselves. Um we reduced our incident response time uh by 80% using Datadog data observability.
[24:52] Uh that means that we we were actually able to cut 90 minutes of responses down to just a couple of minutes cuz we were no longer guessing on if this was actually a problem or not. We cut unactionable alerts by over 50% which means that our engineers spend more time
[25:07] triaging actual problems than ghosts. Uh and Datadog ties directly in with our ticketing system which means when a problem does arise, the right team is automatically alerted. We're not guessing who owns what pipeline or who needs to be the the one to respond to
[25:23] this issue. So all of this ensures that across our 4,500 plus weekly jobs, we have zero blind spots. Now, we've talked a lot about a learning alerting. So now Ryan is going to talk about how to investigate after the fact.
[25:41] Yeah, so once you actually get these alerts and you know something's wrong, right? What do you do next? So the first thing you probably want to check is does this matter, right? Like is this actually something that uh has real impact to my business, uh to the consumers of the data. And so this is where uh data lineage does come
[25:56] in, right? So this is a view of like Datadog's data lineage. Um you can see here on the right-hand side that these are a bunch of BI reports that uh the business is using from, you know, a Tableau, Sigma, Looker there. Uh and you can also see upstream where that alert was on that raw orders table.
[26:13] Um and so in this case we can say, "Okay, it it actually is impacting some pretty critical reports for the business." After I know that it does actually matter, I then want to figure out what the root cause is, right? I actually say, "Okay, what what went wrong?" From there, the lineage graph is also helpful to actually see upstream where was the
[26:29] failure. Here I can actually see it's this shopist raw hourly Databricks job that was actually failing and that actually led to those those uh alerts on the downstream data set. But I can go deeper, right? So okay, why did that job fail? Um so what we're
[26:44] looking at here is a pretty detailed uh view of a single run of a Databricks job. Um and uh we're basically showing the full execution graph of each step of that job. So, the It's hard to see on the screen, but like right what I just highlighted there is it represents a single Databricks
[27:00] task. And below all the Databricks that Databricks task is all the different uh, Spark queries that are running underneath. Um, and but right here I can immediately see like what was the actual stack trace, what was the error, what was the line of code in the notebook that failed. And quickly get like an answer
[27:15] to that question. Um, for also for more complex issues where uh, you might need to kind of look through the logs, we have you can just pivot to this logs tab. You can then see all the standard error, standard out, and Spark logs, Spark cluster logs as well. Um, or as is increasingly the case, we
[27:31] also have an AI easy button where you can just click investigate with Bits, which is our you know, our which is our AI, and uh, it can just analyze all of the telemetry and information we have on the left-hand side and let you know what actually went wrong.
[27:48] But, I'll turn it back to Travis to talk a little bit about how they use Datadog to investigate issues. So, yeah. So, here you can see um, I have received an alert that one of our tables is not updating as frequently as we expect. So, from the monitor alert page, I am able to very quickly I uh,
[28:04] pull up the lineage as you see here, find the impacted table, and quickly zoom out to see all of the downstream impacts. This let helps let us know um, you know, which customers might be affected uh, because of the data set in question.
[28:20] Um, so this all this helps us find triage much, much quicker. We don't have to guess what job failed, we know exactly what table uh, has an issue, we know what job is attached to that table, and so we can get to the root of the problem very quickly. And so now Ryan's going to give us a
[28:35] little bit of insight on how you how you can use Datadog to track one of the those last pillars of data observability, cost. Yeah, thanks. So, obviously first order of business is making sure all the pipelines are healthy, good day good quality data is
[28:50] being delivered. Uh if there are issues, we fix them. But, this last pillar is more about understanding the cost understanding and optimizing your cost and usage around the data platform and data. So, what Databricks can help with first of all is just give very good cost
[29:06] visibility into your data your data platform. So, here you can basically track cost of Databricks, but then also actually the underlying cloud it might be running on, too. So, if you're running on AWS, Azure, Google Cloud, you can actually see the Databricks cost as well as all those
[29:22] other costs in kind of one view, so you can make sure that they're not exceeding budget or even just do some forecasting to make sure to to get a sense of what you think those costs will be at the end of the month. But, this is just like this just gives you the visibility. Where we go deeper is we actually give very specific
[29:37] recommendations on how you can save money and eliminate waste, right? So, what we're looking at here is a recommendation on a specific Databricks job where we're basically taking all of that data that you saw in that um that screen that had the the execution
[29:52] flow of a job run, right? Like all those different the task and the queries and and all the performa- the logs. We're able to kind of look through that and actually determine, "Hey, this job is actually running inefficiently. You actually should filter We see in the query that there's a massive join happening shuffling on
[30:08] many different nodes of the cluster. If you actually filter some of that data first that you that you actually are not using as part final result, um you can actually really kind of reduce that query time by 70% and actually reduce cost as well. So, this is just like one example, but we also have similar
[30:23] recommendations for clusters, so like being able to identify clusters that are over provisioned um and where you can actually we provide very specific configuration change recommendations. And then also SQL warehouses as well is something that we're working on and coming soon.
[30:39] Um and oh, sorry, going back. One thing that uh is uh unfortunate. Yeah. Uh you can see there's a box at the bottom there that is a button that is that basically would lead to this page where you can actually just uh in a single click go and uh basically make a
[30:54] PR for that code change, too. So, um we do it's not really helpful if you get a big recommendation and then you have to go sort figure out actually where do I go make that change. You can actually just integrate with GitHub or wherever your source code repositories are and actually in a few clicks just make that adjustment to your notebook.
[31:11] Um and as is increasingly the case, too, um you don't have to use Datadog's UI for this. So, we have a lot of um a lot of our users are starting to use our MCP servers where you can actually just take go from your coding agents, pull in all that performance data that we have around your jobs
[31:27] uh into the agent context, and then actually use that to make the code changes right there. Um but then back to the usage question, right? So, I think this is what I showed earlier around, "Hey, is there impact?" Um lineage is helpful in showing like what is impacted, but then another layer
[31:43] of the question is is anyone actually using this, right? So, there could be all these dashboards, but what if no one's actually looked at them for like a month? Um and so, the last thing that we can also help with is understanding actually the usage patterns of the data, right? So, we actually bring in the full query history
[31:59] from Databricks, and with what we do with lineage and mapping the tables, we can you can do a deep dive into specific tables and see if they're actually being used or not by different users and who is using them.
[32:14] And I think to to close, I'll turn it right back to Travis to talk about some of the key learnings from our talk. Yeah, so we'll kind of give an overview of everything that we talked about. Um and hopefully, if you don't remember anything else from the session, hopefully you'll remember these four things. First, serverless is not an
[32:30] all-or-nothing deal. Find creative ways to use serverless. You know, across our 4,500 weekly jobs, um there's only a handful that are 100% on serverless. But, that doesn't mean that we can't use it in other cases. I talked about pre-data checks. You can do post-data
[32:47] cleanup, whether it's removing temporary files, optimizing tables. Um the sky's the limit. So, just because you require custom tooling or something else where you're still relying on a traditional job cluster, doesn't mean that you can't use serverless in other areas.
[33:03] Second, define the criteria for your alerting. It's really hard to know what and when to respond if there's not any defined impact. We talk We opened this talk with we had 90-minute response times at AccuWeather.
[33:18] That's because we had no idea what an actual what constituted an alert. Third, data observability is more than just data quality. As Ryan has shown, may just monitoring data quality is insufficient. You want to monitor what's
[33:35] happening earlier in your pipeline, like we did with jobs, as well as understand the data usage, downstream impact, and costs. And this gets more and more important with AI changing data consumption patterns and creating new business risks.
[33:51] And lastly, going back to my DevOps background, treat your monitoring like you would any good software or pipeline. Datadog has the tools for it. We talked about the APIs, the SDKs, the MCP servers to automate setup and creation.
[34:07] And I can guarantee you that your Databricks pipelines are going to change, and so you should update your monitoring to change with them. Spend time to tune your monitors. Refine them over time as you see how they perform.
[34:25] Yeah. Yeah, and if you are interested in learning more about Datadog data observability, you can just take a picture of this QR code or scan the QR code and yeah, schedule a meeting with us. Thank you so much.

Learn more about the Databricks Data and AI platform.

The information provided herein is for general informational purposes only and may not reflect the most current product capabilities or configurations.