Skip to main content

Databricks Serverless Migration: Scaling from 2,000 Jobs to 40% Adoption

Summary

  • McAfee migrated 40% of their 2,000-plus production jobs to Databricks Serverless using a three-bucket job qualification framework, achieving a 50% cost reduction and 83% runtime improvement for migration-eligible workloads.
  • Classic compute left 30% of clusters idle and forced platform teams to spend time managing infrastructure rather than governance; the migration playbook covers job profiling, platform guardrails with Unity Catalog, and batch rollout with rollback strategies.
  • Dashboard-driven monitoring and resource tagging built executive confidence throughout the migration, with the full playbook designed so that other organizations can replicate the approach regardless of their job count or cloud environment.

Databricks Serverless Migration: Scaling from 2,000 Jobs to 40% Adoption

Watch: Databricks Serverless Migration: Scaling from 2,000 Jobs to 40% Adoption
At McAfee, the data platform challenge was universal: scale velocity and reduce costs without breaking 2,000+ running jobs. Classic compute left 30% of clusters idle, and platform teams spent time managing infrastructure instead of governance.
Learn how to build an enterprise serverless migration playbook: job profiling and the three-bucket qualification framework, platform guardrails with Unity Catalog and Databricks Serverless cost controls, real-world patterns for batch migration and rollback strategies, and how dashboards and tagging earned executive confidence. The result: 40% of jobs now running on Serverless, 50% cost reduction, and 83% runtime improvement for migration candidates.
🤝

Chapters

FAQs

What results did McAfee achieve after migrating to Databricks Serverless?

McAfee achieved a 50% cost reduction and 83% runtime improvement for their serverless migration candidates, with 40% of their 2,000-plus production jobs now running on serverless compute. Their classic compute environment had been running with approximately 30% idle clusters, representing significant wasted spend.

How did McAfee decide which jobs to migrate to serverless?

McAfee built a job profiler and qualification engine that categorizes jobs into three buckets based on their compatibility with serverless compute. This systematic approach identified the best migration candidates, those needing preparation, and those that should remain on classic compute — avoiding the common mistake of migrating all jobs at once.

What does McAfee's serverless migration playbook include?

The playbook covers job profiling with a three-bucket qualification framework, establishing platform guardrails using Unity Catalog and Databricks Serverless cost controls, executing migrations in batches with rollback strategies, and using dashboards and resource tagging to track results and communicate progress to leadership.

How did McAfee use Databricks Genie Code as part of their migration?

McAfee used Databricks Genie Code for query optimization during the migration process, identifying inefficient SQL patterns in jobs being moved to serverless. This complemented the infrastructure migration with query-level improvements that contributed additional cost and performance gains beyond compute savings alone.

Full transcript

[00:08] Good afternoon, everybody. Thanks for joining this session with us. Um let me make this session a bit more interactive to begin with first of all. Couple of questions, nothing too difficult, just yes or no's. Raise just raise your hand. We just want to make sure that we know our audience.
[00:24] Um so, how many of you manage Databricks platform directly or indirectly? Wow. Good. How many are on Unity Catalog? All right. Okay.
[00:40] How many of you have serverless enabled in your environment? Pretty good. Um okay. Now, how many of you have seen um a lot of compute clusters sitting idle in your environment? Let's say more
[00:55] than 20%, 30%. Right? Yeah. Now, we are starting to get closer to the problem that we want to talk about. Um and how many of you hear from your developers that their code is running slower, their job is running slower, it's not as efficient as
[01:11] it is? Right? Yeah. Okay. All right. Thank you. Um So, if you if your hand went up for any of those questions, you are definitely in the right room. Because here is the question we had to answer. Wouldn't it be great
[01:30] to make your code run faster and save 50% of your cost? That's exactly what we did at McAfee. And by the end of this this presentation, we want you to walk out with some specific action items that
[01:45] help you improve your platform. All right. My name is Mandar. I work for McAfee as director of data engineering. I've been with McAfee for over 6 years. Uh prior to McAfee, I spent majority of my career in cybersecurity.
[02:01] Amisha? My name is Amisha Singh and I'm a delivery solutions architect here at Databricks. And I've been at Databricks for about 3 years. Been a cloud infra architect at AWS in my previous life. And I've been working with Mandar and the McAfee team for a major chunk of my
[02:17] time at Databricks. And today you're going to hear about the story that we very much lived through together while migrating to serverless. Okay. So, let me give you a quick rundown of what we will cover. First, we will talk about some of the common problems most data platforms
[02:34] run into. Then Amisha will walk us through some options on the table, the playbook we followed. Then I will cover the McAfee journey. And that will include some of the problems that we all ran into with uh uh
[02:56] All right. So, for most of the platforms, the mandate looks like this: increase velocity and reduce data bits cost without stopping any services, right? We are no different. We had this mandate, we still have. And now, let me give you a sense of the scale we were working because I think it
[03:13] will help shape the story a little better. We have over a thousand users that we support on a monthly basis, directly or indirectly. We have more than 30,000 tables in our environment. We have more than 2,000 jobs that run in our
[03:30] environment. So, that's the real pressure, right? Because none of those jobs could go down. We are likely not going to get any head count. That's definitely not a solution to the problem here. So, how do you modernize infrastructure
[03:46] at scale when the team head count may not grow? One thing that was really in our favor was the fact that we were already at about 90% Unity Catalog adoption. So, the governance foundation was there.
[04:03] We weren't starting from scratch. We are ready for the next step. But, that's just the mandate, the scale, the challenge. But, this is really what kept me at night. And this might resonate with a lot of people who
[04:18] um manage the data platform directly. Thousands of jobs running on the platform, and if even if one of them goes down because of a change that we did, it's on me. The downstream impact is real.
[04:37] On top of that, you were asked to deliver cost savings. We were not even sure they were possible at the time. Not trim a little, but cut costs by at least a third. And we had no proof it could be done.
[04:52] All we had was a hypothesis. So, the assignment was basically modernize everything, break nothing, and remain lean. No pressure, right?
[05:08] So, this was my reality. The question was, what could we do to solve this problem? And honestly, that's where having Amisha alongside us changed things for us. She has seen teams exactly in this position before. So, Amisha, what were the options before us?
[05:24] All right, let's look at them. I think I heard someone laughing there as well, so I feel like you understand the problem uh McAfee was going through. And strangely, this is a very common problem that enterprises fall into, that the platform is growing, but the way
[05:40] that you usually run it doesn't keep up. So, before recommending anything, we put all of the options on the table, and we ran through the same three questions across all of them. Does it remove idle compute cost? Does it remove cluster
[05:56] management? And does it remove the Does it actually give you true elasticity at scale? So, starting off with right-sizing your clusters by tuning the right instant by tuning to the right instance types and enabling auto-termination or moving to
[06:12] spot instances, you can generate real savings, but your plat- platform team is still that bottleneck, and they're making every sizing decision. When the demand spikes, there's no dynamic response. You've made the current model
[06:27] a little bit cheaper, but at what cost? Your underlying problem is still the same. With custom auto-scaling, gives you elasticity, that's great. Compute can grow and shrink with demand, but you're still building and maintaining that auto- auto-scaling logic yourself, and
[06:43] the engineering effort still goes up. So, one of the boxes in there is checked, but the other two still remain unchecked. Now, moving on to serverless, so that's what truly helped us close out all three of this. We're only paying while a job
[06:59] is running, so the idle cost disappears by design, and Databricks is managing the compute entirely. So, of course, for McAfee, this was the obvious choice, and from what I could tell, it was the obvious choice for most of us, um since I saw a lot of hands go up when
[07:16] we asked you if you guys are so are on serverless or not. So, let's look at what what are the real value drivers for serverless across all of these pillars. So, on the compute side, before serverless, clusters were sitting idle between runs, and someone
[07:32] on the platform team was manually sizing every one of those 2,000 jobs across all of the different workspaces. But after, compute is available in second to scale the to the actual demand of the workload itself and costs nothing
[07:48] when ideal. On the software side, where teams often get surprised, classic compute means that you have to own the Databricks runtime and with every new DB release that you get, you have to do compatibility testing, upgrade windows,
[08:03] coordination with your teams and across all of the environments. But, on serverless, the runtime is versionless and Photon enabled and the platform team is now out of the TBR upgrade business and they they don't interact with that as much. Then, on the
[08:19] admin side, you have all of these cluster policies, utilization metrics, provisioning requests that are coming in through tickets and scaling decisions that you meet need to make and all of that was eliminated and the entire category of work goes away. So,
[08:35] Databricks manages all of that for you and the platform stops being an infrastructure operations function. Platform The platform team stops being an infra- infrastructure operations function and starts actually doing things that are expanding the capabilities on the platform.
[08:51] And for end users, you have developers who would file tickets, wait for clusters, but now they can just run their jobs within the spend guardrails that the platform team has already set.
[09:07] Now, this is one of the direct testimonials from our platform admins. Um and this alludes to just that. Of course, we saw the cost go down. That was the main thing that we were trying to achieve with this migration. But, what they really saw was this shift in
[09:22] productivity. The platform ste- uh the platform team stopped fielding those provisioning tickets and started spending time on the work that actually moved the platform, which is data quality, governance, and discoverability and searchability on the platform.
[09:37] And that same shift flowed through to all of the developers as well. So, now you don't have to wait, you don't have to think of like runaway costs, but this is truly self-service with governance built in. And at over 1,000 users that McAfee previously had, by eliminating
[09:54] that, that's what drove up the active users on the platform. And now we have different problems at McAfee, but that's a talk for another day. Um so, now of course we've talked about moving to serverless. It can make your life much simpler.
[10:10] But um you might know that serverless isn't the answer for every job. I think uh have you guys struggled with serverless migration? Please raise hands. Um but uh the most important thing that
[10:26] you'll do before you even touch a single job is figuring out which ones actually belong here. And the clearest candidates are going to be your Python SQL workloads or those that are short-running jobs or high frequency uh running at a high frequency
[10:41] where the spin-up time is real and the clusters that are already over-provisioned or under-provisioned or sitting idle between runs. And there's one requirement that's non-negotiable, which Mandar already touched upon, which is Unity Catalog. So, so for to be on
[10:56] serverless, you really need to be on UC. If you still have jobs that are running on legacy Hive Metastore, this uh Unity Catalog migration needs to happen first. For McAfee, fortunately, 85% adoption was already there with UC and that governance foundation was uh what made
[11:13] this whole migration feasible for us. And it's equally important to know the blockers up front. Um Scala and R aren't supported today with serverless. JAR libraries and notebooks are blocked, although I think we have JAR tasks in public preview now.
[11:28] So, it really depends on what you're trying to achieve and then making sure that you're aware of all of the limitations and which category that they fall in, and you need to evaluate against those limitations and then chart out the best path forward after careful evaluation.
[11:45] So, the question now becomes if you're an enterprise customer, how do you apply all of this across the 2,000 or so jobs that you have or more, and that's why we developed this qualification engine to take that raw
[12:02] portfolio of 2,000 jobs across the six workspaces and then turn it into a prioritized evidence-based migration backlog. So, every job that we had ran through these three filters that you see in here in parallel. The first was the
[12:18] infrastructure and cost profiler, which would look at the compute economics. So, anything with low cluster utilization, comparison of like on-demand versus spot VM types, we surfaced those jobs where uh we could get that information and
[12:33] then compared whether they would be a good fit for classic or serverless. Then, we had the workflow blocker profiler, which would scan the job definitions for known serverless hard stops. So, that's your compatibility signal, finding those blockers uh before
[12:48] the migration actually takes place. And lastly, we also built a custom script with McAfee where we added four precise filters on top, which would evaluate the average run time of the jobs, whether they were under 4 hours or not, whether they had any libraries, if
[13:06] they had any Python libraries. We don't We um We don't have support for wheel and jar dependencies yet. So, anything with a Python library, we would update the script to create the environment and the dependency settings for those jobs. And the jobs should have been in a successful running
[13:23] state previously, and only those active jobs were the one that we tackled for the migration. And any of the jobs that failed any of this this criteria, um they would not warrant for migration in the successive phases. Now, after we ran all of these filters
[13:39] and profilers, the class of based on the classification results, we then put the job into one of these three buckets that you see here. So, the hard blocks mean that you stay on classic. So, we all of the serverless
[13:54] limitations that you saw with like RDDs, Spark configs, or caching, those were the hard blocks, and we marked them to tackle them later on. And we continuously continue to review that list of jobs that we have.
[14:10] The unblocked ones, which are the SQL Python workloads, which are clear fits, are the one that we actually proceeded with. Um and we made the changes where we needed to, and these are the ones that you should tackle in your first phase of migration to gain confidence in
[14:26] the migration that you're doing. The soft blocks mean that you need to assess those jobs first. So, these probably need some sort of remediation or refactoring before they can proceed. So, in a lot of these jobs, we had to work around with all of the init scripts
[14:42] that we had or like older DBR versions which needed upgrade. We needed to refactor the code. If you're working with streaming jobs, then you need to handle that carefully as well. But once these were remediated, these were slotted for the second phase of the migration to follow up with.
[14:58] And one of the key decisions that you're going to make in this entire process is also going to be deciding the type of compute that you want to move to when you move to serverless. So, either you can move to the performance optimized mode, which can offer sub-second latency
[15:14] and are good for like really um SLA sensitive jobs. And standard mode on the other hand can give you about like 60 to 70% discounts if you're willing to take that take that slightly longer startup time which was right in the case for McAfee
[15:31] because most of the jobs that we were dealing with were bad jobs and we could move them to the standard mode and that was the right fit for us at that time. And knowing these jobs, knowing and categorizing these jobs, running through this qualification guide, of course this
[15:48] is great, but how do you put all of the governance in place so that you make sure your platform is ready to receive all of the serverless jobs that you're going to run. So, let's take a look at that. Now,
[16:04] before we moved a single job, we wanted to make sure that the platform was ready and the guardrails were set in place to ensure a smooth transition. So, start with a plat with the platform guardrails in place. So, for McAfee, that meant a security review with the
[16:19] infosec team and given that with most enterprise customers that that's a common ask, you should make sure that you're following through with your infosec teams. And also, when you move to serverless, if you have any private connectivity requirements, you need to also work
[16:35] with NCCs to make sure that you're solving for that before you move to serverless. One other critical item is making sure that you have a rollback strategy in place, not as a formality, but because these well-documented rollback steps are
[16:50] what what is going to give your stakeholders the right confidence to actually begin with. Next come the cost and access controls. And this is where we put the serverless usage and budget policies and tagging in place. So, before the migration even started at McAfee, every
[17:08] serverless job was using a usage policy tagged by the team and the cost center. And that before baseline job, before you began the migration and you baseline the cost and performance. Without that, you can't really show for any improvement. So, make sure that you can
[17:24] also allocate that into your whole system as well. And one thing that we actually like that surfaced when we were trying to put the cost and access control in place was that you really need to roll these policies out carefully. So, we built custom queries to actually surface any
[17:41] untagged jobs which were missing these auto-attached policies. And we audited those policies with every batch that we were trying to migrate and not at the end. Now, the last piece is the environment standardization. So, for the jobs which
[17:56] for the relevant jobs, we created a shared environment.yml and specified the Python version and the packages and dependencies which were inherited by the relevant jobs across all of the workspaces. And that single decision was also what made the migration more reproducible and repeatable than
[18:13] manually working with each of the jobs that you're working with. So, put your infrastructure in place before the migration infrastructure um before the migration actually starts. Now, look Let's look at the entire
[18:29] process end-to-end and look at what was the playbook that we followed. So, for newer jobs or new workloads, it's very easy. You can directly move them to serverless. There's no uh run runway for that. But, for existing jobs with every
[18:45] batch across every workspace like specifically with McAfee, of course, you need to have that framework in place. So, this was the framework that you need to follow. So, first, you qualify your job. So, the profiler tool that we looked at and the three-bucket framework produces your backlog.
[19:01] Start with your top 10 to 20 candidates. Keep it small. The order really matters here because you don't want to run everything in one phase. Um phase one, try to tackle the unblocked jobs um that we categorize as unblocked in the previous uh slides to help you gain that confidence in the
[19:17] migration. And then tackle the ones that need refactoring. And you can you can use the same playbook across the waves of the migration as well. Then plan and secure. Get that security sign-off, configure that NCC, set that budget policy, apply tagging,
[19:33] lock in the base environments where needed. And before anything touches serverless, document the costs, the performance so that you have your baselines in place. And that before number is going to help you prove the after as well. So, that's really, really important.
[19:49] Then batch in waves of 10 or fewer. 10 isn't a magic number, but we did hit technical issues. But the whole point of having small batches is to find that common problem and fix it in a job in a batch of like 10 jobs versus like a
[20:05] batch of 100 jobs. So, work with your business units on resolution as needed before you scale the next wave. And then continue to make sure that you're monitoring and comparing the last couple of runs. So, for McAfee, we compared the last 10 runs with the latest 10 runs on
[20:20] serverless. And we excluded the day of the migration to make sure that we were avoiding any skews in comparison. And we also built custom dashboards to surface that cost variance, that runtime delta, and reward the candidates as needed so
[20:35] that every batch had a clear baseline to compare against. And we were able to reward when needed. And the guardrails are are what helped us. So, we set alerts on the costs of the jobs. And if it triggered, we rewarded immediately
[20:51] and investigated the job after. A rollback strategy should should not be looked at as a failure because that's what the system was built for. We wanted to catch those jobs early. And that reward mechanism is actually what gave McAfee these enough confidence to let the entire
[21:07] migration program run through in the first place. And then, once you have that down, um you can just repeat and rinse this entire cycle. And this looks really simple on paper, but now I'll let Mandar talk through all of the challenges that we ran into.
[21:24] Okay. All right. down to 378,000. That's almost 364,000. We cut it almost exactly in half. And that's not a projection.
[21:41] That's measured across almost 15,000 job runs that we observed during that time frame. This is what we we saw we noticed. Now, um
[21:58] the compute spend numbers are good, but the number that really gets me is the runtime drop, if you look at it. 83% drop. Now, this is not all the jobs, specific high-impact jobs that we noticed. So, not all of your jobs that you
[22:15] migrate to serverless are going to have the same runtime drop. So, I really want to be real here. 83% drop is a best-case scenario, but I think that gives you an idea where the ceiling is,
[22:30] right? And Now, here's the part I really want you to take away. We did this by migrating about 184 jobs within that time frame. We did not go after all the jobs. That's
[22:47] not advised. We wouldn't advise it based on our experience. We moved only the right jobs. And Amisha talked about the profiler. Use the same playbook, profile the jobs that really qualify. And then after the migration, we watched
[23:03] e- every batch against a baseline and kept rollback ready all the way. That's the takeaway here. Cost will drop. Do the selective migration, dashboard-driven monitoring, and guardrails.
[23:19] The method The method is what is repeatable here. And it's the method we are sharing with you today. So, I hope really it helps. Now, let's take a look at the messy part.
[23:34] This is the part where the playbook is not going to prepare you. This is the slide I wish someone had shown us when we were doing the migration at the time. Here is what we learned the hard way.
[23:50] Serverless doesn't forgive what classic clusters were hiding. Here is just an example, right? It's not on the slide there, but I'll walk you through. I think most of you have this example in your environment. Let's just say there
[24:06] is a column that's supposed to be a number, but it's stored as a string. On classic, Spark just quietly converts it. Job runs fine. Everybody is happy. You move it to serverless,
[24:21] it's strict SQL. The job stops cold. It goes back to some of the limitations that Amisha was talking about. And it's a type cast error that we found in most of our jobs. So, our first reaction when we saw the
[24:36] failure, our first reaction was serverless broke our job. But, it did not. The mismatch was always there. Classic compute had just been more forgiving all this time.
[24:51] Nothing here was serverless feeling. It was serverless just turning the lights on. That's it. So, our advice is that find these issues before you migrate, not after.
[25:06] It's better when these issues show up in your staging or testing than when somebody's actually waiting for a report. That's not going to be a a good message or email that you end up receiving.
[25:23] Now, here is another messy part. Sometimes a job does not break. It just gets more expensive. And this is the one that also scared us very early in our migration journey. We moved a job over and it came back um
[25:39] more than three times the cost of the um runs that it was um on when it was on classic. So, our two weeks alert triggered, we got the alert, we migrate we rolled it back to classic, no harm done.
[25:54] But here is the thing. That wasn't serverless being expensive again. When we profiled this job very closely, worked with the owners, tried to understand the code, the sequel it turned out that the sequel code was just inefficient.
[26:11] And on classic cluster caching and other features that were just quietly hiding this cost. Serverless was billing us honestly for the waste that was already there. So, we worked with the owners. We optimized the queries and jobs. We
[26:30] remigrated the job and it dropped down again and it was cheaper than what the classic was um costing us. So, we have these four common anti-patterns that I have listed here on the screen.
[26:45] And whether it's full table scans, bad joins, redundant work or partitioning or small table file size type of issues. These are the most common usual suspects in our environment.
[27:01] If any of these resonate with you, when you develop the playbook, make sure you tailor it tailor it to identify these types of issues. When you start identifying them, hunt them down before you migrate. And most of your cost surprises will disappear.
[27:22] Now, when we did the migration, we did not have Genie code. You all do. You can do a lot of this by just working with Genie code, and it will make your life much easier. So, today instead of
[27:39] digging through the execution plans, you just ask, "Find the scans that aren't pruning." So, just a broadcast join here, and it does the work for you. You may have to refine it and tailor it to your own requirements, but still it will end up saving a lot of time for you.
[27:54] And that's the whole point of this slide. It's not about the tool that we are trying to brand or market here. It's about how much time your team can gain back and how much more efficiency you can also experience. And to be clear,
[28:15] if you start profiling all of these jobs, and you start identifying a lot of these issues, you are going to start identifying common patterns that will also help you educate your own users of the platform.
[28:32] So, the sky's the limit. You can take it as far as you would like in terms of how far you want to take the responsibility and how far you want to go to share that responsibility with the code developers as well. So, use Genie code. Um it will find
[28:48] issues for you and possible fixes. You decide what you ship and when you ship, right? All right. Amisha talked about this and just in the interest of time, I'm not going to talk about every single line here.
[29:06] But, this is just an example of the operational checklist that we followed at the time. And just imagine a job either failing or becoming more expensive. As soon as you get the alert, your team
[29:23] is going to be equipped with reverting it back and you avoid all the confusion and the drama that it might cause later on. So, this is the slide where I just want to show you how it was planned. Right? And
[29:38] that's the one key takeaway from this one. Build the dashboard first. If you can't compare serverless against old classic baseline, you can't really prove anything. And you can't defend the work that your team is going to be doing.
[29:57] So, one, the dashboard, daily cost per job. We had this dashboard up and running and that was the first question we asked ourselves every morning when we were keeping track of the migration. We compared, as Amisha said, we compared the last 10 serverless runs against the
[30:14] last 10 classic runs. And just honest comparison. Second one is the tagging. Every workload tagged by team, cost center, environment, because what you can't see, you can't really manage. And the last one is the guardrails.
[30:30] These are super important when the cost of the job goes above, right? And I think both of us have talked about it. Again, you can use Genie code for this one, for implementing the checklist, and again, you can tailor it to your own environment. Same idea as before. You
[30:47] describe what you want, and it will build it for you. And for a lean team like ours, this is extremely good. This is a really good win for us. I really wish we had this when we started our journey about a year ago.
[31:08] All right. So, I talked about a lot of numbers. This is the current state today. As of last month, 72% of our queries run on serverless.
[31:29] And if you look at the number of queries that still need to be fixed, it's about 900,000 that we still um have, and those are the fixable blockers. And we know exactly what needs to be done there, because it's been part of our playbook.
[31:44] We need to work with the core developers. We need to work with the privacy and infosec people to make it more approvals. We know exactly what to do. So, those are the fixables. But, there is that 18% on the far right side that you see. That's where the hard blocked um work
[32:01] loads are. So, Amisha talked about these. There are a lot of streaming jobs that we have. A lot of jar uh custom jar uh based jobs that we have. And that will take time. But, we are happy with where we are
[32:18] today. 72% especially with the direction that Databricks is going towards. I think it's it's giving a lot of efficiency and productivity for our teams.
[32:33] So, that's our story, where we started, what broke, and where we are now. Before I hand over to Amisha, I just want to say that modernizing infrastructure and cost controls are really a shared responsibility. So,
[32:49] real quick, a big shout out to my team, the platform team, um and the broader data engineering team at McAfee, who did the actual work. I just happened to be the lucky guy who's presenting these good numbers, right?
[33:05] So, now, uh the obvious question is, how do you start? Amisha? All right. So, thank you, Mandar, for walking us through McAfee's journey. So, what do you This is something that
[33:21] we want you to leave with. Um when you think about moving to serverless in your own environments, think of it this way. The first week, just think about and sit on it, what's the reason for you to migrate? Is it idle compute cost? Is it developer
[33:36] productivity? Is it platform infra overhead? The reason for your migration is going to drive this whole journey forward. Once you have that reason down, then run through the serverless profiler, um do
[33:51] the whole decisioning framework, and get that list of top candidates, and make that decision about whether it's worth for you to do the migration or not. In the next couple of weeks, my advice would be take your top and clear fit
[34:07] jobs, and run those steps. Run through the infosec reviews, configuring NCCs, budget policies, and so on, and create that dashboard that will help you drive that decision forward. Set those rollback triggers before you even start.
[34:23] And do not do not promote any jobs to production in production to serverless before you test them out in dev. Um and then for once you're done through with that, that's when you think about scaling that framework, right? Um, if
[34:39] the first 10 that you tried out in dev helped, then repeat that same playbook and do that same concept of like batching it in certain batching it and then basically following that same discipline when you try to move the other jobs as well. And bring the business units with you in
[34:56] those conversation and work with them to do the actual migration. The playbook itself is not going to change, but the way that you migrate each of the jobs might. So, at McAfee, we started with 10 jobs and now we are at like 40% of the jobs now running on serverless and growing.
[35:13] And another piece of advice that I have for you is do not do this alone. If you work with Databricks, this is a joint partnership and that's the reason that this worked at scale and at the pace that it did at McAfee. I remember talking to a lot of different
[35:28] enterprise customers that were still struggling with this when we started out with this a year ago. Um, but you are the ones that can bring the tribal knowledge about your enterprises to us, the workload context, the business unit sequencing, and the organizational judgment that you have.
[35:45] And what the business can really absorb when you do this migration. And that kind of internal knowledge can only come from you guys. What we can help you with is that qualification tooling, making sure that the cost visibility infrastructure is in place. You have the guidance around how
[36:02] um around how to do the migration itself, and that fast feedback loop. Because as you may know, serverless is changing every day. So, for example, when Mandar mentioned that we had the typecast error, we realized that it was because of the ANSI mode be ANSI SQL
[36:17] mode being enabled on serverless and we were able to fix that quickly because we were in the room with them. And when a job a bit uh showed like ballooned up costs, we were able to figure out what was the uh reason behind it and if it's something that's tunable.
[36:33] So, that real-time velocity is what kept the project moving forward and prevented it from stalling. And this matters going forward because serverless is not a static target right now. There are new capabilities being shipped every day with stronger cost controls,
[36:50] expanded environment management, better governance, uh which are getting shipped continuously. I think just last week we released a set of skills that can help you with the serverless migration. So, make sure that you partner with your Databricks account teams and your DSAs
[37:05] to make sure that you're ahead of the curve and you're not just like catching up. Uh don't treat this as a solo project. We're here to help. And then lastly, I just want to caveat with saying that we just we didn't just in like do a serverless migration. We're
[37:21] making sure that McAfee is aligned to the direction that Databricks is already moving in. And that direction is serverless first. So, that you guys are doing doing more productive work, managing less infrastructure, and you have that
[37:36] consistent governance with you see and better visibility of your platform, and you're modernizing in the right direction with us and not just improving your current state, but you're shortening the path to whatever comes next on the platform. And this work, I promise, is going to compound every job, every warehouse that
[37:53] you move to serverless is going to make sure that your whatever you're doing is compatible with the future capabilities on the platform so that you're able to absorb whatever's coming. And McAfee is a clear example of what that looks like um when it is done with
[38:09] intent. And with that, thank you so much for joining us and being here today. Um I don't think we have a lot of time for questions, but you can come up to us after. But, if you want to follow either of us, there are LinkedIn QR codes. And again, thanks so much for joining.
[38:26] Thank you.

Learn more about the Databricks Data and AI platform.

The information provided herein is for general informational purposes only and may not reflect the most current product capabilities or configurations.