Migrate to Databricks Lakehouse Serverless: 60% Performance Gain and 20% Cost Reduction with DBT
Summary
- GetYourGuide migrated from its legacy Revolus system on Databricks job clusters to Databricks Lakehouse Serverless paired with DBT, achieving 60% faster execution times and an unexpected 20% cost reduction even at the scale of a leading European travel platform with over one billion euros in revenue.
- Databricks Lakehouse Serverless outperformed job clusters for SQL-heavy workloads because auto-scaling eliminated manual resource estimation and cluster tuning, which had been the primary source of developer friction and wasted idle capacity.
- Integrating Databricks AI Query with Unity Catalog context reduced documentation hallucinations to nearly zero, enabling complex SQL chaining, ad-hoc analysis, and governance across the unified data platform.
Migrate to Databricks Lakehouse Serverless: 60% Performance Gain and 20% Cost Reduction with DBT

Maintaining custom transformation tools at scale creates technical debt and developer friction. GetYourGuide's legacy Revolus system on Databricks job clusters required constant tuning, resource estimation, and version management. When the team standardized on DBT and Databricks Lakehouse Serverless (formerly known as DBSQL Serverless), they achieved 60% faster execution and unexpected cost savings even at scale.
Learn how GetYourGuide implemented this migration and why Databricks Lakehouse Serverless outperformed job clusters for SQL-heavy workloads. Discover how auto-scaling eliminated resource guessing, how DBT's integration with Databricks simplified pipeline management, and how AI Query with Unity Catalog context reduced documentation hallucinations to nearly zero. This architectural pattern demonstrates the power of leveraging native Databricks capabilities for faster, cheaper, and more maintainable data platforms.
🤝
Chapters
00:00Introduction and GetYourGuide Company Overview02:49Legacy Setup: Revolus and Databricks Job Clusters04:07Challenges: Resource Management and Cluster Tuning06:14Project Goals: Improve Developer Experience and Modernize07:35Target Solution: DBSQL Serverless and DBT08:24Results: 60% Faster, 20% Cost Reduction09:11Performance Analysis: Why Serverless Was Faster10:14Cost Analysis: Auto-Scaling and Execution Efficiency11:35Documentation Challenge and AI Query Solution with Unity Catalog13:27AI Query for Complex Analysis, SQL Chaining, and Governance14:31Key Takeaways and Best Practices
FAQs
Why did GetYourGuide migrate from Databricks job clusters to Databricks Lakehouse Serverless?
GetYourGuide's legacy Revolus system on Databricks job clusters required constant manual resource estimation, cluster tuning, and version management, creating technical debt and developer friction as the platform grew. Databricks Lakehouse Serverless with auto-scaling eliminated the need to predict resource requirements, and the migration ultimately delivered 60% faster execution and 20% cost savings.
How does DBT work with Databricks SQL Serverless?
DBT integrates natively with Databricks SQL Serverless as a SQL-first transformation framework that manages pipeline dependencies, version control, and environment promotion without custom infrastructure. In GetYourGuide's migration, DBT replaced the custom Revolus transformation tool and combined with DBSQL Serverless to create a more maintainable and performant data platform.
How did Databricks AI Query with Unity Catalog context reduce hallucinations?
By connecting Databricks AI Query to Unity Catalog's metadata and governance context, GetYourGuide provided the AI assistant with accurate, up-to-date information about their specific tables, schemas, and business definitions. This grounding reduced incorrect or hallucinated responses in SQL assistance to nearly zero compared to using a general-purpose AI without catalog-specific context.
What performance and cost results did GetYourGuide achieve after migrating to DBSQL Serverless?
GetYourGuide measured a 60% reduction in execution time and a 20% cost reduction after migrating from job clusters to DBSQL Serverless, even though the team had not anticipated cost savings. The auto-scaling behavior of serverless compute was identified as the primary driver of the unexpected cost improvement, as it eliminated idle billing on underutilized cluster resources.
Full transcript
[00:08] And now, uh first of all, welcome to the data summit Databricks. I hope you are all having a really great time here. Also, not only here in the summit, but also San Francisco. It's my first time here myself, so I'm very excited. Uh a little bit about me, I'm Giovanni Corsetti, uh Italian-Brazilian, so you might pick
[00:24] up my accent. I actually live in Germany right now, where I work for GetYourGuide as a data platform engineer. So, the topic of the presentation is how we actually migrated from our legacy pipeline setup using job clusters to
[00:40] something more modern with DB SQL serverless plus DBT. It's more on the basic level, perhaps slightly intermediate, so don't expect any crazy AI AI breakthrough, just like we heard in the keynote. Sorry for that.
[00:55] And I wish you all can get something useful at the end of the session. So, before anything else, uh for the ones that are not aware, uh what is GetYourGuide? GetYourGuide is actually a leading travel platform for booking tours, excursions, or any type of activity that
[01:12] you want to do when you travel, or well, if you want to explore your own city. And it's actually the number one experience platform in Europe with 1 billion plus revenue for the fiscal year of 2027 and 2025. So, for the European folks, probably you
[01:27] heard about that. Uh if you're not European, I also hope you heard about GetYourGuide. And GetYourGuide tries to answer the question of what to do. So, the same way that when you have the question of what should I buy, you think about Amazon, or where should I stay, you think about
[01:43] booking.com or Airbnb, GetYourGuide tries to fill the gap of what you can do when you're traveling. A little side story here is that before coming to San Francisco, I actually went to Vegas. You know, the company's paying the flights, so why not take a little bit of holidays?
[01:59] Then in first day in Vegas, went to casinos. Pretty fun, but then lost the money. Second day, I did not know what to do. So, I opened an app. In this app, I got a tour that picked me up in the hotel, drove me to the canyons, spent quite
[02:15] some time in the canyons, took nice pictures, explored, had a lot of fun. The same shuttle drove me back. And yeah, that's actually the app of GetYourGuide. This type of experience for people that want to explore the place they are at. So, yeah. The talk's not only about is
[02:32] is not coming only from employee, but also from customer. And I hope it also adds some value there. Anyway, uh less sales pitch, and let's talk about more of the technical part. So, I will just briefly talk about the setup we had before for some
[02:49] contextualization. And at GetYourGuide, uh we had this custom building tool called Revolus. Revolus was essentially a Scala repository where people could go there, they could write SQL transformations with templates. Then, during compilation
[03:04] time, those templates would be rendered. Airflow would pick them up, would add some parameters like start date, end date, and so on. Then, with a Databricks uh operator, it would trigger that uh on a job cluster.
[03:19] And this would essentially run our transformations. So, if only by listening to that, you already feel overwhelmed, just imagine working with that. So, it was very hard first to onboard people on that, because well, uh no one in the market has experience with
[03:34] the tool you build at GetYourGuide. Second, from a platform perspective, we really struggled to maintain things working, because as well, any kind of security update or new features that people are asking for, the platform team has to go there and implement ourselves.
[03:52] And third, once you're only trying to keep things running, you never have time to effectively improve the system, right? Uh a little bit of side note here, this setup was actually created in 2022. So, back then, Databricks SQL serverless was not a thing.
[04:07] The engineer that developed that, he did try to use DBT back then. But then, uh getting a little bit technical here, DBT needs uh JDBC endpoints to connect with Databricks clusters not provide. But then, he was able to like hack his
[04:23] way through by backing a Docker image with DBT Spark, then exposing the Spark context, connecting via DBT, and he made it work, DBT and Databricks clusters, but it was very complicated, and that's why he actually created Revolut from scratch.
[04:39] So, just a little of a side story here. Anyway, uh going back to the main setup, so for every Databricks cluster that we had, we had to manually specify uh the number of resource. So, how many workers would have, uh the driver type, the node type.
[04:56] And of course, we had some defaults, but things are very tricky there because if you give very tight defaults for Databricks cluster, we you end up with a scenario where you are always undersizing the transformations, and then you're going to have jobs running maybe two or three times longer because they are resource
[05:12] hungry. You're going to have jobs failing because of out of memory, and you're going to have those very annoying Datadog Datadog alerts popping up saying that, "Hey, you have to increase the resource of your Databricks cluster." Second approach is that you are very
[05:28] generous in the default, but then, probably only Databricks get happy because you end up paying more, but then your pockets get empty. So, pretty tricky to find a very good sweet spot there. And then, because of the nature of the business, so, you know, the company is
[05:43] growing, we have some transformations that they just grow with We have to automatically adjust them. Other tables, they are just dynamic by default. So, for example, a booking related table. Maybe during weekends, Easter, or school
[05:59] holidays, you are going to have three three or even like five times the load that you usually have during the weekdays. And it's just not possible to go there and properly estimate how many resources should allocate for that.
[06:14] Then, uh um with that said, let me talk about what were our goals. So, first, uh we wanted to improve the developers' experience. That's because the field of traveling is already hard by the way it is. And analysts or bank
[06:31] engineers, they should be focusing much more on the business logic over fighting the technology. Second is that especially now in the era of AI, with like cloud, uh uh oh oh and our code base, uh Fable, they wanted to expand in the whole
[06:47] world. It's very easy to build things. Everyone's a builder. But what really matters is not what you can build, but if what you build can actually survive the test of time. And in our case, we did notice that in-house, we felt that our technology was really struggling the test of time
[07:03] because we always need someone to keep there uh making things work. And finally, with all that said, we are pretty legacy. Now, with this whole, let's say AI fever, everyone's using AI, we also wanted to, you know, uh surf the wave. Let's say it that way.
[07:20] And for that, we needed to somehow also modernize the setup. Then, what was actually the target solution? So, what did you want to implement before, uh actually even having a solution in place? First is that we wanted to get rid of double
[07:35] clusters and move fully to a serverless warehouse. That's because developer experience is much simpler, and initially we actually expected things to be slightly expensive. I will go back on this point later. But for the developer experience it was worth the price.
[07:52] Second is that we are just tired of maintaining a transformation tool in-house and we wanted to move to probably the most well-known transformation tool in the market, which DBT. Also because we saw that it has at least now very good support with Databricks.
[08:07] And finally, after everything is done, we did not want to just shift and lift, but also try to leverage some of those Databricks AI functions that they are also advertising those keynotes that we listened before. So, I will directly jump to results and then
[08:24] I'm going to deep dive into them. So, first one is that for performance it actually got much much faster by 60%. We were expecting things to get faster because well, that's the nature of serverless. Serverless is faster by default, but not by that that much.
[08:41] And what really surprised us is that by using DBC SQL serverless over job clusters was actually cheaper at scale. That was definitely one of the biggest shocks because when you were doing some POCs, we got a very different result
[08:56] than we got in production. So, when you talk about performance, the first one is that what made things so much faster, right? First is that at scale it's just impossible to fine-tune our jobs.
[09:11] And the moment you have some undersized clusters, they are going to damage your performance. You are going to have your jobs running two, three or even five times longer just because it's lacking resource. Second is because when you are managing job clusters, you always have to take
[09:28] care of this process of upgrading the DBR version. And for serverless that's done by default. You don't have to worry about that. Serverless always has one of the latest versions of DBRs that also comes with upgrades in the software, which
[09:43] makes it faster by default. And then, if anyone here ever used job clusters, you know that it has this very annoying spin-up time that can be 1 minute, 5 minutes, 10, 15. Maybe you're unlucky and there are no resources available. And yeah, uh once
[09:59] you have a lot of transformations, those minutes, they become hours, and things really stack up. But then, the most surprising part is that what made execution so much cheaper, actually? So, first is that auto scaling really
[10:14] shows its value at scale. So, for the oversized transformations, we always had this little bit of fat that uh was exposed, but you know, it was maybe like 5% wasted. But then, once you have hundreds of thousands of
[10:30] transformations running weekly, this extra part really shows its value. You really feel that the workload um balancing is doing its job. Second is that serverless, by nature again, is just faster than job clusters. So, even if you're paying, let's say, I
[10:45] don't know, $1 uh per hour in a job cluster, and serverless is $2 per hour, but one finish in 10 minutes, the other finish in four, you still are uh gaining something there. So, for SQL, our experience is that serverless has a much better performance in general.
[11:02] Um finally, as I mentioned before, by always having those optimizations happening for serverless with new DBR versions and so on, you are in a way always leveraging and always harnessing the uh utmost release of Databricks.
[11:18] So, with all of that said, okay, nice. We are able to, let's say, lift and shift. We are not using wheels anymore. We are not using job clusters anymore. So, what now? What comes next, right? So, at GetYourGuide, we have like
[11:35] hundreds of thousands of tables, and some tables have hundreds of columns. Documenting all of them is just inefficient at scale. So, what we tried doing is that we have our DBT file, we have our schema file, we have some internal documentation,
[11:51] and we tried leveraging some like LLM models like Cloud or Codex and so on by ingesting all of those files, telling the model to document things, and checking the results. It worked really well for upstream transformations,
[12:07] but as I go down, we noticed that most of the documentation becomes a lot of word salad. So, you see a lot of words there without a meaning or obvious hallucinations. So, things that are wrong. So, for example, I work for this thing called CDP, core data platform,
[12:23] but then for the AI to us, uh customer data platform. So, this type of hallucinations. So, we want to mitigate that, and to be honest, even without much expectations, we tried the AI query feature of Databricks. And for our surprise, it was much much
[12:40] much better than just using those raw LLMs. Reason being that AI query actually leveraged the internals of Unity Catalog, meaning that all this information of like the lineage or metadata of our tables, it's all there. You It's all
[12:56] there for you to mine. And the same problem of documenting tables with AI query proved much much more efficient. We got rid of hallucinations because the model is much more aware of what was upstream, and
[13:12] it's SQL native, meaning that we can use that for much more than than that. So, for example, I won't go too deep here because it's not much of my area of expertise, but we do have some analysts that now they're using AI query for some very complex
[13:27] cities. So, I don't know. They have one CTE that's, you know, pure SQL. Second one is already calling AI query for some summarization. Then third one is getting the results of the second one. So, it's very fun because you can see people doing very complex, you know, we
[13:42] LLM chains only by writing SQL. Best part of all of that is for the AI query feature, it already inherits the governance from Databricks. So, I Again, if you attend the keynote, they have this AI gateway.
[13:57] It's more or less that. You don't have to take care of the access control just because it's all inside. Meaning that from a platform perspective, we really don't have to worry about, you know, making sure that analysts cannot touch uh PII data
[14:15] because, I don't know, the AI model hallucinated and it tried to read from table it cannot read. All of that by default. And yeah. Do I do have to say from the whole migration, the AI query was the most unexpected and the best surprise we had.
[14:31] Then, with all of that said, uh just a very quick retrospective of what should I take away from that? First is that from get-go get experience, serverless um DBT uh sorry, serverless uh actually works very well
[14:46] for SQL. And so far, we don't have any complaints. Second is that Databricks and DBT, they do have a very good integration with the DBT Databricks package, which actually in constant development. So, I know that in the new version, the
[15:03] 1.12, I guess, they have semantic layers of code, which is also a very big topic in the next days because uh agents, they are much better optimized for semantic layer than for like documentation general. And finally, if you're interacting with
[15:18] like table-like information, it's much more likely you're going to get better results by using the AI query over any type of vanilla like model like cloud or GPT. Then with that, uh the main conclusion also that at scale the basic serverless
[15:33] is going to beat job clusters um for pure SQL. And it also has like a better developer experience. Then I will conclude for now. Uh Thank you so much for your time. Thank you so much for your time. And a reminder for people to actually fill out
[15:49] the survey in the app. So thanks so much everyone.
Learn more about the Databricks Data and AI platform.
The information provided herein is for general informational purposes only and may not reflect the most current product capabilities or configurations.