Skip to main content

Lakehouse//RT: Real-Time Analytics at Scale with Databricks

Summary

  • Lakehouse//RT is a new Databricks capability powered by the Reyden engine that delivers millisecond query performance directly on the lakehouse, addressing the long-standing challenge of low-latency operational analytics without specialized external systems.
  • The Reyden engine targets real-time BI dashboards, Databricks Apps custom UI serving, and emerging AI agent workloads that require thousands of concurrent queries, and is planned to be incorporated into all Databricks compute options over time.
  • By building real-time performance natively into the lakehouse's separated storage and compute architecture, Lakehouse//RT eliminates the need for separate caching layers or specialized cubing services.

Lakehouse//RT: Real-Time Analytics at Scale with Databricks

Watch: Lakehouse RT: Real-Time Analytics at Scale with Databricks
Lakehouse//RT brings real-time performance to Databricks, addressing the long-standing challenge of executing low-latency operational analytics directly on your lake without specialized external systems. Shant Hovsepian, Distinguished Engineer at Databricks, explains how the new Reyden engine enables millisecond query performance at massive scale while maintaining the flexibility and cost-efficiency of the lakehouse architecture.
Learn how Lakehouse//RT powers real-time BI dashboards, Databricks Apps custom UI serving, and emerging AI agent workloads that run thousands of concurrent queries. The Reyden engine eliminates the need for separate caching layers or specialized cubing services, unifying high-performance analytics with lakehouse flexibility for the next generation of data applications.
🤝

Chapters

FAQs

What is Lakehouse//RT and what problem does it solve?

Lakehouse//RT is a Databricks capability built on the Reyden engine that brings sub-second, low-latency query performance to the lakehouse architecture. It solves the long-standing challenge that operational analytics workloads requiring real-time results previously needed specialized external systems outside the lakehouse.

What is the Reyden engine?

The Reyden engine is the core query engine underlying Lakehouse//RT, optimized for fast operational analytics on a separate compute and storage architecture.

What use cases does Lakehouse//RT support?

Lakehouse//RT targets real-time BI dashboards, custom UI serving through Databricks Apps, and AI agent workloads that generate thousands of concurrent queries. These workloads require consistent sub-second response times that were previously difficult to achieve on a general-purpose lakehouse.

Why did Databricks build Lakehouse//RT rather than relying on external caching systems?

Databricks built Lakehouse//RT to unify high-performance analytics with lakehouse flexibility in a single system, eliminating the need for separate caching layers or specialized cubing services. The separated compute and storage model of the lakehouse architecture provides the technical foundation for achieving real-time performance without adding external dependencies.

Full transcript

[00:20] I am pleased to be joined by Shaunt, who's one of our VPs of engineering. I'm actually not a VP of engineering. I'm just an engineer. Just an engineer. Yeah, the VP is too too Well, no, it's just like if if the president dies and the VP has to take over, it's a lot of responsibility.
[00:38] But but Shaunt, you've been with us for a while and your your greatest hits have included what? Databricks SQL? Yes. Yeah, originally led all the data warehousing stuff in Databricks SQL and we have some fans out there. Hey guys. Uh and like the AI BI and the visual data
[00:54] AI BI, bringing the whole you know, dashboards into Databricks and now recently you've kind of taken on a new a new project. Do you want to do you want to talk about that a little bit? Yeah, so we all know that the best data warehouse is the lakehouse. Yes. Um and the problem was we wanted to, you
[01:12] know, hit that like last little bit of low latency real-time workloads that were always a little challenging to do in warehouses. And that's what uh lakehouse//RT uh is. I think people thought for a while RT meant like requires therapy.
[01:29] You know, re-trying tasks, which is something that Spark likes to do quite a bit. But it's real time. Real time. Just because I think it was always weird to us that people thought something about the lakehouse architecture just made it slow or like hard to do fast
[01:45] things sub-second. Data warehouses could never do it. They're always these specialized solutions. Uh but there was really no technical reason why you couldn't do it with the this separated compute and storage. So we set out to kind of really build a cool system that optimized that entire flow
[02:00] um and that's what Lakehouse//RT is and we're super excited that it's available now. Can you explain a little bit more on what exactly use case we're trying to solve or trying to position this new engine on? Like what exactly is Yeah, so it's honestly it's the intent is, you know,
[02:16] eventually over the next year or two and just like kind of what we had done with uh Databricks SQL in the past with Lakehouse is we're sort of starting with this one use case around real-time operational analytics uh but this core engine, the Reyden engine, which is what Lakehouse//RT is
[02:32] based off of, will actually make it into all of the compute options. So, you know, it's it's going to be everywhere cuz like uh who doesn't like fast, right? Like uh Like every time I use uh Gemini or ChatGPT and you have to like pick the models, you know, you know, do
[02:48] you which model do you pick when you use Gemini, Jason? I just I just pick Gemini. Like Yeah, really? Yeah, I always go for Pro cuz it's like I want give me the best thing. But then it's slow. But it's like it's like a horrible trade-off. So, we we really we want the best option to always also be the fastest option. And
[03:05] so, it'll be everywhere. But for now, it works great for a lot of these uh sort of BI serving type of workloads and use cases, things where you would have to extract data into another system or use like a special cubing service uh or anything where you had like a caching layer for doing lots of lookups uh and
[03:22] this just comes up a lot for like any kind of real-time dashboards, slice and dice workloads, um kind of serving a web application and actually it's been very popular so far from Databricks apps where people are building uh you know, custom UIs and widgets that want to show data uh directly from the the lake. And
[03:40] so, it kind of fills that area um but eventually it's going to be kind of doing everything. Now, you you started off you called it RT and then you just said Reyden. Like which one is it? Everyone wants to know. yeah. So, you saw the keynote, um we love naming and renaming things
[03:55] sometimes. But, what what always happens is there's like the engineering team, um and then there's the other parts of a software company, and they never tend to agree on the names of things. Uh so, in this case, Reyden uh is sort of the name of of the engine.
[04:13] It's sort of like the thing that we built at the core of this. Uh and when the team first came out with the name, uh everybody thought it was Reyden, like um The Mortal Kombat guy, right? god of thunder and lightning, who's also a Mortal Kombat character.
[04:28] Okay. Which totally seems weird, but it kind of made sense cuz if you're familiar with Databricks and all the work we've done, there was always there was Spark. Right. Uh then you had Enzyme, which is like uh part of SDP. Yeah. And then there's Catalyst, which, you know, was in Spark.
[04:43] Yeah. And then we had Photon, which is like the the other thing that makes it fast. So, there's this sort of electricity theme. So, everyone thought this was like such a clever idea in naming this thing Reyden cuz it's like, "Oh, yeah, the guy from Mortal Kombat that controls electricity. Of course, this new
[04:59] orchestration engine is going to be named Reyden." Perfect name. And then the depressing part was uh yeah, the engineer who picked the name uh had no idea what Mortal Kombat was. Oh, really? Yeah, yeah. He, you know, he's
[05:15] uh the year he was born in this the most significant digit of that year starts with a two and not a one. So, he sort of missed That says a lot. Yeah. missed Mortal Kombat. Uh and he was really shocked everyone was saying that. But, it was actually stood for and it was kind of a a joke. We called it
[05:30] Reyden was a Reynolds Dream Engine. So, it's sort of an acronym cuz Reynold Xin, who's our co-founder who presented it, he was the one who was always on our case about doing a you know, try to build something cool. Um and this engineer, actually his previous project Uh
[05:46] who's had a great code name that I won't mention. Uh but it never shipped. It got killed. Okay. So, I think he was like, "Okay, yeah. Uh this is This time I learned my lesson. I'm going to name it after Reynolds." And then he can't kill it. Um
[06:02] That's a very very good idea. Yes. Yeah, it sort of worked. Yes. It's like, you know, it would be suicide. And uh Reynolds wouldn't do it. Uh and then if somebody else tried to kill it, then it'd be homicide. So, it was like it was it was safe. It was perfect. So, he learned from that lesson. Uh I have a question for you. Is Lakehouse
[06:18] RT relevant to AI and agents? And if yes, how? Yeah. So, you know, I one thing that we have seen um is that like agents it's it's like hard to predict what the
[06:33] exact behavior patterns will be in the future. But from the customers and the usage that we've seen so far, uh like agents you know, Jason, you have been doing data warehousing for a long time. You know how to write a SQL query. Yes. Turns out these AI agents kind of aren't
[06:49] aren't very good at it or they like they do all sorts of things. So, do really really complicated ones. some really complicated one-shot things. And what we were finding is a lot of these AI agents, they run a ton of queries. Excuse my language. Yeah. Uh Technical term. And like each one isn't as valuable as,
[07:05] you know, if you and I were writing those queries. So, it's like there was this really big shift where there was way more volume. And it was hard to like make or attribute like the same amount of value to each query. Mhm. And so, by luck, we were seeing this trend with a lot of the early customers we have
[07:21] that were trying out um Lakehouse//RT. And it it was like kind of perfect for them cuz like you want it to be fast and handle the high concurrency, which is what happens when there's just all the uh gigantic workloads kind of trying out a bunch of combinations of slicing and dicing. But you also needed an engine
[07:37] that was not like super opinionated. It it to be flexible. It had to handle like the small little query, but also those crazy complicated. So, we have no limits on like things like joins or complex aggregation, window functions, like uh It's like human beings like to use CTEs.
[07:53] Uh turns out AI agents love using CTEs. just like a large volume of queries from these agents, too, right? So, Oh, I'm being told that we're we're just at time here. So, I tend to ramble. I'm sorry. Never. Never. I love time, but I am not. I love it. I love it.
[08:08] But, um it was so great having you on the show. Um we'll have you again next year, too.

Learn more about the Databricks Data and AI platform.

The information provided herein is for general informational purposes only and may not reflect the most current product capabilities or configurations.