Skip to main content

Bridging AI Ambition and Operational Reality: An Industry Forum - My pick

Summary

  • Databricks co-founder Reynold Xin joins leaders from Meta, Rivian, Workday, ModMed, and Afresh in a 90-minute industry forum to discuss the foundational data and governance challenges that enterprise AI leaders must solve before moving from experimentation to production-scale intelligent systems.
  • A fireside chat covers LakeBase architecture, Genie for natural language data access, and OmniGen innovations announced at the summit, with Reynold Xin sharing which new capabilities he is personally most excited to use.
  • Panel discussions emphasize that data foundations must come first, that successful enterprises choose simplification over sprawl, and that governance works best when treated as a paved road that enables innovation rather than a compliance checkpoint that slows it down.

Bridging AI Ambition and Operational Reality: An Industry Forum - My pick

Watch: Bridging AI Ambition and Operational Reality: An Industry Forum - My pick
As enterprise AI matures beyond experimentation, leaders face the architectural tax of fragmented stacks. This 90-minute forum brings together leaders from Meta, Rivian, Workday, ModMed, and Afresh alongside Databricks co-founder Reynold Xin to explore the foundational challenges that now define market leadership. Move past hype to address how top companies eliminate latency between data and applications, slash compute costs, and establish governance frameworks required to scale intelligent systems.
Discover why data foundations must come first, how successful enterprises choose simplification over sprawl, and the real-time architectures powering agents at scale. Learn from companies tackling governance as a paved road rather than a blocker, managing AI tool proliferation, and thinking about interoperability. Includes fireside chat on LakeBase, Genie, and OmniGen innovations, panel discussion on production-grade AI practices, and strategic insights on winners and losers in the AI era.
🤝

Chapters

FAQs

What is LakeBase and why is it significant for AI applications?

LakeBase is a Databricks innovation discussed in this video's fireside chat with co-founder Reynold Xin that brings real-time, low-latency transactional database capabilities into the lakehouse environment. For AI applications and agents that need to read and write operational data at low latency, LakeBase eliminates the need for a separate OLTP database outside the Databricks Data and AI platform.

How are companies like Rivian and Meta approaching data consolidation for AI?

Rivian's presentation in this video covers how the company consolidated its data stack onto Databricks to reduce the architectural tax of managing a fragmented set of tools and systems. Meta similarly presents how transforming its data infrastructure eliminated latency between data and AI applications, with both companies treating simplification and unification as prerequisites for scaling AI.

What does it mean to treat governance as a paved road rather than a blocker?

This video's panel describes governance as a paved road when it is designed to make compliant behavior the default—embedding access controls, audit trails, and policy enforcement into the data platform so teams do not need separate approval processes. Companies represented in the forum describe this approach as enabling faster AI deployment because teams trust the guardrails are already built into the infrastructure they use.

What is OmniGen and what role does it play in the Databricks platform?

OmniGen is a Databricks innovation referenced in this video as one of the announcements made at the summit, discussed during the fireside chat with co-founder Reynold Xin alongside LakeBase and Genie. Specific details about OmniGen's capabilities are addressed in the fireside chat portion of this video as part of the broader discussion of new platform innovations.

Full transcript

[00:10] Hey everybody, thanks for joining us. Um welcome to our Wait, it didn't advance. There we go. Our Tech and AI Industry Forum. Um we're we're here to talk about, you know, throughout the course of the day and we'll get to the agenda, but basically, or not the day, next couple hours,
[00:26] um both AI and then how data is so important and fundamental to getting the AI right. Um Okay, I guess now I introduce myself. Hi, I'm Manish. Um I am the Tech GM for Digital Natives.
[00:42] Uh that is the part of my our business where we work with, you know, basically all the customers you're going to hear from today. So, you're going to hear from our friends at Rivian, our friends at Meta, our friends at Afresh, um and a number of others. So, you know, very technical,
[00:57] forward-looking, AI-driven companies. Um I do want to thank our sponsor, Entrata. Uh thank you.
[01:12] And here's the agenda. So, we're going to have um a fireside chat with Reynolds, who's kind of waiting over by the side, uh our co-founder and chief architect. Um we'll then have a keynote from Meta, keynote from Rivian, um an exact conversation led by my counterpart, uh my partner Sheridan,
[01:28] um and then, you know, some closing remarks from the one and only Jason Reed. Um because we're going to have a co-founder on stage, we do have to put a caveat up. Uh anything he says, you don't get to come
[01:43] back in like a week and say, "Reynolds promised us this." Um so, all right. With that, Reynolds, now you can come up on stage.
[01:59] All right. Where should I be sitting? Uh any anywhere you'd like. Okay. All right. You're good? Yes. Okay. So, the first question is going to be a little bit of a softball. Um then we'll get to the hard ones.
[02:15] So, you're a chief architect. Uh you are behind a lot of the innovations that have been announced this week. Um but you also actually use our platform. Like you are also, you know, an engineer. So, what are you most excited to like
[02:31] actually use? Yeah. Um for the those in the audience, just a little bit of a context on this event. Um they sent me a bunch of questions they might ask. And then, 5 minutes before this walking on stage, we look at the questions together, we said, "Ah, let's
[02:47] ask something else." So, everything is fresh. I didn't know about what's going to come. All right. Um the So, I think from a personal perspective, my um what I see the most exciting um because I might be able to use it quite a bit myself is actually uh two things that
[03:04] might not actually have maybe the biggest business impact, but one is like OmniGen, the elements genie. Or genie might have massive business impact. But the reason is OmniGen is kind of how um I've been using uh what I've been using to write code. Um I don't write a ton of code lately,
[03:20] but uh occasionally I spend some time writing code. And about a month ago, there was like an intense 1-week period where I spent a lot of time building the new sandbox product um with just two other engineers. And then, uh at the time, we kind of really struggled with um
[03:36] sort of the existing harnesses out there. And uh in particular, because I'm kind of doing the window and Cloud was like becoming super slow. And I started using Codeax a lot. And then, there's certain things I've done in Cloud that doesn't work that well in Codeax. And just getting them to work with each other
[03:51] um was kind of painful. As part of that also realized hey if you ask a bot Codex going to review your output, please don't embarrass yourself. It's actually going to work extra hard. And vice versa. So, Amgen was a great
[04:08] thing we've been building at Databricks internally for addressing such use cases. How do we get a different harnesses to actually work well with each other? It has a lot of other enterprise control features too that are very important I think from a control point of view but just from a personal point of view I'm kind of happy with Amgen. I don't have to
[04:25] carry my laptop around without shutting down the lid. I could actually use my phone to ask it to do anything and I could easily synchronize the different sessions across the different harnesses that I've been using. Amgenie is more of an exact persona like do I ask a lot of
[04:40] questions about hey how our product are trending, all kinds of different telemetry, what's the productivity of engineers across different teams. How do you compare and contrast? In the past kind of I either have to spend an enormous amount of time myself to do
[04:56] those analysis or I ask somebody who reported to me and then they ask somebody who reported to them and then that person asks another person to report to them and have the result come back in a week. Now often I can get something pretty accurate just within like a few minutes on Genie. Yeah, I've noticed that cuz you used to
[05:12] slack me asking me like hey who are our top five customers that do something. And now you slack me saying like hey what are these five customers doing? And Exactly. I don't I have to buy I can't buy time anymore by saying like oh let me go find out. Um
[05:27] I got to actually know. Yeah, so those two are kind of changing my everyday life quite a bit. Okay. Um I think the L tap announcement was pretty exciting for everybody. Um I think two questions I've I've been getting are one is like
[05:46] walk people through like the actual magic there. I think some people were a little bit like, wait, like how is this different than like what we already have with like base? And then there's Raiden, and then there's sequel, like how does it all fit together? Why is it actually different? Yeah. Um the
[06:01] Maybe I'll walk through the story of how it happened. We're talking about it. Um JP was asking me, it sounded like magic, what happened? So, what happened was uh we we acquired Neon, and we sort of built together this lake base architecture where we put the We we sort of separated Postgres storage
[06:17] from compute. So, we sort of tear out the uh compute stateless part of Postgres. And then we uh separated underlying storage into an open storage, eventually hitting the lake, right? Um And uh I think it was right after last data and AI summit, um Ali and I were
[06:33] actually debating about, hey, whether it's going to be possible instead of storing the data in Postgres row-oriented format, why don't we store it in a columnar format, and now analytics access go read it super quickly. And we're debating and debating. I was on the end of this is not going to be possible. You have massive performance degradation. And
[06:50] Ali's kind of like, yeah, maybe it'll be possible. Um and we're just stuck on debating. Um I think towards the end of last year, we hired somebody who's another new engine I'm not going to name who it is. Uh he joined, and then he looked at the problem. And then he's like, oh, uh
[07:07] this weekend I uh I finally got on boarded and figured out where all the code bases are, and I built a prototype that just transcoded the data. Um and I also looked into the actual storage layer. There's massive amount of CPU idle. Um so, what I told on stage was actually
[07:23] the real reason. The storage fleet had massive amount of I think it was like 20% CPU utilization average. So, it was like the extra 80% CPU right there that could be used to transcode the data. And once you transcode the data, it compresses
[07:38] better. Uh because columnar formats tend to compress at very least kind of 10 to 1 compared with our raw format. Um so the data volume shrink and now you can actually write the data directly to the lake without penalty without any performance penalty.
[07:54] We try everything analyzing them to see if there's decrease in any dimensions. We just couldn't find any. Um so it happened. Um now obviously that's more of a prototype. It takes a lot of actual effort to build the real system. Um and
[08:10] I think a lot of you are curious, "Hey, when is going to come?" I would just say it. Um I think if you ask for access now, you can probably actually get it very quickly in a matter of like 2 weeks. We might even make it generally available just in a few weeks. All right, I'm I'm going to go back to that forward-looking statement. Don't
[08:26] all come back in 3 weeks asking for access. Um Why is this a Why is this a big deal? Yeah. So I think the reason this But it's kind of funny because we should have shaped this whole keynote with a
[08:42] theme of agentic data foundation. But honestly, this between you and me um everybody talks about agent these days. Everybody ignores So then of course everybody talks about agent, but is it is it real? Like what what does it have to do with agent? Um and I actually personally didn't
[08:57] believe in like this agentic foundation that much until last night. Uh I was having dinner with a customer. And then the customer was telling me, "Hey, I have all this like log analysis and business signal I was trying to investigate and I unleashed the agents and they actually work reasonably well
[09:13] once your data's in the right place. But they don't have that damn data I need the most which is my actual application state. Um I see product telemetry changing. I see maybe the services have degraded SLAs. I don't know what the hell is happening in those databases."
[09:30] Um Do you And that person wasn't at the keynote so he just we're just having dinner. And then at the time I thought, "Oh, wow. Okay, Delta solves exactly this." Every single one of your table in your OLTP database would just appear
[09:47] as a data set in the lakehouse without you having to do any work. And we can make it every single table available because there's no pipeline, there's no CDC, there's no underlying hidden pipeline. It's just how the underlying storage works because now we can have a single unified storage.
[10:02] Um and with that, actually, the agents can correlate what is happening on the actual transactional databases with your product telemetry, with your uh services logs, and all that. And it's going to make it much more powerful. Um so, I actually think it's it's going
[10:18] to be pretty revolutionary. We obviously will take a while to adopt and sort of get everything in sort of the right place and build the right applications. But, I think once it's built, it's kind of difficult to beat cuz one of the biggest problem in databases for the last like 40, 50 years, since the beginning of
[10:34] databases, is how do you analyze your actual transactional data? Yeah, and I remember when I was uh when when I went to the University of Toronto for my undergrad, um every I think 1:00 a.m. to like 4:00 a.m., they'd bring down the entire sort of student information system
[10:50] uh because they had to ETL that data out. And that very ETL job um actually degrades the performance of the live databases. Uh we will solve that once for all. All right. Now, a question we're also asking I'm hearing
[11:08] is, all right, we're making it easier for agents to access all of this data, get faster access to the data. Um people are already talking about like token efficiency, tokenomics, all that kind of stuff. But, now we're also going to see an increase in infra cost.
[11:24] Um how should people be thinking about that? Yeah, so we actually recently did a similar analysis ourselves, and we realized, hey, aside from our biggest problem wasn't even token cost going up. It was actually our leaving the data topic for a moment. Um
[11:41] in Databricks engineering, one of our biggest problem was actually CICD cost going up substantially. Because of higher token cost, people can be writing more code and submitting far more PRs and changes, and the resulting CICD cost from those are actually much
[11:57] higher than the token cost. Um so Okay. The On one hand, so we kind of did analysis. Hey, so what Now, obviously, engineering product team like productivity is always
[12:12] very difficult to measure unless in sales. Sales is like how much revenue you bring in. But uh for engineering and for product teams, it's always very difficult to measure because what exactly do you mean by productivity? There's something that makes sense of it. Uh what we are seeing is sort of like commit count going up
[12:27] substantially. Um like those are objective metrics. And then there are some more uh so byte metrics of hey, are we shipping faster? Are we getting more? Um and we are seeing that. Um and I think we're probably going to see a similar issue with data also
[12:43] because we like in the last um I saw this already been published. Like I think in the last four quarters, every quarter Databricks revenue uh relative to growth have gone up um compared with the quarter before. Usually, as the company become larger and revenue becomes
[12:59] larger, your your growth rate decelerates. We've been accelerating four quarters in a row. Um and the re- I think a big part of that reason is because uh because of thanks to AI and agents, um it's becoming so easier to query data.
[13:15] Yeah. Um and so more people are querying data, and the same team are querying more the data more. Um the I think this is a trickier analysis, but one of the things is you kind of have to balance out uh cost going up,
[13:31] gaining more productivity, um and at the same time, you have finance teams that are trying to figure out what is my gross margin, free cash flow. Even if you just have productivity say doubling or tripling, it doesn't mean you can actually have the cost go up uh beyond whatever they were trying to forecast.
[13:47] Um so I don't I think the world is hitting this problem. Um like everybody's suddenly waking up realizing we're hitting this problem at about the same time. Um I I don't I don't have sort of a solution for you, but I do think um you have to take into account all three
[14:03] different sciences. Cost, but what what is ROI to that cost? And sometimes even just having enough ROI might not be sufficient because finance team have their own budget and uh some numbers they already communicated to investors. But the other on the other hand, I I do think we are trying as hard as possible
[14:20] to reduce cost from a product development point of view. Like one of the main reason for doing Rayden is we've observed, "Hey, if you're going to be unleashing armies of humans using armies of agents against the data system, um a lot of the the pre-existing
[14:35] technologies are not quite there. It becomes too expensive." So, with Rayden with really Castle T, we can actually shrink the cost quite a bit. Like if previously you want to run 12,000 queries per second with the very simple queries um against for example DataBricks SQL, you probably need to launch a
[14:51] 40 need a 40 cluster 40 separate clusters just to sustain that load. But Rayden, you can do it in like one cluster. All right. Um I'm going to throw you a little bit of a curveball here. So, something I've heard uh kind of through the ether is you have this
[15:07] comment that um everything in data is Um could you explain to all of us what that means? Yeah, yeah. Um I think what I said was uh everything's a gigantic pile of Um not uh
[15:23] but a pile The And what I meant was That's That's much better. It's a giant pile of not just We need more caveats in just forward-looking statements. It's a backward-looking statement.
[15:38] Um the So, what I meant was this. If you look at the most widely used or even not super widely used data systems out there today, they were all almost universally created 10, 15, 20, 30, 40 years ago.
[15:55] I mean, if you look at, for example, Databricks. Right? The core of Databricks is Spark. Um and Spark was created in 2009. It mostly matured around 2015. So, it's like kind of maybe 10-year-old technology. If you look at Snowflake, Snowflake was started in 2013. I think
[16:11] they started becoming super mature around around the same time. It's probably also 10-year-old technology. And if you look at ClickHouse example, maybe someone in the audience here is. It was also started around the same time. I think it was open source a little bit later, but it all sort of started. And all of this tech and I'm
[16:27] not going to go into like IBM DB2 and all Um the but all of this technology was kind of created initially for a subset of use cases. And then as they become more and more successful, more use cases start to show up. And so, the team started
[16:43] engineering around those use cases to support them. Um but the system was kind of not designed from the get-go to support all these different use cases. So, we make a lot of kind of trade-offs, assumptions, and abstractions, and design the abstraction in a way um that are not super amenable to the
[16:59] new use cases. Um so, after probably like a decade of organic evolution, everybody's trying to expand. It typically ends up being, "Hey, it works This system works reasonably well or sort of super well for like whatever
[17:15] core use case we're trying to support in the beginning, and we're kind of okay for everything else. Sometimes not even very good, but it's kind of like maybe possible. It's enough to support some revenue growth or whatever team that's like building that. Um but because there's so much revenue
[17:31] at stake, um and it's the systems become so complicated, nobody dares to let's rewrite it from scratch. Um there The industry is littered with projects that try to the the like rewrite some very successful system from
[17:48] scratch, and it failed. By the way, you can look up Wikipedia. There's a term called second system syndrome, which basically says uh you build the first system, it becomes successful, and you try to build a second one, you become overly ambitious, um and it fails. Um and there's the examples after
[18:03] examples. Um so, what ends up happening is as all of these companies or the teams become super successful, they're just stuck with that gigantic pile of um that were created 10, 15 years ago.
[18:18] And it's so important that you can't really do much to um and Spark's that way. It works pretty well for data engineering, but kind of sucks for a lot of low latency workloads. Uh so, I'm just trying to I'm not going to on others. I would just be
[18:34] on ourselves here. The But everybody else have the same issue. Um so, I was motivating our team um in the Is Is it motivating to be told that you're you're working on a giant pile of
[18:50] then team in particular, and I said, "Everything is a gigantic pile of Um you just have to suck less, and uh build something that's not a gigantic pile of shit." But then, it's actually in some to some extent easier to build. Um the gigantic pile of happened
[19:05] largely through organic evolution, and happens through not thinking through what are the whole total set of use cases you want to support. So, you make those design trade-offs abstractions, and then you end up any given point faster to support by hacking around your existing design.
[19:22] Right? So, if you actually start considering the uh maybe the total set of use cases you want to support from the get-go, you can actually design stuff in a way where the abstraction and the infrastructure makes sense. Now, the second system syndrome fails mostly because people are too ambitious,
[19:38] and they try to do everything at once. There's nothing wrong with designing the abstraction and infrastructure to do everything at once. What's wrong is try to do everything actually doing everything at So, with Ray then um we sort of became very careful.
[19:53] By design, we want to be able to support everything. But, in order to uh actually we wanted to focus on incremental delivery. Like, for example, our first uh release with Lakehouse RT is read-only queries. Um and on up to a certain scale factor.
[20:09] I'm still pretty large scale factor, but we're not going to ask you to throw like a petabyte of data against Um and then also there are certain guardrails on like very complicated joins. Again, we can support them reasonably well, but if you're going to throw like a 40-way join, we're going to forbid it.
[20:24] Um and we're going to try to limit, for example, scale uh there are very kinds of guardrails we're adding. But, then over time we'll expand it, we'll actually relax it. There's no fundamental reason why the system can't support it. It's designed to be able to do that, but we don't want to unleash the uh all those workloads up front.
[20:40] This is kind of my point. I think it's how and we're hoping this would be a it's a very successful example of how we can break the mold of um you you sort of that you start a company based on one system, and then after 10, 15 years it becomes a gigantic pile of
[20:56] that nobody wants to touch. Engineers, I don't mean the customers. Engineers don't want to touch cuz it's so easy to break them. Um and then as a result we're stuck in that sort of local optimum. But, rather let's like create a new system that could do it all, but deliver incremental
[21:11] in a way that actually so all of you don't have to deal with the giant pile of All right. That was something. Um
[21:26] All right, last question. Um What what do you see as the next big problem coming in the next 6 to 12 months that I think people may feel looming above them, but like aren't actually talking about yet?
[21:47] Um I mean, there's so many. I I think one of the thing is all right, I mean, obviously this self-serving, but I do believe deeply, which is uh in order to unleash the agents, uh I think there's so much focus on just the AI part, but honestly, I I feel the reasoning capabilities of the frontier models are
[22:02] pretty damn good now. Uh maybe maybe you get a little bit better for the frontier coding, but honestly, for vast majority of use cases, they're pretty good. Um you don't even need like OpenAI or Anthropic's models, like use even all the open-source models are
[22:17] getting pretty good. I've been playing with a lot of them. Um the But but now, I think if you feed it garbage, or if it doesn't have access to the right data, it's not going to be super useful. I use Gemini largely for questions I can answer on
[22:33] the open internet. Um and that's incredible. It it benefits me my personal life tremendously, but um once I tried to ask it about certain things I care about from a business point of view in Databricks, it's completely stuck. So, I think a lot of it, and obviously
[22:49] we're building products to make that better, but honestly, none of our product can help if the data foundation's not there. Um if your data's scattered across different places, there's many duplicates, there's 13 different customer ID databases, there's nothing Data Bricks or say the AI can do.
[23:06] Um and I think so much of the focus is on just, hey, build agents, build agents, build agents. I think there needs to be more focus. To in order to build those agents, we have to actually think about what's our data infrastructure. And by here, I don't mean like which vendor do you use,
[23:22] but rather, hey, um solve that data debt. It's kind of time to confront them because they're becoming far more expensive than the past. All right. Um that's time. Thank you very much, Ronald. Really appreciate it.
[23:38] Uh Round of applause for Ronald, please. Next up, we have JP from Meta. Um I was going to make a pile of joke, but you know.
[23:55] Thank you very much. I'm very, very excited to uh to come and talk to you guys about the transformation we've been uh we've been going through. Um and uh how we have moved away from paying the tax um of architecture thanks to the phenomenal uh pile of that the team has been building
[24:11] uh to owning outcomes. So, um we'll see how we actually are trying to trying to walk the walk and make it real. Um so, I joined I joined uh uh Facebook um now Meta about a decade ago. Um and uh
[24:26] when I joined, I joined from Cisco. Uh and we had just finished an implementation of SAP Hana, which at the time, I felt like I was living in the future. Um and I come into Facebook, and within a few months, uh we're running 20,000 models with the infrastructure that we
[24:41] have every single day for uh forecasting revenue. Um we we have daily data sets of 2.3 billion rows that we can store without without the infrastructure blinking. Um we had a tough quarter at some point, and we had to run that stuff on a on a weekly basis. Oh, sorry, on an
[24:58] hourly basis. And so, that I know worked like a charm. And so, I felt like I was living in in the in the future. That felt permanent. That felt like this was the the answer to all of our problems. Every query sub-20 seconds,
[25:13] analyst, no matter how much data you wanted to pull. Fast forward to 2 years ago and we still have the same infrastructure. And it's starting to become an architectural tax. We start seeing some challenges, some forward
[25:29] forward usage of AI AI coming and we uh we recognized that we were sitting on a on a timing bomb that was that was not going to help us scale for our businesses. There were three major forces that we
[25:45] saw at the time that would that would compound to making it a problem for us. One was data freshness. And even though we had our data that could update on an hourly basis if we really really forced it, no matter what the scale was, what we realized is that our latency of
[26:02] hourly data or daily data was was always hidden by the analyst that was answering questions. Exact asking a question, analyst taking their sweet time, couple of hours, days. And so, our own latency behind the scenes was hidden. And in a world where
[26:17] we started seeing exacts having the potential of asking questions themselves on the data, we saw that we we could not hide behind the analyst taking time to answer questions anymore. The second part was governance. And in a in the world of of
[26:35] analyst governed kind of data, when you have multiple metrics that have the same name or multiple metrics that are evolving, the analyst would make the decision in general on which data to use or which metric definition to use. In a world where we were relying on
[26:51] machines to make the decisions, that became a lot, lot harder. Um, because machines could only use triggers like usage, and usage is typically on the thing that has been used for the longest time, which may not be the most relevant one. And if you start hard-coding that, you know, use the newest, that that new
[27:08] definition of a metric may not be actually as as uh as relevant or not yet relevant. Um, so we needed a new way of making governance uh machine readable. And the last piece is on latency. Where we essentially expected people to start asking hundreds of questions, or
[27:25] somebody uh somebody asking one question, the agent behind the scenes running hundreds of queries, um, maybe even in a sequential fashion. And with a 20-second latency on on queries, um, we essentially saw that this was going to take minutes to answer a question every single time, which means that we were
[27:41] essentially going to break. So, we essentially started on our journey about a year and a half ago, um, and I believe in this case that the sequence really matters. Uh, I think Reynolds mentioned, right? We need to start focusing more on governance, and
[27:57] this is exactly what we did a year and a half ago. Governance, uh, where we started working with uh with our business partners to define hundreds of metrics, 400 plus across the domain of enterprise that we support, uh, owned by the business, built a tool for this, uh,
[28:13] business data portal, which is quite equivalent to what we have in Unity now. Um, we built 200 plus semantic models, uh, recipes, which are ways that our data scientists can encode encode uh uh explain explained uh analysis steps
[28:29] to take for uh specific analysis that we want to reproduce on an ongoing basis. And then cookbooks, which, you know, bring everything together and connect those recipes to the semantic model in a in a in a meaningful way. After we implemented all of this, uh, we started working on our people and evolving our people. And this is one of
[28:46] the things that has been the uh the most fulfilling and to be honest one of the one of the piece that I think would matter the most in the long run. We started transforming our data engineering community from owning the the pipelines and the data sets and the quality of those and the timeliness of
[29:02] of lending to owning the business outcome and owning the quality of the response that an agent could answer based on the full stack of data pipelines, semantic models, and recipes. And then and only then we started
[29:17] implementing uh at speed the uh the uh the technology. Um technology where we moved from um daily or hourly updates to uh microbatching, then microbatching for to uh to real time uh with Zerobus uh
[29:33] ingest, and then we started moving also in terms of the uh of the speed at which uh at which the technology allows to uh to have our data in the gold layer updated, which is subminutes, and uh and the speed at which the queries get answered.
[29:54] At the end of the day there are these four three transformation that we drove in parallel. Um one transformation is is uh is around the uh the business transformation. Building trust, building trust with our business partners, building trust with all of the uh all of the uh information that we can
[30:09] provide to the machine so that the machine give us good good answers. I believe that in the future every domain that we support, HR, supply chains, finance, anything in the enterprise support, um will essentially have uh information available by the
[30:25] agents. I think Reynolds mentioned as well, right? That the the agents are like the uh the the base model, the frontier models are getting better and better at knowing those domains. Um all we need to really do is provide context around our own data, um and make sure that the that those those models
[30:41] understand and connect to uh to uh to our data in in in a smarter way, which allows us to have scaled governance. Ultimately, we have built 200 plus semantic models, but a lot of this now can be assisted by the knowledge that the the uh frontier models have, and the
[30:58] knowledge of our data, and the lineage, and the transformations that we apply from the bronze layers to gold layers in in uh in Databricks. Finally, our people, and I want to emphasize that in that you know, I know there's a lot of uh people that fear
[31:14] about the roles of data engineering and engineering in general in the world of AI. Um and I actually genuinely believe that the tools that we have at our disposal now uh give us an opportunity to elevate that role instead of instead of reducing. As I said, I think owning
[31:29] the business outcome gets us closer to the business, and it essentially pushes from owning just the technology stack to really having the opportunity through tech assisted by agents to own the business outcomes that that the uh that the business actually needs to do.
[31:50] I'll give an example of our of our uh world right now and closing in in closing in that in that we have transformed our supply chain uh models with the ability to have real-time information that allows us to to uh to connect uh in uh in uh in a more meaningful way meaningful way with our business. We've
[32:07] accelerated with Genie the ability to create semantic models in a much much faster way. So, we're essentially seeing that as much as we utilize the technology that is becoming more and more available, we're able to accelerate all of that governance that we've put in place that becomes the foundation for
[32:23] everything else that we deliver. The the outcomes that we're driving for a business that now executives can ask real-time questions on pretty much every single one of our domains, and we're able to govern that information with the agents helping us in in uh in that
[32:40] transformation. This is This is not necessarily a demo anymore. It's AI that actually works in production. Be it with Genie or being with the with the real-time sub 1 second query capabilities that allows us to now
[32:56] have hundreds of queries running at the same time when we want to do detail analysis and agents. I'll leave you with this. If you haven't started to Reynolds points on the transformation and on unifying, start
[33:11] with the governance governance first right now. Create your semantic models. This is what works and this is the work that compounds ultimately. This is the work that with the technology will take you up into the and into the right. If you have already started scaling,
[33:26] work with your people. Work on your people. Move the move the goal post from from owning the data and managing the data and curating the data to owning the outcomes and the answers that the business needs answered through agents thanks to all of the data on the governed
[33:42] governed semantic models that we actually can build. And if you're worried about whether your team's out the future, I genuinely believe that the opportunities we have in front of us are much much greater if we take advantage of the governance and the capabilities that these that the technology gives us
[34:00] to to own the business outcomes for all of our businesses. And that essentially means that if we build the right foundation, we will ultimately compound the outcomes that we get in our favor.
[34:17] Thank you. Thank you, JP, for threading that needle at Meta. Please welcome to the stage Mikey Flynn, director of core data Rivian. Hey, everybody. So glad to be here. Delighted to see you all and yes, my pleasure to get to chat
[34:33] with you a bit about Rivian and and our data and AI journey. Um so I've been there about four, four and a half years. So when I started we were just starting to make vehicles. Now it's at the point where you can walk outside, you'll probably see a vehicle crossing the road. Um and it's been a thrill to see this amount of scale and this growth
[34:48] behind the scenes and under the hood there's actually been the inverse journey in the data world. We have had a very scattered and very disparate data landscape and an AI landscape. And the journey that we've been on uh behind the scenes has been very much one of convergence, simplification, really hardening on patterns that we feel like
[35:04] are scalable. Um and so it's just been interesting to see this juxtaposition. So I'll walk you with a bit through that journey, some lessons we've learned along the way, and give you a sense for where we are now, the sort of things that we're working on that we really think are going to take us uh into this next era. And if you're not convinced already, um here to convince you that
[35:20] your your AI problem is actually a data and a data architecture problem. But oh, quite the setup I've gotten from the previous people before me. So I think you'll see a lot of coherence in that message going forward. Um so um Okay, I've gone through here. So when I uh began, the intense focus was on the
[35:36] factory. How do we get the vehicles off the line? Um and over the course of a number of months of of really intense work, we were able to sort of build that production um that production capability. It started with an open empty shell, then we begin to sort of land equipment, we start collecting
[35:52] equipment from the factory. And once the first vehicles roll off the line, you have this sort of big relief like, okay, job job well done. Uh and then we have this realization, oh shoot, now there's vehicles that are out in the wild and they're streaming data to the cloud and the service centers are operating and there's recalls and there's things that
[36:08] we need to understand about the telemetry of the vehicles in the field. So we were far from being done, it was very much just the beginning. Uh a big part of the journey uh that we went on on the data side of things was reconciling the data ecosystem that existed for the factory with the data ecosystem that existed to support the
[36:24] telemetry data from the vehicles and then the ecosystem that supported that just running the business function. So, when I inherited this scope after we sort of got the factory up and running, we were spread across multiple clouds, multiple data platforms. There were dozens of tools being used for analytics and data science. Every
[36:40] tool and every team had a different pattern for ingestion, for transformation, for governance. And so, the surface area there was just a tremendous amount for us to control. It's really hard uh to get a handle over that. So, we stepped through it logically over the course of a few years, and we collapsed the the
[36:56] complexity of our our data platform into Databricks. We got very opinionated about how we bring data in, how we govern that data, um on and on and on. So, these first three boxes of platform and governance, ingestion and transformation, and the patterns and tools that people use for data science and ML, that became very
[37:12] clear. Once again, we had one of these, "Okay, job well done" moments. You know, if we've we're victorious, things are good, we've hit the finish line. Uh and that's just about the time that all of these AI tools start to proliferate and start to come in. And once again, we realize, "Okay, we're at the beginning of yet another journey." Um so, sort of
[37:28] dust off that uh that one victory and move on to the next thing. So, the good news here is we've gotten really good at converging. We've gotten really good at figuring out what are the patterns that are going to scale. We know where the story ends, right? The story doesn't end with a bunch of different tools with sort of
[37:43] bespoke needs scattered about with each different team using their very esoteric tool, right? This this journey ends in convergence, and it ends in simplification. Um so, that's the good news. The The challenging aspect of this is the AI dimension brings a bunch of wrinkles that are going to make it more
[37:59] challenging to do that convergence into that simplification. I'm sure you're probably feeling some of these sentiments right now. Um In In the manufacturing space, uh there's a lot of decisions that need to go right for the factory to be successful, for the vehicles to be successful. Uh one of the ways that you
[38:15] protect yourself against things going wrong is something called a failure modes effects analysis, an FMEA. And it's basically a risk uh uh a risk assessment. And the three dimensions of a risk assessment are how bad is it if it goes wrong, so the severity, how likely is it to go wrong, and are you
[38:32] going to catch it? What's the detectability? So, when I think about the sprawl of of AI tools, it's quite concerning across all three of those dimensions. It it's never been easier to have a new tool or a new pattern introduced. It's never been easier for people to sort of spin up and get started. So, the likelihood
[38:48] is quite high. When we start thinking about the severity of what can go wrong, there's been a lot of conversations around sort of cost going out of control. I think we've sort of gone through the the honeymoon phase of of large language models and and tokenization. I think people are really appreciating there's real cost behind that. There's also a really meaningful
[39:05] governance exposure in terms of these tools proliferating. Every surface you have is another governance and control surface you need to take care of. And then lastly, the detectability side of things. As these shadowy AI tools pop up, as a central platform team, it becomes really hard to
[39:21] manage all of those. And so, on one hand, we have the playbook, we know where this ends. On the other hand, there are some wrinkles we need to pay close attention to this time. Um we're getting sharper that AI is not the goal. AI is a means to the goal, and so
[39:37] I really appreciate the perspective of focusing on the outcome, owning the outcome and not the data. Could not agree more with that. And I think I'll I'll I'll leave you with one statement here just in terms of the proliferation is every time you bring your data to an external system or
[39:53] you move your data to the application, you're creating a governance problem, you're creating a cost problem, you're creating a telemetry problem. Every time you simplify and you converge, you are eliminating one of those problems, right? So, the forces are definitely moving us towards simplification. So, I think now is time to take a hard look at
[40:10] the stack and figure out what's really going to scale with us. And those same uh that guided our convergence for sort of multiple clouds, multiple platforms into one, multiple ingestion and transformation patterns into very few, multiple analytics and data science tools into very few.
[40:27] This case one. Um that same trajectory is going to play out uh in the AI world. The other thing that's changing here is that the goal posts are moving a bit and this I think resonates with some of the previous talks as well. Uh and so another sort of job well done, well not
[40:43] quite. There's there's a bit more in front of us. Um there's no way you can get around having a solid data foundation. So I'll echo that statement that's been said before. If if you're not starting on that journey or if you're just beginning on that journey, stay there. That's the most important sort of no regrets work
[40:58] for you and and your team to to go into. You'll always pay dividends for fortifying your data ecosystem and that's the boxes down in the bottom. You know, are your pipelines tested? Can you, you know, qualify the the quality of your data? Do you have a hardened data governance surface? Um can you
[41:13] explain with semantics and with context sort of what that data means? Do you have lineage to trace that data back to its source, the sort of data provenance? All of that good data engineering and cataloging work is going to really pay dividends um for the sort of previous goal posts and even more so in in the AI era. Um we're quite excited though about
[41:32] what comes beyond that. So now that we've invested the work to sort of lay that foundation, sitting above that uh that sort of gold layer, so to speak, um maybe you call it platinum or diamond data or diamond data ecosystem, uh is is all these exciting capabilities that that we've been um
[41:47] working on in partnership with the the product team and and I'm just so impressed with with what's been delivered by Databricks here. Um the new standard, the new goal posts that we're going after um are not just the gold data, but layering on top of that genie spaces to wrap around a subset of that data ecosystem with a bunch of context
[42:04] and semantics and understanding um to leverage AI Dev Kit and and Unity AI Gateway to give you that governance, to give you that control and that visibility, and collapse this distance between where your data lives and where you're consuming and using that data to
[42:19] minimize the the tax again that you pay. The tax takes the form of inefficiencies for your team, it takes the form of token utilization, and I think some of the more concerning dimensions are it takes the the the form of of these governance surface that that become very
[42:34] hard for the team to manage. So, firmly believe that Unity AI Gateway is going to do for your AI ecosystem what Unity Catalog has done in in the data space. It's just the next extension there. And finally, just to make it punchy
[42:51] here, you know, models are commoditizing. I think they have already commoditized. I was I was excited to see, you know, when Fable came out, it's like, oh, I can, you know, model a 737 on my laptop in 5 minutes. But how often are you actually like modeling a 737 on your laptop? Most of the time our tasks are much more well-scoped. So, I do believe
[43:07] that we really are there for the vast majority of the cases. Trust in your data, trust in your foundations is not something that you're going to be able to buy your way out of, right? That takes the work to invest in your data quality and your data foundations. Going beyond that, once you have that foundation,
[43:24] it's this existential threat of exposure from when you have too much surface area. And so again, you have cost exposure, you have governance exposure, and you have efficiency exposure on your teams as they start to work through a fractured ecosystem. So, so convergence and simplification is definitely the
[43:39] name of the game. And then I think that the real call to action here that that drives urgency for me and I try to impart on my team, and maybe I'll leave you with this, is this capability gap is going to continue to grow. You know, that the teams that have made the effort to bring their data into a central ecosystem, to govern it
[43:55] well, to wrap it around with Genie spaces, etc., they're going to be the most nimble to move forward as the industry continues to expand. And so we're just starting to see the divergence between sort of the legacy way of doing things and these rocket ships that are going to start taking off. So, yeah, strongly
[44:11] encourage you to lean in on your data foundation and and recognize the goal posts have moved. You start at the same place, but we got to go a bit further than gold data and go up above into that genie dev kit AI gateway space to really excel into the future. Otherwise, you'll just keep paying this this tax over and over. So,
[44:27] that's it. Thank you so much. Thank you, Mikey, for making the case for simplification. Keep the applause as we welcome to stage Sheridan and McDonald. Thank you. All right. Thank you, Mikey, JP,
[44:45] everybody. Uh I get to welcome up Yeah. Amazing group is Ronnie here yet? Amazing. Present. All right. Okay, here's one. Thank you.
[45:09] Good. Good. Good. Good. Great. Awesome. So, hey everyone, I'm Sheridan. Um I am Manisha's peer at Databricks. You need some IT support? Sorry, sorry. Come here. All right. We're here. And you can see who I have here with me. Um some amazing customers and uh we're
[45:26] going to do kind of a Q&A, get their take. Um I've met with all of these guys in the last couple weeks to talk about this and they have some really interesting viewpoints. Um so, why don't we start if Ronnie, you want to start and uh introduce yourself. Who are you? What you look after? Where do you work?
[45:42] Ronnie Johnson, CIO at Workday. I'm responsible for our global IT and for that that means for us we we run a customer zero program, which is Workday on Workday. Responsible for our core infrastructure, responsible for our architecture and AI, so that's corporate AI engineering, as well as our
[45:58] go-to-market systems and strategy and operations. Uh I'm Dan Cane. I'm the co-founder and co-CEO of ModMed, uh formerly known as Modernizing Medicine. We're a healthcare technology company that builds software solutions for ambulatory surgical and medical specialty practices. That's a
[46:14] mouthful of basically saying private practice physicians, uh dermatologists, ophthalmologists, orthopedic surgeons throughout the country. Uh we've been a Databricks partner for 9 plus years, uh and I could couldn't imagine a better data platform for us to have made that bet, made the transition, and uh
[46:30] been such a good partner as we build out on our AI journey. And I'm Christina Kemmer. I'm the CTO of Afresh, and we're the AI-powered operating system of fresh food in grocery. And so, what that means is that we're helping grocers make decisions
[46:46] around, you know, how many bananas to buy, things like that. And our whole goal is reducing food waste. And um as CTO, what that means is, you know, I have the teams that own ingesting all of that data, that's super messy real-world data, and serving it up to
[47:03] the models that actually like produce the ordering decisions and recommendations to our customers. So, everything in between there. I love it. And um these three companies are great because most of the people in the room are Workday users. I mean, they're every day. Um so, improving the the product and
[47:21] experience there is amazing. And then in the healthcare space, Dan told me this morning that the healthcare interactions have an NPS score of minus 45, which is quite worse than anytime you get on the phone with uh Comcast customer service.
[47:36] And so, we're rooting for you. And then obviously, Christina um on the uh food spoilage, I mean, it's just um amazing kind of like mission that you guys are behind. So, let's start talking about um AI prototyping and AI production. And
[47:51] everyone in this room probably got uh a very angry email from your CEO about 2 years ago saying, "What are we doing with AI?" And then we all started running around trying to build AI. Um so, Dan, I'll start with you. Um from
[48:06] the concept of piloting um and moving to production, um what's the single biggest thing that has uh surprised you about the gap between a working demo and actually getting something live? Yes, and because I am the CEO, I sent myself that very nasty email. Um and
[48:24] then I uh brought in a true story. I I didn't send myself the email, but I brought in a co-CEO so that I could really focus uh kind of founder journey on reinventing ModMed for this agentic world we're we're living in now, which is exciting. Um I mean, we made we we had the data. So,
[48:41] 16 years ago we made the correct decision of saying, "In order for us to build the software solution and platform of the future for the AI-powered practice." We didn't call it that back then cuz no one knew what AI was, but that we were going to need data, right? Data is the the what feeds your greater AI. Um and so, we made sure that we had
[48:58] secondary use rights on our de-identified data. And in healthcare, there's a lot of restrictions on what you can and can't do with data. And so, we need to be really good stewards uh and have really good governance of how we use that. So, the the data was there and and today that's almost 1 billion patient encounters worth of
[49:13] data. So, it's a tremendous amount of data that we have to be good custodians of. But the biggest surprise is um like the first generation of our AI products uh we built some AI-only SKUs, we built some AI features to enhance existing products, was that you can't just take data, train model, ship it. Um
[49:31] that that failed miserably. Um and it never really took off. We had to figure out how to decompose the problems into uh orchestration of different agents uh and sub-agents. We had to figure out how to uh have the the visibility uh into what they were doing. And we had
[49:46] to figure out how to iterate much more quickly than we were. So, the first generation of AI for the first year we did two two years ago when we started on this journey. Uh we about once a month we're releasing new models uh to our customers. Now we do a couple hundred
[50:01] evolutions in a week. Um and it's it's all measured, it's all tracked. We've got, you know, our our our bronze uh silver, gold, and now platinum data sets. And so we're constantly evaling. And we're getting to the point now where it's it's really going to be more reinforcement-based uh than anything
[50:16] else. And so the hardest trick that we didn't didn't have and didn't have the ability to do um wasn't so much how do you do inference at scale. Databricks helped us with that when we had problems with that. It wasn't so much like how do you measure success? Well, you have your your data sets and your
[50:33] evals to be able to do that. The hardest part was changing the iteration uh capability of our products. So we were a scaled agile company kind of shipping products every 2 weeks. And we had to go much faster than that in the age of AI to get the product to where it needed to
[50:48] be. And Ronnie and Christina, similar question. I'll start with you, Ronnie, and then Christina. Do you think that this piloting moving to production, do you think this let's just ship it and see season is over? Um I love what you told me last week.
[51:04] I I I think that season is over. And I'll actually kind of share a kind of a similar um experience. Like I um I didn't get an angry email, but I actually started at Workday the week that GPT-4 came out. Um and before I knew where the bathroom was,
[51:20] someone was telling me I'm presenting to the board around what is our AI strategy. So uh Uh with with that it caused me to understand that it's the board's not asking like what are you doing? They're asking what are the outcomes you're about to deliver. Um and so we in the beginning let a thousand flowers bloom.
[51:36] We are now mowing the lawn and we're being very very very thoughtful about how we're cultivating the garden going forward. And so for us that means we're actually prioritizing our use cases extremely thoughtfully. So, we we look at the intersection of feasibility and impact. And for us, feasibility is
[51:53] things around um is the tech ready? Is it actually ready to deliver the outcomes? Is the org ready? Are the workflows ready? Is the business process ready? Do we have ready to go business stakeholders to actually do this body of work? Um and then what is the scale of the impact we can deliver? And so, we prioritize those use cases around impact
[52:09] because of like we've already made, you know, like the flowers blooming that was fun, and you know, everybody is way over that piece of the the work. Um they're truly in focusing on outcomes. And so, I loved in my first year um when I got to go give that board presentation and they let me like drop a lot of seeds down and see
[52:25] what happens. Um they are now asking like what are we doing? Um and so, we're actually having to actually um produce ROI against those agents and the those investments that we made. And we've actually really legitimately started re-platforming
[52:41] areas where some of the results weren't what we wanted, didn't meet the hypothesis, or where we found that certain platforms just didn't scale to actually deliver the value over time. Question on that. So, being able to track the ROI and the impact, how much debate or contention is
[52:57] there internally? Because I would imagine the owners of some of these flowers are going to be very very dogged that there is ROI. So, so that the first year when we didn't require an ROI analysis, um we we were just over that fight. In some cases, we say if you don't give us that
[53:13] data, we're going to turn off that capability. Um now, we actually make them express a hypothesis around those outcomes, and we want to understand who is the like the DRI or who's the owner of those outcomes because they're meant to on a quarterly or a bi-annual basis give us those to give us that data.
[53:29] Otherwise, again, we will turn it off. And so, you have to prove that it's still delivering for us to continue it to graduate those investments. Love that. Christina, what about you? You know, A Fresh is um probably quite smaller than both of your organizations. And so, for us, experimentation
[53:45] definitely isn't over, but we have to be really, really purposeful around where we do that experimentation. So, um, our customers are grocers. The change management at grocers takes a long time. And so, kind of the rip it
[54:00] and ship it is like not ever in our DNA, um, as far as from a production, um, possibility, but you know, we've made purposeful architecture investments to say, like, as an example, rebuild our model infrastructure so that it can be
[54:15] really easy to swap out the models that we're using or to experiment with different models with different particular use cases like complex promotions or seasonal items, things like that. Those are hard problems in our industry. And so, where we've invested, you know, is to say,
[54:32] like, okay, it's not acceptable if a recommendation doesn't get to a user or decision doesn't get to a user, but the quality of that decision or the way that we get to that decision, we should play around with and see what's working and see what, you know, could be better
[54:48] there. And so, I think for us it's been much more around the where do we experiment and where can we not and be very solid and and be purposeful about that. Okay. So, let's talk about the actual architectural tax. So, Reuben and J- and and Meta both highlighted this. Reynold,
[55:05] uh, his last answer, uh, very much highlighted this. Um, Christie, I'll start with you. Like, what do you think is the what's how's what's the architectural tax look like at your company? What is like one of the most expensive pieces of plumbing that you want to rip out or
[55:22] have ripped out? Yeah. Yeah, I think for us, um, a really expensive part of our plumbing was in really the connective tissue. So, basically, the the pipelines and the glue code that took data from one place and transformed it so another system
[55:39] could use that data and having that all over the place. And so with fresh data, you know, fresh data is very messy. Um it's not very clean data that we're getting from our customers. It doesn't, you know, align and make sense. So we have to do that. And so the pipelines can't be just dumb, right? They just
[55:55] can't be passed through. That's a lot of the value that we're providing. But our customers don't value us for the pipelines, they value us for the outcome at the end of it. And so for us, um we made a huge investment for us to really
[56:11] consolidate systems. So Mickey talked about this around consolidating from we were on Databricks and Snowflake, consolidated all the way back to just Databricks from a foundational perspective. Um so that there, you know, all of our teams are in one spot. When we're making improvements, it's in that
[56:27] same system. And then we're also migrated to Lake Base. So that was a really interesting thing for us because you know, we um want to make sure that our product operational data, things like that, are in that same system. The the
[56:44] time and the pipelines and the fragility that we were seeing between like, okay, our warehouse and our database and finally to the product, things like that. That was the the most fragile part of our architecture. And so we really invested in that piece to really drive that cost down, but also increase
[57:00] that stability and reliability. So, let's talk about governance. And Ronnie, you told me you have a compliance team, you have a security team, um you have a governance team. Those teams in the past have always, uh
[57:16] and I'm not, if there's governance and compliance people out there, bear with me. Those teams in the past have almost seemed like a blocker to innovation and moving fast. Um how have you seen them actually be an enabler for Workday? So we try to treat those teams like they
[57:34] are like a paved road versus a a brake pedal. Um and a paved road for us means we've got lanes, we've got rules, that means we can start to go faster because we have a rule set. We know when there's like an accident, call it a data privacy accident or security accident, or even
[57:49] let's just say an accident of over token usage and something costs too much or there's waste. Whenever there's an accident, someone's going to way over rotate and put way too many rules. And so, we think about governance as kind of laying that paved road before we get started so that we have an understanding of who's moving so we don't have duplicative agents, we don't have
[58:05] duplicative data, we don't create something that would create an accident that makes everything go way slower. And with kind of with that kind of shared understanding, we actually do take make the investment to build that paved road so we can all go faster. Um that said, I think one of the things the governance is not just your yeah, we have a tech AI
[58:22] advisory board, we have an exec AI advisory board, it's not just the security and the compliance and the responsible AI teams. We actually build governance in into tech technology. We actually leverage Databricks actually to build a really cool thing that our our team I'm looking at one of my my team members here who actually built it. Um we have this um environment that is
[58:38] essentially the QA of our agents. Um and it we built this because back in a 2023 when we were first building agents, we were all saying, "Hey, human in the middle." None of us are really saying that now, but how we got over that is that we actually built an agent um or frankly um a model that monitored our
[58:55] agents for accuracy, um bias, toxicity. And think about this, for Workday, it's really important, even tone. If you're asking a question around, you know, let's say bereavement leave, um there's a tone that we want to take when we're responding with an agent. When you're asking a question that maybe something
[59:10] like wildly inappropriate, um there's a different tone we take with that, too. Um and so, for for us, we really needed to build something that actually unlocked the ability for us to deploy agents that were actually um Workmate facing for us at Workday or even customer facing. And so, we want want governance into our agents and into our
[59:27] our agentic strategy, so that we actually can again, that paved road for us to go faster. I love that. So, Workday has empathetic agents. We do, actually. I'd love to see how you monitor and measure empathy in an agent. Um and by the way, your team member, would that be Phoenix, the best-dressed man in tech?
[59:43] Always. Right there. Okay. Uh Dan, what about you? I mean, on the on the the tax piece, the the hardest thing right now is that we we are the while we have great governance, um we unleashed our developers throughout the organization,
[59:59] and and they're because it's a bigger organization, there are pockets of technical people, not just in product development, where we had a lot of good instrumentation and best practices. So, um the the the picture from the keynote this morning of like the wild west of agents all doing what they're doing without us knowing is the reality right
[01:00:16] now at ModMed. Um now, they all have, because it's all hooked into our authentication and identity management, like we have traceability, but not where we need it. So, we've got individual Claude code terminals with MCP running directly against this, that, and the other, and it's the wild west. Um so, we
[01:00:32] were in the process of building our own AI gateway. So, thank you for building one, um because now we can stop working on that. But, trying to bring some MCP governance, some traceability to all of our conversations. Remember, we're a healthcare company. So, we need to understand we've got PHI in all of our
[01:00:49] systems. It's in our Salesforce, it's in our Workday, it's it's in our product, of course. It's in Jira tickets. And so, we need to be able to monitor all of this, even though we know it's going to be mostly agentic. We need a a single place to be able to go, monitor, log, track, uh refine, query. Um and so, I'm
[01:01:07] really excited about, you know, Omnigent is going to be amazing for our development. I think it'll really reduce our costs, because for a long time we were telling people to use the highest quality models, cuz they gave the best answers. And then, when we saw the cost shoot through the roof, we're like, maybe you don't need to do that as much.
[01:01:23] Um, but it was it was the right answer three months ago, right? It was to use the most powerful model as much as you needed to get the best answers. Uh, now we're at the point where we're using multiple different agents uh, concurrently checking each other and and making sure we go back, using lower cost
[01:01:38] agents for things. So, the more we can streamline that, I think the better off we'll be. Um, one trick on the compliance and um, legal compliance and infosec is if they're the first group you go to with your AI automation, uh, they actually get cuz they learn how it works, they learn the rules of the road,
[01:01:54] they get to learn how to make these these like guardrails, uh, and then how to have agents help enforce guardrails rather than they have to enforce guardrails. So, all of a sudden, the team of of lawyers who I didn't even know we had who just review marketing, which I didn't know was a thing, um, but
[01:02:10] everything that went out from ModMed had to go to like this review process. And I'm like, can't we just you guys create the agent, we'll show you how to do it, and then everything marketing does will go to that agent, and the agent will say whether green light or yellow light you need review. And so, it it was like super helpful. Once they got how it
[01:02:25] worked, they were much better and much more willing to go with you on these journeys. Um, and for them, it's really just around reportability and and observability and making sure the more you can plumb it into something like um, the AI gateway, I think the happier everyone will be um, because they know
[01:02:41] they've got one place they can go for that source of truth for how all the all the agents are working, how all of your MCPs working and your skills and and so on. All right, so let's talk about interoperability. Um, you know, Databricks feels very strongly that
[01:02:57] open source is like core of our DNA. We try to be humble enough to not think that we will own the entire data landscape tip to tail forever. Um, Reynold was kind of talking about that as well. And um, just so you guys just so everyone in the audience knows,
[01:03:14] the reason why Ali continues to open source things that we built like a year or two before. If you ask him, his answer is is very simple. It's to hedge against the innovator's dilemma inside at Databricks. Because you can imagine there's a lot of
[01:03:30] debate when we open source a product that has incredible IP, that has great traction in the market. Ali's perspective on this is, "No, I want to open source it cuz I don't want the engineers to get fat and lazy, and I want them to have to go build the next
[01:03:46] thing." Because, you know, it's like in in the pharma world, you have how many years uh you have 7 years until the drug gets generic? Um and so, let's talk about interoperability. I'll start with Dan and Ronnie. How is that a part of your data strategy
[01:04:02] and platform strategy? What does it look like in your guys's organization? Okay, I'll go first. So, for interoperability, I mean I'm let's be honest, I'm we're not going crazy. Um we're we're thinking about interoperability anchored on a core set of platforms and principles. And so, for
[01:04:17] us, um we've made some like key investments. We decided like well, here's these are the systems of record that we're using. Here's where we're going to put our our AI and ML workloads. And so, from from there, interoperability then is okay, we're fit for purpose. Um so, we've decided these
[01:04:33] workloads will handle this type of work. Um we make sure that all agents actually have to register their purpose, and that all of our AI tools are recorded. Um and so, when when we're we're thinking about modular additions to our our core infrastructure, our stack, um it's
[01:04:50] really areas where there are gaps. Um the other thing that's interesting about what how we think about modularity, too, um anywhere we're doing something that's not in our I'll call it our core, um we think about that in maybe one to two-year terms because we think
[01:05:05] that the the industry is moving so fast that that modularity, that interoperability that we're seeking, too, might be something that's built in by this kind of those core platform providers that we're actually staking kind of our our our our key bets on. Um another thing that I think is important to support interoperability is having
[01:05:22] the right architectural model and frankly review boards for your organization. So, for us, we needed to be really really clear and actually redefine what it meant to be kind of at the top inner we call it enterprise architectural layer. Um our domain layers are where we actually focus the domain knowledge plus
[01:05:39] the kind of the architecture and then at the solution architect solution architectural layer. Um and we get clarity around the decision authority for those groups and we make sure that they're subscribing to a set of clear and key um common architectural principles so that when we are doing something interoperable, we actually can
[01:05:55] also modularly pull those pieces out um if something you know more performant comes to provide us more capability. And so, for us, interoperability is really a it's it's it's an architectural philosophy, frankly. From modernizing medicine's perspective,
[01:06:11] um the data that we have is is our customer's data and we make sure that our customers get it back in the form in which they want it. And Databricks has been phenomenal and integral to that process because as a true single-instance multi-tenant SaaS solution, our customer's data is
[01:06:26] massively commingled with every other customer's data. And so, teasing that apart is is actually a drawback of modern cloud architecture, right? Is people have this vision in healthcare, probably isn't true in all of tech, but they're they're imagining a database that's theirs that we just are somehow
[01:06:42] not giving them, you know, a SQL blinking SQL cursor access to. And so, we had to go through quite a bit of work to make it safe for them to be able to access their data, export their data, ingest it into their own versions of whatever data lakes they were doing. And so, we we've been on that journey for a
[01:06:58] very long time. We're very good at giving them back their data and and we absolutely believe that this is their data. We're being good stewards of it. It's the patient's data. It's the provider and data. Um, on the interoperability side, healthcare is frustrating. Uh, and for those of you who've ever come into the
[01:07:13] healthcare industry, um, there are many, many, many competing standards for trying to be interoperable. Um, the latest standards are called fire interfaces, f h i r, and they're on version three or four at this point, and I'm going to say they've maybe got 10%
[01:07:29] of the way that they need to get to. Um, there's there's not a lot of real-time interoperability, which is why when you go into one clinic and then go to another place, like the information doesn't transfer. Um, it's incredibly frustrating. A huge amount of data transfer right now in healthcare is still faxing. Um, because it's HIPAA
[01:07:46] compliant. It's HIPAA compliant cuz it's analog. It's actually a fax modem, right? And so, that's the only reason it doesn't come into like the wrong contact. Like, it's it's the most backwards and frustrating thing, but that being said, as more and more companies, um, embrace these
[01:08:01] technologies, and as more customers get savvy with what they want to do with their own data, and they want to build their own agents against their data. And so, giving them really fresh data, um, with the right data dictionary, with the right ontology behind it, I'm super excited about what Genie's going to be able to do so that we can help them
[01:08:16] annotate the data in a way that when they ask their questions, it's going to give back the right answers. And without the ontology engine, we've been playing around, like we we do a lot with Genie. Um, Genie was was good at a fallback. Genie was not the best at getting to the right
[01:08:31] answer the right the sort of on the first time. Um, because in our system, even when we had Genie spaces created correctly, um, there's a lot of ways you can answer that question. Without being able to have the annotations and the metadata, uh, and for it to be aware of that, sort of the data provenance,
[01:08:48] it would tend to go to the wrong place a lot of times, or it would take an inordinate amount of time realizing it went to the wrong place, and then sort of try again and get the answers. And so, we created hundreds of MCP tools to sort of short short Genie, uh, which now with ontology, I I don't think we're going to have to do that as much. Um but
[01:09:04] by the way, that's a really good short circuit. If you're developing these applications, AI, chat apps, whatever, and it's not getting what it needs, you can stop beating your head against the the the wall there. Uh just create a little MCP tool and give it the description of what you want it to go get and how you want it to get it, and
[01:09:19] it will just short circuit that and get exactly what it needs, and then it can work with genie to get it and combine it with other data and do things. So, what what we're going to do with genie ontology is what we're going to be able to do with genie one is really I think game-changing, and we're going to have to figure out how do we expose our customers to that. Um because we'll have
[01:09:35] access to it internally, obviously, but I think there's a lot of power in customers going on that journey with you. Wasn't on my bingo card to hear about fax machines at Data and AI Summit, but yes, healthcare is frustrating. Um so
[01:09:50] one part of Interop is like openness, right? And, you know, I think if you guys came to Summit 2 years ago or even last year, it was a a lot of talk about Delta versus Iceberg, which is kind of like saying should you store your pictures in PDF or PNG, right? It's just
[01:10:06] a kind of a nonsensical debate that unfortunately the market had for a long time. So, Christina, like open formats, Delta, Iceberg, open databases, how do how do you view that, and how does it affect even your vendor choices? Like is
[01:10:21] proprietary lock-in ever acceptable? If so, how do you think about it? Yeah, I think for us, we think about like especially at the data and the storage layer, open format is so much the preference because then your data's not trapped. Um and you can make it
[01:10:38] allows us to make different decisions about compute or tooling on top of that later if we need to without this giant lift and shift migration. Uh we've already done enough migrations. I joke with my team that we end a migration, and then we start another one in a different part of the organization. So,
[01:10:53] if we don't have to, that'd be great. Um but you know, I think that's a really important piece of this and I think, um, same going back to Lake Base, that decision was pretty easy because it's Postgres underneath, right? Our team knows how to work with that. You know, they understand, um, it's not some brand
[01:11:11] new thing that we have to go and understand how to operate and and maintain things like that. So, I think that those are really important decisions for us to keep it simple, not get too weird with our technology and keep it simple so that that reduce our cognitive overload and our costs and
[01:11:26] things like that. So, um, you know, I think the third part there for us is like really at that model layer. So, being able to route to, tune, choose different models. If you lock into a model today, it's great for today, but then, you know, in 2 months
[01:11:44] when some other Greek mythical, you know, creature version of a model comes out, you like you're already behind. So, I think having that flexibility is super important, um, probably for all of us, honestly, um,
[01:11:59] but the is it ever acceptable acceptable to have proprietary lock-in? Um, we do in some things, but we make purposeful choices around that because if it gets us to speed, you know, speed of delivery, um, being able to prototype
[01:12:15] something where it's like, hey, we're we're kind of going into a new world and actually using this particular tooling it is great. So, one example, we're using Databricks model serving, right? So, that's kind of a lock-in in that in that sense for us, but it helps us to understand and learn and say, okay, is
[01:12:32] this the right decision to make and keep going on or is there a different kind of solution that we need to take? So, um, I think those are the choices that that we make because sometimes it's just cheaper for us to try it that way. So, Love it. And Ronnie, you know, you're not getting off this stage without this
[01:12:47] question. Uh, everything we're talking about, the question is how has that changed how you restructure your teams and how you think about the roles and responsibilities? Okay, so Sheridan's put me on the spot cuz I made up this word. It's called PIMBASAQUA.
[01:13:07] Um I now Real quick, real quick. When you told me this, I was quick Googling so I didn't look like an idiot cuz I had no idea what it was. He asked me. He was asking me the etymology of this word and I was like from my brain. Um
[01:13:23] So um a PIMBASAQUA is a product manager, business analyst, solution architect, QA. It's what I expect now of a product manager. I'm expecting them to go full stack. And so what's really changed is
[01:13:38] the expectation of what you get out of the roles. And so it's changing like frankly AI and and especially when you actually start to get to getting true like universal standards for how you can actually deliver AI successfully. It's changing the expectation of the value delivered. So maybe you don't get a full
[01:13:54] stack PIMBASAQUA, but then you're getting a product manager who can be a product line manager. It's either going to go full stack or it's going to go wider, but the expectation around what you can deliver and how you can deliver it is wildly changing. And so you know, you all can borrow PIMBASAQUA, but I think that it it is it it is the new
[01:14:10] frankly the job definition that you're getting more out of your teams. And so you you know, we're having to let go of what have been traditional roles that don't serve as well in the AI world. Just so everyone can write this down. It's PMBASAQA.
[01:14:26] Rolls off the tongue. Yeah. Okay, last stretch. Let's talk about big bets and predictions. Let's talk about winners and losers. Um Dan and Christina I'll start with you guys. Single infrastructure investment that you think
[01:14:43] will make the most impact. The question is supposed to be over the next 12 months. Although the next 12 months kind of feels like the next 48 months these days. So, what's what's the big bet that you think for the next period of time that you feel confident giving a prediction?
[01:14:59] Uh I mean we we've done a great job in getting all of the data into Unity already. But um and so we've got good governance on data. I think it's it's the AI gateway uh is going to be key, Genie ontology, Genie 1. Um some of the
[01:15:14] other things around um security that we saw. I think there was a so much announced this week that um I'm very willing to data bet on because uh it is based on open standards and it's based on things that I can get behind and it works so well in our
[01:15:30] existing infrastructure for our model serving and our inference at scale and and how we've developed it. Like we've already built, sold, and shipped high-performance AI-only applications uh through Databricks. And so I've I've great confidence that that works. Our next frontier is actually a lot more
[01:15:47] internally, going through the org design changes for our product teams and how they work and upskilling people to be more across and I don't think that that necessarily means like all product managers will will be great at all of the other aspects, but they will absolutely touch it. But the same thing
[01:16:02] I think is true of our our more like hardcore developers will have to learn how to be more product owners. Uh they've always we've never separated product and QA. That was always the product person owned it. Um but I absolutely agree, those lines are going to blur. Team sizes will shrink tremendously because when your code
[01:16:18] velocity uh is so much higher, if you have a team size bigger than frankly two or three, they step on each other's toes. Um and then you're just waiting on each other for for constantly or or worse. And so I think that there's there's a huge transformation from what was sort of the old SDLC to more of an AI
[01:16:34] software development life cycle, different team sizes, different deployment strategies, a lot more code gets pushed, which can be double-edged, which gets all the way back I think, philosophically to the first thing you said, which is you need to lead with the why. Like, what's the ROI? What is it you're trying to do? Is it increasing
[01:16:50] TAM? Are we reducing churn? Like, is it just you built this because someone asked you to build it and a customer thought it'd be cool? We need to always understand like, well, what are we doing cuz we're able to do so much more so quickly. We need to understand, okay, what's what's the cost? What's the benefit? What's the why in everything we do, both
[01:17:06] externally with our customers and our products and internally with all of our modernization. I think, you know, thinking about what's important to do today. And AI is this world is crazy because I feel like you constantly feel
[01:17:21] like you're behind. Like, I I know my team feels like that, I feel like that. I think that you come and talk to people doing amazing things and it's like, oh my god, I'm behind, you know? And I think, you know, that's not true. I think that everybody's kind of in the same boat. We're trying to figure it
[01:17:36] out. And so, one of the things I think that you heard today over and over and over again, at least in in these sessions, was like, "Hey, invest invest in your data foundation. Invest in that, you know, in the governance around it." Um that is so critical. And if you haven't done it yet, that's okay. You
[01:17:54] know, there's a saying about planting a tree, right? Best time was 20 years ago. What's the next best time? It's today. So, do that today and then have your org benefit from that in the coming months because, you know, the the windows get shorter and shorter and shorter as far as the rapid changes
[01:18:10] that's happening. And so, what's great is that if you have that foundation to build from, the stuff on top can kind of swap in and out and that's okay, but um I would just encourage that. So, Ronnie, let's talk about winners and losers. I feel like right now in the
[01:18:26] market, we're watching some companies have this like very, very hockey stick rapid like TAM capture. And we're also seeing companies um on the opposite side. What's your take on the infrastructure
[01:18:42] necessary to win versus lose right now? So, I mean, the infrastructure, I think, of a of a Oklahoma loser, um it's it's those those companies that maybe have just handed out productivity tools. I mean, it's not enough to say, "Here, Copilot. Here, Gemini. Or here, Claude." and go.
[01:18:58] Um you'll see those losers just have a ton of sprawl and no ROI. Um I I think the winners are those who are actually making sure that they're um investing in the skilling of their workforce so that they can actually get the not just a personal productivity, but actually
[01:19:14] getting to business outcomes. It's where you know, the the winners have discipline in how they're actually evaluating and selecting the tools and building strategic partnerships and in in really truly um enabling each of their business teams to deliver outcomes. You'll see in those business
[01:19:30] teams that they're actually uh very thoughtful about how they're making their investments and how they're choosing those platforms to actually graduate those investments off of. The The infrastructure, I think, of a winner is really just a disciplined organization. It's It's got some clear principles that's not moving in silos and is just really um
[01:19:47] thoughtful about how they cultivate their little AI gardens. I love it. Hopefully, this is super helpful. Ronnie, Dan, Christina, my team loves working with you guys. You push us and you push my product teams. Appreciate it. Thank you, everyone.
[01:20:05] All right. Next up, uh to close us out, we have Jason Reed. I will let him introduce himself if you don't know who he is. Um Speaking of Interop, he is uh in charge of all Interop at Databricks.
[01:20:21] All right. Mics are working. Yes, mics are working. I think I got I'm miked, so we're good. Uh yeah, thank you, Sherrod. This is fantastic. Click next. Where are we at? Yep. No, not that next. Here we are. Okay, that's me. Um really great to hear from everyone on
[01:20:37] the panel, so thank you to Reynold and the rest of the crew. Uh I'm just really going to be here my goal for like the last 5 minutes to get you guys on your way. It would be hopefully to uh get you out of any state of paralysis you might be in. And it's wholly understandable if you're in a state of paralysis in the sense that
[01:20:54] the amount of information coming at you, not even just like just this week, the amount of information coming at you, and the new products and new capabilities, all the technology, it's coming like fast and furious along with everything else that's happening in the AI and genetic space, it's like really easy to feel like I don't even know what I should do next. And so hopefully after
[01:21:11] everything you've just heard today, uh you can walk out of here with some like, "Okay, I actually do know what to do next. I can take the next step." Uh it doesn't have to be all the steps. To Reynold's point, we don't want to do everything all at once, but at least we have a design framework for where we're going, and we can take that one incremental next step. And so in order
[01:21:27] to take the next step, for me, I was like looking back to see like where was the last time I was in a similar situation because like yes, the world's moving super fast, but these are things that we have been through before. Uh I started my my career in data like 15 years ago when I when I joined Netflix, and we were just going through
[01:21:44] our like cloud migration. Like it was another another wave of like technology migrations that we did. And we were moving out of data centers, moving to the cloud, and we we ditched Teradata, I mean we did everything. And this is like pretty early days, this is like 2010, 2011. We're like, "We're going to decide to put our entire data
[01:22:00] state uh into S3." Wasn't called the uh the lakehouse yet, but it was the beginnings of what we now have as lakehouses. And it was all this promise of what was going to be great about that architecture, which is, "Oh, let's throw all the data in there, we'll collect everything, we'll make sense of it later, it'll be fine, it's like schema
[01:22:15] on read, it's going to be great. Uh look at this big Teradata bill we got rid of." And of course the realities came pretty like quickly as well, not as quickly as they're coming now. Like this still took years, and And that we're finding these realities out this months. so everything is compressing. But, you know, we figured out that this it wasn't the Nirvana we thought it was. And we
[01:22:32] had to go back to the drawing board. And we realized that we needed a couple of key pieces, right? To we needed some like foundations. The the sort of the theme of today has been we needed better data foundations. We needed SQL, we needed transactions, we needed data we could trust. Like that's the very bottom
[01:22:48] layer down there. That's where like I live all my days as like the very bottom layer of the stack. You know, Iceberg was something that we developed at Netflix. And here I am a decade later like still like helping that thing evolve and making sure that all the companies across the ecosystem can take advantage of open data formats, like
[01:23:03] really strong foundations that are also flexible and modular. But the other thing that we needed has been more recent was like we also need governance. We need a catalog, we need a providence, lineage. We need all this metadata. This is part of our strong foundations. Yes, open table formats, but governance as well.
[01:23:19] And so, you know, we heard a lot about that today. Like like there's no wasted energy in improving your governance story. Always be better. So, if you do nothing else after you walk out of here, go back and figure out what's the next step you can take on laying yourself better data foundations. And technology is not the answer, right? Like Unity
[01:23:36] Catalog is a tool to help you build the foundations, but you all in this room have to go out to your organizations and you have to work with your people and your process and your culture to build those things. Uh we're we're here to help you give the tools, but like alone Unity Catalog is not going to solve your governance problem. I think you all know
[01:23:52] that. Okay. So, that's good. But this is also our reality. Right? There's the the tidal wave of of agentic and AI workloads. We have all of our agent co-workers now. They are incredibly data hungry. It puts a lot more pressure on those
[01:24:07] data foundations. And then it adds. It It adds. We saw like, "Hey, the gold layer isn't enough. There's another layer of stuff we have to think about governing." Uh otherwise, you know, we end up in this picture, right? This is like going back to the data lake days of like, "Hey, we got all
[01:24:22] this data in the lake. What are we going to do?" Nobody knows what's in there. I don't know. We're just racking up S3 bills. So, you know, in order to like if you're already like down this path and many of us are, um you know, the reality is it's it's not like too late, right? Like I've seen
[01:24:37] this movie before. We've been here. We know that like those governance programs matter. The technologies aren't going to solve them for you. Right? We have a bunch of announcements we made this week. A lot of them around adding Unity AI Gateway. That's like that next layer of governance tooling,
[01:24:54] but it's just the tooling. And just like Unity Catalog is not the answer, uh AI Gateway is also not the answer. It's a means towards you all need to go out and like build your your governance models. Simplify the surface area, right? Reign in the surface area of which your companies are doing so you
[01:25:09] can wrap your arms around it. Uh and then, you know, leverage the tooling, put the programs in place, and you know, it's like a little bit of like go slower to go faster, right? Like we need to like put these these foundations in place, uh lay the the roadwork. I think that was Ronnie's metaphor which I love like lay the the the lanes, the
[01:25:26] groundwork, and that actually allows everybody to go faster. Once we understand the rules of the road, we can all go faster and we're not like, you know, just running into each other everywhere. So, the tools are here, uh but your next action should be go out to your organizations, figure out how to put this stuff into practice uh sooner than later because the longer
[01:25:42] it goes like the bigger small the bigger mess potentially to clean up. So, uh you know, the time to start was yesterday like the great analogy with the trees planting like it's time to start governance programs was, you know, 5 years ago. The next best time was like this afternoon if you can. Um and so yeah, that's my sort of
[01:25:58] uh my takeaway for you all is I got to do those things. Uh and just a reminder like this is not as much as it feels new and scary, uh we've been here before. Many of us in this room have we've been through these challenges before. We will get through them again. A lot of lessons you learned the last times around are still very applicable.
[01:26:14] So, So lean on that experience. Uh and lean on the people in your organizations who have that experience. And and they will help you get through it. Okay. Uh, just a reminder outside of this, which was an amazing session. So thank you everyone who was part of this and
[01:26:30] organizing. Uh, lots of the things going on this week for, you know, you and your cohorts of, you know, the people out there on the on the forefront of technology and pushing the boundaries and making healthcare better. Thank you. Uh, I very much appreciate that. Um, and
[01:26:45] also one like plug, shameless plug. I have a talk this afternoon if you want to hear about Phoenix and how he, uh, did unleash his agents at Workday and all the great work that we're going to double click into like how that actually happened. What was the foundations that
[01:27:01] Workday needed in order to do that. So Phoenix and I are going to give a talk this afternoon at 4:00. You can come hear the details about that. And thank you very much.

Learn more about the Databricks Data and AI platform.

The information provided herein is for general informational purposes only and may not reflect the most current product capabilities or configurations.