Skip to main content

AI-Native App Architecture: Building for Agentic Workflows and Real-Time Data

Summary

  • Zillow and Workday explain why traditional ETL-optimized cloud architecture breaks under agentic AI, and how replacing point-to-point integrations with a unified lakehouse foundation converts the M×N complexity problem into an additive M+N architecture for 1,200+ engineers.
  • Lakebase provides a unified foundation for OLAP, OLTP, and agent workloads, with Zillow using database branching to run parallel agent evaluations in isolation without risking production data integrity.
  • Offline evaluation at scale is presented as the key to confident model upgrades, enabling teams to verify replacement model performance before switching production traffic when OpenAI deprecates endpoints.

AI-Native App Architecture: Building for Agentic Workflows and Real-Time Data

Watch: AI-Native App Architecture: Building for Agentic Workflows and Real-Time Data
The 20-year-old cloud architecture optimized for ETL silos breaks under agentic AI. Zillow and Workday are reimagining the foundation: Lakebase for unified operational and analytical data, Databricks Apps for rapid experimentation at scale, Agent Bricks for coordinated multi-agent workflows, and Unity Catalog for governance that doesn't bottleneck innovation.
this video covers how to convert the multiplicative complexity of point-to-point integrations (M x N agents to N data systems) into an additive architecture (M + N). Learn why a single source of truth eliminates data duplication, reduces trust problems, and lets 1200+ engineers safely deploy AI features. Hear how Zillow handles branched development with Lakebase for parallel agent evaluation, why memory management must be architected from day one, and why offline evaluation at scale is the key to confident model upgrades when OpenAI deprecates endpoints.
🤝

Chapters

FAQs

Why does traditional cloud architecture fail for agentic AI?

Legacy cloud architectures were optimized for ETL batch pipelines, creating isolated data stores that do not support the real-time, bidirectional data access patterns AI agents require. Point-to-point integrations between agents and data systems grow multiplicatively—M agents times N data systems—creating a maintenance and governance burden that becomes unmanageable at enterprise scale.

How does Lakebase support parallel agent evaluation at Zillow?

Zillow uses Lakebase's database branching capability to spin up isolated copies of operational data for each agent evaluation run, allowing teams to test new model versions against real data without risking production state. This pattern enables safe parallel experimentation across multiple agent variants simultaneously, which is essential for Zillow's AI readiness mission across its data mesh platform.

Why is memory management critical in agentic application design?

Memory management must be architected from the start because agents accumulate state across interactions, and without a governed memory layer, agents can surface stale or conflicting information to users. A single source of truth in a unified lakehouse eliminates the data duplication and trust problems that emerge when agents maintain their own isolated memory stores.

How do teams handle model endpoint deprecations in production agentic systems?

Zillow addresses model deprecations through offline evaluation at scale, running new model versions against large evaluation datasets before switching production traffic. This approach lets teams confirm the replacement model matches or exceeds the performance of the deprecated endpoint, avoiding live traffic experiments on time-sensitive real estate data products used by Zillow customers.

Full transcript

[00:08] Hello everyone. Um, I hope all of you really enjoyed the the keynote presentation uh today and looking forward to what I think is going to be an amazing conversation with uh two of the most innovative leaders I've had the pleasure to work with. Uh so thank you thank you for joining us and maybe you should start by just giving a little bit
[00:24] of introduction of uh who we have here the company you are working at and the fun question of uh what is one thing about you that people would be surprised to know about okay thank you thank thank you a uh good afternoon everybody uh so I I am Phoenix
[00:40] Majamur uh I work for workday uh just completed my two years mark uh here um my uh primary responsibilities are a few folds, right? I run our uh data and AI platforms team. I run our AI engineering
[00:55] organization and most recently I had the pleasure to expand my remmit and start working with our data science team and building some of the most sophisticated AI and agentic applications. Fun fact is a hard one, but uh many of you may or may not uh resonate with it, but I
[01:13] growing up when I was in college, I used to play guitar in a rock and roll band. All right. U my name is Jaady. I'm from Zillow. I've been at Zillow for um year and a half, a little bit over that. I
[01:29] lead the data mesh platform at Zillow. So our mission is to like enable um Zillow's data to be AI ready both for like customers and our uh internal employees. So whether that's like real time, you know, streaming or graphical
[01:44] APIs or, you know, just dashboards and insights, we want to just like enable everyone to self-s serve themselves. So that's been our mission. Um, and I can't beat that fun fact, but uh I am an avid hiker. I've been hiking
[02:00] through all the national parks in United States. I've am I've completed 46 out of 63. Wow. So what are the four ones that are left? 46 out of 63. So there are many left, but I mean yeah, like all the ones in Alaska, the eight of them, that's some of them are hard to get to. Some of
[02:17] them your favorite actually I don't know like I don't know if I should say this but B in Canada is Yes. Yes. But and I think Death Valley has been very unique for me like I think I enjoy Death Valley a lot. Yeah.
[02:34] And thank you again for coming here. Um I think it's uh something that we all can agree right we are living through what I think is one of the most transformational era of our careers. Uh I don't think we have seen the kind of innovation the kind of uh uh paradigm shifts that we are seeing in the industry. Uh often I
[02:51] myself get surprised to realize that the whole AI LM era is only four years old. I think charge GPD came out in November of 2022. Mhm. Uh but what has been extremely exciting for me personally is actually to see the innovation that's been going on for the last couple of years where the way
[03:07] people are building software, the way people are building applications and the playbook that we all probably grew up learning uh in our career is now starting to feel a little bit more outdated. Um so maybe taking that as a segue would love to like start by saying
[03:23] uh Phoenix when did you start realizing that the the way applications are being built uh has to change in your uh environment you know that the realization happened as I observed transition on my the kind
[03:38] of emails and slack messages I start get started to get right I was getting like before the era was okay we have this problem statements let's build the technology around it. But then the Slack messages changed. It was about I want to experiment with this frontier model. I
[03:56] want this particular capability AI capability enabled. We have so many hypotheses we want to validate. And these are not technologists. These are individuals who are product managers. These are individuals who are like uh
[04:12] finance personals. These are individuals who are like uh people personals, right? like human resources and so on and so forth office of the CFO and it was a profound realization that as AI proliferates it is important that we create an
[04:28] experimentation framework and that's when we came up with this uh acronym sale sandbox for AI innovation and learning but the whole idea was right can we encourage the organization to test and experiment with AI and do fast
[04:44] enough and validate prove or disprove their hypothesis at very early stages right creating prototypes creating MVPs validating hypothesis became important previous world right like even to validate our hypothesis was hard to do
[04:59] because it's there is so many applications you have to integrate you have to consider like engineering resources on the pool AI changed that for us experimentation became lot more possible what about you J so uh Zillow went from optimizing for
[05:15] classic ML to agent AKI in a matter of months. Wow. Um like before like Zillow grew through acquisition. We had like inherited a lot of companies with different tech stacks. So Snowflake, Trino, Tabler, they name it, right? Like we used them all. So we were already on a journey to consolidate
[05:31] and we just like migrated data bricks was really going well. And when the opportunity came uh to work with partner with open AI on like building the first like app in open AI like uh with charge JPD for Zillow, we just like all rallied
[05:47] around it. We embraced it like you know Zillow has a culture of like being the first to adopt when you know mobile apps came out like Zillow was one of the very early apps out there. So we rallied around it and like we realized that we needed a database that would uh you know really work with like vector search. We
[06:03] needed like really you know high performance uh you know uh throughput. We also wanted to like we realized that we would need all this data offline to evaluate and like you know improve the quality. So lakebase was just announced right right around the corner and we were already thinking about exploring it and we jumped in like I know it was not
[06:20] even you know generally available at that time but we saw that as an opportunity to really explore and innovate and move fast forward. So that's what we did like I think with um lakebase building um you know the first app with openAI using lakebase as our backend like uh for agentic memory and
[06:37] you know going and evaluating and enhancing from then. So that was like where we were like okay we need to really change how we think about platforms and how do we support agent AI in the new era where things change so fast we need to be fast adopt and like
[06:52] try new things out. So that's been you know our journey so far since then. Yeah, I think the hunger for innovation is the common denominator here, right? That's that that's what drives us and the pace I think people are like hoping to do things in days when it would take months or years. I think that's I mean we see the database as
[07:08] well like suddenly I heard from both of you is like there's a shift left happening where now there are more and more people empowered to experiment otherwise it always used to be engineering which was the bottleneck for the lack of better word and then obviously the products are evolving. Yeah. uh and you know you need a new
[07:24] stack. So um curious as you were facing these challenges you know you mentioned about Slack messages changing from we have an idea let's go through this to I want to do this right now. Uh what shifted like what kind of strategic shifts you had to make within your organization or as a company to start catering to these needs.
[07:42] uh a few factors right so first and foremost right it it became quite apparent to us that we we should allow experimentation but it has to be governed and guard railed right it's because obviously like any
[07:57] agentic application AI application it is so data heavy but we also have a high degree of sensitive data so like putting the right governance parameters at the very beginning ensuring like the the right individual uals like personas have right access to the data that became
[08:14] quite important. The other significant factor there was how we are we received requests from our end users previous world right and end user would write a 10-page PRD that here you go
[08:30] read the PRD create a prototype for me tell me when you can build it the world today is very different right we are getting a created prototype and MV MVP hand it over to us and more often that so right the PRD is created after the
[08:46] MVP, right? The product already exists in some shape or form, but the expectations are also changing, right? Like take example, right? Like someone in the office of the CFO may have created a finance application, they have iPoded it. Now, they may have done it in two weeks, but their expectation is can
[09:03] I deploy it, can I scale it, can I productionize it in two weeks and very different thing, right? Like zero to zero to one, right? Prototyping is possible but taking it 1 to 10 needs a very different skill very different scale as well. So it's a technological
[09:20] advancement for sure. We have to rethink the way we deploy technology. In fact we are thinking about like deployment mechanisms to support vioded applications. We are thinking about how do we bring in bring in right harnessing techn technologies right databases but it is also educating the
[09:37] organization and managing this change and that's the people part of the things right like it's like somebody can vibe code it but they have to also understand and realize that it is not a linear journey right so you could make the 0ero to one very fast but there is a the 1 to 10
[09:53] journey has many other considerations security considerations considerations around scaling token consumption like app viability and many many other factors. So we have to be just be pragmatic of the fact that individuals and organizations they should allow
[10:10] experiments non-technical people will build apps but they need to be equally educated around the idea of what SDLC's are and considerations around STLC as they plan their projects and programs. Yeah. Yeah. Like plus one to what
[10:25] Phoenix was saying. I think for us um like honestly adoption was the easiest part of like just using self-service technologies. We were both shifting left and shifting right. You know shifting left as in um you know our product engineers never used data bricks before
[10:41] are like you know building silver level gold level tables in an hour like someone who's never done the barrier to entry is so low. Similarly on our uh you know sales ops marketing side folks who've never even had access to these systems uh or knew how to like build dashboards on their own are now like
[10:57] deploying apps and data bricks like you know so the uh adoption and self-s serve is not was not the problem the governance and the trust and no and giving to folks the confidence that they're doing right thing has been the challenging part for us. So it's more about like how do we make sure that
[11:12] there is curation there is business semantics there is metadata everything to help guide you know uh AI to make the right decisions but what data set it uses how it scales the eval around it and how do we kind of continue with that's been the hardest problem it's still work in progress I think um you
[11:28] know like the latest announcement on geneontology we're excited about that but I think that's like where it comes down to like from scaling is like how do we go and actually trust what people are building and is it starting to um I'm curious that are you starting to see some
[11:44] paradigm shifts in the the the stack or the primitives that we have been so used to using uh are they missing features are things that you know are finding were meant for when cloud was preai versus now when we want cloud to support us are the things that are becoming more
[11:59] important or less important yeah like um I can at least talk about some of the changes that I'm seeing um like lake based definitely you know I can talk a little bit more about that
[12:14] but where um there are like three different ways where how you're using lake based today one is like your um you know kind of an alternate for your traditional operational databases like Aurora MySQL Dynamob so there's like
[12:30] emerging use cases AI use cases where we're you know very proactively using it piloting it in actually production use cases. So that's like one shift that where that is helping us is not having to then constantly build pipelines to get the data offline, right? And we're
[12:46] also like and I'll probably come to you talking about lakebased because that's one thing I'm most excited about. But um uh the other area we're using it is we do have our own homegrown system where we move use it to move data online offline like if that if you think about it in the retail world you're always
[13:02] constantly moving data from real time to offline offline to real time and all of that stuff. So that's another area where it's becoming easier for uh you know our teams to like bring their offline data online and make it available in a very scalable manner. The third one is obviously the you know vibe coding you
[13:18] know um apps everyone is building it's so much easy to uh you know kind of just create your OLTP database and schema and like get your app running against it and if you want that offline yes it's easy to do that with lakebase and if you want offline data to be part of your OLTP
[13:33] database for whatever faster access or whatever live queries you have that too uh just like a testament to like how much we're can I you know excited about this um recently our founder built like uh cloud skill where you know you can
[13:48] build your wipe code your app and like you know use it to deploy to data bricks it has the option to like create a oldp database for you like using lake base and even a you know syncing automated syncing from like lakehouse to lake base and all of that and it's just like oneot prompt in cloud to do that so yeah
[14:06] that's yeah so that's like kind of where we're like most excited about lake base so I I'm seeing that as like really like a gamechanging technology for us you Yeah, I think I I I I would actually 100% align with um all of all of that. I uh if I were to really add more to it
[14:23] is semantics was hard even in the human era, right? Yes. It gets even harder in the AI era, right? Because obviously many of the contextual information is sitting on our head, right? So and and and that that was the hardest
[14:39] part, right? getting the information out of someone's head and documenting it and uh the taxonomy through which you document. Now in the previous world right um we had isolated applications pre era era AI era applications may talk
[14:54] to each other but it's not the isol the applications could fairly at least 90% of the time operate independently but in today's world right we would have a pyramidal structure of agents like mother agent chalk talking to child agents agents that are like interystem
[15:10] right like we may have a network of agents running in data bricks but that network of agents may have to talk to Salesforce, right? And they have their own agents, right? Obviously like A2A communication, right? M MCPS are becoming more and more of a norm. But these all these systems,
[15:27] these agentic workflows are probabilistic in nature, right? And uh we do not have a horizontal semantic engine, the contextual engine yet that can operate across the horizon. I think that that is the next innovation and evolution that
[15:44] we should target. That's our northstar in my opinion right like when when we are able to establish that agentic communication and agentic interaction using the context and the semantics that's uh that's that's the fundamental shift compared to right traditional cloud
[15:59] engineering space. Yeah. So I mean look uh I think a lot of us work in large companies and as I said like this is probably the biggest change we have seen in the industry. uh curious what kind of challenges internally because change is hard. We all know that changing how people operate, how people
[16:15] work is hard. Any challenges you faced internally as you were trying to educate the organizations about this shift? Uh any advice that you can share as we all are trying to uh you know enforce that change and organizations to adopt this in a safe and secure manner.
[16:31] Mhm. Uh I'll take a I'll I'll shift from a perspective of like implementation of technology to how do we build trust around AI right it goes back to my previous point around like organizational change and organizational education is needed right for AI to
[16:48] truly work. So we took the opportunity to build a new program which we are very early stages right in the in the program we call it like AI trust engineering or AI trust office. What really it is is like how do we establish trust across
[17:03] large complex organizations and improve the confidence quotient of the organization to consume AI and it's not always technology there are three pillars to it um the first pillar is the who pillar right like leaders who should influence like these are influencing
[17:19] leaders who should set the strategy around what are the guardrails governance parameters operating principles through which we should op use AI for example, right? You could have an AI that is built into your Zoom app or you could have an AI that is
[17:36] completely custom created using data bricks apps and we should not monitor and evaluate them in the same parameters right it's not one sizefits-all and that takes to the second pillar establishing policy parameters establishing rules of engagement as you
[17:52] consume AI in the organization extremely important right like take a quadrant-driven approach like is the how complex your AI is and how what's the risk parameters associated with AI, right? Like it's a complexity and risk parameter mesh of four quadrants, right?
[18:10] That's how we are establishing like evaluation of our AI tools and technologies because at at the end of the day, if we do not have a right evaluation parameters like we don't know what the AI is doing and that leads to the third pillar of observability and
[18:26] monitoring around AI. While we may unanimously agree that there is no set of defined industry parameters or KPIs, metrics through which we should observe and monitor AI, it might be different for different organizations, right? Like regulated organizations like banks and
[18:41] financial institutions may have a higher bar and unregulated institutions may have a lower bar. But we all should come up with parameters through which we will monitor our AI. And then are then there are organizations who are purely in the business of AI observability and monitoring like we have been speaking
[18:56] with some of them but we at the same time we are analyzing logs we have these monitoring dashboards we are we have these uh these uh indicators through which we we can operate to use AI with a higher degree
[19:13] of confidence quotient. So it it's not again going back to the point that the change was not about yeah we have the technology but change was also about that we can observe monitor and give guiding principles around how to operate the technology and we are very early
[19:30] right I mean we are we are learning as we as we as we move but it it part of it is like hit and trial experiments uh for us like I think with uh data mesh where um we've had to invest a lot in is
[19:46] how do we make it easy for people to do the right thing and that like included like when you're building a you know a data product or exposing your data product how do we build governance into it so that um at the time of authoring you know you're able to enforce these
[20:02] things because it's much harder to enforce things later um and also as part of that it's like the golden paths or pave paths is like here's how is your path to production for anything that you're whether you're building a a a subgraph on GraphQL, whether you're
[20:18] creating events for like you know real-time triggers or whether you're like building dashboards or insights or anything. What is the path to production and what is the pave path for that? So we've been investing in that a lot especially with data mesh shifting things to the left like where we're trying to help our developers is you
[20:36] know you model your data once define it once reuse everywhere no matter what access pattern it is. So that is like how we are scaling is like reduce the work for them to like make it easy that you do this work once and represent everywhere and on top of that like to
[20:51] honestly in the like world where everyone is able to use any tool to build anything it's like how do you still curate like how do you still like govern so we've been building registries registries of like your building skills like okay let's register let's really articulate what when it should be used
[21:08] and who should use it and same thing for MCP tools we're building is like the registry to like still bring the governance not only to data products but also like to how you access data. So that's like where like honestly like where as we're moving beyond just like you know getting everyone to self-service and VIP code is like to
[21:26] really scale that effectively where we're investing. Yeah. And um going back a little bit like to to your favorite talk lake base uh I think early and in the keynote also addressed this new pattern because now data is becoming the fuel for applications. Mhm. And historically we had spent a lot of
[21:41] time moving data from OLAP to OLTP. Curious with innovations like lakebas and other technologies, how are you seeing the world go forward? Are you starting to see that it'll be all one big single system or uh what's architecture that you're trying to
[21:56] aspire to? I mean yeah I mean that's the dream like you know like no more different operational systems one system one source of truth because um if you like just think about like where the trust problem comes in is like multiple copies of the same data
[22:12] and like then you have like it gets stale as long as soon as you copy it. So if you can avoid that copy and you have a single sort of through proof then like that reduces the risk and of like data corruption and data quality issues and that's like like if you think about our production like sites and production
[22:28] applications they don't have the same data duplication problem you have like in the data like lake world or you know analytics world right so because there's so many copies that exist so I'm really really excited about um if I think about it um like lakebase with their real time
[22:45] lake codes and um you know the LTAP that's most exciting for me cuz then that means you can reuse the same data in whatever you build whether it's production applications or analytics. So I'm really really excited about that. Yeah, I think it's a it's a it's a
[23:01] mathematical problem, right? So, if we look at pointto-point connections as AI scales in the organization, pointto-point connections is m * n m applications talking to n applications, right? It's not scalable.
[23:17] Now if we have uh 10 agents maybe we can scale it 20 agents we can scale it 100 agents absolutely we cannot scale like it's the simply the permutation mathematics doesn't work at one one point things will fail so technologies like lake base technologies like iceberg
[23:34] com layers in fact even like centralized mcp registries that converts that m* n problem to an n + n problem linear right and I think that that is that should be like our our goal as practitioners as engineering leaders
[23:49] that what efforts we could make to convert a multiplicative problem into an additive problem right I think that that's that's the in my simple mind like whenever I think about the problem statement is like is it an am I solving a multiplicative problem and making it an additive
[24:06] problem and how the technology is helping me to achieve that it's very nicely put um changing gears a little bit right we touched a little bit more about wipe coding and now sly is becoming way easier to do 0ero to one and we talked about how governance is is a big part of how to cover the last mile
[24:23] or the last five miles. uh but one of the things that I think you guys briefly touched is also how do we make sure the quality of the applications the trust uh and how do we iterate on it because while I think this technology is extremely innovative uh very rarely have I seen it work to the level of quality
[24:39] that we expect particularly for sometimes like financial systems or uh curious how how what systems have you put in place internally to like enable your application developers to iterate quickly on the quality and how do you define that this is not good enough to
[24:55] be productionized. a few parameters like I spoke about the three pillars of uh AI trust that that we have thought about but most importantly like the way we are we have operate we are operating is uh almost a reverse funnel right so very top of the
[25:12] funnel is quite wide right very little guardrails and controls like obviously there are guardrails and controls around like who has access to what data who has access to what kind of entitlements in the in the organization but we are allowing a lot of experiments at the very top right maybe hundreds of users
[25:29] will go and experiment if they want to but the experiments cannot live forever right so there is a there is an expiry date for the experiments right like the every on a 90-day rolling basis we review is the experiment still viable
[25:45] or we take away the data bricks workspace right and and it is by design another thing on top of the technology another another aspect which is by design is like any experiment It has to be supported by an executive sponsor. It has to be important enough of an
[26:01] experiment for an executive or a PNN leader to to back it out, vet and validate that this is a worthwhile experiment to spend time on. Now experiments can go costly. So we do have a fullyfledged PHOPS operations on
[26:17] top of it where we would actively inform the experimenters that what's the cost of the experiment. M we have a mechanism through which we are looking at like cost of AI versus ROI on the AI right it has to be obviously the ROI has to be equal or greater at least
[26:33] and I I would agree that if if it is equal it's worth worthwhile because sometimes it gives us other advantages over it's not always the greater factor right then there is the middle of the funnel that's where the MVPs are are there so
[26:48] MVPs are not experiments anymore these are these are budding applications. They may or may not become applications but we are now exposing them to their real world situations. We are uh scaling them. We are uh pressure testing them. Uh we are exposing it to
[27:05] even like end consumers who like let's assume like president of sales wants to consume a particular applications or uh seuite wants to consume a particular application. getting the feedback from senior leaders, getting the feedback from individuals who will use it on a day-to-day basis. Also pressure testing
[27:22] it from a perspective of token consumption, pressure testing it from a perspective of like security. Um, user experience is an important part of it, right? Like it's 0ero to one, you could create a concept, but you know that doesn't mean that the app is ready for end users. How a user is interacting
[27:38] with it. And almost like 50% of MVPs will get purged at that middle layer. Then there is the final layer which is the bottom of the funnel. The applications or prototypes that become enterprise
[27:54] applications truly scalable support is garnered. we could commit SLAs's around it, right? with like we could u provide right kind of sur and chaos engineering support around it and and going through
[28:10] that funnel maybe from starting from 100 experiments to possibly 20 MVPs to five applications at the bottom of the funnel and doing it over a period of like 3 months 6 months maybe even a year and repeatedly doing it that that framework
[28:27] is working well for us I'm curious like who owns the MVP layer is the your teams that because the top funnel is probably anyone who has access to a wide coding software can do something. Yes. Exa Exactly. So, so that's a good question, right? Like how the operating model is shifting. The top funnel you
[28:44] you give the tools, right? You go and go and go crazy with the tools, right? As long as like they are not breaking certain guardrails of the organization. Middlefunnel, it becomes a hybrid model, right? is like both the participant the product team or the sponsoring team is participating along with like technology
[29:00] organization engineering organizations like mine. Bottom of the funnel it shifts back entirely to the engineering organization to manage because now the sponsors the business users the customers they are they they just remain as a user. They
[29:15] can they can they can be an influencing party they could be an informed party but they are not a responsible and accountable party. one more zero. So at at Zillow when I just think about like quality and um you know how we're
[29:30] iterating especially for our agentic experiences, customerf facing agentic experiences um uh we depend a lot on offline evals. So that means like you know all our conversation data we like you know we're using already link based so that's available uh we pull in evals
[29:47] and traces data in there and we are then mapping it to your frozen snapshots to like reproduce what the agents saw or did that time so that like that feedback loop back to improving our agent experiences was like the five effect that we are like kind of like maximizing what we're also like investing currently
[30:04] like heavily exploring is like the lake based data you know database branching which is super easy to set up, but it's not just about database branching, but also like how do we evolve our application layer and how do we evolve our own development process to be able to iterate quickly and experiment like
[30:20] with all the like leverage the technology where you can like branch and have like all these experiments running in parallel. So that's our next step there like to uh take it from like where we are today to accelerate even further. Um so like that's like primarily like the eval is like we are or like you know
[30:37] because at at a consumerf facing level you're not going to be you know kind of leveraging UI if you don't trust the answers. So that's like so much important for us that we are actually helping the customer move to the next you know phase in their like house buying life cycle and are you guys seeing a shift like we
[30:52] databicks also we we uh are starting to see early signs of but curious like traditionally there was like offline evals you build an application runal if it looks good deploy it and you could afford not to change it for most part but with the innovation speed happening there's a new model coming out
[31:08] or there's a new technique coming out uh are you seeing a shift towards more online real time or continuous iteration of the quality even after it has reached the the bottom of the funnel. Yes. Yes. Absolutely. Absolutely. I think the eval frameworks are not static frameworks
[31:25] anymore. Like it's easier said than done, right? How do you make a dynamic eval framework like which means more work to be done by our research teams but that's just a grounding reality. Yeah. Yeah. And are the tools that you guys are finding useful or this is still
[31:40] early to I I think like for us like the challenge like towards moving online evals or like in general like scaling evals is you want that state machine of like what was the state of an entity at that time and be able to like use that to really
[31:55] evaluate. So that's like where we're like investing is how do we build that state machine make that available real time to like kind of scale like overall evaluations both online and offline. You see so from a technical stance I I I I don't think we are there yet where I could confidently say that we have a tool a
[32:11] technology specific like service provider who has taken us to that dynamic online state but yeah I mean I'm very open to the idea right opportunities to partner with even data brick share nodes to get to that that level where we could take it more dynamic and real
[32:32] um cool uh one to wrap it up one last question uh If you know we all carry a lot of uh legacy technical stack and we always have built in the era of cloud uh if you had to build your entire app stack uh from scratch uh and wanted to optimize for agentic applications and
[32:47] your applications what's the single most important architectural uh choice you would make that will enable you to move faster? I mean, I'll probably sound like a broken record here, but uh but honestly, like I think when you think about um you know, how can you uh scale
[33:04] AI with confidence, it's the trust that's the factor that's holding us back, right? Like it's how like like unleashing agents is like the only thing holding me back is like I'm confident the agent knows like how to get the same exact result every time it like gets the data, right? So kind of kind of I you
[33:22] know ref you know talked about this earlier but what gets in the way is that you don't have a single source of truth and then you spend so much time copying data from here to there and all that like that impacts engineering velocity that risk data quality and then it also like creates like this curation problem
[33:39] that's all this audit complexity so actually I was just like you know uh internally talking to my teams is like we should do a pilot we should do a pilot with um LTAP and lake house real time and see what we can do with it and prove it out and where can we use it because I really want to cut down all
[33:55] that work that I'm curious how much of your engineering hours are spend on this integration plumbing system a lot yeah like data engineering pretty much has like the same job the most fun part of the job I I would 100% agree with J right like
[34:12] creating that single source of truth uh we made an attempt at it last year uh time will tell how successful we we are using data bricks iceberg right so that that that is helping us but if I were to really reset everything uh one of the areas I would be more deliberate is like
[34:29] thinking about short-term and long-term memory systems from the get-go we thought those systems came afterthought in many situations um we obviously retrofitted things but if yeah if if we could I could really get a time machine that that would be
[34:45] something different that I will take an approach and perhaps ability to share memory in a governed way across agents and exactly exactly because we we retrofitted a lot of like memory management systems afterwards. Cool. Uh I think that's most what we had the time today but uh uh this is
[35:02] obviously one of the many sessions that we are hosting in terms of helping all of you understand uh how the entire agentic application ecosystem is evolving. Uh there is a a lounge in Moscone South that I would encourage all of you to uh spend some time and
[35:17] network. uh we also have uh tech industry main event which is the forum on Wednesday at 1 pm and another uh interesting panel and a talk that we are doing which is enabling and uh helping people build high velocity custom applications. We heard a lot about uh
[35:33] wipe coding but how do we use uh databicks app apps to build custom applications which are both secure and also uh ready for productionization. Uh so hopefully you will be able to attend those sessions and uh uh and you enjoyed this conversation as well. Thank you so much uh Phoenix and J for joining us.
[35:51] Thank you everybody. Thank you

Learn more about the Databricks Data and AI platform.

The information provided herein is for general informational purposes only and may not reflect the most current product capabilities or configurations.