Skip to main content

Custom Models at Scale: Reinforcement Learning for Enterprise Agents

Summary

  • Jonathan Frankle, Chief AI Scientist at Databricks coming from the Mosaic acquisition, explains that custom models trained with reinforcement learning on enterprise data can outperform general-purpose frontier models for specific agent tasks and can be built in weeks using Databricks infrastructure.
  • Databricks shipped three custom RL-trained models: AI Parse for document parsing, Knowledge Assistant 2.0 with 3x faster response times, and a specialized model for Genie — each built by the AI Research team as customer zero for the platform.
  • Reinforcement learning success is 90% dependent on data quality, requiring high-quality training examples, synthetic data generation, and the infrastructure to iterate quickly — all available to customers on Databricks.

Custom Models at Scale: Reinforcement Learning for Enterprise Agents

Watch: Custom Models at Scale: Reinforcement Learning for Enterprise Agents
Foundation models are fast, but custom models trained on your enterprise data are faster, cheaper, and more reliable. Databricks has released three custom models this week built with reinforcement learning: AI Parse for best-in-class document parsing, Knowledge Assistant 2.0 with 3x faster response times, and a specialized model for Genie. Each was built in weeks using infrastructure available to all Databricks customers.
Jonathan Frankle, Chief AI Scientist at Databricks, explains what makes custom RL work: it is 90 percent data quality. Success requires high-quality training examples, synthetic data generation, and the infrastructure to iterate fast. Hear how to build and fine-tune custom models on Databricks AI Runtime, capture inference calls automatically in Delta tables, and scale your agent ecosystem with specialized models that outperform general-purpose alternatives.
🤝

Chapters

FAQs

Who is Jonathan Frankle and what does the Databricks AI Research team do?

Jonathan Frankle is the Chief AI Scientist at Databricks, joining from the Mosaic acquisition, and leads the AI research team responsible for taking on technical risks and delivering the 'magic' that makes products like Genie fast and cost-effective. The team also acts as customer zero, building ambitious agents on Databricks before customers attempt them so products are tested and ready.

What custom models did Databricks build with reinforcement learning?

Databricks built three custom RL-trained models: AI Parse for best-in-class document parsing, Knowledge Assistant 2.0 with 3x faster response times than its predecessor, and a specialized model that powers Genie. All three were built in weeks using Databricks infrastructure and serve as proof points for what customers can achieve with the same platform.

Why is reinforcement learning described as 90% about data quality?

According to Jonathan Frankle, the success of an RL training run depends overwhelmingly on having high-quality training examples and well-constructed synthetic data rather than on model architecture choices alone. Without curated, accurate training data and fast iteration infrastructure, reinforcement learning will not produce a reliable, specialized model.

How can Databricks customers build their own custom models using RL?

Customers can capture inference calls automatically into Delta tables through Databricks AI Runtime, then use that logged production data as the foundation for fine-tuning models through reinforcement learning. The same infrastructure Databricks used to build AI Parse and Knowledge Assistant 2.0 is available to all customers on the Databricks Data and AI platform.

Full transcript

[00:19] Welcome back. So, we are joined with Jonathan Frankle. Jonathan, you came from uh the Mosaic acquisition, I think a couple years ago, right? I've been at Data Bricks longer than I've been at Mosaic at this point. So, you know, congrats. It's crazy, huh? Yeah. No, February was a crossover point. It was kind of nuts. Yeah, the first I mean the first data AI summit in 2023. I remember us chatting right after
[00:36] the mosaic acquisition and now you know three data dases later. Yeah, mosaic was like a a big turning point for data bricks. It really established a lot of credibility for us um with um generative AI and I just want you to talk a little bit about your role. you lead up our our AI research
[00:51] team and so can you explain a little bit about what that is, what you do, what your team does and um and what your contribution is to data bicks because I think it's really unique value prop that data bicks has that you don't really find in a lot of these other companies that are uh claiming to be able to do AI. So
[01:06] yeah, it's a complicated thing explaining to people what a research team does. So there are a bunch of different ways I look at it. One is that you know we're the tip of the spear for data bricks. We are supposed to take on big risks and try lots of big new things and some of them are going to fail. Some of them are going to work out really well. Many of them got announced as
[01:22] products in the past couple of days. Um and many of them will never get announced because they didn't work. Um but my job is to take on a lot of risk cuz you know as a company we should do that. Another way of looking at it is that our job is to provide the magic that we need to build new products that couldn't otherwise be built. Um you know
[01:37] things like Genie where we need some magic to make it really fast and really cost effective. And so cool we brought in some magic from the research team. things like AI parse. We brought in some magic from the research team to pull that one off. So, you know, our job is to do that. Another way of looking at it is that we're supposed to be customer zero.
[01:52] We need to do big ambitious things on data bricks that involve agents and AI before anybody else tries to do it. Break everything, help to fix everything, and then our products are ready for customers to go and, you know, build their own ambitious agents, try to get ahead of things. Yeah. Awesome. And I I know one of your favorite topics which is custom models
[02:10] powering our agents. That's one of the things you're super excited about, but I want to hear all about that. I love that you're using the phrase super excited. You're really, you know, talking right now. Um, yeah. So, we I think if you look throughout, you know, scattered throughout the announcements this week,
[02:25] you'll see I think I can name three custom models I'm allowed to talk about off the top of my head that were discussed this week. The first was an AI parse. We have a custom parsing model that is bestin-class in terms of both cost and quality um developed by folks on our team. The second is our updated knowledge assistant product which has a
[02:41] new custom model we've been working on um that I believe is like 3x faster and you know has better quality answers than you know the previous model we were using I think claude sonnet um for that particular product. The third one is, you know, we kind of teased this a little bit during the keynote today, so I can finally mention it. You know, a
[02:57] model for Genie, um, that has been customtrained and is both faster and cheaper and better, um, than other models that we've used previously for Genie that are, you know, from close providers. And this is just the beginning. This is just scratching the surface. Um, you know, that model for Genie was built in a few weeks. The
[03:14] model for, you know, KA 2.0 was built in a few weeks. And so we're going to be turning out a lot more of these models to come to make sure all of our products are great and highly specialized for what our customers need. But that's not the end of it because we use data bricks to do this. Oh, I guess they scored against they love it. No, they're
[03:30] cheering. They love the models. There we go. Um, no. I guess Croatia scored England. This just in England won. So I wasn't enthusiastic about England or Croatia until just now. We're we're dual Summit Live data bricks
[03:45] and sports. Exactly. But the I think the coolest part for me is that just like in the mosaic days, just like we've always done, the same technology we're using to build these custom models is available to our customers and our team is working handinhand with a handful of customers to help them build custom models for their products to also reduce their
[04:01] cost. We've all seen this move from kind of token maxing in the beginning of the year to the idea that we need to make sure things are cost effective and get value. And one of the best ways to do that is to train your own model. and we're here to help with our AR runtime, with our custom reinforcement learning stack, with our model serving, um, with all of our amazing data, you know, data
[04:18] stuff that we've had forever. It's an amazing way to do custom reinforcement learning and we're having a blast. I think that's a really good point because we're seeing this a lot with customers, too. They they start building an agentic app of some sort and to get to to get to basically market really quickly, they're going to use one of
[04:33] these foundation models, which are, you know, native in data bricks. Um but then over time you know if they get enough in insight you know they can basically get enough insight from basically going through our AI gateway capturing all those inference calls into a delta table which we do automatically for them and then they can basically use that to
[04:49] basically train or fine-tune a new model to focus on just like a a smaller component of that overall agentic system. So I wish it were so easy as to just you know take everything from inference table and train on it. Yeah. Um it's not just one button click. Uh if only it were one button click and reinforce learning just worked I'd be on
[05:05] the beach right now. We wouldn't need a research team at that point. My whole team would be just hanging out right now instead of working their asses off on this stuff. Well, now with Genie one, you could be on the beach and I mean that's the thing like actually one of our team members gave a great presentation on using Genie to figure out how to improve Genie. Like you know Julia who just joined us
[05:22] from Quotient was showing off how she was using Genie code to analyze Genie code traces to figure out how to improve Genie code. Um which is give you some good ideas. You never got to try it. Yeah. But I think the amazing thing about RL is that like everybody thinks it's the algorithm. Everybody thinks it's, you know, about the GPU infrastructure and
[05:38] all of that is really important. We've worked hard, but I have to tell our team members, it's 90% data. It turns out that, you know, good data quality, the right examples, all of that matters. So, yeah, you can have all your inference tables, and that's great for inspiration and figuring out what people are trying to do, but getting really high quality data. It's kind of classic
[05:54] data brick stuff. You need Spark, you need DBS SQL, and we've written a lot of custom infrastructure and data bricks to do synthetic data generation for reinforcement learning. And, you know, that's how we got that awesome model for Genie. That's where we got that awesome model for um AI pars. That's where we got that awesome model for KA 2.0 and a
[06:09] bunch more that I'm unfortunately not allowed to talk about yet. But you know, I was going to say what's the next model? But uh you know, I I like to speak through my work. That's been a consistent theme over the years. I'll tell you about what we've done once we've done it and it worked. And until then, you know, that's a that's a common like chess
[06:24] thing. I'll let my play my game speak for itself. Go try Genie. Go try KA 2.0. And I think the product speaks for themselves, especially AI parse, like the folks who worked on that did an amazing job. That speaks for itself. And and so that the AI pars is that like parsing PDFs, documents,
[06:41] parsing PDFs, best in class in the world. Um, and that was not us a year ago, but we worked really really hard. I I can't take any credit. This is other folks on the research team. They did an amazing job. They worked that problem incredibly hard with a lot of customers for a long time and built something pretty incredible.
[06:56] Yeah. Now I know um I I I travel around the world and people are very appreciative of you personally but then the the whole organization since you know uh companies in general could like just listen to customer feedback which is very much like priority here to
[07:13] improve products but then sometimes you want also want to get ahead of the market and do your own research to push the limits that's how spark itself was was created no one was working on that saw a problem saw potential so very much appreciated. And uh how do you figure
[07:29] out like what what your team should be working on? I think it's it's always a collaboration. You know, we we talk to customers a lot. Like I would say I do multiple customer meetings a week. It's this is probably the only company in the world where I could be a research scientist and spend a lot of time training models and working with, you
[07:45] know, folks who, you know, are incredible caliber researchers, just world class, and then also go spend a lot of time with customers. And nobody's confused about why I would do both. Why are you a researcher if you're talking to customers? Why are you talking to customers if you're a researcher? No, the answer is like the only true test of
[08:01] research. It's not some eval. It's not some benchmark. It's not anything but did we make a customer happy? Did we solve their problem? And if we haven't done that, we're wasting our time. Yeah. I think uh one of the things that's really unique about data bricks that I don't know any other company that does this is we you know the company was founded by seven PhDs basically and
[08:18] Mattea is still a full-time professor at Berkeley. Uh and so we've always had like one foot in academia and research and another foot in commercializing technology. So we're able to see like what customers are using today, but also what might be around the bend. And I think you guys actually do a lot of interaction with uh different research
[08:34] um around university as well to like try new white papers that come out to test them out to see if it's you know going to be useful or not. Yeah. And I like I think the people who we hire on our research team love to talk to customers. Like it's a it's something we actually screen for. like someone who literally just joined our team last week. He's here this week and
[08:50] he's been nagging me all week to come to one of my customer meetings and so he just came to a customer meeting with me. Had a bunch of great suggestions for them and wants to spend a bunch of time with them. That's awesome. That's the best kind of researcher in my view and that's how a lot of the key data bricks products actually came into being like Delta Lake by working really closely alongside customers and helping
[09:06] solve their particular problems. Awesome. Well, we have about 2 minutes left. Uh you're saying you're hiring a few people in the team. How does that look? Is it all in New York? Are you growing? Should they send their resumes directly to you? I mean, you know, you're welcome to send your resume to me if you really want to.
[09:21] Um, but I can't promise it's the fastest way in. We're hiring, honestly, in a really slow, picky way. Um, because like finding the kinds of people who are both amazing scientists and who love customers is actually really hard. It's not something you get trained for during your PhD. It requires kind of, you know,
[09:37] good empathy and a lot of good like soft skills in addition to just being an amazing scientist. It's not something you find every day. And we really work hard and work carefully to find those specific people. So I think just for like my subp part of the team that I'm focused on for reinforcement learning, I think I've hired one person the past
[09:52] year. Oh wow. Um because we have an amazing core of people and I like to keep things small and you know that startup atmosphere and you know finding the right people is actually really hard. So if you think you know you've got a PhD from a top place and you think you're an awesome scientist who's worked deeply in reinforcement learning and you spend a
[10:09] lot of time with customers, you know, shoot me an email. Um but it's a like it's a very small group of people in the world who would make it a data bricks in that way. Well um your PhD thesis the lottery ticket hypothesis I think we could we could end by saying data bricks won the
[10:24] lottery with you and mosaic you won the lottery here. So uh absolutely I I do have to say to my team like you know whenever the lottery ticket hypothesis comes up you're allowed to take a drink. So you know we I I brought it up without knowing that everyone take a drink. whenever a customer brings it up, you know, the
[10:40] lottery I I always say about the lottery ticket hypothesis like it was a lot of fun back in 2018 and you know it's been God and I feel so old. It's been almost a decade since I did that work and you know the whole world has changed since then like you know nobody knew what an agent was nobody knew what an LLM was
[10:56] like. So you know the field moves on. Excellent. Well thank you so much Jonathan for joining us today. I really appreciate it and it's uh it's always great to get to work with you and your team. Thank you so much for having me.

Learn more about the Databricks Data and AI platform.

The information provided herein is for general informational purposes only and may not reflect the most current product capabilities or configurations.