Skip to main content

Self-Serve Data Foundations for AI Agents: OpenAI Case Study on Databricks

Summary

  • OpenAI's Head of Marketing Operations built a trusted data foundation on the Databricks Data and AI platform after years of failed AI projects caused by siloed marketing data, manual CSV workflows, and a broken sales contact form incident.
  • The architecture implements medallion architecture with bronze, silver, and gold layers, integrates Hightouch for activation, and uses Databricks Genie for data validation, enabling marketers to self-serve without SQL expertise.
  • The data foundation delivered $400,000 in monthly storage savings, reduced marketing form setup time from days to 5 minutes, and provided the governed data layer that turned previously failed AI projects into reliable production systems.

Self-Serve Data Foundations for AI Agents: OpenAI Case Study on Databricks

Watch: Self-Serve Data Foundations for AI Agents: OpenAI Case Study on Databricks
Jeff Canada, Head of Marketing Operations at OpenAI, shares how breaking the sales contact form led to rethinking data infrastructure. Faced with siloed marketing data, costly storage spanning 1 billion users, and failed AI projects, OpenAI built a trusted data foundation using Databricks as the system of record, Hightouch for activation, and the medallion architecture (bronze, silver, gold layers) to enable self-serve access.
Discover how to move from manual workflows and CSV files to automated, governed data pipelines that power AI agents and marketing teams. Learn real-world patterns: activating from the lakehouse instead of external CDPs, implementing medallion architecture, integrating Hightouch, and using Genie for data validation. Results: 400,000 USD monthly savings on data storage, forms spun up in 5 minutes instead of days, and marketers empowered to self-serve without SQL expertise.
🤝

Chapters

FAQs

Why did OpenAI's AI projects fail before they built a proper data foundation?

OpenAI's marketing AI projects failed because they were built on siloed, ungoverned data. Without a trusted data foundation, AI systems could not produce reliable outputs that could be handed to marketers or integrated into production systems. The speaker notes that the AI was doing what it was told — the underlying data was the real problem.

What is the medallion architecture and how did OpenAI implement it?

The medallion architecture organizes data into bronze (raw), silver (cleaned and conformed), and gold (business-ready) layers. OpenAI implemented this on Databricks as the system of record, with each layer serving progressively more trusted and curated data for downstream marketing and AI use cases.

How did the Databricks and Hightouch integration help OpenAI's marketing team?

Hightouch enabled activation directly from the Databricks lakehouse rather than requiring a separate customer data platform. This allowed OpenAI to power marketing workflows from the gold layer of their medallion architecture, replacing manual CSV exports and giving marketers access to governed, up-to-date data.

What were the three key lessons OpenAI's marketing operations team learned?

The video presents three key lessons drawn from OpenAI's journey: trusted data is the prerequisite for reliable AI, moving from manual workflows to governed automated pipelines unlocks both cost savings and self-serve access, and building a shared data foundation enables AI agents and marketing teams to work from the same authoritative source.

Full transcript

[00:06] Cool. Um welcome everybody. Super excited to be here with the like die hard fans who stayed to the second to last session of all of Data Bricks Summit. Um I appreciate it and I I hope to have a fun, you know, like 20-30 minutes with
[00:23] you. Um You know, every time I do one of these I always like to start with like a hook, right? So, the hook today is don't make any decisions based on the things that I'm saying and make sure to fill out your surveys.
[00:41] Did it work? Are you hooked? No. Um My hook today is 70%. Gartner just released a study at the beginning of this year saying 70%
[00:57] of enterprise AI projects failed to make it to production or show any demonstrable ROI. That's a lot of percents. I actually think it was 72, but I changed it to 70 cuz what's a couple percentage points amongst
[01:12] friends? Um My name's Jeff Canada. I run marketing operations at OpenAI and ladies and gentlemen, I have a confession to make. In 2025, I was part
[01:29] of that 70%. You might have thought I was going to get up here and be like, I have all of this AI figured out. Here's exactly what you need to do. That's not the case. Last year at this time, every single AI
[01:44] project I had was garbage. I couldn't give it to my marketers. I couldn't give it to production systems. It just wasn't good. So, today I want to walk through a couple of things that I did and how I messed it up and then how I
[02:01] went back and fixed it and some lessons I learned, so hopefully you don't make the same mistakes. I want to talk about two main things. The first one is the day I broke the sales contact form. Yes, the sales
[02:18] contact form, the one that's responsible for billions of dollars of revenue. I broke it. I was trying to solve a problem and I created an even bigger one. And the second thing
[02:34] is the time that AI failed me, but listen, I'm going to be honest with you cuz you stuck it out to come here today. It's really the times that I failed at AI. AI was actually doing the things that I told it to do. I just kind of
[02:49] sucked at AI at the time. Cuz I'll be honest, I'm learning this at the same time you are. I was just talking to Whitney and we were talking about how the playbooks that we had before are kind of thrown out the door. I don't know if you guys have learned something really cool this
[03:04] week that you can tell me or if you have a playbook that you can share, I'm all ears. I actually learn a lot just talking to people at shows like this. Um a lot of the stuff that I have put into production has come from learning from you guys cuz we're all in
[03:19] this together. So, the quick agenda I like to ground in like what my mandate is, where I work, um the mess that I would love to say I inherited, but I I created it.
[03:35] Uh cuz everything Excuse me. Everything is kind of being built from the ground up at OpenAI. Uh we'll talk about the fix that I implemented for the mess that I created, uh some of the results and the lessons.
[03:51] I always like to anchor these first in why I'm here. And why Open AI is here. And Open AI has a mission. And that mission is to create AI that benefits all of humanity.
[04:06] That is a very large total addressable market. I asked ChatGPT before I came into this presentation cuz in my head it was 8.8 billion people. It's actually 8.3 billion. I had it think really hard, too. I did like a long term like a reasoning. So if I was going to be late
[04:24] cuz I was like, "Come on." But it's 8.3 billion. But to anchor on what I'm looking at and the volume and the scale that I look at, ChatGPT, Open AI, and our API products have 1 billion weekly active users.
[04:42] Of that, there are about 2 million business customers as part of that 1 billion.
[04:58] One, two. Just two. I was a team of one for a year. I'm actually 12 years old, so I aged very gracefully. Um Anyway, so let's talk about the mandate for those two people, myself and and Steve, who's great. He was supposed to be here
[05:14] today, but he's holding down the fort. It's hard with two people. I'm responsible for building the foundation for marketing, for agents, and in that case, for agentic marketing
[05:30] at Open AI. One of my core things that I look after is rewriting the playbook for marketing operations. How do I make marketing operations AI first? I've been doing this for longer than I care to admit.
[05:46] And like I said, my playbook is out the door. So, I'm figuring it out. I'm talking to people like you. I'm talking to our vendors, our partners to figure out the best way to do this cuz this is all new. Everything's changing all of the time.
[06:02] So, let's figure it out together. Really, the core thing, the core responsibility that I think is my responsibility is creating a frictionless ecosystem where our marketers can
[06:18] operate. Whitney here is on our marketing team at OpenAI. My job is to make sure that she can do her job without worrying about like data, systems integrations. When she gets back to her desk on Monday, all of the leads from this show
[06:34] that she talked to should just be perfectly in the systems without having to wait and wait for me to do it and fiddle around with SQL queries and make sure the data is clean and part of that is making sure we have clean data.
[06:49] That's compliant with all of our global regulations and not messy. Speaking of, let's talk about the mess I made. And the day
[07:05] I broke the sales contact form. The funny thing is is that I was trying to solve a problem when I broke the sales contact form. But, like what happens a lot sometimes, you create more problems when you solve a problem. So, what was the problem I was trying to
[07:21] So, I was trying to solve? Marketers weren't able to access information easily. Marketing data was siloed in one one pocket. Product data was in another. If somebody filled out the sales contact form,
[07:36] we didn't really have an easy way to see if they was if those were ChatGPT users. There's a lot of privacy and compliance regulations. There's a lot of data that I can't access and I shouldn't be able to access, but it's all over the place.
[07:55] Storage for a billion users expensive. Very expensive, and we'll come back to just how how expensive that might be. There's a lot of operational risk. I talked about that. There's privacy and
[08:10] compliance issues with with marketing systems being able to access certain data. I don't want marketing systems to access certain data. I can't access certain data.
[08:26] And that created a bad experience for customers because we were having to wait to engage with people. We were having to wait for systems to catch up. And even the most basic personalization was really hard. If I wanted to send an
[08:42] email to you because you haven't logged into ChatGPT in a week, I had to download it on a weekly basis onto my computer, manually upload it into my marketing systems, and send an email. But if I did that, downloaded it on Monday, sent the email on Tuesday,
[08:58] and you logged in on Monday, you just got an email that says, "Hey, we haven't seen you in a week." And that ends up on LinkedIn or Reddit or the thing formerly called Twitter.
[09:15] And then I'd have to remember to actually delete that data from my computer. Cuz I manually downloaded it. It was sitting on my computer. If I didn't remember, I actually had a calendar reminder on Friday afternoons that said, "Don't forget to delete the data before
[09:30] you go home." I don't know why I talk like Goofy in my calendar invites, but I do. Um Also, so I wanted to solve that, right? I wanted to create some data pipelines, so I did. And I built an API from Databricks
[09:47] into our marketing system, and I turned it on and I went about my merry my merry way. Not knowing that I completely and totally overloaded the APIs in my email system, including the APIs
[10:02] that pulled in the data from the sales contact form to get that data to the sales team. That just stopped working, and I had no idea. All I was trying to do was just make it so I could know that if you wanted to talk to sales, I could say, "Oh, yeah, they are a ChatGPT user."
[10:20] Little did I know what I was doing was actually preventing them from getting any sort of communication whatsoever. That was not a Slack message that I wanted to receive, being like, "Hey, we haven't seen any leads today from this form."
[10:36] And I was like, "Oh, shit." So I needed to fix it, right? And like a good at OpenAI employee, I wanted to like throw some AI at the at the at the at the problem. So that brings me to when AI failed me. Or actually, we we talked about this.
[10:53] It's the time I failed at AI. I want to take responsibility for that. So the problems that I was trying to solve was that those disconnected systems were creating a lot of in- inefficiencies. We went through that manual workflow that I was just talking about.
[11:09] I didn't have the ability to put AI in there cuz I didn't really understand the process. I didn't have it documented. I didn't have things like naming conventions or what I was actually trying to do written down. There was no context. I was a team
[11:26] of one at the time, so all the context was up here. And finally, and probably most importantly, the data was siloed, was duplicative, was conflicting.
[11:44] And all of this, man, like I I really couldn't effectively run AI on top of these these processes. And even capturing capturing new data, what was a like a relatively simple process like creating a form, became a really long multi-step process.
[12:01] I had to get the web edge team involved. I had to go and create new fields in a in our marketing systems. The sales ops team had to create salesforce fields. The I had to make sure everything was configured and syncing. That is like if I timed it from end to end, probably a
[12:18] few hours, but in reality with meetings and slack requests and all that kind of stuff, it took days. Just to spin up a new form or make changes to an existing form. So I was like, okay, well, I want to create some AI to do this. So I go to
[12:34] Codex and I say, "Hey Codex, build me a new form." And it did. It built me a new form, but it didn't connect it to anything. It didn't route it to the sales people.
[12:50] It just did what I told it. It's kind of like running up to an intern and being like, "Here you go, run a marketing campaign. See you." and running away. Like I needed to give it the context and I needed it to have the data to support it. I needed to fix this
[13:06] because that was doing nothing for me. So how did I fix it? All roads were leading to data. My data was disjointed. My AI couldn't run on it. My marketers couldn't run with it. I needed to fix it
[13:21] and I needed to provide marketing and agents with easy access to customer data and actually I would add in here to context data. And maybe even like change customer data
[13:36] platform, we'll just keep it as CDP but like C squared DP which is customer data and context data platform where the United profiles aren't just people but they're also business processes. Naming conventions.
[13:52] Other things so that my marketers know what's going on and at the same time my agents know what to build and what to run off of. Now I told you I'd come back to this. I did some calculations on what I pay for my marketing system. If I were
[14:07] to put a billion people in there it would cost me a hundred million dollars annually. Right? Does anybody have a hundred million dollar marketing tech budget?
[14:24] Cuz I don't. And actually it's it's kind of wasteful. It's very wasteful. I don't need all of those people in my marketing database. There's a form on our website where if you want like a cool Codex sweatshirt
[14:40] you can fill it out. They don't really need to be in the sales pipeline marketing funnel. I need that data cuz I need to send them the cool sweatshirt. But I don't need it to be over there and I don't need to join it with all these other things. So I was I was going out
[14:58] and I was like okay I need a way to to connect all of this data together. I spent I've been at Open AI for about two years and some change. I spent six years before that at a CDP or a company. And I was looking at them and then I
[15:13] realized I was creating basically the same problem that I had before. The CDP was becoming this source of truth, but actually the source of truth for the rest of the organization is Databricks. Is my lakehouse. That's
[15:30] where sales is running reports. That's where data eng is running things. That's where finance and HR, I mean, everybody lives in the data warehouse, in the lakehouse. Why don't I live in there?
[15:46] So, that led us to Hightouch, which really just sits on top of Databricks and allows us to pull in data from all of these various sources and then activate them in different destinations. Now, again, I like to be honest with
[16:02] you. That didn't just fix everything. I was running into problems again because Hightouch was sitting on a bunch of disparate data systems and kind of making some assumptions. I was like, "Well, I Sorry, I wasn't going to
[16:17] say Um now I said it twice." Every time I tell my mom I'm going to speak at a conference, she's like, "Jeff, don't say fuck." And then I'm like, "Okay, I won't." And then I end up saying it three times now. Um I'm only allowed to say it three times. The fourth one, they'll pull me off stage.
[16:33] Um anyway, I digress. Uh Hightouch was there, but my data was still messy and Hightouch was having some problems with it. So, I went and talked to my friends at Hightouch. I talked to my friends at Databricks and I was able to actually talk to the marketing ops team, Elizabeth Hobbs, at
[16:50] the at Databricks, who is brilliant, and introduced me to something that I wasn't super familiar with. Um so, before this, I was like I was running through the slides sitting out on the floor over there and I was like,
[17:05] "Okay, I want to talk about medallion architecture." Is this something that everybody knows already? And ChatGPT was like, "Well." And kind of like linked it out and like, "Databricks and data analytics people for sure, marketing ops maybe, business
[17:21] users probably not. So, definitely talk about it. Don't spend a lot of time." And I was like, "Okay, good. Whew." Cool, because this for me was actually this was new. I've spent my career in marketing technology working in my silo little marketing tech.
[17:37] Not really thinking about how to pull in raw data in a bronze table, join it in a silver table, and then have these gold level tables to run my marketing off of. This was new to me and it like I just remember sitting in the office
[17:54] having this explained to me being like, "Woah, this is so cool. Now I get Databricks and data warehouses and everything's great." So, I used this and the first thing I used was actually forms. So, form data
[18:10] was piping through Hightouch and Hightouch events into a bronze raw table that's then joined into a silver level table with like the actual forms. Here's who who submit these individual forms. Here's the people and how they
[18:25] submit the different forms that then is joined into this beautiful gold table that I can now empower marketers agents to run off of. Took some work up front.
[18:41] But, I only had to really get the architecture built one time. And now I've enabled some really cool things to run off of this data schema in Hightouch. New data that comes in, a new form that
[18:56] gets spun up is now instantly in Databricks. I don't have to do anything. Those those people who fill out that form can be queried and built and an audience can be built in Hightouch and an email can be sent to them.
[19:12] Marketing can self-serve and build audiences and journeys. They don't have to come to me to write SQL, to download things on my laptop, to break the sales contact form. They can do it on their own.
[19:28] And I'm no longer in this huge bottleneck for people. I'm I'm an enabler because I came back and fixed my data. Marketing campaigns can launch super, super quick. I no longer am doing all of
[19:43] this querying. I do the that I actually like to do, which is like solve problems and configure how things get from point A to point B and build models. And then the marketing team and every time I say marketing, I'm also talking about agents that we're
[19:58] building. They can access those audiences, self-serve those audiences, and create really personalized information personalized emails, personalized journeys,
[20:15] and personalized experiences for people that I wasn't able to do before. This kind of a thing I wasn't able to do before. I didn't know that Ashley also worked in finance and was in the enterprise and was a technology employee. I couldn't connect that to the
[20:31] fact that she engaged with our emails before, but now that I have all of this in a shared data layer, I can put personalized content in front of Ashley based on all of this information that I couldn't do before because it was a CSV file sitting on my desktop
[20:50] or on my laptop. Who has a desktop anymore? No, no offense if you do have one. It also enabled me to have more time to do some cool stuff. So, I created a form form,
[21:05] which might not sound super cool to you, but I think it's pretty cool. And what this was doing, this gets into the next part of the context piece. Because I want my agents to be able to build a form and have the context as to
[21:20] what they're trying to do. What is this form actually trying to do? Is it just giving somebody a sweatshirt? Or is this Does this need to be routed to a sales person? And this data is now collected in a form, which I have pipe coded this.
[21:41] Um and piped into the data warehouse, where agents can now have access to why this form is being built, along with the workflows that I've documented that looked kind of like that five-step slide that I showed before of like, this
[21:58] is how I used to do it. WebEngs would do this, I would do this. That's in there. The context of why the form is being built is in there. And now AI has access to this full suite of information to build out a process
[22:14] that used to take teams, plural, systems, plural, hours, if not days to do, that now can be spun up off the back of that form form being submitted
[22:30] in minutes. I'm in the middle. I double-check everything. I do the same thing with Steve on my team. I have Steve double-check my work. We're, you know, making sure it's not just completely off there running doing
[22:45] its thing. But all of this is now happening in an automated fashion, and I can spin up I mean I no joke, watched a form get built right after I asked ChatGPT about medallion architecture. I got a ping
[23:01] that was like, "We need a new form for Tuesday." And I spun it up, checked it, approved it, checked that it got into Databricks, checked that the audiences were available in high touch in like 5 minutes.
[23:19] And for me, that's priceless. My time is my money. And Lord knows I'm really good at spending both of them. In inappropriate ways. Not inappropriate, but I'm just good at spending money um and
[23:35] time. So, now I've got this shared foundation. And it's not just a shared foundation for my team, which was like, I remember like back when I started this and the first time I I did a presentation in front of people, I was like, "If it's not in Salesforce, it doesn't exist."
[23:52] Well, that's that's that's not the case anymore. It's if it's not in the date in the lakehouse, it really doesn't exist. And if it's not accessible from the lakehouse, it's useless. So, now I have this foundation in my
[24:08] lakehouse in Databricks. I have these different agents. I have these different capabilities sitting on top of it. And then I have an activation layer. And if I I know I was saying there aren't any playbooks, but if there was a playbook, it would be this this kind of like data layer, intelligence layer,
[24:25] activation layer, which I think is the future of marketing. And then we can go in and and change in and out different agents. Marketers can jump in wherever they want to along the route. And we can start doing things at the scale of 1 billion people, of 8.3
[24:43] billion people that would probably have blown my mind in my past roles when I was like working in a database of a million and thinking that that was crazy big.
[24:59] So, what this unlocked was really a complete view of a trusted data foundation. So, I know if a marketer goes and runs a query using AI, which they love to do, and a few months ago they do that and I'd be like, "Yeah,
[25:15] sorry, no. That's looking at the wrong thing. That's looking at the wrong It's joining to the wrong account table." Now, marketers are running queries using AI on top of a trusted data source. So, I trust what they're doing.
[25:32] They trust what I'm doing. We all trust what the AI is doing. AI can self-serve. Marketers can self-serve. And I no longer have to go and fiddle around with the SQL query. It's all kind of built in to the
[25:49] structure of how things are operating now. And this is where like the cool agents and all of that stuff. I I started off trying to build huge and build all this agentic stuff and realized it's all
[26:05] garbage unless my data's right. It's that age-old adage. Garbage in, garbage out. But for this time, guys, it's real. Like you can't fake your way around it and just be like, "Oh, I'm just going to change this column in this one system because it's all now
[26:22] across the board. And now we can save time and money. So, we can go spend it on things that we want. So, the results really talked about this. That new form, I just
[26:38] spun one up in 5 minutes before this before this talk. It doesn't involve teams. It's just It's just us double-checking that the agents that I've built are doing what they're supposed to be doing and you know, not making things up.
[26:56] She still do sometimes. Um And I know I was talking about that 100 million annually. That's could be a real figure, but this $400,000 a month is what I'm saving on data
[27:11] storage costs because I'm no longer piping every single form into my marketing database. I'm only piping the stuff that's relevant to marketing and sales into marketing into my marketing database and therefore saving millions of dollars
[27:27] for the organization. On top of just like the headache of like accidentally emailing my mom about an enterprise feature and then getting a phone call from her. Or actually I have to call her. She doesn't call me cuz she doesn't want to
[27:43] interrupt. You know how moms are. Um And then there's a couple of lessons I learned along the way. We should activate from the lake house. That's the place where our data lives. That's where everybody lives. That's
[27:58] where things should start and end. You have to fix your foundation first. It's like a house. You try to build it on a shitty foundation, you're going to have a shitty house. And that single source of truth is that
[28:15] foundation. That's what enables my marketing team. That's what enables my agents. And that's what enables me to trust the things that it's doing because I know that the data that it's built on is up to date, is gold.
[28:33] And that's all I could ever hope for. Um I I think that's all.

Learn more about the Databricks Data and AI platform.

The information provided herein is for general informational purposes only and may not reflect the most current product capabilities or configurations.