Building an AI-Native Enterprise: How Versant Media Unified Data Across 11 Brands
Summary
- After separating from NBCUniversal, Versant Media unified content metadata, audience analytics, and behavioral data from 11 brands serving 65 million households on the Databricks Data and AI platform, using Delta Lake and Unity Catalog to establish a single governed lakehouse.
- Versant accelerates AI-native development by using Genie for natural language queries across gold and silver data layers and generative AI tools including Claude Code to accelerate pipeline development, enabling blended AI-human engineering teams to move faster with fewer hand-offs.
- Strategic data products including customer identity resolution drive new revenue through personalized content targeting and smarter marketing campaigns, with Versant's Insights Portal exposing Genie APIs to business stakeholders across all 11 brands.
Building an AI-Native Enterprise: How Versant Media Unified Data Across 11 Brands

After separating from NBCUniversal, Versant Media faced fragmented data across 11 brands serving 65 million households with 14 billion hours of annual content consumption. Building an enterprise data platform meant consolidating content metadata, audience analytics, and behavioral data into a unified lakehouse using Databricks while maintaining granular compliance and privacy controls. Versant uses Delta Lake for reliable batch and streaming ingestion, Unity Catalog for governance and lineage, Genie for natural language queries across gold and silver layers, and generative AI tools like Claude Code for accelerating pipeline development. By shifting from dashboard delivery to rich silver data layers, blended AI-enabled teams, and strategic data products like customer identity resolution, the company drives new revenue through personalization, improved content targeting, and smarter marketing campaigns.
🤝
Chapters
00:00Introduction to Versant Media and Data Vision03:21Four Dynamic Markets and Business Portfolio06:17Unified Data Platform: Lakehouse and Products07:36Key Challenges: Data Fragmentation and Silos13:15Solutions: Databricks, Delta Lake, Unity Catalog16:42AI Acceleration: Genie and GenAI Toolkit19:05Operations: AI Agents and Workflow Automation23:05Team Structure: Blended Teams and Workforce Evolution25:52Strategic Data Products and 360 Data Sets27:14Versant Insights Portal and Genie APIs29:24Customer Identity Solution and Data Platform33:31Business Outcomes: Personalization and Revenue Growth
FAQs
Who is Versant Media and what data challenges did they face?
Versant Media is a media company that separated from NBCUniversal, serving 65 million households across 11 brands with 14 billion hours of annual content consumption. After the separation, Versant faced significant data fragmentation across brands and needed to build a unified enterprise data platform from scratch while maintaining granular compliance and privacy controls.
How does Versant Media use Databricks Genie?
Versant uses Genie for natural language queries across both gold and silver data layers, exposing these capabilities to business stakeholders through the Versant Insights Portal via Genie APIs. This allows non-technical users across all 11 brands to explore audience and content data without writing SQL, accelerating time-to-insight for marketing and content decisions.
How does Versant Media use generative AI to accelerate development?
This video describes Versant's use of Claude Code as part of a generative AI toolkit to accelerate data pipeline development, enabling engineering teams to build faster with AI assistance. The company is structured around blended AI-human teams designed to minimize hand-offs and inefficiencies, treating AI as a core productivity tool rather than an experimental capability.
What are Versant Media's strategic data products?
Versant's strategic data products include customer identity resolution, which links anonymous and known user profiles across brands to enable one-to-one personalization, and audience 360 data sets that consolidate behavioral and content data. These products power personalization, improved content targeting, and smarter marketing campaigns that drive new revenue across the company's portfolio.
Full transcript
[00:09] Thank you very much. I super stoked to be here. Super stoked to be at Databricks Data and AI Summit. I hope everyone's having a a good summit so far. The keynotes are hard act to follow as well, but we'll do our best. I'm going to be talking today about Versant Media, our new company, and how we're really trying to reinvent ourselves as
[00:25] an AI native startup and using data and analytics as a way to to fuel of that. Before I get to that though, a couple of things about me. Um I'm SVP of data products and technology at Versant. I've been in the space for way too long and embarrassingly long
[00:41] amount of time. And most of that in data and analytics for telecoms, media, and entertainment. Prior to this role, I was chief architect of Peacock and our other NBC Universal streaming platforms around the world. And prior to that, I was head of solutions for media and entertainment
[00:56] and telecoms at AWS. Um A couple of things though that are kind of personal to me that I'm going to use as themes as we run through this. One and thankfully they're nothing to do with obnoxious heavy metal music or my love of horror movies.
[01:12] One is is cars and I draw a lot of analogies between the cars that I love to tinker with and Databricks as a platform and how we leverage it. I'm a big fan of American muscle cars. The older the better, the bigger the block V8 that you can fit in that thing
[01:27] the better. And an and a big engine is a great foundation for for a for having a fun time on four wheels. But it depends what you do with it. You can then add a lot of value by tuning that. You can add new cold air intakes and turbo chargers and super chargers. You can drive it
[01:42] differently. You can have different electronic systems around it. And that's kind of how I see Databricks as a platform that's a good foundation, but it depends how you drive it. It depends what you build on top of it. And it depends what data and fuel you put into it. So, that's kind of point number one is kind of taking that solid foundation,
[01:59] that powerful platform, and building on top of it. Point number two is my dislike of inefficiency, and I think as we enter an AI-native world and really reinvent the ways we work, you know, things like hand-offs, like motes around teams, really slow everything down.
[02:14] Anything that requires signatures in triplicate just kills me unless it's a mortgage or a will. So, I'm trying to stamp those out. And we're not only going to talk about the platform we use the data that we put into it, but also how we use that platform and try to drive the most efficiency that we can out of it.
[02:33] So, what is Verson Media? Who are we? You may not have heard the name, but you've probably heard some of our brands. We're home to a diverse collection of 11 well-known brands. Our portfolio includes TV networks, streaming services, digital transactions, software
[02:50] solutions, and our brands play an integral part in the lives of our fans because live is still the domain of the television network, and Verson is there at the moments that people truly care about, whether that's a stock market opening bell, whether
[03:05] that's election days, exhilarating moments on the pitch, or the most talked about red carpet events. And we spend we spent the past year separating from Comcast NBCUniversal, our parent company, but also looking at how can we reinvent ourselves for the future and
[03:21] really transform the way that we operate and how we how we turn up. Now, we operate in four dynamic and separate markets. We have business and financial news, which is led by CNBC. Political and opinion news, which is led by MSNBC. We've got golf and athletics
[03:40] where we're combining live golf coverage through the Golf Channel with golf tee time bookings through our GolfNow business and educational resources in our GolfPass business to combine those into a true holistic golf portfolio. And
[03:55] we have sports and general entertainment, everything from our television network through to purchasing movie tickets via Fandango, rating and reviewing movies in Rotten Tomatoes, and streaming content on Fandango at home, now called Fandango Stream.
[04:11] And I'm not going to read our whole vision statement to you, but a couple of things that I do think are important to call out. We do see ourselves as an industry-changing force, and we aim to be disruptive, innovative, and entrepreneurial. And to do that, we need to turn up differently. We work across a
[04:26] broad set of brands, content types, and genres, and we have a relentless focus on our customers and the communities that we serve. And talking of those customers, a few headline stats, um in 2024, we had 14 billion hours of our content consumed
[04:43] across all of our nascent brands. We served 65 million households in the US, which is around half of the US population. When you compare that to popular streaming service, that puts us in front of every streaming service with the exception of Netflix. And our
[04:58] digital platforms processed over 140 million transactions per year. That could be movie ticket purchases, it could be a golf tee time bookings, or other transactions that we handle. So, we we have a lot of scale and a lot of customers on the platform.
[05:14] And how we how we intend to grow is really fueled by data and analytics as well. Obviously, we have content on the left, content is king, but which content do we invest in? Which content do we acquire? Which platforms should we partner with for the distribution of
[05:29] that content? And how should we combine that content to create a compelling slate? We also intend to broaden our audiences, but we need to understand more about the millions of customers we have on the platform today, understand more about them, and meet them when they're at where they're at. And we're looking to
[05:46] build new digital personalized experiences for those customers, but meeting them where they're at requires a lot of experimentation. Maybe they want different content types, different ways to consume that content, and we're going to be experimenting actively, which generates a ton of data and the need to
[06:02] analyze that data very, very rapidly. So, in short, we've been presented an opportunity. We were born from Comcast NBC Universal, a company that has over 100 years of experience in the entertainment industry, but we now we now have the opportunity to reinvent
[06:17] ourselves as a true AI-native startup. So, let's talk about how we actually do that when it comes to data and AI. So, our vision is intentionally straightforward. We create a data platform that offers trusted, reliable,
[06:34] high-value data sets. On top of that data platform, we then create data products that are reusable, intuitive, and simple to use. We then ensure that these products are laser-focused around business growth and value. And then we democratize those products
[06:50] to our key stakeholders in a self-service fashion. But when you think back to the day of the the previous role of data products and analytic and analysts, um that role was typically around creating dashboards and reports. So, a business user would
[07:05] ask for a new dashboard, an analyst would go through, create the requirements, we'd build data sets and pipelines, and eventually after a few months, we'd give that business user a report, and they'd say, "I guess it kind of fits what I need." Uh we'd email it to them probably on a daily or a weekly
[07:21] basis, and eventually that gets filed into their junk email, and and they've forgotten that they requested that report in the first place. I'm not sure that uh is the most valuable use of the time of an analyst or a data product person. What I think the way this is evolving is
[07:36] that we're now focused much more on building really rich data the If you're familiar with the medallion concepts, really rich silver layers that are granular that join things like content, audience, product analytics, and then let analysts go nuts in there. Ask the questions that the business cannot
[07:52] answer today, and generate net new insights, and and really add value to the business. So, that's where this vision is is essentially taking us. But, we have headwinds, and I think the first set of headwinds that we'll approach are ones where we we we're
[08:09] consuming data, we're collecting data. So, when it comes to content metadata, each of our brands is storing that content metadata in different formats, in different technology stacks. So, the first job is just to bring that together into one single content da- data lake.
[08:25] The second area is around audience analytics, and it's really similar. Whether it's Nielsen data related to cable TV viewers, whether it's event tracking for websites or digital apps, or whether it's through third-party channels like YouTube or Roku, we bring all of that into one data lake for
[08:42] consolidation. Similarly, on the product side, we gather how a customer uses our apps, what they did or did not view within the apps, and the quality of their experience while they were in the apps. Again, bringing that all into a single place. And for revenue, we track direct billing
[08:59] relationships, third-party billing, ad revenue, and the success of our marketing campaigns. Lastly, we bring in third-party data such as social media analytics, box office forecasts, competitive analysis, and we use that to enrich our
[09:15] first-party data. But, we're still faced with a number of headwinds around this. So, viewership and content data is fragmented across our platforms, and it and it's really difficult to map. So, revenue calculations are increas- increasingly
[09:30] complex in a modern digital landscape, especially when you're using partner channels and third-party in-app purchases. Uh we also inherited silos of customer identities. So, if you imagine we have silos of identities for Fandango, for
[09:46] Rotten Tomatoes, for our golf businesses, for our entertainment brands. And bringing all of that together can be really challenging. Also, within that world, we have some brands where we know a lot about our customers. If I buy a movie ticket on Fandango, Fandango probably knows my email address, my phone number, my
[10:02] address, my billing information. There's a lot of rich data there. But, what if I'm just watching content on cable TV? Is you know, do do does does VIZIO know anything about me in that context? What if I'm watching content through a website and I'm not logged in? Or I'm a
[10:17] user on Fandango who's unregistered but watching content that's funded by advertising? There's kind of a sliding scale where we know a lot about some customers. For others, we may just have a device ID, an IP address, a zip code, or nothing at all. So, really how do we bring that together and do something
[10:34] usable with that data? Also, as we're launching new digital experiences, as I mentioned, there's a need for experimentation. And that experimentation requires richer product analytics data. We really need to know how a customer entered the platform, how
[10:49] they traversed the platform, how they interacted with the platform and for how long. What did they watch? What did they skip? And we need to be able to analyze that very rapidly. So, more data, more volume is a is a problem. And then when it comes to um
[11:05] uh governance and compliance. Governance and compliance can mean a number of things. We've clearly got regulatory compliance. We have customers opting in for different brands in different ways. Just because I opt in for marketing for Fandango does not mean I'm necessarily opting in for Fandango, sorry, for
[11:21] marketing from CNBC. So, we need to manage compliance and privacy at a very granular level. Um but we also then have to enrich our semantics layer and sure that that's up to date. We have to track data lineage for our auditors, and uh we need to
[11:37] manage cost. And I think cost has always been a concern for data platforms, but now in the world of AI, the last thing we want to be is uh the the next headline about uh a billion dollar spend on gen AI because we didn't keep a finger on the pulse of cost. So, cost is
[11:53] is only more important uh these days. Then when we talk about data science and AI, I think there's a hunger. You know, personalization is no longer just about having a rail that says, "Hey Simon, here's what we recommend you watch next." I think there's a a desire to
[12:09] personalize every way that we interact with the with the customer, and that requires not only clean data in order to drive that personalization, but especially as AI becomes a piece of the puzzle, we then need to uh look at the semantics layer and ensure that that's
[12:25] actually feeding the right context into AI models. So, that becomes another challenge that we need to solve. And then data democratization, really we want to get out of the business of answering every single question that every single business stakeholder has about our data and be able to
[12:41] democratize those insights in a way to our to our stakeholders, but we need to do that in a way that still maintains trust and integrity. If we lose the trust in our data, we lose the trust in our platform and and people stop trusting the the analytics function as a whole. So, there's nothing more critical
[12:58] than the integrity and the trust of our data. So, let's doom and gloom more about how we solve these, and it would be weird for me not to be stood here and talking about uh using Databricks tools to solve these. So, what are some of the tools that we're bringing to the table? Well, to solve the fragmented data
[13:15] issue, it's really a combination of two things. We're bringing the data together in a unified lakehouse. I think that's a concept that's been around for quite a while, but just bringing the data together is helpful to us, and we'll talk about how we do that in a moment. Uh but we're also as we're bringing that
[13:30] data together, we're leveraging AI to enrich that data. So, one of our first forays into that world is using AI to scan articles that we're bringing in. They could be text articles, video articles, and then creating rich tags. So, in the news world, this may be that
[13:46] we have articles that mention Donald Trump. It's a political piece. It mentions the White House. It mentions the conflict in Iran. So, that can be used then for analytics purposes so we can understand what the appetite of our customers are for for reading different articles, but it can also be used for
[14:01] personalization, for richer search experiences. So, leveraging AI and using Databricks as the engine that kind of drives that tagging is really important to us as we're bringing that data in. Then when we talk about disconnected customer identities, again we're doing
[14:17] two separate things. One is using Databricks as the landing place for those identities and for the behavioral data related to those customers so that we're at least having those in in one central place to slice and dice. The next thing is we're leveraging a Databricks premier partner to go out and
[14:35] actually activate on that data. And I'm going to talk through the workflow around that in in just a moment in a lot more detail. When it comes to volume and velocity, again we're kind of solving that in two ways. Uh Delta Lake I think is really helpful for the reliability and the
[14:51] quality of loading the data into Databricks uh both when we're doing that in in a batch workload point of view, but also for real-time streaming as well. But I think what's really important to us is we don't necessarily need to um have have a full-on migration to Databricks to take advantage of the
[15:07] platform. If there are large data sets, of which there certainly are when you talk about data sets like Adobe Analytics, Google Analytics, and Particle, they're pretty significantly sized data sets if you're if you're running a platform that supports millions of customers, um you don't
[15:22] necessarily need to load that into Databricks. You can actually federate that data if it's in something like Redshift or an S3 bucket or an RDS database. You can federate that data into Databricks and then make it available through a zero copy mechanism. So, that data appears to be immediately
[15:39] available in its raw form and you can then build silver anonymized layers and then gold business logic layers on top of that. So, it's been a great way for us to move fast and then mitigate the cost of some of those larger data sets. Now, I think everyone's probably
[15:55] familiar with Unity Catalog as a semantics layer. The I guess the the hidden benefits of Unity Catalog uh as we're adopting it uh allow us to use it more for governance as well. So, we can use it for granular access control. We can integrate it with
[16:10] our single sign-on solution to determine who has access to what uh workspaces, schemas, tables. And we were were also able to use it for um uh data lineage as well. Whether that data resided in Databricks or not, we can trace the data lineage. So, that
[16:26] keeps our our auditors happy when it comes to audit season. Um, and we can also use Unity Catalog to uh to both keep control of cost and have visibility of cost, but put budget constraints around the cost of the the data within Databricks as well. Uh
[16:42] and particularly as we're looking at AI workloads, Unity Catalog now has an AI gateway that can actually govern the the costs of the the uh the AI that's in use within Databricks. So, we're also really leaning hard into AI on Databricks as an accelerator. And
[16:58] I think this comes in two forms. Uh one is Genie and I think most people are familiar, particularly if you're here, that you can go into the Databricks console, you can use Genie and say, "Hey, what were the biggest Fandango movies at the box office over the past weekend?" And how does that compare to
[17:13] the previous 3 months? Give me a table. Give me visualizations. And it would do a reasonable job of creating bar charts and trend analysis and all of that good stuff within the Databricks console. Um what we're actually now looking to do is leverage that and democratize that
[17:29] out to more customers. And I'll talk about how we're doing that in a moment. But one of the reasons that we're doing that, even before today's announcement of Genie ontology, which I think is going to be huge, um we did a bake off between uh other tools like Claude code, like Codex, and Genie to look at the
[17:46] level of accuracy and fidelity if we're querying data within Databricks. And Genie outperformed all of them significantly. And the reason for that is simple and and makes sense. Genie is connected to Unity Catalog, so knows about our tables, our column
[18:02] definitions, knows about the business logic that's related to our data. It's not trying to figure that out on the fly. So, just because we've seen Genie provide such higher accuracy results, we're really going to be doubling down on Genie as a way of democratizing that
[18:17] data and offering it up in a natural language uh model. The other thing uh that we see GenAI within Databricks as being a huge accelerator is that it offers a GenAI toolkit. So, along with Unity Catalog AI Gateway, we can use tools from the GenAI toolkit like um
[18:34] like Claude code, like Codex. Uh we can use Claude code within our engineering teams to speed up the development of new pipeline code, new tables. We can use Codex within data products to be doing data modeling and data schema design and building more of a strategy. So, it
[18:49] offers different tools that that excel in different areas that we can then bring together without the need to provision any new infrastructure or request anything from IT teams. So, again, just allows us to move very, very quickly. But it's not just about the platform
[19:05] itself. I did think it was worth just mentioning how we operate. So, along the top here, you can see a very traditional way of working and this is our data products and technology way of working. Could just as easily be a product team requesting new features for an iOS app
[19:20] for a streaming service. It's very very similar very familiar to all of you I imagine. On the left you've got intake a business stakeholder requests something. We gather some requirements. We we refine those requirements. We produce a rough level of effort and we start building Jira tickets and and all of
[19:37] that good stuff. Then we get into into prioritization and roadmap. We look at what is the business value of that feature that's been requested and how does it stack up against the rest of our roadmap and we get into governance then around things like data governance, semantics layer governance,
[19:54] regulatory control, then into designing the data model, the schema and building it building out the pipelines, building any required dashboards or reports and productionizing and operationalizing it. So hopefully none of this looks too wacky or unfamiliar to anyone. What we
[20:11] then looked at was how can we use gen AI tools to speed all of this up? Look at each of these areas in isolation. How do we use tools to do that? And as I mentioned we found we started early with GitHub Copilot. We saw some minimal benefits, but as the tools have evolved
[20:26] tools like Claude Code have really taken an exponential leap forward. So have really helped us within you know data engineering write faster code, validate that that code and test it and roll it out faster. Similarly within data products Codex is looking really promising for helping us do data
[20:42] modeling and schema design. But all of those have been used more in terms of a person with a browser or an IDE that's calling that that that AI capability on the side asking it a question implementing it back into their workflow in a highly manual way. What
[20:59] we're now looking to do is roll out AI agents across each of these job families. So we'll have a project manager gen AI agent that's going to be doing a lot of the intake. It will do basic things like is you know, does a request already exist for this same feature? Does that feature even exist
[21:15] within our platform? And and ensuring that we're we're deduplicating our requests as they come in and doing some of that admin work. It'll create early tickets for the work. Then we get into a data product manager role where it'll be looking deeper at the requirements,
[21:31] having a conversation with the requester to refine those requirements, looking at the business value of the request when we've got the requirements a bit more fully baked, and then reprioritizing and communicating our road map based on that. Then we get into governance and and
[21:46] governance is not an area where I would feel that comfortable relying on AI at the moment. You know, there's the stakes are too high there. So I wouldn't necessarily rely on AI for things like regulatory compliance checks, but where I do think it adds a ton of value are
[22:02] things like populating our semantics layer. So it's pretty great at scanning tables and schemas and and future designs and then building a semantics layer from those that a human operator can then go in and and improve upon over time.
[22:17] Then we get into design. As I mentioned, schema design, data modeling, and the rolling out of those. And the same on the engineering side of things. And there's no reason, even though we're building out these agents, there's no reason why we can't use either things like Claude Code or
[22:33] Codex behind the scenes that are being instrumented by those agents, or Databricks native functionality like Genie as well. So really just ways that we're we're speeding up our workforce. And then last on the operations side, we see a ton of of promise in AI as an
[22:49] operator as a support mechanism both for old-fashioned data pipeline succeed and fail type checks, but even more when it comes to anomaly detection, data quality checks, validation, that kind of thing. So, how does that change the way we
[23:05] actually turn up in terms of project teams? Well, I think we're already working as blended teams. I think the days of siloed organizations slowing us down have hopefully ended. So, um what we tend to end up in is a world
[23:21] where we've got one or two analysts. So, these are the subject matter experts who are experts on a specific business domain. They may be experts on Fandango movie tickets or on CNBC stock tickers or on MSNow's opinion news.
[23:36] From there, um and and they I guess as I mentioned earlier, these are folks who probably in the old world were focused on getting a request and building a dashboard. Right now, their their their role is so much richer than that. It's uh querying rich broad silver data
[23:52] layers, finding out what uh questions the business is unable to answer today, and then asking those questions and coming up with new with net new insights that really show growth. Uh they're backed up by data product managers. So, these are the people who
[24:07] are leveraging AI to where we need new data sets, where we need new uh new uh artifacts like dashboards and reports, these are the folks who will be owning the road map of delivering that, but heavily leveraging AI to turn those requirements into a reality.
[24:25] Then we have a a full-time on-staff data engineer, and this person is again responsible for building that institutional knowledge. We We feel that the the role of people in this equation is much more that. It's the the people who truly understand our business, who understand the the business logic, and
[24:41] have that acumen. And then they're responsible for managing our consulting teams, managing the AI agents, and ensuring that we maintain the optimum efficiency and productivity out of those groups. And those uh those people are backed up then by consulting teams that can be a
[24:58] blend obviously of onshore and offshore and uh, who are responsible for QA validation of anything that AI is producing and the AI agents themselves who really just act as virtual team members to uh, to actually go accelerate. So, the whole model here is
[25:16] we have consultant data engineers on each project who are delegating down to AI to actually go write the code, write user tests, write do the validation of that code. The consultant data engineers are responsible for implementing and checking that code and productionizing
[25:32] it. Then it all rolls up to our our in-house uh, engineer who is uh, hopefully sitting on that that uh, business acumen along with our product managers and our our analysts. So, really should allow us to be much more agile, much uh, move much faster and maintain a good velocity.
[25:52] So, this then allows us to offer an enterprise intelligence platform that is trusted, accessible in a self-service manner, and accessible to allow our customers to glean insights in new ways. And on top of that, we then build rich 360° data sets that are views of our content,
[26:09] our customers, our digital products, and our business revenue. And we join all those at a granular level and provide that to our analysts to be able to gain the maximum uh, value from those data sets. And then we use that data in strategic data products. So, I'm going to walk
[26:26] through a couple of examples. One is our enterprise-wide insights portal uh, that provides our executives and stakeholders with real-time uh, up-to-date business information. Another one is our customer identity solution where we pull together those
[26:42] those fragmented IDs and activate upon them. Uh, but we also have a semantics layer that we provide to our analysts so that they can find the the data that they need and the business logic that they need within the context of our enterprise and tools for monetizing that
[26:58] data. So, we provide some data to other partners as B2B opportunities to monetize it and obviously there are B2C opportunities for monetizing our data as well. But, let's take a deeper dive into a couple of these products. So, the first one is the Verizon Insights
[27:14] portal. So, this is a a single source of truth for high-value insights, KPIs, dashboards, and reports across all of our brands. And we have granular access controls in place that integrate with our SSO platform and determine who can and cannot access information for
[27:31] specific brands and sensitive information within those brands. So, we have granular access control there as well. And for each line of business, we surface KPIs um for the performance of those brands across linear cable TV, digital distribution, and streaming
[27:47] platforms. And we support live dashboards in tools like Tableau, Domo, Looker, Omni, including things like filtering and email subscriptions. We also support the upload of PDFs and PowerPoint documents so that we can have custom reports in
[28:03] there as well. But, I know what you're thinking. Um you're probably thinking, "Simon, this just looks like you've built some iframes around some dashboards. What's the big deal?" And And it's fair, but remember our company's only been alive for 6 months. So, it's we're still on the on the start of
[28:18] our journey, but I think where that journey takes us next is a lot more exciting. And on next on our road map is leveraging Genie APIs. So, again, I'm sure people have experimented with Genie within the Databricks console. Genie
[28:33] also offers APIs from which you can you can send uh queries, you can pull back results. So, again, because Genie's been so promising and the data accuracy has been so high, we're now going to be offering Genie APIs into the Verizon Insights platform for business users and
[28:51] key stakeholders, that's primarily going to be sourced from the gold layer, which is our highest accuracy curated layer that has business logic in there, where we can be most confident of the correctness and the accuracy of our data. For analysts and data product managers, we also then have the silver
[29:07] layer, that rich broad granular layer, that can also be offered up through Genie. So, we're going to be offering both flavors depending on the job role there. And then I think that the value of this portal exponentially increases. Moving on to customer identities and
[29:24] customer data platform. So, again, we we ended up in this world where we inherited a bunch of brands that stored different identities in different tech stacks, different different formats, and some brands that had no identities whatsoever, or just things like a a
[29:41] mobile device ID or an IP address. So, what we're doing here is we're pulling together a number of pieces of information. We have tracking data that could be from Google Analytics, Adobe Analytics, mParticle, the usual suspects. We bring that in so that we have some rich
[29:57] information about where a customer came in from, their device that they were using, those types of data we bring into Data Bricks. We then bring in compliance and privacy information so we understand did Simon only opt in for marketing comms for Fandango, or did they opt in
[30:14] for for other comms from other brands, so we have a granular view of privacy and compliance. We then bring in customer profiles, so that could be from CRMs like Salesforce where we're bringing in profiles and know a little bit about our customers. It could be other data sets that sit on
[30:30] S3 and Redshift and other other data repositories, but we consolidate all of that. Um and then we we bring in our other customer IDs from the various source environments for each brand. That all comes into a raw bronze layer where we keep things as raw and as as
[30:47] original as we possibly can. From there, we start to graph it and create an identity spine. So, this is where we would use deterministic mapping where we know something about our customers. If we have their email address, phone number, address, we bring all of that in and and start to do
[31:03] deterministic mapping across our brands. So, we'll know, for example, that Simon was a CNBC pro subscriber, a Golf Now subscriber, and bought some tickets on Fandango. We also then bring in that partial anonymous data for customers who have
[31:19] just watched uh ad-funded content or browsed our websites, but we do have a device ID that we can tie back to those customers. So, what we end up building there is a household picture. So, we know within that household the customers who live within it, the devices that they use within it, and the the types of
[31:35] information that they're viewing. We then purchase third-party data from services like Experian to enrich that first-party data with things like average household information, uh sorry, average household income information, and other demographic information to to
[31:52] flesh out flesh that out. So, we end up with this very rich uh repository of customer information that we can draw down from. From there, we expose that in a silver layer. We anonymize that information, and we make it available to train data
[32:07] science models on. So, that's where we can start experimenting with different segments, and we could look at users from different backgrounds, from different income levels, from uh different interests. So, we can have female NFL fans ages 21 to 25, and use,
[32:24] you know, use those segments to to target for for marketing and ad sales campaigns. Those segments get pushed into a gold layer, which is where our curated segments will live. And those uh segments can also be distributed both internally and to external partners
[32:41] using clean rooms. From there, we use a Data Data Bricks premier partner. I'm not going to name them because we're kind of early on this journey and uh don't want to give props until we see how it goes, but um that Data Bricks premier partner will be used by our marketing and ad sales teams
[32:58] to take those segments, link those back to the customer IDs, and then actually target those customers through ad sales uh relationships, for example, through FreeWheel, our ads partner. Uh we could be using tools like Sailthru for email
[33:13] marketing uh or any other really uh marketing or ad sales platform that we want to integrate with downstream. And then we keep a a real-time log of those campaigns in in back into Data Bricks.
[33:31] So, this allows us to change that old paradigm of being the peddlers of dashboards and reports. That's That's no longer the business that we're in. In this new world, we can add business value in new ways, such as ensuring that our investments in original or acquired content are optimized with data and
[33:46] insights. We want to know more about our audiences and connect them with our content in more meaningful ways. We want to provide more authentic personalized experiences that drive greater engagement and less churn. And we want to grow new revenue streams
[34:03] through smarter ad decisioning, marketing campaigns, B2B monetization, or upsell opportunities. And we want to drive business strategy based on more sophisticated forecasting, segmentation, and modeling.
[34:22] So, bringing this all together, the uh the big block V8 engine in our car is is Data Bricks, but what that allows us to do if we drive it right, if we build the right uh if we fuel it correctly, and we build the right components around it is that uh we can collect data, ingest it, and map it in a more cohesive,
[34:38] consistent, and reliable fashion. We can offer a trusted, accessible, and intelligent data platform on which we can build enterprise-wide 360° rich silver data sets for analysis and experimentation. And then we can offer Oops, excuse me.
[34:56] And we can offer strategic high-value data products to our enterprise and external partners and drive significantly more business value from content and audience insights to new customer experiences and business growth. But most importantly to me, it
[35:11] means that we can turn up in a scrapier, more innovative way as an AI native startup and stamp out some of those inefficiencies that drive me nuts. Um so that is all I've got for you today. We have 5 minutes left and I'm happy to take any questions if you have any.
[35:29] Thank you very much. Thanks for coming along. Thank you.
Learn more about the Databricks Data and AI platform.
The information provided herein is for general informational purposes only and may not reflect the most current product capabilities or configurations.