Skip to main content

Emotional Congruence at Scale: Using Databricks AI to Match Ads and Content

Summary

  • NBCUniversal built an Emotional Congruence engine on Databricks that quantifies alignment between ads and surrounding content across three dimensions—themes, tones, and values—to keep viewers in the moment and improve advertiser outcomes.
  • A two-track pipeline extracts multimodal signals from both content and ad creatives, validates tags against human-labeled ground truth with F1 metrics, and deploys via MLflow 3.0 with Unity Catalog for config-driven production promotion.
  • Testing confirmed that emotional alignment between ads and content drives measurable lift in viewer intent and conversion in double digits, making emotional context a quantifiable and actionable signal in NBCUniversal's contextual advertising stack.

Emotional Congruence at Scale: Using Databricks AI to Match Ads and Content

Watch: Emotional Congruence at Scale: Using Databricks AI to Match Ads and Content
Contextual advertising has long focused on placing the right ad in front of the right person at the right time, but a critical element remains unsolved: emotional context. At NBCUniversal, we built an Emotional Congruence engine to quantify whether alignment between emotional signals in ads and content drives viewer intent and conversion, and how to scale it using Databricks foundation models.
Learn how to build a taxonomy-driven tagging system using themes, tones, and values to capture emotional profiles of content. Discover a two-track pipeline that extracts multimodal signals from both content and creative ads, validates against human-labeled ground truth with F1 metrics, and uses MLflow 3.0 with Unity Catalog for config-driven production deployment. Testing shows emotional alignment drives measurable lift in double digits, now embedded in NBCUniversal's broader contextual advertising.
🤝

Chapters

FAQs

What is emotional congruence in advertising?

Emotional congruence refers to the alignment between the emotional tone of an advertisement and the emotional context of the content surrounding it. NBCUniversal found that when ads and content share emotional signals—captured across themes, tones, and values—viewer intent and conversion improve by double digits, making emotional context a measurable advertising variable.

How does NBCUniversal's Emotional Congruence pipeline work?

The pipeline runs two parallel tracks that extract multimodal signals from both content and ad creatives, then matches them against a shared emotional taxonomy built from themes, tones, and values. Results are validated against human-labeled ground truth using F1 metrics, and the system is deployed via MLflow 3.0 with Unity Catalog for config-driven promotion to production without code changes.

How does Databricks AI power the emotional tagging system at scale?

Foundation models served on Databricks generate emotional tags by processing multimodal signals from both content and ads, applying the taxonomy across NBCUniversal's catalog at scale. MLflow tracks every evaluation run and prompt variant through a four-step evaluation framework, enabling the team to select the best configuration for production before deploying.

How does MLflow lineage connect evaluation to production in this system?

MLflow 3.0 tracks the full lineage from evaluation experiments through to production deployment, recording which model version and prompt configuration produced each set of emotional tags. Unity Catalog stores and versions configurations so the team can promote changes without modifying code and trace every production output back to its evaluation data and ground truth labels.

Full transcript

[00:08] Uh, good afternoon everyone. Um, thank you for being here. Uh, I know it's a late slot. Um, and I'm probably the only thing keeping you from your last happy hour or dinner um, and drinks and so I'll try to keep this as as brief as I can. Um,
[00:24] I I realized like literally just now as I was up here that I forgot to put a title slide. So, we're just going to pretend when I'm starting to get into the preamble that it says emotional congruence at scale. Um, before that though, just a reminder to complete your survey um after the course or after the
[00:42] uh the session. Um, okay. So, um, when's the last time that you guys were were watching a show or or a movie and you were fully immersed um and then an ad break hit and it broke that spell? Um, you could
[00:59] imagine watching like a tense thriller and then a commercial for a beach vacation comes up and you're like totally pulled out of whatever headsp space you were trying to escape into. Um, that disconnect's not annoying just for you. It's also wasted money for an
[01:15] advertiser who just spent money at best on an ad that you won't remember and at worst on an ad that that frustrated you. Um, and so at NBC Universal, we're using data bricks AI to to try to fix this. um to align ads uh with the content
[01:30] around them to keep viewers in the moment and to help make every uh ad placement more effective for advertisers. My name is Alex and I'm a director of data science at NBCU in the ad sales division where I focus primarily on audience science and
[01:48] applied ML and support of the various products that we offer our advertisers. Um this talk uh the next 30 or 40 minutes is going to be about a system that we've built using data bricks AI. So MLflow Unity catalog foundation models served on data bricks uh to solve
[02:04] a problem that I think is is just now being addressed in in the TV ad industry and that's um emotional context matching between an ad creative and a piece of content. Um, I'm here primarily to tell you how we're doing this. Um, to give you
[02:22] insight into the tools that we're using on data bricks and maybe you can apply them to your own use cases. Um, but I can't really do that unless I tell you the why. Um, and so before I get into that, let me explain to you what I even mean by emotional congruence, uh, emotional context and why that matters
[02:39] to NBC. Um, historically effective advertising has been about two things. um getting an ad in front of the right person at the right time. Uh the industry spent a lot of money on on doing this effectively. Um even in a cookieless world, there's
[02:55] identity graphs uh that map device IDs to a household. There's programmatic delivery. There's dynamic ad insertion, frequency caps. There's an entire ecosystem that's really mature and really good at solving um the problem of
[03:11] getting an ad in front of the right person at the right time. Um, but more recently, I think there's a third component. Um, and that's context. We've heard a lot about it. Um, and and that's more about not only the right ad in front of the or sorry, not only the the ad in front of the right person at the
[03:26] right time, but the right ad in the right moment in front of the right person at the right time. Um, and that I think remains largely unsolved. Um, and it's where I think that the biggest opportunity is. And I think it's it's a problem legitimately that AI can help us uh solve. Um, context here isn't just
[03:43] about like brand safety. Um, it's a fundament more fundamental question of does this ad belong here? Um, does the content set the viewer up to receive the message in a way that makes the ad hit harder? Um, historically, genre and
[03:58] keyword matching might have filled this void and it kind of gets you into that neighborhood, but they don't really capture how a show makes you feel. And that's really what shapes your mindset um when when you're at hits. Um, emotional congruence. What is that?
[04:14] Uh, at NBC, we define that as the degree to which an ad's emotional signals align with the content surrounding it. Um, I'll explain what those signals are in a little bit. Um, but know that that we're not doing this just to do it. This isn't an exercise in futility. Um, there's a
[04:31] lot of research that suggests that, you know, when a content primes a viewer's mindset or or or or emotional state to match an ads that the ad performs um performs better. And that makes sense. If you go back to the thriller and beach commercial example, um, the less that
[04:47] you have to shift your mindset, uh, the better, the more memorable that ad's going to be, the better recall you're going to have around it. Um, and so we're trying to operationalize that at scale. Um, and if we can do that, we'll have unlocked a new layer of targeting around emotion. Um, I mentioned genre keyword
[05:05] matching kind of historically serving this this purpose. Um, you know, if a show's classified as a comedy, the assumption is that you're in a light-hearted and receptive mood. If it's news, then you're attentive and engaged. Um, it's a helpful curistic, I think, and it's it's worked, but it also
[05:21] is blunt and it lacks nuance and it's coarse. Um, two shows can be classified as dramas and feel completely different. Um, Breaking Bad and and This Is Up This Is Us, for example, or both dramas, but I'd argue if you're familiar with the shows, you're probably in a really different state emotionally or mindset
[05:39] when you're watching them. And the ads that run in each should should reflect that genre doesn't necessarily capture that that nuance. Um, so what does uh what signals emotionally am I am I talking about? Um, at NBC, we've settled on three um, three
[05:57] main dimensions that that we think together capture a fuller emotional profile of a show. Um, and those are themes, tones, and values. Themes answer what a story is about. Is it uplifting? Is it dark? Is it suspenseful? Um, uh,
[06:13] sorry. Themes answer what what a story is about. Is it romance? Is it coming of age? Is it personal struggle? Tones answer how you feel when you watch it. Is it uplifting? Is it dark? Is it suspenseful? Values answer more broadly what it says about the world and by extension um the audience that watches
[06:30] it. Is it indulgent? Think I don't know Real Housewives and some of the unscripted Bravo content. Um is it reflective of family and community? Um of ambition, success, of safety. Um and and we didn't like just craft
[06:46] this taxonomy out of thin air. It was a thoughtful exercise where we started with a really broad library of hundreds of keywords. We embedded them using a burst style model and then we let unsupervised clustering tell us what the natural groupings were. We used pretty
[07:02] standard um metrics of of uh cluster evaluation such as silhouette scores to help us, you know, help guide us and tell us how many clusters there might be. uh and the results of taxonomy that's grounded in rich semantic meaning and not just editorial opinion. Um we
[07:20] validated the clusters manually and gave them definitions so that we could apply them consistently throughout throughout our tagging pipeline. And so with that said, let's get into what this pipeline looks like, how we tag this, and specifically where we're using data bricks.
[07:37] Um the pipeline consists primarily of two tracks. um one for content, one for ads, and it runs symmetrically on both sides of the match. So content goes through signal extraction, and I'll I'll talk through what signals we're actually using in a bit. Um synopsis
[07:53] summarization, and then metadata tagging with our emotional cigarette signals using um foundational models served on data bricks. Um that produces a dense embedding uh of each show's emotional profile. um ads work. The ad tagging
[08:09] works in in a very similar way. The difference here is that we're actually using the actual ad the the creative video. Um and ad spots only 30 to 60 seconds. So it's it's fairly cheap to do. Um so the difference here is that instead of synopsis sources, we're using
[08:25] frame. We know we're extracting frames from the ad and audio transcription. We're combining that into a synopsis and tagging that. Um once we have the the embeddings and tags, we'll just match those across both sides um to surface the most congruent uh show placements.
[08:41] We'll take those titles or those ids, we'll package them up and then we'll send them downstream for activation. All this is governed in uh Unity catalog and and and ML flow. Um terms of the signals that we're using here um for content we use a multisource
[08:59] uh approach that includes first and third party um synopsy sources. Um those all get fed into uh a prompt and a foundation model that's been tested and chosen um to produce a structured synop synopsis that's optimized for tagging
[09:16] not just a description of the show. Um for ad again it's multimodal. Um, so we'll track frames at one to two seconds. So maybe 15 to uh to 30 per per per ad, transcribe the audio, package that up um into a synopsis, and then
[09:34] that goes through again a model and prompt that's been chosen specifically for the ad tagging use case. Um, and then we just we we match them. Um, all of this runs as a data bricks job uh in a pipeline. And it's a
[09:51] multitask job that runs uh linearly. So each each task is linearly dependent on the prior. Um so first you have synopsis generation and then content tagging and then uh and then creative matching. Um this job runs on a standard cadence. Um
[10:07] so that we pick up any new shows that have come out in say the past week or month on Peacock or any other uh endpoints uh that NBC has. Uh we also have a set of evaluation runs. The evaluation runs help choose which prompts and which models we're using for
[10:23] tagging. And I'll walk through that framework in in the next slide. Um those run kind of offline though because we're not needing to evaluate the prompt or the model every month. Uh we run those roughly quarterly. Um and then we'll
[10:39] just promote if if a new model or prompt is chosen, we'll just promote that uh to our production job. Um but MLflow tracks every prompt configuration uh every run, every evaluation run, every model version. Um the winning models and prompts get promoted to Unity Catalog.
[10:57] Um and then we update a configuration table in Unity Catalog. So if we ever need to change the model or change the prompt, we just change the the table in Unity Catalog. Um and nothing else changes. We don't have to change the job. There's no code changes necessary. It's just all all pulled from Unity
[11:13] Catalog. and it's made it really easy for us to do this because everything's config driven and fully observable and and reproducible. Um the evaluation itself framework um is
[11:28] largely a four-step process um but it runs identically for every stage of the match. So we have evaluation of prompts and models for synopsis generation. We have one for content tagging. We have one for ad matching. Um it starts though
[11:44] with uh selecting prompt variants. So we might have three to four prompt variants um within a stage say for synopsis generation. Um and the differences between those prompts aren't just like small wording changes. They're fairly um
[12:01] there's a strategy behind them. And so um one prompt might just very simply hey give me the synopsis for this show. Another might employ more chain of reasoning like first you know what emotional state might the viewer be in and then give me tags for this show. Um
[12:17] another might be uh from an advertiser frame of mind. We'll ask the the model to adopt an advertiser um um you know frame of frame of mind. So we have different prompts for uh within each
[12:33] within each bucket. Um and the prompts are designed to uh extract the the structure that we that we want from from the signals um that we have in in in in uh in each um in each bucket. Um we'll
[12:49] then test each of those prompt variants against a set of foundation models. So we might have three to four prompt variants. We'll test those then against three to four uh models. Typically they're the foundation models, the big three. It's GPT, Claude, Gemini. Um we'll test them at pretty deterministic
[13:06] temperatures so that results are reproducible um across runs. Um, but the result ultimately is a is a winning combination of a prompt and a model that's specifically designed to do what we want it to do, whether that's
[13:21] generate a synopsis or tag a piece of content or tag an ad. Um, what defines the winning is um our it's assessed against a human uh a human labeled uh ground truth. So, we're pretty fortunate
[13:37] at NBC to have um some resources that do content tagging. And so, um there's other contextual efforts beyond just this. And so, um we have a fairly rich set of titles that have been human tagged and verified. And so, um we'll
[13:55] assess the models and prompts against that human tagged ground truth. Um fairly standard um metric. we use F1 score um a 75% threshold which I think is relatively um normal in in the in the
[14:14] in the literature um uh and so every run every evaluation run is logged to MLflow we log the prompt artifact uh model parameters the evaluation metrics um so we'll have a complete lineage from uh from you know
[14:31] from eval to registered model to um to production output But again, if we want to update or refresh the production model, whether that's because a new frontier model version comes out or uh because our taxonomyy's evolved or we just want to try a new prompt, all we
[14:46] have to do is run the eval and then update the config table that's in Unity catalog and um that reflects in the in the production job. Um no, no notebook edits and and and no redeployment.
[15:06] Um, and so the result here then is uh basically uh a a a every ad creative uh or every title is tagged and every ad creative we have is is also tagged and we can match those match those at scale. So for example, take Yellowstone. I think it's a fairly unambiguous show.
[15:22] Everybody here is probably really familiar with it. Even if you hadn't seen it, I think you can attest to what it what it's about. Um, so, you know, I think it makes sense that it would be tagged with, uh, gritty, tradition, um, suspenseful, right? Um, and you'd
[15:39] imagine that the top creative matches that might want to run in that show would be maybe a truck brand, uh, maybe it's uh, an alcohol brand, maybe it's a whiskey, u maybe it's a workware brand. All of these have very similar
[15:55] ethoses, brand ethoses that you would again imagine would would run running in Yellowstone. So, it makes sense intuitively. Um, but I think the more important uh piece here is the the scale at which we're able to to do this. Um,
[16:11] we've tagged over 10,000 titles with relative ease. Um, and there's no code changes that have have needed to be made um at all. We've we've set we've set our notebooks up to run and we haven't had to change those in in months
[16:27] because everything is deployed um via Unity catalog. Um so, you know, that's that's where we are now, but I'm really excited about where we're going to go uh in the
[16:43] future. I think um right now a lot of what we're doing is deploying foundation models via um UDFs um uh at NBCU we have a fairly strict governance team and so a lot of the
[16:59] stuff that's available in uh public preview we can't quite use yet. So once it becomes generally available uh we will start to shift this architecture over to use information extraction which is an agent brick um or AI extract AI
[17:14] classify. These are all um these are all native sort of SQL um functions AI functions that are good at batch process processing. Um, if you guys are familiar with them, um, I hear they're pretty powerful and I can't use them yet, but
[17:30] um, uh, so the goal is to to shift this process over to to to to run in those. Um, and ideally, I think that's going to really make it much easier than it is now. Not that it's hard, but um, right
[17:48] now we're batch process. Right now, we're running these these these jobs in they're UDFs. And so every title has to there's a foundation call for every title and every every every instance that we that we that we need to to to to
[18:04] tag there's a call to the model. And so if we shift this over to batch processing, it it it it you know, we don't need to to worry about that anymore. So um it's going to significantly reduce our our cost. Um, that's it. Uh, I I told you
[18:22] guys I would try to keep it short and sweet, and I think I did.

Learn more about the Databricks Data and AI platform.

The information provided herein is for general informational purposes only and may not reflect the most current product capabilities or configurations.