Near Real-Time Media Analytics: Spark Declarative Pipelines and Lakebase
Summary
- Conde Nast, publisher of Vogue, The New Yorker, and Wired across 37 markets and 26 languages, replaced a fragmented analytics stack with 15–30 minute latency by rebuilding on Databricks Spark Declarative Pipelines and Lakebase, achieving 30-second data freshness with P50 latency of 140 milliseconds.
- Spark Declarative Pipelines enforce consistent business logic across 10 petabytes of first-party tracking data while eliminating manual DAG writing and orchestration complexity, and Unity Catalog enforces governance, PII masking, and role-based access for 275 concurrent users across all brands and markets.
- The new architecture enables editors to make real-time decisions on content performance, subscription conversions, affiliate revenue, and ad monetization from a single unified platform, replacing the multi-system approach that had delayed editorial response to trending stories.
Near Real-Time Media Analytics: Spark Declarative Pipelines and Lakebase

Editorial decisions in global media need instant data. Conde Nast, publishing 115+ year old brands like Vogue and The New Yorker across 37 markets and 26 languages, faced a critical gap: their analytics stack delivered insights 15 to 30 minutes late, costing revenue on trending stories and delaying subscription conversions. This talk reveals how Conde Nast rebuilt media analytics on Databricks to achieve sub-second insight into content performance.
You'll learn how Conde Nast unified 10 petabytes of first-party tracking data with Spark Declarative Pipelines that enforce consistent business logic at scale while eliminating orchestration complexity. See how Lakebase provides low-latency operational serving, how Unity Catalog enforces governance, PII masking, and role-based access across 275 concurrent users, and how declarative patterns eliminate manual DAG writing and reduce operational overhead. Discover how the new architecture delivers P50 latency of 140 milliseconds and P95 of 380 milliseconds with 30-second data freshness, enabling editors to make real-time decisions on content performance, subscription impact, affiliate conversions, and ad monetization all from a single unified platform.
🤝
Chapters
00:00Introduction to Conde Nast and Data Mandate02:37Today's Agenda: Editorial Analytics Solution03:23Editorial Needs: Freshness, Dimensions, Outcomes05:16Solution Landscape: Unifying Data and Embedding07:08First-Party Data Strategy and Scale09:20Old Architecture: Fragmented, High-Latency Stack10:58Pain Points: Operational Complexity and Governance13:56New Solution: Spark Declarative Pipelines and Lakebase14:45Architecture: Streaming to Medallion to Serving Layer17:26Benefits: Orchestration, Quality, Governance Unified19:22Unity Catalog: Role-Based Access and Consented Data22:05Results: P50 140ms, P95 380ms, 30s Lag, 300 Users22:54Lessons Learned: Cost, Culture, Testing, Governance24:28Future: Genie, Audio, Video, Anomaly Alerts
FAQs
Why did Conde Nast need near-real-time editorial analytics?
Editorial teams need to act immediately on trending stories to maximize subscription conversions, affiliate revenue, and ad monetization. Conde Nast's previous analytics stack delivered insights 15 to 30 minutes late, meaning editors made decisions on stale data and missed revenue opportunities when content began trending.
What role do Spark Declarative Pipelines play in Conde Nast's architecture?
Spark Declarative Pipelines enforce consistent business logic across all 10 petabytes of Conde Nast's first-party tracking data, eliminating the need to manually write and maintain DAGs across 37 markets and 26 languages. The declarative approach also unifies data quality checks and orchestration in a single framework, reducing operational overhead.
What performance benchmarks did the new Conde Nast platform achieve?
The rebuilt platform delivers P50 latency of 140 milliseconds, P95 of 380 milliseconds, and 30-second data freshness, supporting 275 concurrent users across all brands and markets. This compares to the previous stack's 15 to 30 minute lag.
How does Unity Catalog support Conde Nast's governance and privacy requirements?
Unity Catalog enforces role-based access control and PII masking across all brand and market data, ensuring that editors and analysts see only the data they are authorized to access. This is especially important given Conde Nast's first-party data strategy and the need to manage consented user data across 37 markets with varying privacy regulations.
Full transcript
[00:09] Hello. Uh good morning everybody. I think it's uh very excited to see a lot of people in the third day of the conference. So hopefully you're not too tired from the party yesterday. Um yeah, so without further ado, we'll get started. So I'm Prasa. I lead uh business
[00:24] intelligence and enterprise analytics at Condai and I have Arun also with me. Hey uh I'm Arun. I'm director of data engineering in condai. So together we're going to present uh about what we did as
[00:40] a near real time and why it was a near real time and I'll let presenter to get started with the problem that we faced in the business side of it. Sure. Yeah. So together uh we belong to the data group. What it entitles is what is the mandate for us is to democratize
[00:57] u data at continast for every employee irrespective of the brand or the market they belong to. uh because in the last two years two last four to five years since the since Roger came as a CEO the biggest mandate for us is to consolidate the operations across the globe for
[01:12] Continast um so so that's the mission that we work towards and we both uh are representing a lot of people uh who work behind us uh along with us uh uh to be here. Yeah. With that note, uh let's quickly start with the brand itself
[01:28] because a lot of people are familiar with um uh with the individual brands that Continent uh encompasses within itself but Continast itself may not be familiar for a lot of you. So we um uh own a lot of the forefront brands like uh Wired,
[01:47] Vogue, New Yorker uh who are kind of trends setters in their own way uh in the in the category that they belong to and uh we are over 100 years old. Um I think uh um so that tells you like the resilience for the company like uh how
[02:03] it has renovated re reformed itself uh over the over this period and we operate uh in around 37 markets and uh we deliver our content in uh 26 different languages. So what it uh should convey
[02:18] to you is the breadth of the uh audience that we serve towards and um the challenge that comes with it and the opportunity that gives for the data folks like us right so so yeah quickly about today's agenda um we we
[02:37] want to talk precisely about a solution that we picked up for our editorial group which is at the heart of what Kai is about um and how we went ahead with a solution in a in a
[02:52] in a not uh very orthodox way I would say and where it led us and how um we went with data bricks and uh improvised the solutioning and what are the key lessons that we learned with it and how it has set the platform ahead ahead for
[03:08] us right so that's the agenda for us today now I want to just uh set the business context uh and then leave it to Arun to talk about the architecture and the solutioning itself. Um so when we spoke
[03:23] with the various editorial groups across the globe um and brands as well within each market there were a lot of needs as you know um any company that produces original content there is a lot of independence for the editorial units. Um
[03:39] so they think in a very um um um independent way and they give us um varied needs. So we generally bucketed into three different categories like um the key important factor for them was freshness of the data. They can't wait
[03:55] typically our data solutions were almost like a day delayed. Uh so the key concern for them is like they're not able to activate in a lot of their distribution uh for with the delay. So they wanted it um near real time if not real time, right? because there's cost associated with being real time or near
[04:11] real time. And then they just don't want um um one metric uh they want to cut it by different slices of dimensions for example different traffic sources or author units that are resonating with the audience right now or uh be it the
[04:28] device where it is being activated properly um on paid versus organic. So things like that and the last one that is very important I feel is in the last two to three years the way we have taken the editorial group in the data journey
[04:43] they are very matured right now they're not just looking for a topline conversion of uh page views or unique visitors they definitely want to see the downstream metrics uh that impact uh downstream metrics to really find out the impact of a content right so how much subscriptions I'm converting uh
[04:58] what is the a percentage of ad units that I'm able to monetize in my engagement right and how my content is doing on the commerce conversions right so we try to um that was a big need from the stakeholders so now when we take that into a solution landscape what what
[05:16] we decided is like okay so it's very important for us to get the traffic and all associated impact data sets into one pane so that it makes our life easier for the solutioning and also um um we want to ensure like the solution
[05:33] um meets the customer where they are. So we know like they frequent the CMS tools every day. So we want to make sure like our solution gets a good embed within that uh CMS solutioning so that every day they go into a CMS to curate the content and publish it they can also see
[05:49] our solution giving the rightful insights to make a make the right decision. I'm going to quickly uh double click on that a little bit. I'm going to go from um right to left. So the embed I spoke about the internal CMS uh solution that
[06:04] we use is uh copilot. So our solutioning we want to embed it within that so that uh editors have an easy way to access it. Um and then we talked about unifying all the content relevant to traffic and other uh impactful metrics in one place.
[06:21] So, so yeah, this is a good opportunity for us to get all the siloed business uh rules uh in different platforms to get it into one place, right? That's data bricks for us. And then the key thing um is about relying on the first party data
[06:37] sets about two years before the company made a very cognitive decisioning to to reduce our dependency on the third party tracking like Omnature or Google Analytics and go to first party tracking. So it's a very important pivot for us uh because we realize like more
[06:53] and more third party uh um tracking is becoming less reliable. So we want to ensure like u our data capture is uh pristine uh using our first party tracking tool and it gives us the flexibility to track whatever we want
[07:08] right um yeah so that's so this tool that we are building is purely on the first party uh data sets uh so so yeah there's no dependency on the third party one quickly on the sale of scale of the data lake itself um of course it holds uh 10
[07:25] pabytes of data um on any given day we easily get over a terabyte of data. Where it's stemming out is on an average every day we deliver about 30 to 40 million impressions to users and in a month we have over 100 million visitors
[07:41] uh visiting our sites and uh on any given day we easily deliver somewhere between 8 to 12 million uh page views and in our tracking we don't just track the page views all the scrolls all the conversions uh we also have multiple content units within a within a page
[07:57] like audio units video units we also track a lot of events related to it so that's where the volume is coming from and one other uh unique demand that came from the editorial team is like in any given month there are thousand plus users from the editorial unit but that's not the mandate from them on any key
[08:12] 10pole events they might have more than 300 concurrent users so we want to make sure like our solution should be able to support without any u um without any inconsistency for 300 two 200 to 300 parallel users. So that was a big
[08:28] demand. Now this might look very simple need but then our budgets are very limited. So we have to deliver all these things with a with a small budget or no budget right. So now that's where Aron comes in to see like uh to say how it happened. So as you saw like uh the data volume
[08:47] that we are looking at for this particular data product is like 10 pabit volume and also the concurrency of the users and the time of uh delay for this data what we are looking at is also should be within your budget right so that's the exact way that we are looking
[09:02] at so when we started off so what that we did so our collector uh is we use snowflow SDK for collecting our events our realtime needs and all the user behaviors and the clicks and click views and all the conversions. So we use that
[09:20] uh injection flow from that collector and then we ingested into our uh time stream database right from which we were able to aggregate and then during that process of time there were delayed datas. So you cannot you can just imagine like data coming from different
[09:36] regions, different time zones, different markets, different uh languages. So it has got also the failures wherein you'll have your records. So what we actually thought of is like to have a first real time near real time. I wouldn't say like
[09:52] it's a real time on the batches which will be taken in from our first party data. Okay. And then flowing it through the uh ritzy and then feeding it into our uh UI. Okay. So basically beacon UI is what we call it. uh but it is
[10:08] actually a JSUI ReactJS UI which actually gives you the view of all these numbers right so the initial half which runs um without its synchronous with the backend data so what is the backend historical data so we also had a request
[10:25] from the business team to say like they need two years worth of data to analyze along with the existing real-time data that we getting so what happened here for us so we really need to sync up with all these three things. So what are those three things? The historic two years worth of data and the metric that
[10:42] we get on a near real time and also there is a co-pilot uh injection which is happening. Okay. So now the heavy load goes to your u aggregation at the API layer right so all these metric that you get and then you try to ingest all these data and then your
[10:58] final aggregation should be with all these three layers. So it was really a tedious one. Okay. And there are a lot of pain points. So as I said what are all the pain points? I would say in two different aspects. So one is the architectural aspect other one is the
[11:14] operational aspect. So taking in consideration uh architectural aspect. So it is an a synchronous fragmented way as I said like there is one which is going to be near real time the other which is going to be your uh batch mode right. So then the API will taking the
[11:32] consolidation taxing. Okay. So the API is doing a heavy lift. Okay. So that layer have to get all the the fragmented uh metrics and it has to consolidate in your UI side. Okay. It was heavy lift in the API side. So that being said the
[11:49] scaling is going to be a problem. Okay. You have click which is the UI which going to be a batch base and also you have uh time stream DB which is again going to be your batch and when you have a concurrent uh users of 275 then your report fails it is not going to be very
[12:05] happy situation when someone is not able to see a metric during events like let's say medgala or like your prime day sale or when someone is writing other uh events. So it's it's going to be a very difficult scenario for them not able to see what that article is performing like
[12:21] right so that being said we had limited cost okay because as we know like this is the first time implementation or the implementation itself has a budgeting we really don't say like okay we spend and then we really say what is the throughput so we had a limited budget
[12:36] that is for sure okay so that's another uh architectural overhead as well so in the operational overhead what happened so we were not able to uh uh have the freshness of the data the way we looked at because you have your streaming data coming on the top and then your batch
[12:52] data coming on the bottom and then you again have another co-pilot which is again going to be a injection so error handling was difficult okay whenever there is a issue there is a failure you are backtracking your data trying to identify where the issue was and where the flow was that's a difficult process
[13:07] for us to and it was even tedious for us to do right and then the debugging cost as I said like the 0.1 error handling and debugging both are adding up to our time and also even to the restart of everything. So it's going to be difficult uh for us to do right and
[13:24] finally the governance gap. Okay. So the governance of your data to handle your PII to handle your um data for different GDPR and other other governance on top of it. That's another thing which was having an heavy lift because you have
[13:40] different places where you need to apply different uh rules. So your governance is going to be overhead in a detached architecture is what I would say. Okay. So, so that being said, so all the difficulties, all things that we saw in this existing one, what is that we moved
[13:56] into? What is that we looked up for, right? Um, so we wanted to get all these uh different tools into one single layer, okay? Or one single umbrella. That's where we wanted to get into the
[14:12] Spark declarative uh pipelines. And then we wanted Unity catalog for all the governance and other um things to handle. And then we also looked at lakebased as an option wherein for the UI to have a consolidated aggregated uh API that can be done and then stored in
[14:29] your meta layer and then it can be uh viewed in your um UI. Okay. So now coming back. So what is that we wanted to do right as uh we initially did. So we really have to ask our source okay as
[14:45] I said like uh our spruce or the snowflow collector that we had it has to have a streaming loader. So instead of using their uh near realtime batch loaders we just asked them for the streaming loader. So that is a uh key
[15:00] thing that they have to enable us with. Okay. And then we had all our historical data and revenue and engagement data which is again another set of thing that we have and then all our uh CMS and metadatas that we had. So were unified into one single uh DT or like SDP that
[15:18] we call uh right now into like bronze, silver and gold. Okay. So what is happening now since we have declarative pipeline? So the cleaning of your data with the use of expectations which is again like that's one of the key part
[15:33] where you have your quality can be done and then your goal layer aggregation everything is going to be on a real time right now I have all these three managed now I am going to have my serve layer which is going to be in leg base and then the realtime UI and the UI call is
[15:50] going to be done in the API and then the co c co-pilot embedded okay so it was really a seamless way to look at how we wanted to operate in this model. So this was good okay in the papers and everything is monitored or governed by
[16:05] your unity catalog. So the lineage axis on what not that you wanted to look at. Okay. So how did this go right? So what we did so we had our collector which is now a streaming loader and then it goes
[16:20] to the STP. We use um spark declarative pipelines and then we place our data and even in that situation we did have the uh what do I say uh lag okay 30 minutes of late arrival of data so for that we
[16:37] used uh a running near realtime uh airflow so that what happens is as you get your real time even the airflow runs and then you will get a full realtime in every 30 minutes cycle okay so what are we looking at currently will actually be a realtime one But it will get to a
[16:53] change of minimal number of records depending on the late arrival of data that is really happening and it is it will happen in many cases for us. Okay. Then our metric is then it moves to the red uh red where it is cached and from the caching layer you
[17:10] have your uh numbers then move to the uh UI the J uh react UI. So this is the new way of what we did right. So what is the advantage where did we end up with? Okay. So first on the declarative side
[17:26] as I said so what are the benefits when we started using the spark declarative pipelines the orchestration layer is completely removed. So a huge uh DAG and their implementation those kind of things were removed. Okay. And then we
[17:44] also had all our uh dependencies with respect to our back fills and uh reduction in terms of operational overhead failures everything was able to done uh able to be done in a one single environment. Okay. So then your
[18:00] expectation right so you have your quality when you have to check your data quality. D offers the expectations testing. So you can do your testing and make sure your QA uh is also written on top of it. And then your back fills again as I said like you you can have either your continuously running um
[18:18] every 30 to five 30 minutes airflow batches which can backfill your lineage and then you can also have your u things avoided as I said like dag writing your code for more than 3,400 lines and uh then the retries which were having a
[18:34] delay. So those are all avoided and then you have a drift uh anything which were having like a changing schemas and all those issues were even addressed in the current one and also what happened is the current pipelines were all uh
[18:50] failures or any delays anything that we have we have messaging in Slack which actually helped us to make sure like okay what is a failure how we will be able to handle it so those failures were all slack messaged okay so that was also handled so This helped us to keep the
[19:05] pipeline running up and running in a near realtime basis. Okay. So coming to the unity catalog how it improved us. Okay. So the lineage as we spoke okay first and then audits and then even the governance. So we have PII to be implemented. Okay. Column masking was
[19:22] enabled at your uh in your realtime DT table. So that is a key thing which we are able to do it. Okay. And we also have something like consented postconented data. Right. So we collect all these user information. Uh that's one of the key challenges we face.
[19:38] Initially we just have to collect and then have our tables aligned accordingly and then have the pipelines designed in a consented and postconented manner. But having Unity catalog so you really have your role level access to say like okay uh these are all the role level details and these users have access to the
[19:54] certain data and these regions of users don't have access to your certain data and you'll be able to make sure like uh the user level access is applied on them right that really helped us okay and in terms of like what benefit that we were
[20:09] able to harvest and what we avoided we weren't using our IM policy tools to go and then say like okay certain accesses have to be done and it was tedious to use IM uh policy tools and then all the lineages that we were tracking in confluence right so that was avoided
[20:25] okay and reconciliation with respect to the cons uh consent flag okay so we are flagging as I said like we are flagging it preconent and postconent and then trying to take it so that is avoided and we using uh unity catalog to help us out with respect to that and then any kind
[20:41] of access with respect to uh tickets one by one so that for the user level access was also avoided using the groups. We were able to add them to the group. That was uh one key thing which we were able to avoid. Took away the access security layer from
[20:57] click and it brought them to the unity catalog. Yeah. Oh yeah that's true. Okay. So now so now having like this so now we started off with the scope. What was the scope for us? We wanted to have more than 275 users in nine different markets
[21:16] have to be concurrently able to access the dashboard seamlessly right and we have more than thousand users as Pasa said previously we have more than thousand users who are currently uh active but in any given point of time it should be 275 users at least and then
[21:32] the P95 for the latency should be 10 seconds okay that was the expectation the delay can be at the 10 seconds manner so that we can do it. Okay. And that being said, so what is the architecture we adopted? So Postgress compatible API for uh for the UI teams
[21:49] and then the autosync uh from the core layers and also replica without the cluster operations like scaling up and down. So then the governance with the lakehouse. So this is what we did. Okay. So which actually helped us and moving
[22:05] in this way. So what is that we were able to achieve, right? So we were able to achieve at least like the P50 was like uh 140 milliseconds, right? And the P95 was like 380 milliseconds. So everything got into milliseconds but we
[22:20] wanted somewhere around like 10 seconds was our goal. Okay. And the freshness the lag that we were able to do is like 30 seconds with more than 300 users still we are able to serve their uh uh UI metrics. So which was really a win-win for us. Okay.
[22:37] And that being said, so we had a certain lessons learned in this uh activity, right? So when we did from where we were in the old architecture, cost was our main thing. But eventually we learned there are better solutions
[22:54] with lower cost. What is that we are able to achieve? The thing is declarative pipeline usage is a cultural shift. Okay, I know like people still want to have their own realtime um pipeline and they want to measure it
[23:09] with respect to that. But I think I'm not just think like we did implement this and we were able to get like it's a culture shift and it is not just an API thing that you want to do and then you get it right and leg based changed how
[23:24] we shift things to the UI. Okay, so we're able to shift or like we were able to move uh in a faster pace and also the lag was even latency was very very low and we were able to add a new metric in a very quicker and faster manner with DT
[23:40] and then leg base getting into place and then testing is very cheaper with the expectations uh functionality in your DT and also the way you are trying to test and keep the quality up to the mark was really good. Okay. and adding up the governance was
[23:57] really good with respect to the unity catalog access audits and again coming to your PII masking. So it was really able we were able to get it to the top notch. Okay. And if some of your uh data analysts want to have the actual data
[24:13] from your uh back end they want to query it even for them we will be able to do a role level access and then you'll be able to do your PII masking as well. Okay. So these are the lessons learned. Okay. And now that we have everything so
[24:28] the business is happy. So they have different plans. I'll leave presenter to take the next one. Thank you. Yeah. So I think um uh we we have shown business like what we can uh deliver um based on our first party data because the premise here is not just to
[24:44] uh give another alternative solution. It's also to change the uh mindset of a business to move from a third party solutioning to a first party solutioning. Um so I think we we have a very solid platform right now. Um and we are looking at u uh some key additions into the platform. The first big thing
[25:00] is obviously conversational analytics. I think uh the editorial community um has repeatedly uh said like uh they would like something to to explore in a more non-technical way. So we feel like uh a genie on top of a kind of a supervisory agent on top of uh what we already built
[25:17] would be a very good add-on. And then um the next big thing is uh now that we have done this predominantly for long form content um we we we have introduced narrated audio into most of our long form content. We're also doing um uh u we are also converting a lot of long
[25:33] long form content into uh video prototypes. So I think with AI in the center more and more content would be interoperated uh across different content types. So it's a right time for us to expand into video and audio analytics as well into the same platform. So that is also something we
[25:49] are uh actively working on. And then uh anomaly alerts is a big thing uh because uh to give you a few examples um when tech team goes through a template release or something if there is an impact on ad density we don't want to know about it a day later we want to everybody to know then and there uh and
[26:05] if a content is getting viral in a social channel we also want to know about it quick. So we have we have devised some uh quick slack alerts uh based on real-time data so that uh editors and the concerned teams uh know about uh any uh anomalies that happen in
[26:21] the traffic and then the last one is uh contextual benchmarking. Every editorial unit gets a target uh and it gets uh given to individual editors uh at the beginning of the year. So right now I I think they do it but then it's only
[26:37] on the topline KPIs. they don't do uh there is no opportunity to do it in a more intelligent way and editors don't have power to say like what they can commit and commit not commit for so I think u a benchmarking across content category would really give them u um a
[26:52] good tool in their hand to say like hey I can commit for this I I don't want to commit for this or have a good conversation with their editorial unit um and have a have have a very productive uh content curation right in the in the in the future so yeah I think We are really
[27:08] excited to do all these things. Uh it would have been nice to have a demo as well but then we are a private company so we have a lot of restrictions into what we can share and what we cannot share. Hopefully we will work with data bricks to put on a white paper or something to really show you like what we have built internally. Uh but yeah uh
[27:26] that's pretty much it. Uh we can open it for Q&A if you guys have any questions. Yeah. Thank you.
Learn more about the Databricks Data and AI platform.
The information provided herein is for general informational purposes only and may not reflect the most current product capabilities or configurations.