Skip to main content

Magnite's Petabyte-Scale Iceberg Migration: Interoperability and Zero-Copy Data Sharing

Summary

  • Magnite migrated their petabyte-scale SpringServe dataset to Apache Iceberg on the Databricks Data and AI platform, achieving a 33% reduction in ingest costs and 50–70% improvement in query times.
  • Zero-copy data sharing with clients was enabled through Unity Catalog, eliminating the need to physically copy data and reducing operational complexity at petabyte scale.
  • The video covers the technical tradeoffs between Iceberg and Delta formats, optimization strategies for large-scale workloads, and the inflection points that signal when Iceberg adoption makes strategic sense.

Magnite's Petabyte-Scale Iceberg Migration: Interoperability and Zero-Copy Data Sharing

Watch: Magnite's Petabyte-Scale Iceberg Migration: Interoperability and Zero-Copy Data Sharing
Magnite processes over 2 trillion ad requests daily, ingesting 1 petabyte of data every 10 days across hybrid cloud and on-premises infrastructure. Managing this scale while providing unified data access through multiple platforms required rethinking their data architecture and format strategy.
Learn how Magnite successfully migrated their petabyte-scale SpringServe dataset to Apache Iceberg using Databricks, achieving 33% reduction in ingest costs, 50-70% query time improvements, and enabling zero-copy data sharing with clients through Unity Catalog. Discover the technical tradeoffs between Iceberg and Delta, optimization strategies that scale, and when Iceberg adoption makes sense for your organization.
🤝

Chapters

FAQs

What results did Magnite achieve after migrating to Apache Iceberg?

Magnite achieved a 33% reduction in ingest costs and 50–70% improvement in query times after migrating their SpringServe dataset to Apache Iceberg on the Databricks Data and AI platform. They also enabled zero-copy data sharing with clients through Unity Catalog, eliminating the overhead of maintaining separate data copies.

How much data does Magnite process and why did scale drive their format decision?

Magnite processes over 2 trillion ad requests daily, ingesting 1 petabyte of data every 10 days across hybrid cloud and on-premises infrastructure. The scale and multi-platform access requirements made interoperability a top priority, which drove the evaluation of Iceberg as an open table format.

What is zero-copy data sharing and how did Magnite implement it?

Zero-copy data sharing allows clients to query data directly without receiving a physical copy, keeping storage costs low and ensuring clients always see current data. Magnite implemented this through Unity Catalog on the Databricks Data and AI platform, which governs access to their shared Iceberg tables.

When does it make sense to adopt Apache Iceberg over Delta Lake?

According to the speakers in this video, the decision depends on interoperability requirements, existing ecosystem integrations, and whether clients or partners need multi-engine access to the same data. They emphasize that there is no one-size-fits-all answer and describe specific inflection points that indicate when Iceberg adoption becomes the right strategic move.

Full transcript

[00:08] Hi. Hi everyone. Thank you for coming to our talk. I'm going to let you start with intros. Sure. Hello everyone. My name is Brian Licester. I go by Lico. Please call me Lico. Don't call me Brian. I won't answer to it. Um that's why you won't ever hear me say that again. Um I am a senior specialist solution architect at Databricks. But prior to
[00:25] Databricks I was at Tabular working on Apache Iceberg for a number of years. Prior to that at AWS prior to that at Confluent prior to that at Red Hat. I'm an open source nerd. I love open source stuff. Um but I've had the privilege to work with and probably hundreds of Databricks customers on their Iceberg journeys. And
[00:41] Magnite is my recent favorite. We got to have dinner last night and maybe I drank too much. That's not the point. The point is I'm here and we're excited to share with you what we've seen. I'm going to hopefully share some I'll call them stories from the field like what other customers have seen because I think Magnite's journey is a
[00:56] little unique compared to what most of y'all will experience. So I kind of want to provide a counter balance to this because there's not a one size fits all journey, right? Everyone's on a different path. So but good to meet y'all. Thanks for coming. Thank you. And I'm Keyvon Raphael. So I work at Magnite. I lead our data
[01:12] engineering effort there. I'm going to talk about our journey as Lico said, our journey through to Iceberg. Just to give you guys a little preview and road map of what I'm going to talk through. I'm going to give you an overview of Magnite, our company,
[01:27] and what we do and what we process. Then I'll talk a little bit about our data landscape. So what systems we have, how those systems and how people interact with our data, and those two together form our requirements. And then from there what led us to choose
[01:44] Iceberg, and the research we did and what we learned through that process. And then I'll talk a little bit about the specific data set in question. I think I'm guessing that was part of the draw. We have a data set that is a petabyte in size, which is pretty cool. So, I'll
[01:59] talk a little bit about that data set and its results from moving over to Iceberg. Along the way, Liko alluded to this, some of the problems that we face and the technical challenges and considerations that we have are because
[02:14] of our scale. I'll I'll touch on that on the next slide. I will try to point out as we go what sorts of things arise from that that are maybe a little unique to us because of our scale and what sort of things, in my opinion, I think are core fundamental things that everyone
[02:31] should care about regardless of scale. So, I'll try and kind of point that out as we go. And from there, let's get into it. Next slide, please. All right. So, we'll get into Magnite. Great. So, Magnite, who are we?
[02:47] We have a video advertising platform, which includes a video ad server. So, if you're not familiar with ad tech, that's an exchange or you can think of an auction platform where we run these auction transactions where people are are bidding on and serving ads.
[03:03] We also have a large display SSP, that's a supply-side platform. That's a platform that helps publishers who have ad inventory sell their inventory. Some of the industry or business challenges that we face that we work
[03:20] with, we handle video, display, and audio format ads. We service both sell-side clients, as I mentioned, and also buy-side clients. And we reach about 99% of US streaming
[03:36] supply, so like video streaming supply. This all means we generate a lot of data, a lot of traffic, a lot of data. We need to process that, make it useful to power our platforms. So, a little bit about us, we operate a hybrid architecture. We have a very
[03:52] large footprint both in the cloud and on prem. To give you an idea of scale, this year alone 2026 that we're in, we will spend a total of $60 million just in CapEx. So, that's on prem hardware.
[04:08] Uh over the past 3 years we've spent $130 million on CapEx for that on prem hardware. And to complement that in the cloud, can't share specific number, but to give you an idea, our cloud scale in terms of uh compute resources is
[04:23] roughly comparable in size to our on prem footprint. So, we're running two fairly large footprints. We process over 2 trillion with a T ad requests each day. Uh we have one
[04:38] roughly 1 petabyte that like flows into our pipelines the beginning of them and we have to process through. That results in about 400 terabytes of data that we persist to output data sets across our various platforms. That's every day. Uh we have to manage producing transaction
[04:56] level, aggregate, and real-time data sets. So, to power various different use cases. Uh and just as an example, um this large-scale data set from the title of from the talk title is from one of our pipelines. That single pipeline, one
[05:13] of three main ones, processes 30 million events a second on its own. So, these are the sorts of scale issues we're dealing with um that sometimes face unique challenges for us.
[05:29] So, getting into our starting landscape, this is kind of the what of what our what our world looks like. We have multiple different internal platforms within our company. Each one has their own data pipeline, their own data sets, their own backing
[05:44] data platforms, again across cloud and on prem. The thing that makes that a little challenging is that we have a very frequent need to blend the data across these platforms. So, we need to provide a holistic view of this
[06:01] data both internally, so you could think of if you have an account, if you're an account manager, you have an account that has data across these platforms, you need to see the full picture, um for example, or if you're data science, you want to work across the whole platform, not just subsets of it.
[06:16] Uh and we also need to prevent that uh present that holistically to clients externally. So, again, same thing. You're a client, you work on our different prod or you use our different products, you would ideally like to see one view of everything you've got going on.
[06:32] We also need to support data sharing to our clients. So, that is both export, so, you know, you're a client, you want your subset of data each hour, each day, we export that out to your S3 bucket somewhere, old school way, but also more importantly, direct data
[06:48] sharing through some of the data sharing uh mechanisms that platforms have now. So, like Databricks has what is now called open sharing, used to be called Delta Sharing, Snowflake has Delta uh direct data sharing as well. Um so, that's a very key tenant of of
[07:04] the what of our ecosystem. And to highlight like what why this starting landscape wasn't working perfectly and why we need to make it better, uh we wanted to bring the company towards one data format and ideally also one data
[07:21] platform to make a lot of these working across boundaries a little more seamless, a little easier. So, that was our goal. I'm going to get into the how a little bit. So, that's that's our our landscape. This is now how systems and people interact with that data.
[07:36] So, we have consumer access uh through a lot of different channels. We Snowflake that we run things in. We have open source Spark on prem that we also run. We have a lot of things in Databricks, both uh Spark jobs, Spark
[07:52] notebooks, notebooks, things like that. Also SQL, ad hoc SQL, things like that. We have a few custom Java applications powering reporting and similar use cases that use JDBC connections. And this is to variety of different backing data
[08:09] sets. All of a lot of this is to a variety of different backing data sets. We also have people doing ad hoc uh personal exploration and analysis locally using things like DuckDB, PyIceberg. On top of that, we're also testing out a Trino-based system on prem
[08:24] as well. So, to see if there's room for an engine like that to, you know, utilize some of our on prem capacity for certain workflows. On the Excuse me. On the On the producer side, uh we have data coming into our data sets from Databricks jobs.
[08:41] We have some custom applications that output their data to S3 that we then need to load into uh our various data sets, the final data sets, both in formats of gzip CSV and parquet files as the starting point. And we also have on prem open source
[08:58] Spark that writes directly to their output data sets. Uh And I'll give you a sneak peek of like final end state, but that used to be entirely on prem writing to on prem. And now in this like future state we ended up in, we have on prem writing to on prem and
[09:13] on prem writing to Iceberg in the cloud. Little little sneak peek. Uh The one point of friction I'll call out cuz this is a this is a a little bit of like a through through line and and topic throughout all of this. Um direct data sharing to our clients.
[09:30] The consequence of one of our pipelines that does streaming event processing is events are processed as they come in, which means that an accounts data is interleaved all the data is just kind of a big puddle. It's not originally sorted by
[09:46] account or partitioned or clustered or anything. So, some of our clients, which are very low volume, faced a lot of friction and frustration because their data is interleaved amongst terabytes of everyone else's data and they had to do really large scans just to access their
[10:03] specific data. I'll I'll talk more some I'll talk some more details about that a little later on, but that's just a contributing use case of like something we have to solve cuz we don't like it when our clients are unhappy. We don't we don't want them to have a subpar experience when they're accessing our data through direct data sharing
[10:19] methods. Cool. Great. So, why Iceberg? Um how did we end up how did we end up here? Off the bat, extremely wide widespread support. Um
[10:35] almost everyone supports it. You get maximum interoperability. That was top of mind for us. Uh I'll talk a bit more about this as we go, but basically you get all of your consumers that are able to read your data in Iceberg format
[10:51] from a single source of truth, which is great. And to call out as well, consumers yes, but also producers. So, it would be a little useless if you had maximum interoperability for querying your data, but you still had to do a lot of rigmarole for producing it. So, also,
[11:06] you know, everyone can write to Iceberg now. It's great. And uh just I'll highlight again that like this also includes ad hoc personal usage. You're someone on a computer and you're like, "I want to bleep bloop around and look at some data." You can do that very easily. DuckDB, PyIceberg, etc. Things like that.
[11:23] From that, it also maintained excellent query performance. So, um looking around when we were doing our research, there are that indicate that sometimes if you move away from a proprietary format and a proprietary system, you might take a performance hit. I think that is true in
[11:39] certain cases, but overall we were seeing excellent query performance. We actually um saw better performance, which I'll I'll talk about when we get to this example data set as a bit of a consequence of data layout and things like that. But as a baseline, cool, we
[11:54] can maintain we can maintain the SLAs that we have to. The really nice point coming into our multitude of consumers is the zero data zero copy data architecture. So, this was a really fundamental
[12:11] consideration that I think I'll point out as one of these things that is universally applicable regardless of scale. I think even if you have you don't have data at the scale of our size, you probably have different consumers uh who want to access your data. You have different platforms you want to use, different tools, etc.
[12:28] Even if your data's fairly small and manageable, it's still a bit of a pain to set up these copy jobs and sync jobs and more moving parts and maintenance and now you divert from a source of truth and hey man, this thing says 10, but that thing says eight. And you're like,
[12:44] okay, now I have to figure that out. That's all That's all frustrating. I'm guessing Actually, let's make this interactive. Who has dealt with that? Great. Uh zero copy became really really helpful cuz now everyone you just cool. You want
[13:01] to read our data, great. Iceberg, here it is. Point to the source of truth. Everyone can just read one source of truth, no more flopping data out. This leads into our last point uh on here, which is that also made us very well positioned to not only continue, but also expand
[13:18] our direct data sharing with our clients and partners. Why? Well, before the conversation was when client was like, "Hey, we'd like to, you know, do some data sharing. We'd like to retrieve our data from your platform, put it in our own." Conversation was cool.
[13:33] What platform do you use? Maybe we use the same one in a line. Otherwise, I have to now set up this thing to copy some data to the place you want it so that you can consume it from there. And like, that's annoying and expensive. Now, the conversation is, "Great. Where do you want your data? We can support
[13:50] that. We will just stand up if we don't already have it, we'll just stand up at a mini account there, pointed to our source of truth, and boom, you've got your data." Like, that's a lot nicer. It's a lot easier. Um that was that was huge for us.
[14:06] I'll share some other points that are on the slide. I'll just talk to a couple points of uh things that we learned throughout our research journey. Um case this is, you know, helpful considerations. Delta, also very popular. We started this at a point where uh the explicit talk about Delta and Iceberg
[14:23] converging to basically be one hadn't yet fully started or wasn't as for uh like top of mind. So, Iceberg appeared to just be the clear one ahead. There was almost universal support for it. Uh it was a bit easier to just go with that. Uh we
[14:40] also discovered Delta's better for streaming workloads due to the difference between Delta and Iceberg in how they write metadata. Iceberg will rewrite your um manifests and your snapshots each time you you do some new data. Delta will just create, as it's named, Delta
[14:56] updates. So, we decided, "Okay, we don't want to shoehorn some part of our pipelines into a format that doesn't really make sense." So, we decided, "Cool, we're going to say Iceberg is the format for all of our gold data sets." And I know everyone has a different definition of gold, but for as we define
[15:12] it, we said, "Cool, the final ones, the important ones that have lots of lines coming out from them." And things in the middle can be what makes the most sense uh if they're just purely internal steps in a pipeline. So, we have for example some Spark streaming jobs on Databricks. Those output Delta, easy peasy.
[15:30] Couple of other benefits that we liked. Customer owned S3 buckets. So, a lot of things in the past, I know this has changed relatively recently, but a lot of things in the past and some things still are in proprietary vendor owned storage locations in some
[15:46] format. Iceberg, it's all in our S3 buckets. Even in Databricks, you can you can still open your AWS console and say, yeah, here's my bucket, here's all my data. Great. It's ours, we can move it around should we want. That leads to less vendor lock-in, which is great.
[16:02] Our account rep is here in the front row. She is sick and tired of me saying every time she's like, here's a cool new feature. I'm like, cool, but does it lock us in? So, that was a nice that was a nice benefit. And then Iceberg with Databricks as well,
[16:19] just moving towards one platform in addition to one one format. You get central governance via Unity Catalog. It's a nice benefit to move for us to move to a more like hub and spoke model where we have one central thing and then we can branch out from there to other tools and platforms that want to consume
[16:35] from there. When you mentioned you're sharing with partners, I'm curious about um like you said, oh, we'll get it there. We'll set up a mini account. What does that kind of mean from a technology point of view? So, like is it perhaps I'm in uh US West 1. I hope no one's in US West
[16:52] 1 on AWS by the way, but hypothetically, let's say you're in US West 1. Would would you be like, yeah, I can surface your data there like or how does that kind of play out? Yeah, so the the like classic example would be if we're going to cross region is a bit
[17:09] of a separate thing cuz you have to deal with now some egress and things like that, which is becoming easier. There's now some things that are on the road map to make that a little more seamless, but if we do a little scale back example of like I'm on Databricks, you're on Snowflake.
[17:24] Sure. Before again, it was like cool, I have to would have had to copy my data from Databricks and set up like a ETL job and export it and ingest it into Databricks and into Snowflake and pay for that extra loading and whatever. Now,
[17:40] you can just cool, I'll make an account in Snowflake. I will go through the steps one time to like set up an iceberg integration, create a catalog, point to Databricks, make sure the credentials are set, make their view for their data sharing set up one time and I'm I'm done. And I just manage my one table in
[17:56] Databricks and that's one consumer pointing in and we have that set up for clients and Got it. Easy peasy. Okay. Yeah, so and it's all automated, I imagine, right? You put rep that with scripting of some sort and you're all good to go. Ideally, yes. Yes, you have infrastructure as code, you have things Terraform, so that's easy and
[18:12] reproducible. Yeah. I I By the way, we did not plan the questions, okay? Like I don't assume I'm asking him things he knew he was going to get asked or We were prepping and he's like, I have some questions for you and I'm not going to tell you them. I was like, great. Love that. All right. Keep going. I'm good.
[18:28] Send me more. All right. Well, they're coming. All right. We're going to talk now about a specific data set as one example. One of our platforms is called SpringServe. We call and our internally we call transaction level data, we call that log
[18:45] level data. So, we have this log level data set for our SpringServe platform. This records individual event level rows for all traffic coming through this platform. This is, by the way, the platform that has that pipeline that does 30 million
[19:00] events a second coming through. This is a 10-day rolling retention. Um so, 10 days of data. This loads in about uh depends peak and off-peak, you know, we have very cyclical traffic. So, anywhere from like
[19:15] 3 to 7 TB per hour that we load into this data set. Uh that is approximately 25 billion rows. Um so, this is this is a fair amount. These are coming in from S3. So, we're loading in We have a custom application stream processor lands GZIP
[19:37] We need to put those into a table. So, that's a challenge to begin with. On top of that, it was previously prohibitively expensive for us to run clustering operations to optimize the data layout for this data set. So, I touched on this before, you know, we have stream processing, so it's just events as they
[19:54] come in. Uh all that data is interleaved. It's not segregated by account. Um it's, you know, roughly blocked by hour, which is nice, but within that, it's all just a big jumble. We could not run optimized commands
[20:09] again to recluster this before. That was causing some friction. In terms of the uses use and consumers of this data set, so internally, it's it's kind of one of our um uh like I I had a word for this that
[20:24] I've now forgotten, but it's like our golden child data set. If people go to it for debugging, for analysis, the data scientists love it because they get raw unfiltered data with all of the little attributes that they use for all the cool stuff that they do that I don't understand.
[20:41] And then, it's also very important for client-facing. So, we do uh some of these exports to clients, but also direct data sharing. So, we have that set up on this data set because clients also want that transaction-level data to see what's happening and be able
[20:56] to use that for their own for their own sort of use cases. Um this I didn't call out, but as you could have surmised, this is about a petabyte in size for 10 days of retention. Uh I'll tell you fun fact, when we started this data set, when when Liko and I were
[21:12] starting for this talk, it was like 900 TB a couple months ago. I was like, okay, we'll fudge it, we'll round up. And then a few weeks ago when we were prepping, I was like, let me just check in on it. I went into Databricks and I clicked the thing and I went on the details tab. And it was like, size 1, 012
[21:30] TB. And I was like, great. It's now actually a petabyte. I It's my fault. By the way, I knew you would You knew it'd get there. Yeah. So, it's actually a petabyte now. So, this is our starting point. This is This is what we have. This is what we have to deal with. Um
[21:45] I'm going to go to the next slide. I'm going to show you some results. Before you start reading what's on the slide, I'll just call out and I These are not benchmarks. Like, these are not People love saying benchmarks results and it immediately means benchmarks. This is a data set
[22:01] that we migrated to managed Iceberg in October. It's been running fully It's production state since October in this setup and I'm going to show you some results we've seen from that, not just from testing. Cool. Screenshot this one. It's a good one. Go
[22:17] ahead. Get your phones. Go ahead. It's good. Off the bat. So, on the producer side, again, reminder, we're loading in gzip CSV files from S3. Off the bat, 33% reduction in loading cost uh compared to previous system. The kicker for that is that includes
[22:35] Databricks auto clustering your data on write when ingesting this data set. So, before, very expensive to load, was not clustered. Now, cheaper to load, includes clustering, so we have data data layout optimal according to our
[22:51] liquid clustering column definition. I will touch on that. You'll see a little later down why that becomes important. On the consumer side, we saw a 50 to 70% reduction in P95 query time. So,
[23:08] um internal, external, all around, you name it, because data was now optimized by hour, clustered by hour, clustered by account. I think we have a couple of other columns in there that are frequently filtered on. Uh we we saw great performance gains.
[23:25] I'll call out that second bullet point under consumers. We're coming back to this through-line topic of direct data sharing to to clients. Those clients that were complaining to us of like, "Man, we need to run a warehouse that's five times the size that it should be because we have to
[23:41] sift through all of your data to get to our data." 90 to 98% reduction in the amount of data scanned in their queries, and they can now run right-sized clusters that are appropriate for their data volume. So, again, this was huge. I think the the thing I want to call out
[23:57] on that, too, when when Kevin updated the slide, and I saw that bullet point, I was I was floored and thrilled because in Iceberg, we worked really hard on getting scan planning right, so that you're filtering as much as you can before it ever comes down the line. Yeah. And seeing more of the real-world results, cuz you can imagine like,
[24:14] depending on what kind of dev you are, I'm a very local dev. I'm old school. Sorry, not sorry. It's all like my my test is my laptop. That's not their scale, right? Like, what's the best I can do on my laptop? So, to see real-world results like that, I thought it was super
[24:30] powerful, especially cuz I feel bad. I kind of like, he called them low-volume clients. I kind of called them the little guy. They're still You still love them, right? You love all your customers. And and so, it's like to tell a guy, "Oh, well, you're small, so you don't deserve good per- Come on." So, this like I love that bullet point.
[24:45] Yeah, it's a great call. I It really I mean, it really broke my heart having to like hear clients say that, and I'm like, "I really don't want that to be your experience." Like, this is This is my like this is I see this as my responsibility, and I'm failing you. So, that's a great call. Thank you. Um so, yeah. That solved that problem for them.
[25:02] Uh following from that, corresponding shouldn't be a surprise, but we saw a 70% reduction in the consumer warehouse cost now just due to shorter queries. So, we needed less horizontal scaling to be able to service our existing workloads. Um kicker,
[25:18] last bullet point, no changes to consumer queries. So, uh our direct data sharing is mostly done to clients through Snowflake currently. No one had to change anything. We just swapped a view behind the scenes and we're now reading
[25:35] from a single source of truth uh from Databricks to those clients. In the interest of being honest and transparent, I will say we did try to roll this out and due to a fault on myself, we did break it in the test beta release. So,
[25:51] before we fixed that and then afterwards, no changes all all opaque. Um But yeah, that was that was some great uh great results we saw from this. I'll also just share now a bit about the process of of what like what we tested to get here just to give
[26:07] you an idea of some of the things that again, this is one of those call outs of like is important regardless of your scale. So, we tested all the different permutations when we were testing this load job to get the data into Databricks. We tested
[26:23] copy into. Um Pro tip, if you have a lot of data, don't use copy into. It's not performant at scale. So, it's a freebie. Uh we tested batch spark. We tested autoloader. Uh we tested serverless compute, provision classic compute. All
[26:39] of this with Photon, no Photon. So, all these combinations to see what works best. And something else that came out of this, it's not strictly Iceberg, but uh included benefits with running this on Databricks, I called out that we got
[26:55] that liquid clustering on the right. Um we also So, that's a great benefit. We also got the benefit of predictive optimization. So, it's just uh for those of you don't know, it's the background worker that does like file compaction and re-clustering and optimizes your
[27:11] data layout and files and all of that stuff. Um we were skeptical cuz again, it was prohibitively expensive to do that automatically or manually before, and they're like, "Try it out." We tried it out. Surprisingly reasonable cost. Um so, that now just
[27:27] runs in the background, and we don't have to think about it, which is great. And the last thing I'll call out again, which I think is like a universally applicable consideration, is it's just got to remember it's not one-size-fits-all. Just because one thing works one way for
[27:42] one workload, it does not mean that that same thing will work for a similar but slightly different workload. So, to give you a a concrete example of that, in this case, when we tested serverless compute for this low job, was like many
[27:59] times more expensive than what we ultimately ended up going with, which was provision classic compute for this low job. Um but then we've now since moved other workloads onto Databricks, onto managed Iceberg, that are smaller
[28:14] smaller relative for us. They're still like almost terabyte per hour. Um but uh smaller aggregated data sets that we load more frequently, and serverless actually turned out to be the most cost-effective option for that workload.
[28:30] So, like great takeaway, you know, we couldn't just assume that this this would apply the same solution would apply to everything else. Uh another example of that, another concrete example, uh I mentioned Photon no Photon. We use Photon for this data set because
[28:47] uh one of the benefits of Photon is that it can do splitting and parallelization of processing gzip CSVs. That's not something that's available in open-source Spark. That's like a Databricks secret sauce. At the scale we were at here, that made
[29:02] a huge difference. But at a scale of some of these smaller data sets that we load in, Photon did not make the job materially faster to make up for its increased cost. So again, just two examples of still have to go through the various permutations and testing to make
[29:18] sure that it is the right solution for that specific problem and specific use case. I I think we try really hard to hit the sweet spot where you get bang for the buck or cost efficiency at every turn. So when he talked about predictive optimization, the goal is that this will
[29:35] only run when needed and only if it's going to benefit you. Like we don't want it to run just because, "Hey, look, it's Tuesday." Like that's not a reason to run. But if you have the right number of updates or whatever else, and now it's starting to take into account There was a separate session I don't know if it's happened yet or not cuz holy cow,
[29:50] there's 800 of them. But there's a separate session talking about how we're also going to start taking into account the client access, in other words, the reads, to maybe lay out the data a little differently or more optimally for those use cases. So it's I The way I think of it from a scale point of view, scale is all relative and
[30:05] that's why Kayvon keeps calling out, "Like, hey, this is our scale, but your scale might be different." Scale means different things to different people. My point is that if you have thousands upon thousands of tables, do you really want to be hand tuning every single one of them? I hope the answer is no. If it's not Yes, the answer is no.
[30:21] Yeah. It you know, like that's why this exists, but you can also turn it off if you're one of those people who likes to go in and tinker. That's up to you. So we give you that flexibility and freedom. I just hope we're doing our job right so you don't have to. Yeah. And I'll just add one other point of color, whether you guys believe me or
[30:36] not is up to you, but I am very skeptical of sales people and vendors when they tell me things. I'm like, there's no way that's true. No way that's true. PO turned out to be the Liquid saying what our account team told us pretty true. It is for us so far very
[30:52] reasonable cost for just not not having to think about any of that stuff. Um there's been a couple of similar things, but yeah. Still be skeptical, try it out, make sure, but try it out. Think the phrase we used at the dinner last night the most often was trust but verify. Like it's fine. It's what you're doing
[31:07] with your AI things that are like coding stuff for you and you know, doing databases in production. So, yeah, trust but verify. Um yeah, that's the end of my section. Liquid, do you want to Yeah, yeah, yeah. want to go through a little bit? So, let's talk about me. What about No, no, no, no. What When I say me, I mean you. Because you're not Magnite. I love
[31:23] you, too, but you're not Magnite. Your scale might be different. So, how does this apply to you? What can you expect? And I want to start by quickly asking, who's using Iceberg today? I know it sounds like an contrived question. I I'm sincere because I know there's also hype, right? Okay, this is good. Thank you. Thank you. Um
[31:39] I I work at Tabular. So, we had the first like reference implementation of a catalog out there. And so, a lot of people worked with us. And I I think that was around when the hype with Iceberg started. And so, a lot of people like, "Oh, I need an Iceberg." And it's like, "But do you?" And it's really weird, maybe hard for me to say, but
[31:55] maybe you don't. Um and so, like that's why um part of what I wanted to do is kind of almost balance what Kevin's been sharing with you is like Iceberg uh and everyone's hammer. And they're like, "Ooh, look at all these nails. I'm going to just keep doing Please consider carefully whether this makes sense for you. Um so, what I want
[32:12] to talk about are the uh inflection points. Where I see and totally worked with over 100 customers. I'm not trying to brag. I'm trying to give you perspective. Um I've worked with over I've lost count. So, I'm guessing it's over 100 customers. And a lot of their inflection points come around where they're like, "Hey, I
[32:28] have interoperability needs." And I'm not just talking like Databricks and Snowflake, but they're like, "Oh, I'm also working with uh a native just environment that needs to read this. Oh, I have my Iceberg in AWS first and I want to get it into Databricks. How would I do that? What do I do? And so, interoperability is key. In fact, that's what Kedar started talking about. It's
[32:43] absolutely what we see all the time. Um, I think the other interesting piece there, um, is data science use cases. I'm also not a data science person, um, but they all want the data. They want it where it is. They don't want to have to wait for it to be um, you know, you're
[32:59] talking about your raw before. I like that's powerful because we can't imagine what we're going to glean from that data until we do it, right? Like that data was useless yesterday, but now today it's like, oh shoot, you know what? This with this with this, wow, it's really powerful. So, I think those other things are happening, and that's been driving a
[33:15] lot of people into our Iceberg offerings in Databricks because they want to use our data science platforms. And so, I think that's a really powerful paradigm, and depending on how deep into this you are, you know, we have two different ways of looking at this at Databricks. We have managed Iceberg and foreign Iceberg. So, how many people write their
[33:30] Iceberg outside of Databricks? It's okay. Raise your hand. Don't be afraid. I love you equally. Don't worry. All right. No, good. Thank you. So, um, the thing I said because we came from Tabular, right? Thing I said coming in is I was like, look, everybody's first Iceberg platform is not going to be Databricks. And they're like, what? And I'm like,
[33:47] how could they do that? Like, you didn't have an offering. You had Uniform, and Uniform was a way to make a Delta table have Iceberg metadata so you could read it in those other platforms, uh, back in 2023. And it was available via Unity Catalog, so there were ways to make that
[34:02] easy, but not everyone was integrating with it yet. When you fast forward to 2024 and when, you know, our the acquisition was announced, we had to bring our catalog along, and we spent basically a year putting that together with what was already in Unity Catalog, and that's when write access became available. So, external Iceberg writers could write, you could write from within
[34:18] Databricks platform. And now we're GA today, which is really awesome news. It's been a journey of our own here, but my point in bringing that up is I think a lot of people, uh, get really confused about all this, and it's easy to get confused, and I'm sorry. And this is where I feel like, um, I I have a subtitle here I put in yet I
[34:34] have failed you. Here's why I feel I failed you. Maybe not just me, some other folks too. We talked about the formats coming together. If you watched keynote yesterday morning, you saw Ryan Blue up on stage with Ali and you know, he's like he's like, "Oh yeah, it doesn't matter. You know, tables don't matter." And I was like, "Yeah, they kind of
[34:50] still do. Sorry, Ali. I'm going to be in trouble." Um the way I say I failed you is this. We talked about the formats coming together and people already assume both can do everything. That is not true. I wish it were. We're working on it. We're not there yet. And so as much as,
[35:06] you know, whatever timelines you hear, whatever punchlines are on a keynote on stage, the reality is that the hard stuff was what was left behind. You know what I'm saying? Like we're working towards we did the easy stuff first. Getting the the underlying data to be the same,
[35:21] arguably easy. Ryan spent a lot of time last year working on Apache Parquet more than he did Iceberg, which is pretty funny to say, but it's true. That's because we wanted to get the data aligned underneath. Now it's the metadata layer. This is the harder part. And this is really where the the real work begins. So,
[35:37] I just I wanted to be clear about that because so many people are like, "Oh, sweet. I'm going to have streaming tables in Iceberg." I'm like, "Iceberg doesn't know what the hell that is." And they're like, "What?" I'm like, "Yes, we don't There is no concept. There is nothing in the spec that says this is what a streaming table is." So, you know, Kevin talked earlier about
[35:52] how Delta can handle these things inherently. Iceberg didn't even consider it, didn't even think of it. And so what we're working on with part of Iceberg V4 is to get the metadata rights to happen differently, to be smaller, maybe not happen every commit.
[36:08] That's it's controversial and it's actually happening right now in discussion wise. So, we'll see where that lands. But the thing I'll I'll also point out is that we are um and V3 right now is what's GA in most vendor platforms today. And V4 is being
[36:24] talked about. What you always see is a speckle come first and then we have to implement the spec. And I think it's going to mean time. And this is why I say like I wish we were closer together with this, but even if we said before was here tomorrow, it's not. But let's say that. It's going to be about a year because if you look at history, if it's tells
[36:40] anything, V3 was ratified in May of 2025, and we only just now saw the GA and Databricks GA and Snowflake. I think Google announced some in April. Like they're only just now coming out. So, it's kind of interesting how you'll see this kind of happen in stages. So,
[36:55] try not to get too excited about what one of the formats is doing. I'm also encouraging you don't wait. I don't want you to wait until Iceberg does something that you already had with Delta. Just use Delta. I'm not saying that cuz of the Databricks shirt on my back. All right. I don't want you to miss out because we
[37:11] could go a different direction. Communities can be very interesting, sometimes fickle. I always use the metaphor of like a neighborhood potluck. You know, everyone comes, they bring their dish to pass, and and we all share ideas, and we discuss these things. It's one of the last kind of pure places on earth where politics are almost civil, almost. It's
[37:28] it's a really nice thing, but it also means that an idea that none of us has even thought of might get brought up. And we might go, "Ooh, that's really good. That's better than that streaming table thing. Shoot, I'm going to do that." And you know what that'll do? That'll delay your adoption because we're going to have to implement that, and it's going to take a while.
[37:45] So, things to bear in mind and and just, you know, that was most of the perspective I wanted to share. Um I think we are at time. I know we started a minute late. I did want to offer the audience an opportunity to ask questions of us, and also I'll try to step outside and be available to you all as well. But
[38:00] likewise, yeah, any any questions for either of us? Do you think we'd buy a room? Uh-oh, we have a plant, I think. I mean, yes, you in the front. Hey, um great talk. Two questions. One is, you said don't
[38:16] use copy into. What do you actually use when when you're doing this fast data ingest? And two, is is there any permissions or entitlements or data access controls on on this data? Like, how does that work? Yeah. So, the questions were one, you said I said don't use copy into, so
[38:32] what did you use instead? And two, do you have entitlements or permissions or anything on this data? For the first question, so copy into you might be familiar with, it's the the way you do it in a lot of platforms where it's a SQL statement that's like copy into table from something in S3.
[38:49] Uh that that's what I mean by copy into. The alternative that we ended up going with is a Spark job on Databricks that writes Spark code that loads in data from an S3 directory into a data frame, does some schema application, uh little bit of normalization, things
[39:04] like that, and then does the Spark data frame.write to your data set. And that's faster? It was, if I remember correctly, it was faster and much more performant and cheaper. So, again, depends on your
[39:19] This is a thing that's greatly affected by data scale. So, we have small platform tables that we just They're very small and we they You could think of it as like metadata about your business. We use copy into for those cuz it's very easy to set that up. It's just a SQL statement, you fire it off. It's
[39:35] like, cool, ingest that thing into this table. If it's tiny, it's great. Um as your data gets a little bit bigger, that's where that performance difference starts to catch up. Second question, access control and entitlements. Uh no, all our data is publicly accessible to everyone.
[39:53] I'm starting at Magnet 2.0 myself. Uh that was a joke, to be clear. It's being recorded. Um yes, we do have entitlements, access control, governance through UC. So, within UC we have um user groups, service principles, all
[40:08] that stuff. We have full permissioning on the data set, so the applicable people internally have it. We use service principles for any platforms we want to data share to, so that is all controlled through again our central governance, which is nice. Does that answer your second question? Yeah, and if I can dovetail into the end
[40:24] of that one, one important callout for y'all to be aware of is again from an expectation point of view. I feel like a lot of times people have grand expectations of what Iceberg can and can't do. The governance elements of Iceberg are part of the Iceberg catalog specification, which is an open spec that anyone can implement. We've
[40:40] implemented it in Unity catalog, others have implemented in their catalogs, that's great. It does not have a common concept of K bond can access K bond the data that gets written to the object store where the data is. And a lot of people are surprised to learn that because that's the way it used to be. Yep. That's not the case yet, it's another
[40:55] thing we're trying to figure out how to sort that in the community. You might have seen me zip through a few slides, that's for the recording cuz this will be somewhere that you can visually see it as well. Look at what I said on some of the points and then email me, I'm Liko at Databricks, it's pretty easy to figure out, I'm unique in that way. Spell it out.
[41:10] L I K O. It doesn't sound that way, but yeah, anyway. Um, but email me if you want to discuss any of the three bullets on the like where do we go from here because that's where I talk a lot more about those things. I'm also going to set up a brain date for later so that we can actually chat about that. If you haven't looked at that in your apps, but fill out your surveys where I get fired.
[41:26] Please don't get me fired. Thank you.

Learn more about the Databricks Data and AI platform.

The information provided herein is for general informational purposes only and may not reflect the most current product capabilities or configurations.