Building a Scalable Social Media Research Platform with Databricks
Summary
- Princeton's Research Accelerator built a multi-tenant Lakehouse on Databricks giving researchers access to 2.1 billion Telegram posts, 250 million YouTube videos, and 400 billion rows of web traffic data to study social media at scale.
- The platform uses semantic search across 100 languages with multilingual embeddings, automated workspace provisioning via Databricks Asset Bundles, and Unity Catalog governance to reduce research setup from months to days.
- After a single runaway query consumed 16,000 DBUs, the team implemented automated cost guardrails, a cost transparency dashboard, and budget tracking to protect shared research infrastructure across multiple institutions.
Building a Scalable Social Media Research Platform with Databricks

Researchers studying social media platforms face a critical infrastructure problem: platforms have shut down API access, data requires massive scale to analyze, and no single institution can afford to rebuild tools for each project. Princeton's Research Accelerator addresses this by building a shared data infrastructure on Databricks, enabling researchers across institutions to study 2.1 billion Telegram posts, 250 million YouTube videos, and 400 billion rows of web traffic data.
Learn how to build a multi-tenant Lakehouse that democratizes access to complex social media data. Discover semantic search over 100 languages using embeddings, automated workspace provisioning with Databricks Asset Bundles, cost management guardrails, and Unity Catalog governance. This platform enables researchers to go from manual data collection to analysis in days instead of months, with examples including multilingual engagement analysis, gambling behavior research, and propaganda reach studies.
🤝
Chapters
00:00Introduction00:23Princeton Research Accelerator Platform01:28Computational Social Science Problems02:50Global Scale and API Access Crisis05:30Data Telescope: A New Model for Shared Research06:03Telegram, YouTube, and Comscore Data07:28Building Shared Research Infrastructure09:42Research Workflow: Onboarding to Analysis10:44Databricks Lakehouse Architecture12:35Three Data Sources: Structure and Challenges14:13Pipeline Architecture: Crawlers to Workspaces15:30Medallion Architecture and Data Organization17:04Semantic Search and Multilingual Embeddings18:26Semantic Search in Databricks Notebooks: Live Example20:23Workspace Provisioning and Automation21:10Declarative Infrastructure with Asset Bundles23:34Cost Management: The 16K Runaway Query24:21Automated Cost Controls and Guardrails25:43Cost Transparency Dashboard for Research Teams27:04Research Results: Geopolitical Engagement Analysis28:22Online Gambling Research and Power Users29:45Propaganda Reach: Measuring State-Linked Domains31:06Building Sustainable Research Communities
FAQs
What is the Princeton Research Accelerator platform?
The Research Accelerator is a shared data infrastructure built by Princeton's School of Public and International Affairs to help computational social scientists study online platforms at scale. It provides researchers with a familiar notebook environment backed by terabyte-scale Databricks compute, with datasets including 2.1 billion Telegram posts, 250 million YouTube videos, and 400 billion rows of web traffic data.
How does the platform support multilingual social media research?
The platform implements semantic search using multilingual embeddings that span over 100 languages, allowing researchers to find relevant content across global social media datasets without building language-specific pipelines. This enables studies like multilingual engagement analysis and cross-language propaganda reach measurement that would otherwise require separate data engineering work for each language.
How does Princeton manage costs for multiple research teams on shared infrastructure?
After a runaway query consumed 16,000 DBUs, the team implemented automated cost guardrails, budget tracking, and a cost transparency dashboard so research teams can monitor their own spending in real time. Workspace provisioning through Databricks Asset Bundles also enforces resource policies at the infrastructure level, preventing any single project from monopolizing shared compute.
What research problems does the platform enable?
The platform supports computational social science studies including online gambling behavior, teen mental health and social media use, and online extremism and radicalization pipelines. By providing shared access to large-scale datasets, it enables research that no single institution could afford to undertake independently, turning months of manual data collection into days of analysis.
Full transcript
[00:08] Um but hi everyone. Thanks for joining our breakout session. I'm Kai. We're with the Princeton School of Public and International Affairs. Um I'm the director of engineering for this group and this is Lily. She's one of our data scientists and a clutch data engineer. Um So we're here to talk about how we built
[00:23] a scalable platform for computational social science research on Databricks. So a little bit about us. We're from the Research Accelerator. We're a team of software engineers, data scientists, and data engineers brought together at SPIA to help create
[00:40] scalable software and data tools to tackle really difficult computational social science problems. So as you can see, we're a pretty small team. We're five engineers, a product manager, couple of ops folks, and that's kind of what this talk is all about. How a
[00:56] pretty small team can build an entire research platform on Databricks. And we are basically running a multi-tenant privacy-preserving scalable research platform and we've leveraged several Databricks tools and technologies to help researchers and present them with a
[01:11] notebook environment that they're all familiar with, but they're able to run terabyte scale experiments on. Okay, so here's a little bit of a sample of what computational social scientists are tackling today. Um so online gambling, I'm sure a lot of you guys are
[01:28] into sports. We have like the World Cup going on in this great city, the NBA finals just happened. Go Knicks. One thing I was really surprised about was the prevalence of sports betting commercials on on both of these events. And sports
[01:44] betting is just like really everywhere. State by state legalization moved super fast and now there's basically a sports book in every everyone's pocket, but there's very little information or knowledge on who's getting hurt by these platforms. Teen mental health, I'm sure some of you guys have heard of Jonathan
[02:00] Haidt's book Anxious Generation, but that kicked off a national discussion about school phone bans, age verification laws. And this is all on the premise that social media platforms and the prevalence of smartphones may have this really bad effect on some of
[02:17] the youngest members of our society. And really the technology is moving way faster than our understanding of it. Online extremism in the third panel, that's another complicated topic that's been coming up quite quite recently in our in our politics. Radicalization pipelines run
[02:33] through recommendation engines driven by AI, messaging platforms and forums. And we're trying to understand like how do these groups form and spread and what would it take to give people the tools to push back. So one thing all of these platforms have in common is they exist on the vast
[02:50] super complex global information ecosystem. And this includes some of the largest records of human behavior ever collected. So think about the fact that most of these platforms have billions of people online sharing billions of ideas, arguing with each other on all these
[03:05] platforms. And so the solution to study these platforms requires really big resources and scale. So ultimately the social sciences faces a problem of collective action. There's about 3,000 social science
[03:21] research groups globally that study issues like this. And each research is chasing their own grants and their own funding and each is building their own tools and infrastructure in a bespoke way to try to solve these problems and to try to collect data. It used to be a lot easier to deal with. So the
[03:37] platforms themselves like Meta, Reddit, excuse me, and Twitter, they once gave researchers generous API access and tools to get data sets and to study and do research on these platforms. But since the 2020s these have been like slowly shut down. So, there was once a
[03:54] group called CrowdTangle that was a part of Meta and they they kind of interacted with researchers giving them access to APIs and tools and working kind of hand-in-hand with researchers even providing some resources. That was like kind of slowly shut down.
[04:09] I don't know if any of you guys know this, but Meta is like one of the toughest platforms to do any type of research on. They are like actively fighting against researchers. Reddit and API and Twitter APIs are basically today like way too expensive for most research labs to afford and the terms of service have
[04:25] just become really expensive. So, in essence, the age of easy API access for researchers, it's over. And what's left is researchers are going through like really extraordinary lengths to try to sample data. So, I really love this middle image here. It's
[04:42] from a research group. They were basically trying to get a census of TikTok. Basically trying to figure out like you know how much content is being created in any time slice. They basically reverse engineered the video IDs so they were able to get like the subject matter in part of the string and
[04:59] they they realized that the other part of the string had a timestamp in it. And so, they built this series of this is Raspberry Pis actually. It's kind of a bunch of Raspberry Pis and they just did all of these web requests against TikTok and they figured out that and this is a
[05:14] 99% complete census that the platform produces about 269 million posts per day and that's about 174 years of video content daily. So, that's it's really incredible. This is really incredible
[05:30] work and researchers are doing this worldwide every day. The problem is is like after something like this gets created and you are creating all this code, building all this engineering effort, spending cloud credits, and spending real money, once the papers are published, it's all gone.
[05:47] We can't afford to work like this. So, what this should really highlight is these experiments have to be run on data coming from some of the most well-resourced um tech platforms in existence with billions of users. And so, I'm going to like talk about two of
[06:03] them that we have on our platform, um namely Telegram and YouTube. So, YouTube has like billions of monthly users, um billions of images and and videos and messages crisscrossing the crisscrossing the globe daily. Um so, there's 250 million YouTube users in the
[06:21] United States. That's almost every single adult that lives in this country. Um there's about half a billion people in India on the platform and hundreds of millions of people on Telegram in in primarily non- um English-speaking countries like Brazil and Russia. Um
[06:38] Telegram is particularly challenging to deal with. Um it's kind of unlike the what the way the rest of the internet works or the way that the rest of um these social media platforms work. There is no like centralized directory. There's no global search. Everything is kind of like a forwarded message. So,
[06:54] you get like clusters of groups that are forwarding messages on and on and on. So, just crawling it is a complete nightmare um to do and what you end up with is a huge tangle of data um and information that is insanely complex. So, we're basically in this like crazy bind where the scale just
[07:11] dwarfs any labs um resources, the data itself is super messy and difficult, and now the APIs and the platforms are actively fighting against us. So, what we thought is that the social media or sorry, the social science
[07:28] community really needs is collective action and their own large-scale uh s- uh science instruments. So, I know you guys are all nerds, like fellow nerds here. Um most of you know what this is. This is the James Webb Space Telescope.
[07:44] Um, and it is like a really awesome example of what collective action accomplishes. So, this wasn't just NASA. This was NASA, this is the European Space Agency, the Canadian Space Agency, 14 countries, hundreds of uh research institutions globally involved to build
[08:02] this one science instrument to study the universe. Um, so there's about 2 trillion um galaxies in the known universe and one of uh the projects uh which I mentioned up here, the Cosmos Web, um they imaged about um 780,000 of them. So, it was
[08:19] like a tiny patch of sky compared to how big the universe is. But, I I think like the idea is like, you know, pretty compelling here. Like, they they don't have to capture the whole universe. They don't We don't have to sample all of YouTube um to to to understand it. They just get a small
[08:36] piece of it, um carefully characterize it, curate it, and then you share it with a bunch of researchers. And from one small, well-chosen sample, you get a general understanding or a general understanding of the universe. So, we don't have an international team
[08:52] of thousands of engineers and we don't have a $10 billion budget, but I think the concept is very useful and it's something that we thought we could use um within the accelerator. How could we Oh. Oh. Sorry about that. Shouldn't have pushed that.
[09:11] Okay. So, how can we How can we um use this idea to build a shared infrastructure, build that community um so that not any single institution is having to bear the cost of all of uh all of this and trying to discover um
[09:26] uh truth about science. So, we borrowed this concept, we built the community, um we have kind of I won't mention all the uh universities up there, but there there's a bunch of them. Um and we built this shared tools and infrastructure. The whole research experience on our platform comes down to
[09:42] these four basic steps. They onboard with a data use agreement. That's kind of the entry point. We sort of taken care of the terms of service, the contracts, the access rights, and everything. They get curated data, which we call the typical activity data set.
[09:57] So, that includes 2.1 billion Telegram posts and conversations, 250 million metadata to YouTube videos, and 400 billion rows of Comscore web traffic data. And this is basically taken from a sample of Americans over the past half
[10:13] decade about their web use and traffic over mobile and desktop devices. And this is all cleaned up and ready to analyze. They get scalable compute. And they can analyze and run experiments. And again, in an environment that they're used to
[10:29] in these like notebook environments with SQL, but now they have like massive compute clusters that they can bear on the problems that they're trying to tackle. So, this is a high-level view of the platform. It's a basically a managed Databricks Lakehouse for social science research.
[10:44] From our custom crawlers to the legal institutional guardrails to the apps and tools that the researchers use, everything is connected through Databricks. Our own crawlers on the left bring in the data and it flows through a medallion Lakehouse. Lake Flow, Auto Loader, and multilingual embeddings, and
[11:01] other ML feature extraction steps managed by MLflow. And on the right, researchers reach all of it through a portal, notebook, and some of the software tools that we've built for them. So, governance is built into the Unity Catalog. So, the platform itself is kind of enforcing everything that we need.
[11:18] And and the platform is sort of like kind of policing itself. And so, what we've built in the end is this fully governed multi-tenant and multi-institutional Lakehouse for really messy human interaction data up on social media. And it's an idea that we
[11:33] think has use cases in other industries, but we're here to share how we are using this for computational social science. Okay, so that's a quick overview. Um Lily as our data scientist and engineer, she's like way smarter than me. She's going to actually go into some of the technical aspects um and also give some
[11:49] examples of the research that we've done on here. Lily? All right. Hello. Can you all hear me okay? Awesome. All right, so like you have I So we've got billions of rows of data from Telegram, YouTube
[12:04] videos, and 400 billions of rows of behavioral data. So all of this is sitting on a lakehouse and it's run by only a handful of of engineers. There's so few of us that you could fit us all in a minivan. So let me break down how we're actually going to be building all of this.
[12:20] So before I show you the architecture, I want to show you what we're actually working with, right? So the data itself is the reason why a standard pipeline isn't going to cut it for us. So you see some examples up here we have three sources as Kai broke broke down and they couldn't be more different from each other. So Telegram, it's
[12:35] decentralized and it's massively popular in non-English parts of the world. Um so if you're not super familiar with like the Telegram interface, there's private one-to-one messaging, but there are also public-facing channels that will post messages to their subscribers and that's the data that we're thinking about here.
[12:52] You're probably a lot more familiar with YouTube. So there are public channels with with a huge volume of video metadata. So you can think of transcripts, video descriptions, channel info, and there's so much engagement going on on YouTube. Think of views, likes, comments
[13:07] especially. So when you think of engagement metrics between these two platforms, they're very different. So Telegram, you have emoji reactions, it's multi-lingual content. You might be more used to something on YouTube here. Uh and finally, Comscore is our licensed panel behavioral data set. So these are actual
[13:25] real mobile and desktop usage over time. These are actual users hitting actual domains with timestamps. Um so we have data from the United States about a 5-year uh worth of data supply, and then it's about 100,000 or maybe about 200,000 users here. So this is about 30
[13:42] terabytes of raw text data for all of our Comscore usage. So you can see an example on the screen up there. So you have a unique hash for a user, and there's a timestamp where they're visiting a domain. Maybe they're going on YouTube to find funny cat videos, or maybe they're going on FanDuel to place
[13:58] a bet against the Knicks. So not a good outcome. Uh so three sources with three massively different structures. The way that you're even able to get these are different. Different engagement metrics.
[14:13] It's all different. So we need to figure out how are we going to put these all on one pipeline, and on a platform where researchers can actually work with this. There's a lot here, but this is the full pipeline. So the data's going to flow from top to bottom, and the end state is a researcher in their own Databricks
[14:28] workspace. So that's in the green down there. So we start with our crawlers on AKS. So they're pulling from Telegram and YouTube. Uh everything lands as raw JSON in our Azure blob storage. So Comscore up there, again, that's our licensed data. We're not actually crawling that. So a lot of the details I'm going to be
[14:44] giving are more in the context of YouTube and Telegram. Anyway, this raw JSON that is our untouched source of truth, and it's going to get picked up by the autoloader and put into a Lake Lake Flow declarative pipeline. So the first step on the pipeline, the bronze layer, is where we're going to
[15:00] handle schema validation and any deduplication. So in the event where we recrawl any of our channels or posts, we're not double counting it. So a little bit of context there, we would recrawl whenever we're trying to look at engagement and track metrics over time. So how large has a platform
[15:15] grown over a couple of years? Um how large have the top channels grown in the in the last few months? So things like that. Silver is where it's going to get a lot more interesting. The data here is organized at the post level, reply level, and channel level. And so,
[15:30] I really want to highlight how interesting it is when you get replies into the mix. Because once you start thinking about replies, like comment sections, for example, you're massively increasing the amount of data that we're working with. So, think of a like a YouTube video, maybe it has 1,000 likes, but there are tens of thousands
[15:45] of comments because people decided to start a fight. That's all data we need to be able to account for. Don't go to the comment section. No, don't. Don't go to the comments. So, all this data is going to be cleaned up. It's got common keys and timestamp metrics so that users can actually do like cross joins and analyze across all
[16:02] these different tables, and it's all governed by the Unity Catalog. And I'm going to talk about why that matters a little bit later. Um another thing that you see up here in purple, there's room for machine learning integrations on our pipeline. So, that's for feature extraction or maybe enrichment. So, an example
[16:17] enrichment here are is an embedding model. Uh I'm going to go into a lot more detail of why we have embeddings in the next slide, but I'm just going to highlight that it's part of it. Um but before I do, I just want to show this green layer at the bottom here. So, that's a researcher in their own Databricks workspace. So, it's
[16:33] completely scoped to whatever their usage requirements are for data or their team. And here they're getting the full Databricks experience, right? So, they're getting notebooks, SQL, dashboard. It's all there. So, the whole point of this is that we're not just giving researchers the data to work
[16:48] with. We're giving them an environment to actually get their work done. All right. So, on top of everything we just built, we're also factoring in a semantic search. So, the whole point of this is that a researcher shouldn't have to know what language their data is in in order to find what they're looking for, right? So,
[17:04] we're layering in semantic search. And so, we start with an index. So, we're taking a stratified subsample of all of our data by language. So, we want a really accurate piece of the data that we're working with. So, there's not It's not too English-heavy for YouTube or too Russian or Farsi heavy for Telegram.
[17:21] Everything gets equal weight. If you're not familiar with embeddings, so you can just think of it as turning text into numbers. So, these are going to be points on a high-dimensional uh graph space. So, I have an example here of some fake data. Um So, you can see these clusters forming around specific topics. So, this is text
[17:39] about sports up there or tech or politics. And everything that's going to be closer together has similar semantic meaning. So, on the query side at the bottom here, you can see that a researcher can just type in whatever they're looking for in any language. So, they're looking for a World Cup, if anybody is watching
[17:54] that, maybe you've heard of it. Um but we encode that the same way, so it gets put in the exact same embedding model. Um this is the BGE M3 embedding model, and it's pretty good at handling over 100 languages, which is really important when our data has such a large multilingual distribution. We don't want
[18:09] anything that's really English heavy like other embedding models. Anyway, it's going to get embedded, and we're going to compare it against everything else on this index on the graph. Anything that's closer together in that cluster is going to be surfaced as topically relevant to whatever that search query is.
[18:26] So, that was a lot of information about the architecture, especially for semantic search. So, I'm going to show you uh what a researcher will actually see in their workspace. So, this is a Databricks notebook that you're all probably familiar with. Um and we have this set up for every single research workspace that gets auto-populated there. So, at the top here, we have a
[18:42] bunch of parameters that they can use to enable semantic search. Go ahead and play this. All right, awesome. So, this researcher is going in, and they can set keywords, they can set as many of these filters as they want. So, they might be interested in specific languages like English or Farsi and a bunch of others. Um they can
[18:59] limit the results they get back, and then they can also filter for date ranges or platform. And then they're going to type in their search query. So, what I just described, it's going to get embedded that exact same way. So, they're looking at cryptocurrency price crashes. And if you're ever really in Telegram, you'll know that this is
[19:15] actually a super popular uh topic of conversation on there. So, this similarity threshold that they've set, so if you think about that graph I described, if I have a higher threshold, I want everything clustered super together, uh super closely together. If I want a lower threshold, I want to cast a wider
[19:30] net. So, they don't have to write any code. They're just going to run these cells, and they get row-level data that matches their search query. This also lets them go back in and edit and maybe they're going to tighten their search or broaden it, but they can click through the row-level data and see, "Oh, this isn't the platform I was looking for. Here's the raw content. If I can
[19:47] read these languages, I can see that this is about cryptocurrency." So, there's some in English if you read English, other languages if you read other languages. It's all there, and it all matches the parameters that the user has set. And because it's a Databricks notebook, the fun doesn't stop there. So, they can keep going, and they can do an
[20:03] additional analysis on top of it. They can do sample statistics. They can look at uh their query results by language, over time, by platform. They can even build visualizations on top of this, put it in a dashboard, export it. It's all doable in their workspace.
[20:23] So, I've mentioned a few times now that we have Databricks workspaces for every single research project. So, Databricks isn't just where we build our architecture, it's where the research is actually happening right now. So, every project is getting its own isolated workspace, and getting them in their workspace is an engineering problem of its own.
[20:40] All right. So, before we automated this process, onboarding a new research project was super painful, and it was super manual. So, there's an engineer, maybe me, who has to go ahead and click through all of these separate UIs, or maybe go through Azure and set configurations there. So, I would have
[20:55] for every single workspace, I'd have to go, "Okay, they need access to these Unity Catalog bindings." So, I would add the bindings, and I would have to manually add groups to the workspace, table-level permissions, over and over again for every single workspace. And I think you all understand that this just
[21:10] isn't scalable. So, we rebuilt this as a declarative pipeline. So, now it's going to start with a single entry in a registry. So, we have a workspace that's been provisioned in Terraform, and now we have a JSON object that tells me that tells the job "Here's the new workspace URL. Here's the group that's going to be
[21:27] assigned to the workspace, and here's the workspace type." Now, the workspace type we actually map over to their data use agreements. So, if you might think of if you're a little familiar with how research projects are are handled they'll have a data use agreement so they can only have access to our
[21:42] Comscore data set, or only have access to YouTube or Telegram. So, because it's declarative, we can have that mapping configured in our in our DAB. So, then this is going to kick off a GitHub action workflow. So, it's going to run the bundle. The bundle's going to take it from there. Um it's going to
[21:59] deploy uh and bind Unity Catalog. It's going to attach workspace group. It's It's going to apply all the grants and permissions, cluster configurations. I do just want to highlight that right now in uh the workspace uh at the workspace level, that's the only things that we can handle it for
[22:15] the DAB. So, anything at the account level, we have to bring in additional steps. So, things at the metastore label label have to be done probably using the API or um other methods. But, the whole point of this is that the core provisioning workflow is the
[22:31] same pipeline regardless of what your workspace requirements are. And because it's declarative, we can repeat it, we can audit it, and we can scale it and keep it consistent.
[22:46] Okay, so a little bit of motivation for this slide behind me. Um when researchers come into Databricks for the first time it's not just their first time in Databricks, this might be their first time in a cloud compute environment to begin with. So, they might be used to running code in their laptop, maybe an Excel sheet that they're having some
[23:02] undergrads do, or maybe their university has provisioned some HPC cluster, and it's basically getting managed and handled for them. So, this mentality or this mental model of, "Oh, I'm running these SQL queries. This is costing me money." may not have landed yet. And so,
[23:17] we learned this the hard way. Um so, early on we had a researcher who kicked off a serverless job. Uh they shut their laptop and then went home for the weekend. They came home Monday They came back to the office Monday morning and saw that they had racked up a $16,000 bill.
[23:34] Let that sink in. So, The console said 100 trillion row output, 196 hours, and they didn't they ignored it. And um it actually didn't finish running. It was scheduled to take 40 days. Yes.
[23:49] That was my reaction, too, when I took a look at it. So, basically, this is the moment that really made us take cost and resource management seriously as as an infrastructure problem, not just as a as a researcher education problem, because there's only so much that maybe Kai and
[24:04] I can do to say, "Okay, you're going into Databricks with these auto scaling compute uh resources. Good luck." We can't really say that. We can't just say, "Be careful." We actually have to build guardrails around this. So, we did a couple of things.
[24:21] The first is what I've called originally the compute kill job, but now I'm calling it the reaper because I think that sounds cooler. Um So, what What this does, it's an automated job, and it's going to run every 30 minutes. It's running from our main uh dev workspace, and
[24:37] it's iterating through every single workspace in our account. It's looking for any running compute resources, and it's going to evaluate them. If the resource has been running for more than I think the threshold is 6 hours, it's terminated. Doesn't matter if it's in the middle of a job, nothing. So, it's going to get logged that it's been
[24:53] terminated, the user is going to be alerted about this, and this basically takes care of that runaway notebook or runaway SQL query problem. Um another thing that we did Excuse me. Another thing that we did was we introduced node limits onto our compute resources. So like I mentioned
[25:09] in the configuration tab workflow that we're deploying, we can handle a cluster configuration. So the type of work that researchers are doing in their workspaces don't really need that crazy auto scaling that serverless allows for. Classic compute with set node limits is
[25:27] just as fine for what their workload is. And you can see there's a little graph that when we set these node limits, when we introduce the kill switch or the reaper, we actually can significantly cut costs for every single workspace. And that's really important if you're working with a really fixed budget. Like your grant
[25:43] only covers so much money, every penny counts. All right. And finally, the dashboard that you're looking up here looking at here, it's rolling up DBU consumption and some Azure VM cost estimates. So this is an automated dashboard that we
[25:59] built and it lands in every single workspace. It's pulling from system billing tables and it's pulling from usage tables and it's giving researchers a better understanding of how much am I spending and where am I spending all of it. What we don't want is a researcher getting their Azure bill at the end of the month and having no
[26:16] idea where all this cost came from. So the idea here is that someone in charge of a research project, maybe there's multiple people in there, can look at activity on here and say, "Oh yeah, I can see this cost. It was from that big job we ran the other day." Or "Hmm, I think I made an app as a demo
[26:32] and I never turned it off and this is costing me 20 bucks a day." So all of this is about transparency. It's about getting users into a cloud compute environment safely and securely. That was a lot of engineering. So now I just want to wrap this up with what is
[26:48] the real research that's happening now on our platform and I and I think you're going to find it pretty cool. All right. So, this is the first one made by our very own research scientist on the team, Shawn Norton. So, what Shawn is doing here is a great example of what you can do when the infrastructure is actually built for
[27:04] researchers, right? So, Shawn here is looking at the top 1% of most engaged with YouTube channels, so all the top channels, and he's measuring he's doing a multilingual analysis to measure viewer engagement with Russian language, Ukrainian language, and English language videos
[27:20] around the time of the full-scale Russian invasion of Ukraine. So, what he's finding here is that engagement with Russian and Ukrainian videos skyrockets during this time. And then there's also a little bit of a drop-off with English here. Um if you want to add that
[27:35] Oh, yeah. Some of you probably noticed that we don't know why. Um but yeah, there We have been actually looking into this. One thing we did find is that YouTube started really promoting shorts around this same time. And the platform got
[27:51] completely flooded with a high number of shorts videos. So, we think that possibly it could be the the denominator just got really inflated, and so the average number of views in English just like went down a lot. Um and it might be also partially explained by just a bunch of bilingual Ukrainian or Russian speakers
[28:07] they're like displacing all of this English-speaking content. So, but it's kind of a conjecture. Yeah, and this is all part of our YouTube data set, multilingual analysis. This other one was also done with Shawn and a team at Rutgers, and this is probably my
[28:22] favorite thing that has come out of here because it's so relevant. If you are alive and have internet access, you know about online gambling, and you know how big of a problem it's been. So, Shawn and this team here are looking at using the Comscore data. So, again, that's that um
[28:38] real domain usage and mobile usage data set. They're looking at 18- to 25-year-olds and their daily visits to gambling sites and apps from 2019 to 2024. So, on that little time graph chart there, those little red lines, those are those are showing when
[28:55] Pennsylvania legalized online gambling and then when New York and four other states legalized online gambling. And there's this huge spike in usage at this time. The chart here is breaking down the same age demographic, 18 to 25-year-olds, by income level. And what they find here is that
[29:11] those who make under 60k a year are drastically more likely to become what they call a power user. And that is not a good thing. You do not want to be a power user. That means you are the top 5% of all gamblers on these sites, like FanDuel and all the other ones. Um
[29:27] So, yeah, this was this was a super interesting analysis because they weren't able to just track if you've been to those sites. It track The Comscore data tracks how often do you go? What queries are you doing on there? How many How long do you spend at on each domain? So, it's all extremely relevant.
[29:45] And finally, this is another um interesting uh piece of research done by uh one of our other researchers, Kevin Green, using Comscore data as well. So, this is a propaganda study. So, we're looking at something kind of different from the other two studies, right? So, this is still the uh this is
[30:02] the actual US reach of state-linked propaganda domains. So, using mobile engagement data with news domains from that same range, 2019 to 2025, um those red lines on the chart are Kremlin-linked domains and the blue lines are the Chinese Communist
[30:17] Party-linked domains. And so, it's kind of hard to see on here, but this graph is actually on a log scale. So, what might look like a small difference is actually huge. So, basically, what you're seeing here is that these domains, these propaganda domains, are basically getting no traffic from US
[30:33] users, even with years of effort trying to target Western audiences. The actual reach is really tiny. Um So, yeah, these are also made possible by the Comscore data, and you can watch it over time and continue to add to this analysis as more of that data becomes
[30:50] available. And so, yeah, that's the research. So, you can see these are three completely different questions with completely different methods, but they're all made possible on the exact same platform. I'm going to hand it back to Kai now. Okay. Okay, thanks, Lilly. Um
[31:06] All right. Well, thanks again for coming. Um so, we started with a problem of a vast information ecosystem and scientists trying to study it who have been systematically locked out um by the industry. And so, one last thing I want to talk about is is university funding,
[31:21] um which has um been an increasing problem um in the last couple years. Um so, what you've seen is a a shared instrument, you know, that many institutions can build together and share. Um just a handful of people that it takes to build and run one of these things and growing that consortium.
[31:38] And that model matters a lot more today in today's funding environment than it did 5 years ago. The way that we've been doing social science, um running the same type of pipelines, building the same stuff, and then throwing it all away, we cannot continue that way, especially in this environment.
[31:56] Researchers at Princeton are doing an amazing job doing what they can, pulling together resources. Um researchers at local institutions here like Berkeley and Stanford, they're doing the same things. And this is really important research. We can't lose
[32:11] this stuff, and we can't just like bear the the cost of this anymore. Um so, whether you're a researcher, a potential partner, or you just want to chat about um the Telegram schema or how to scrape it more effectively,
[32:26] um this is how you can reach us. Um thanks so much, and we hope you have a great rest of your conference. And if there's any questions or comments
Learn more about the Databricks Data and AI platform.
The information provided herein is for general informational purposes only and may not reflect the most current product capabilities or configurations.