Skip to main content

Production-Scale Healthcare Data Extraction: 400K Records from 5M Web Pages

Summary

  • Databricks for Good partnered with the Virtue Foundation and AI for the Earth to extract and deduplicate 400,000 healthcare records from 5 million web pages across 73 countries, creating one of the most advanced open-source medical datasets in the world.
  • The production pipeline used Delta Lake with Photon acceleration, multi-step LLM extraction via OpenAI, Splink probabilistic record linkage with salting to handle hot Spark partitions, and LakeFlow Jobs for asynchronous DAG-based recovery.
  • On top of the unified dataset, the team built a deep research agent and the VF Match platform for medical desert analysis, demonstrating how Databricks handles real production AI challenges that proof-of-concept projects never encounter.

Production-Scale Healthcare Data Extraction: 400K Records from 5M Web Pages

Watch: Production-Scale Healthcare Data Extraction: 400K Records from 5M Web Pages
Building production AI systems means solving problems proof-of-concept projects never encounter. The Virtue Foundation partnered with Databricks for Good to extract and deduplicate 400,000 healthcare records from 5 million web pages across 73 countries, creating the world's largest open-source medical dataset. This talk reveals how Databricks handles scaling challenges that separate concept from production: LLM extraction from millions of pages, entity resolution on hot Spark partitions, and orchestrating asynchronous recovery workflows.
Learn how to architect production pipelines using Delta Lake, Photon acceleration, and Spark parallelism. The team implemented multi-step LLM extraction via OpenAI, Splink probabilistic record linkage with salting, and LakeFlow Jobs for DAG-based recovery. For technical audiences and decision-makers, this covers real production challenges including API quota management and skewed data distributions.
🤝

Chapters

FAQs

What is Databricks for Good and how does it support nonprofits?

Databricks for Good is a program that provides pro bono services to nonprofits, including advisory support and hands-on engineering from forward-deployed engineers who set up repositories, build implementations, and teach teams to maintain them independently. The program also provides free compute or compute discounts to eligible organizations.

How did the team handle extracting healthcare records from 5 million web pages?

The team built a multi-step LLM extraction pipeline using OpenAI to process web pages at scale, with LakeFlow Jobs providing a DAG-based recovery mechanism for asynchronous processing of failed or delayed tasks. Delta Lake with Photon acceleration provided the storage and compute foundation for processing millions of pages efficiently.

What is Splink and how was it used to deduplicate healthcare records?

Splink is a probabilistic record linkage library used to identify and merge duplicate healthcare records across sources where no common identifier existed. To handle hot Spark partitions caused by skewed data distributions, the team applied a salting technique to distribute the computational load more evenly across the cluster.

What does the final dataset enable for healthcare access analysis?

The dataset of 400,000 healthcare records across 72 countries powers VF Match, a platform for analyzing medical deserts—areas with insufficient access to healthcare resources. The team also built a deep research agent on top of the dataset that can be queried conversationally to surface insights about healthcare availability globally.

Full transcript

[00:07] Hello everyone. How are you doing? Good. Great audience participation. Thank you. Um raise your hand if you are a technical person, meaning you write code or do AI. Raise your hand if you're a decision maker. Okay, and then like a manager {{}slash}
[00:22] leader. Awesome. A lot of people wearing many hats. Um cool. So, we'll try to tailor this talk to that type of audience. We're really excited to be chatting about um this project today. And that project is uh the Viraj Foundation.
[00:38] So, taking a step back, my name is Michael. Uh I'm one of the founders of Databricks for Good. We help nonprofits uh by giving them pro bono services. So, we bring in experts for free. And we also give free compute or compute discounts. Um to my left here or to my right here
[00:53] is So, I'm Nico. I'm a data scientist at Viraj Foundation. And I'm Priyanka and I'm at AI for the Earth. Uh and we've been working for the last couple years uh collaboratively with Viraj Foundation to create one of the most advanced medical data sets in the
[01:09] world. And then on top of that data set a deep research agent. Um so, you saw the video at keynote. Um same project. And then also we have a demo downstairs at the Expo Hall if you want to chat with the agent or just learn more about the project.
[01:25] So, uh Databricks for Good. The origin story is I was working in uh the retail vertical as a forward deployed engineer working with some of the largest problems, big data issues, uh working on GenAI as well. Um and I saw the power that the
[01:40] Databricks platform offers. I also saw the power that a dedicated Databricks employee can can leverage uh to make impact. And I've been a mission-driven person my whole life and I was like, "Hmm, well, we can help the bottom line of a bank or a large retail
[01:55] organization, but can we help a lean nonprofit? And so that was sort of the impetus of starting Databricks for Good. Um so, what is Databricks for Good? Well, it delivers three key pillars. The first is pro bono services. Um there's two types typically. One is we bring in an advisor. They'll chat with you.
[02:12] They'll look at your diagrams. They'll uh be on video calls, but they won't actually have access to your environment or typically write code in your environment. And the other type is a hands-on keyboard, which is myself, a forward deployed engineer. What that means is we go in and we actually set up repositories, build
[02:28] implementations, and teach along the way so that you guys are able to take the implementation and run with it and not be dependent on services. Uh next up, this is a uh vertical that we've launched really, really recently, and I'm super excited about it, but we
[02:43] look to give pricing discounts to nonprofits. And sometimes we're able even to comp a lot of the uh DBU threshold. So, not 100% so far, but we're trying to get our execs to sponsor this, and we've been making a lot of headway. Um and then finally, we do work at a
[03:00] for-profit organization. Uh I wish we could just do good left and right, but there are externalities that are required to have leadership approve of these initiatives. So, one of them is doing joint PR and marketing. Um we work very closely with nonprofits to create marketing deliverables and then also
[03:17] have Databricks uh endorse, approve, and share broadly within social channels. So, that's basically what Databricks for Good is. If you have any questions, feel free to come up to me at the end or go to forgood@databricks.com. That's our uh email address. If you have
[03:32] any questions or you're a qualified nonprofit, we'd love to work with you. Over to Nico. Cool. So, we are uh health and nonprofit. We deliver health care in developing countries. And so, why do we have a data
[03:47] project? So, over 20 years of existence for Virtu Foundation, countless sent to different locations across the globe led to the realization that there is enormous inefficiencies and slightly better data, slightly better decision-making could actually have a
[04:04] really large impact um over the resource allocation and people deployment in developing countries. And so, you know, we're at we're in an era where there's like over approximately 150 million people waiting for surgery in low- and low-middle-income countries.
[04:21] And so, this is a problem that uh Virtue Foundation is looking to tackle to this through this great missions, but also through um making the knowledge around, you know, the global situation available. Uh we believe that uh healthcare
[04:36] philanthropy follows market rules. And so, from that premise, you know, we seek to impact the shape of medical deserts through things like deploying tricycle ambulances. So, that's like accessibility and transportation. Or figuring out where to develop new uh
[04:51] capabilities at existing facilities, or developing new uh healthcare centers altogether. So, for example, where to build the the next surgical center. And the vision for the data project is as such. So, the public internet constitutes uh a an enormous source of
[05:08] data. Very noisy data, but you know, dealing with like the lower end of the economic development spectrum, you got to do what you can find. And so, it's like a pragmatic approach. In which plugs our uh foundational data refresh project uh built with
[05:24] Databricks, of course, Michael and Priyanka going to uh dive in depth. And this foundational data, in turn, powers applications that help uh steer incoming volunteers, that help identify development opportunities at a country level, uh at a regional
[05:40] level. And um and this is, of course, I mean, all tied with the analytics layer being able to uh to define and figure out where are the medical deserts. And we'll have we'll talk a little more about that later. Handing it over to Michael.
[05:57] I'm going to take that one, actually. Clicker. Thank you, Nico. Awesome. So, what we're working with fundamentally is a resource allocation problem. Virtue Foundation has medical volunteers and doctors, and they need to be they need to know where to go to make the best use of their specific skill set. They also have the ability to build
[06:13] new facilities and surgical centers. But how do they know where their resources should be deployed? So, the answer to this is that we need a comprehensive view of healthcare infrastructure. So, what does that mean? Right now, if I asked you if there's an ICU within a 200-mile radius of us,
[06:29] that's a pretty easy question to answer. You would just do a simple web search. But what if I asked you to prove that there was no ICU within a 200-mile radius? That question is a lot more difficult because it requires you to have a comprehensive view of all of the medical facilities near us. And that is
[06:45] exactly what Virtue Foundation and Databricks partnered to create, the world's largest open-source healthcare infrastructure data set, powered by pipeline that we call Foundational Data Refresh. So, let's step into it. Foundational Data Refresh is powered by two main incoming data sources. The first is Oversure, which is an
[07:01] open-source maps data set, and the second is Bright Data. Bright Data is a platform that we've partnered with that provides us with web scraping and crawling capabilities at scale. This means that we're able to scrape web pages from all over the internet and then use that data to build our data set.
[07:17] Both of these data sources are ingested into Databricks via standard structured streaming with Delta and Spark, and then transformed into a schema that is compatible with LLM-based analysis. That brings us to the next step, data processing. This is the GenAI heavy portion of our pipeline, and it's where we use an LLM
[07:34] to extract information from all of the web pages that we've curated. What does this mean? So, on a given web page, you can have a facility, multiple medical facilities, or NGOs. We want all of the information about those facilities and NGOs curated in a structured format such
[07:49] that downstream actions like an agent or a platform can utilize that data. So, what that looks like is using an LM to extract information like facility names, contact information, street addresses, what medical specialties are present at that facility, equipment, and more.
[08:05] Now, you might think that the work ends here. But, how do we know that two facilities across Oops, sorry. How do we know that a facility on web page X is the same as this a facility on web page Y? Or that they're different? Or that they're not a facility that's already in our data set?
[08:21] The answer to this is something called entity resolution. Where we have to reconcile attributes across multiple facilities to see whether or not they refer to the same source entity. Once that's done, we have a deduplicated version of our data set that contains a unique row per medical facility or NGO
[08:37] that is then fit for consumption on the Virta Foundation layer. The Virta Foundation has two consumption layers. The first is VF Match, a platform that Nico will speak about in more detail later, and VF Agent, which Michael referenced, a deep research agent that allows you to ask and ask questions about the data set.
[08:53] So, what some of you may not know is that last year Michael and the VF team actually presented this exact architecture blueprint. And over the last 12 months, we were able to scale it to a production grade system. Now, when doing something like this, inevitably we're going to run into some challenges. So, next I'd like to step
[09:09] through three main challenges that we encountered while scaling out this pipeline and how we solved them on the Databricks platform. So, just to set up the problem here, how exactly are we finding healthcare facilities on the web? The answer is we generate search terms and then crawl the web pages that show
[09:24] up. It sounds pretty simple. Let's walk through an example. So, in this case, we have a location marker that we've matched with a medical keyword. In this case, critical care medicine charity Angola. We do this across multiple medical keywords and multiple location markers to generate
[09:40] search terms. We then use Bright Data's SERP API get the top 10 URLs per search term, and then Bright Data's web crawler to crawl those web pages and extract information. But this is only a single facet of all of the medical care that is available in Angola. So, we have hundreds of terms
[09:57] per country, 72 countries supported by the Virtue Foundation's mission, and about 300,000 terms that we searched uh sourced from Overture, which is the open-source maps data set. Also, we perform a refresh of existing URLs in our data set. The goal is to not
[10:13] only surface new medical facilities, but enrich our understanding of existing ones. So, all in all, we're looking at something like this. 5 million web pages per refresh of our pipeline. That's the scale that our data ingestion has to handle.
[10:28] And our current ingestion breaks at that scale because a monolithic Databricks job is absolutely certain to time out when it's waiting for hours, if not days, for Bright Data to return its results. So, when we're processing millions of URLs, how do we solve this problem? How
[10:44] do we enable ourselves to perform everything with broken-up tasks, and with part and with recovery? The answer is LakeFlow Jobs. LakeFlow Jobs maps the asynchronous life cycle of web crawling into a DAG, a directed acyclic graph.
[11:01] This means instead of running one long continuous notebook job, we break it up into multiple tasks, each responsible for an atomic unit of work. So, task one, the submit tasks, submit submits Bright Data um batches to the API, and then retrieves the batch IDs, and then
[11:18] stores them in a table. And job two, the polling job, then just pull um runs a Python loop over those IDs, and checks to see whether or not those batches have finished. This means that if there is a failure anywhere in the pipeline, we're able to just restart the pulling mechanism
[11:34] without having to do our entire job from scratch, which would mean that we lose all of the web pages that we've collected so far. Additionally, what we use here is a Delta Lake status tracking. This means that in between the two tasks, we have a table that a task is reading to and the other task is reading
[11:49] from. And that's because the tasks on their own, two direct tasks that talk to each other, aren't enough to maintain the state of our back of our batch pipeline. Because if task B, the polling task, were to terminate before it was done, we would still lose all of the data. We need an external table here to
[12:05] track all of the batches that we're submitting and what their status is. Next, I'll talk about information extraction. So, information extraction is difficult for reasons that go beyond just the infrastructure. To prove just how hard, let's walk through an example.
[12:20] In a show of hands, how many people in this room can tell me the name of the hospital on this web page? It's not a trick question, I promise. Okay. What if I asked you for the contact information, a phone number, an email address?
[12:36] Okay. Now, what if I asked you all of the equipment present at this hospital, the procedures that that equipment enables, and the medical specialties that doctors would have to possess to volunteer that hospital?
[12:53] Not so easy now, is it? And so, that is exactly the problem that we're trying to solve with information extraction. Given a web page, how do we extract not only the obvious stuff, a facility name, contact information, how do we extract the more complex features, and what are the insights that we can derive from them?
[13:08] So, this level of complexity breaks the single-shot model. An LLM cannot handle the as complex of a task as this one in one go. And that's because as our inputs approach the context limit of LLMs, performance degrades. Beyond a certain threshold,
[13:23] the more information you give it, the worse it's going to do. And this applies to a large breadth of tasks as well. If I ask an LLM to do many things, it is almost certain to do each of them worse than if I had asked it to do them one at a time. And so that is exactly our target
[13:39] architecture. A sequence of focused calls, each with an atomic unit of work that do a certain part of the extraction task. And so here's what our architecture actually looks like. First, we extract organizations. From each webpage, we get every unique
[13:55] facility and NGO present. That's it. That's all we do in step one. Then, we perform information extraction. So, that's the task I had you do earlier. Can you get me the facility name? Can we get contact information? Can we get a street address? Etc. After that, we extract medical
[14:11] specialties. So, the Virtue Foundation team has actually developed a hierarchy with our subject matter experts that map medical specialties into multiple domains. So, we feed that entire hierarchy into an LLM and ask it to classify the medical specialties on a given webpage.
[14:26] Then, we perform free form extraction. So, this obviously, not all fields that we want from a specific facility can be formatted into a structured output. There's just some information that's going to spill over. And so, we extract three main free form fields: medical
[14:41] equipment, medical procedures, and medical capabilities. That aims to cover basically everything in a facility that doesn't fit one of our like pre-templated structures. After that, we extract metadata. So, we essentially want to get confidence scores for each of the webpages that
[14:57] we're scraping. How do we know that the data we're getting is actually good? And in certain low- and lower-middle-income countries, it becomes increasingly difficult to gauge whether or not a website is, for example, up to date. And so, in the metadata extraction step, we extract things like how recently a page was updated, whether or not it has
[15:13] external links, and how many of those. After all of these extraction steps, we're able to have a unified view of each of our facility pages. We also implement content hashing, which means that every time we run the pipeline, we're not re-scraping all of our URLs. We're able to see whether or
[15:29] not the hash of a particular web page has changed and update accordingly. And honestly, I could talk forever about the problems that we've encountered with information extraction. Every time it felt like we'd fix the pipeline, a new problem came up. But, I'd like to step through one of the most memorable issues we encountered.
[15:45] That was OpenAI limits. We hit OpenAI's project-wide 2.5 terabyte storage quota multiple times. And OpenAI's 15 billion NQ token limit. Both of these sound like big numbers. I know they did to me. They're not numbers
[16:01] you ever anticipate reaching, especially not with projects like these. But, that's kind of the whole lesson here. Is that a lot of what separates a proof of concept from a production system isn't clever architecture, it's building guardrails around problems like these. And so, for us, the solution was just write a couple of scripts. Make
[16:17] sure you don't hit the retry or at the token limit before you submit more jobs, and clean up the files whenever you hit the limit. And so, our solutions aren't always as glamorous as the ones you see here, but I think that comes part and parcel with just building a production system. So now, after all of this, our
[16:33] information extraction pipeline is stabilized. We've got extracted rows from millions of web pages. A lot of them though are secretly the same place. So, how do we figure out which records point to the same source truth? That takes us to the last step in our
[16:48] pipeline, primary key resolution. Now, up here, you see two records for medical facilities. We, as humans, can see that they point to the same source, right? They're both pointing to the same clinic. But, if you look closely, there's a couple of discrepancies. The name is spelled
[17:05] differently. The street address is reordered. The locality, one has an abbreviated country, one does not. The phone numbers, thankfully, are the same, but the websites are or And the coordinates are only similar up to a certain degree of precision. So, how can we teach a machine that
[17:21] these two actually refer to the same place? The answer to that is Splink. Splink is a probabilistic record linkage framework. That's just a really fancy way of saying we figured out what attributes two records have to have in common to deduplicate to the same source
[17:36] facility and then encoded those attributes. The key here is that Splink allows us to give each of these attributes a weight. And when those weights add up past a certain total, we can confidently say that these two facilities are the same. This means that each of the attributes
[17:52] contributes a weight and no one attribute can decide. And so in this case, name would get a weak positive because it mostly matches but not entirely. Address would also get a weak positive, but website would get a weak negative. And then phone number contributes a
[18:07] strong positive score. So, in this way we're able to accumulate all of the weighted evidence and decide whether or not two facilities are the same source. Now, if I wanted to compare every facility against every other facility in our data set, that would be about 49 trillion
[18:23] comparisons. If you have a couple years to wait, that's fine. We didn't. And so the solution here is blocking. Blocking is essentially where Splink lets us split up the comparison into two stages. The first stage is the blocking stage
[18:39] where we create a minimum set of attributes that two records have to have in common to refer to the same source facility. So, in our case we decided on three attributes: a name representation of some sort, a location representation of some sort, and finally country. So, if these three things did not align,
[18:56] then the records were deemed not a match even before they were fully compared. Then, once pairs make it past the blocking stage, we can actually do a full comparison and deduplicate. But, blocking has its own problems, of course.
[19:11] And this is because blocking creates partitions in Spark based on our most common attributes. So in this case, a name and location representations are fuzzy matches, but country is not. We encode ISO country codes, and the country must be an exact match for two records to compare.
[19:27] So that means that when we block, all of our Spark partitions are based on a country basis. And naturally, one country is a little bit bigger than the rest. In our case, India. So more than 40% of our data set contains records for India, which means that when we wanted to block, all of
[19:43] those records were forced into the same partition. And so when we ran our primary key resolution job for the first time, every other partition finished in under an hour. India ran for almost a day, and then we made the decision to kill the job entirely.
[19:59] And so the solution here is actually two things. One is salting. Salting is a functionality that Splink offers that allows you to split a hot partition into multiple smaller ones. So essentially, instead of doing all pairwise comparisons in one partition, you distribute them across multiple
[20:15] other ones. Additionally, Photon, the Databricks vectorized execution engine, allows us to speed up columnar operations. So both of these combined allowed us to take a job that initially ran for 23 hours without completing to a job that
[20:30] finished in under an hour. So salting fixes our distribution, and Photon speeds up the work. Now, what does fixing all of these problems actually unlock? Surely, there's a point. And the correct answer here is that it
[20:47] allows us to expand the ceiling on each of our runs. So every time we solved one of these problems, we were able to generate more records and more create a larger data set for Verity Foundation. And solving these challenges didn't actually change anything substantial about our pipeline. It didn't change in
[21:03] what it does, it just changed in and it could handle without failing. And you can see that in our numbers. So, solving each of the scaling challenges I just enumerated across data ingestion, information extraction, and primary key resolution directly unlocked a bigger run for us.
[21:19] The first one added 17,000 records. The second 79,000, and the last run added 290,000 new facilities and NGOs. But wait, we said there were millions of web pages and millions of extracted
[21:34] rows. So, where did that information go? There's two answers to this. The first is deduplication. That's how we know that it's working is we have more rows coming in than we do going out. But also, not all of the facilities that we've extracted from the web are suitable for ingestion on the data on
[21:51] the Virtue Foundation platform. And that's because there's a certain set of criteria that the Virtue Foundation need for a record to be publishable. We require a valid name, a valid set of medical specialties, and the kicker, we require coordinates. That's not something that a lot of
[22:06] facilities publish on their websites. And so, a lot of our records are lost simply because they don't meet the publishable criteria. But at the end of the day, we were able to curate more than a list of 400,000 records of medical facilities and NGOs across 72 countries. Compared to the manually
[22:21] manually made data set originally that had 30,000 records. So, what does this actually look like at a global scale? We have more than 72 countries covered in our data set. More than a hundred medical specialties developed with in conjunction with Virtue Foundation. And 400,000
[22:38] facilities and NGOs. So, the number of NGOs chapters we've discovered is actually somewhere around 20,000. And these are the organizations that are actually doing work on the ground. So, this data set that we just talked through powers everything you're about to see next, and I'll hand it off to Nico.
[22:54] Thank you, Priyanka. So, yeah, so I as I alluded to in the intro, so this data set, of course, is groundbreaking in terms of you know, enabling a number of applications, and one of them is VF
[23:11] Match. So, that's our um matching platform meant to help connect medical professionals with the most relevant opportunities and with the most relevant locations where to conduct medical missions. So, um
[23:28] on VF Match, you can actually browse medical deserts. So, this is something that is derived from FDR and augmented using other publicly available data layers. So, looking at population density, looking at what the capacity of the facilities are like,
[23:43] and basically highlighting on a map the the areas that would have good coverage, and outside of those, looking at the population density. So, that's what you see here in red. You know, on VF Match, we can actually overlay that now with the locations of facilities, the
[24:01] locations of local NGOs, and we can find opportunities. So, that's like kind of the end goal of the platform, to connect and enable the global health market. And one of the novelties that FDR has enabled is to look at medical deserts
[24:18] from a much more granular and specific angle. So, we're not just looking at, you know, a high-level medical deserts, but we can look at it in terms of specialties and groups of specialties. So, for example, you may find a very different the very different picture looking at imaging, looking at women's
[24:33] health, looking at cancer, etc. So, here I picked imaging. And so, if I'm looking at, you know, general surgery, like simpler procedures, actually Ghana is pretty well off. But, if I want to consider more advanced procedures, I see that the north is
[24:49] lagging behind, and mainly the capital city area is has enough coverage. And so, these are some of the insights that the new FDR data sets enables to to derive.
[25:06] Cool. So, this was a pretty cool talk. Um I want to highlight again that this was via the Databricks for Good program. Uh we gave pro bono services to work with Virtue Foundation through some of our best people such as Priyanka at this problem. And it's a hard problem. Like it's really really challenging. I've been at
[25:22] Databricks for 4 years. Again, working on some of the biggest scales, and this blew my mind. Like exceeding the Open AI 2.5 TB limit multiple times in one run. Like that's just ridiculous. It's It's a lot of data. Um so, if you have any questions about a
[25:39] similar for good uh mission that could use this technical expertise, um we'd love to hear about it at forgood@databricks.com. And then also just taking a step back, all of us are wondering why we're at this summit. Um this is hopefully a great example of how the Databricks platform can work at
[25:56] scale uh for you. We have a bunch of interoperability. So, we worked with Open AI, worked with Bright Data, we worked with Overture. We built a lot of bespoke and custom implementations, but we also used out-of-the-box connectors. It was all governed in Unity Catalog. All of our tables were stored in Delta, the Delta
[26:13] format. Um and so, it's really amazing to see again what this platform can do when applied to the world's toughest challenges. Thank you.

Learn more about the Databricks Data and AI platform.

The information provided herein is for general informational purposes only and may not reflect the most current product capabilities or configurations.