Skip to main content

LSEG Financial Data at Scale: Unity Catalog and Delta Sharing for Governed Data Access

Summary

  • LSEG, the London Stock Exchange Group, manages 75 petabytes of financial data and uses Databricks Delta Sharing and Unity Catalog to give a tier one bank's 2,000+ data scientists direct, zero-copy access to live financial data without ETL pipelines or data distribution copies.
  • Delta Sharing bundles shared notebooks, pre-trained models, and expert guidance alongside financial datasets, enabling data scientists to move from data discovery to working AI prototypes — including sentiment analysis and scenario analysis — in days rather than months.
  • Unity Catalog provides centralized governance, auditing, and compliance controls across all shared data, with Genie integration enabling natural-language discovery and exploration of available financial datasets by teams across investment banking, asset management, and wealth management.

LSEG Financial Data at Scale: Unity Catalog and Delta Sharing for Governed Data Access

Watch: LSEG Financial Data at Scale: Unity Catalog and Delta Sharing for Governed Data Access
LSEG, the London Stock Exchange Group, manages 75 petabytes of financial data that must be shared securely with regulated financial institutions. Using Databricks Delta Sharing and Unity Catalog, a tier one bank unified its 2,000+ data scientists across investment banking, asset management, and wealth management divisions onto a single governed platform.
Learn how Delta Sharing enables direct, zero-copy access to live data without ETL pipelines or data distribution copies, while Unity Catalog provides centralized governance, auditing, and compliance controls. With shared notebooks, models, and expertise bundled alongside datasets, data scientists can prototype in days instead of months, fail fast, and deploy AI workflows like sentiment analysis and scenario analysis directly on regulated financial data.
🤝

Chapters

FAQs

What is LSEG and what financial data does it provide?

LSEG, the London Stock Exchange Group, is the largest provider of financial data and manages 75 petabytes of financial data — equivalent to approximately 15 million films in terms of storage volume. LSEG delivers this data to financial institutions through multiple technology channels including Databricks Delta Sharing for direct, governed access.

How does Delta Sharing enable zero-copy financial data access?

Delta Sharing allows the data provider to grant direct read access to live Delta tables stored in the provider's cloud storage, so the consumer can query live data without it being copied or moved. For the tier one bank in this video, this eliminated ETL pipelines and data distribution copies, providing governed access to live LSEG financial data with compliance controls maintained by Unity Catalog.

How does LSEG help data scientists prototype AI faster using Databricks?

LSEG bundles shared notebooks, pre-trained models, and expert guidance alongside the financial datasets shared via Delta Sharing, giving data scientists a complete starting environment rather than raw data alone. This approach allows the 2,000+ data scientists at the tier one bank to move from data discovery to working AI prototypes in days rather than the months previously required.

What governance and compliance capabilities does Unity Catalog provide for regulated financial data?

Unity Catalog provides centralized governance including fine-grained access controls, audit logging, and lineage tracking across all shared datasets, which is essential for financial institutions operating under strict regulatory requirements. In this video, it enables LSEG's tier one bank customer to maintain compliance controls while giving data scientists across investment banking, asset management, and wealth management unified access to shared financial data.

Full transcript

[00:08] Well, thanks a lot everybody for coming. We we know there's some competition out there. So, we're going to start with a number. 75. Think of the number 75. What might that mean? What could it be? I mean, the sun's coming out now, so
[00:23] maybe it's going to be 75° Fahrenheit in a bit. Uh maybe after this conference you're going to have today you're going to have 7,500 emails that you'll have to go through. Or maybe that's the the time of your commute from your home to your office if
[00:38] you're still going into the office. 75 actually for LSEG is the petabytes that we have in terms of financial data. We we collect data as a matter of course.
[00:53] LSEG, which is the London Stock Exchange Group, is the largest provider of financial data uh currently on the market. So, 75 petabytes, if you think of it, if you want to translate it into a different way of thinking about it,
[01:10] it would be about 15 million films on Netflix. Which would take you, if you watch these films consecutively, I think 1,700 years if you watch it 24 by 7. So, you can just imagine the
[01:25] amount of data. And that's what we're going to be talking about today. And actually it was really nice listening to the keynote this morning because data was emphasized so much. So, we want to talk about the data and the technology and solving customer problems.
[01:40] So, let's go quickly through the agenda. So, LSEG and Databricks have a partnership. We're going to talk about that. We're going to talk about AI-ready data that LSEG deliver via multiple channels of technology.
[01:57] And then the heart of the presentation is actually the customer challenge. We want to talk to you about a customer that came to Databricks and LSEG to say, "We have a challenge. We have a problem. Can you help us with this? Can we leverage the technologies? Can we
[02:12] leverage the data?" And then Kriti is going to go go through the actual solution around Delta Share, Unity Catalog, and how we actually built for this customer the solution. And then interestingly enough, once
[02:28] again, if you've heard the keynote, which I think most of us did, talking about the framework around governance and how incredibly important it is, particularly in this AI age. So, and then if we have time afterwards, uh we'd be very happy to have a Q&A. Or
[02:43] we can if if we run out of time, we can always pop outside as well. But very very happy to have a Q&A. So, let's kick off. Let's talk about the partnership first. The partnership between LSEG and Databricks.
[03:00] What's interesting is actually LSEG has been using Databricks for quite quite some time now. But what was happening is that more and more customers were coming to us and asking us about putting data into the Delta Share in the Databricks. So, LSEG and Databricks officially announced in Q3 last year
[03:17] the actual partnership. But LSEG has been a customer for quite a while of Databricks. What's been fantastic is the amount of interest and attention that we've received from this. We have many customers who use Databricks. Customers such such as the tier one type
[03:34] of bank that we're going to go through today, but also a lot of the quant hedge funds. So, it's really interesting. You get the whole spectrum of of of customers using this. And really what it's about is about building a partnership, building a platform that allows for the prototyping
[03:51] and the acceleration of AI processes as well as other downstream applications, pricing engines, etc. The use case that was actually that we're going to present, the tier one bank, the challenge that they brought to us is the fact that first and foremost they
[04:07] have 2,000 data scientists, probably 2,000 plus data scientists. And they're across many of the actual ambitions, if you will, within the bank. A bank often has multiple divisions, the investment banking side,
[04:23] asset management and wealth management. And they have over 2,000 data scientists across the board. And what they've asked us to do is say, "How can we actually build a utility which we can hydrate with actual LSEG data to allow for a much easier path in
[04:40] terms of building applications, in terms of building AI agents and bots?" So, that was actually what what what what was presented to us. We'll go into that a bit further as well as to why this was such a challenge for them. But first what I want to do is just briefly talk about the data.
[04:57] And once again, it was interesting to hear the keynote this morning because we heard about the enterprise context and how everything's very contextual. So, what LSEG have in the first instance is three pillars of the AI data.
[05:12] We call it the first pillar being the trusted data. Trusted data means you can actually it's auditable. You understand the lineage. You understand where it's actually coming from. There's so many hallucinations that we know about
[05:28] and actually whether in particular if you're trading if you're trading off this data or if you're actually communicating to a downstream customer, you really have to have trust in the data and the accuracy of the data. And that is one of the biggest things that we really talk about. We also talk
[05:44] a lot about the fact of the processes and the AI readiness of the data. What does that actually mean? But if you think about it, AI as an actual application or a downstream is actually not so incredibly
[06:00] intelligent. If you throw tons and tons of data at it, bulk data, raw data, it has to really decipher what what to do with it. And as you heard perhaps this morning, you know, the cycles it goes through of course consumes quite a few tokens, etc.
[06:17] What we do with our data at LSEG is it's heavily curated. Curated means it's checked for the accuracy, for corrections that come in from the market. We also add metadata to it so you can understand location, where it's traded,
[06:33] what symbols it might be traded under. Entity data, how a customer might be seen from a symbolic perspective on that. So that is really the the the the bulk of the data. So the tagging and
[06:49] the metadata is a phenomenal part of what we call AI-ready data from LSEG. Then of course we have the infrastructure, how we distribute it out. So it's not just the data, it's the distrib- distribution. A lot of that distribution can be internal to the customer
[07:05] or indeed they can use partners. LSEG is actually agnostic in this respect. We have a a lot of the cloud providers as you can well imagine. We put our data in all cloud providers and we use all, if you will, the cutting-edge technology and that's why Databricks is one of our our our strong
[07:21] partners in this whole space. Let me go through a little bit longer or more in terms of an architectural schema. And once again, I was quite I was actually pleasantly surprised and very happy when I watched the keynote.
[07:37] Um the CEO of Databricks had something to the the the effect of the same type of schema up there. It's around the data and then how you distribute the data. The distribution in this day and age is is we all know around MCP.
[07:54] For AIs for for AI AI downstream applications, if you will, bots, agents, LLMs. So, of course, we Alsec does have their its own our own MCP. Our customers build their own MCP as well. Once again, we're agnostic. We can
[08:10] deliver data to the customer's MCP or into our MCP. But, lo and behold, a lot of customers still want the user interface. So, we have the user interface which we which we call it's a desktop product called Alsec workspace, which has a lot
[08:25] of AI into it as well. Also, it it it probes AI via MCP. Of course, there's agents and then there the actual feeds. The feeds, of course, we we expose APIs to all of our feeds, so you can get to the feeds that way.
[08:43] Most of the customers have, if you will, this hybrid way of consuming data. Either through products that Alsec deliver, through their own proprietary products, or indeed through partners. We're we're going after many type of requirements to satisfy, you know, the
[08:59] multiple requirements. So, that's where we see it. We see that with Alsec in an AI data everywhere, it's very much part of our, if you will, DNA now.
[09:14] Let me walk you through the customer challenge though. And this is the part that once again, you know, through my long career, I I really enjoyed this challenge cuz I thought, you know what? We can solve this. We can actually figure it out. With technology as a rule, you know, technology sometimes leads leads us too
[09:30] quickly. But, technology is to solve problems, actually, business problems. It's not the it's It's for the sake of the technology. But, let's go through the actual challenge that they presented to us. And perhaps this will be very familiar to a lot of you that that that that are, you know, have a have a look at it.
[09:50] As I mentioned, many of the large universal banks, the tier ones, have multiple divisions. Be it the investment bank, be it wealth management, be it asset management, sometimes retail as well, depending on the banks. Within these divisions, there are other divisions.
[10:06] So, for example, in the investment bank, you've got the effects. You've got equities. You've got fixed income. And then you have two fixed incomes. You have credit and rates. And you can imagine how many different types of, if you will, divisions, or as we put it here, as siloed,
[10:21] silos are there. And what was interesting is that we noticed that every single silo was doing the same exact thing. So, I'm going to take a step back. So, I was um We came in from London a couple days ago. So, a lot of my friends and a
[10:38] few family members were asking me, "So, what are you going to present in San Francisco? What are you going to talk about in San Francisco?" And I said AI. And I said, "Well, yeah, okay, AI. Everybody hears AI every single day." So, I tried to come up with an analogy for them, because they're not really in this world of financial services, of
[10:54] financial data, of technology and distribution. And so, I was trying to think, "What what might be a good analogy?" And then it occurred to me, actually, San Francisco is known for its food. It's one of the food capitals of the world, actually, with the fusion of food, the different foods. Obviously, we've got North Beach, Chinatown, etc.
[11:11] There's so many different types of food here, and diversity of food. So, I thought, "Well, here's maybe a good analogy." Is that if you're having an event, if you're throwing an event, and you hire a Michelin chef, Michelin star chef, and you tell that Michelin star chef,
[11:27] before that chef starts to cook, to actually make the food that you want for the event, you say, "First, you've got to build a kitchen." And then after you build the kitchen, you got to go out and get the tools, pots, pans, bowls.
[11:42] Then after you get the tools, you got to go out and get the ingredients, fresh vegetables, meats, herbs, spices. And so, 85% of this Michelin star chefs time is spent
[11:58] actually providing an environment where that chef then can cook and present the food. 15% of the time is actually making the food. But lo and behold, that's what data scientists face on a
[12:14] regular basis. Procuring data, building environments, and then of course, they've got the Catch-22. All of this costs a bit of money. You don't get it free of charge. So, they have to go to the COO of the business unit. They say, "I need this this type
[12:30] of budget." The COO asks the normal question, "Well, what's my return on the investment?" So, "I don't know. I got to prototype it first." "But it's going to cost you this much money to prototype?" I said, "Yes, but I have to prototype it for the return return on investment." And what happens is this cycle
[12:46] repeats itself in every single business unit. And that's what this large tier one bank came to us to say is that we spend so many cycles, we consume cycles building out an environment, procuring data, arguing with COOs and budget
[13:04] holders about the need for a bit of extra cash so we can actually build it out. And what you have to think about is that nobody actually owns the holistic picture because it all sits in the business units. And lo and behold, it might be
[13:19] different technologies as well. So, that was the challenge that was presented to us to say, "How can you guys help us with Databricks as a utility or to build a platform using Databricks and ingesting it with the data from LSEG.
[13:37] And we came up with a solution. But the solution I'm going to let Kriti talk about because she was very instrumental in it. Hello. Am I audible? Yeah. Okay, cool.
[13:52] Uh so we heard the challenges. So the challenges are too many data pipelines and a very long onboarding cycle and the constant pressure on data scientists to prove the business value before they can even see the data.
[14:08] So the solution is here. The The solution as that uh that's exactly the challenges that marketplace solves for you. So think of marketplace as a controlled B2B data sharing solution.
[14:25] So here LSEG publishes data to the marketplace and the consumer subscribes and consumes the data. So on the left-hand side you have LSEG. So they remain responsible for the data. So they take care of curating the data,
[14:42] they manage the data quality, and they manage the freshness. So they publish the data into the marketplace. So marketplace they are they act as the distribution channel. So they actually share publish the data
[14:58] in a secured way to the consumers. The most important thing here is that you don't just get to share the tables, you can share lot more assets. For example, you can share notebooks, models, agents, etc.
[15:13] So let me give you an example. Say if LSEG is publishing a data set, they can also publish a notebook along with it that actually demonstrates a how to derive insights from the data. So, think about the data data scientist experience here. So, the data scientist
[15:30] they get the data and they also getting a worked example along with the data. Or say Alteryx has got a model that actually identifies a particular market signal. So, data scientist in this case, they can actually deploy the model in their
[15:46] own environment combined with their own proprietary data and derive insights right from the get-go. So, here the data scientist is not just receiving the data, they are also receiving the expertise and reusable
[16:03] analytical tools along with it. On the right-hand side, the consumer is able to see all of this data in their own Unity Catalog. So, this enables them to join with their own internal data
[16:19] and apply their own internal governance to that. This makes it a very powerful experience. So, the value here is very straightforward. You're sharing a live data. No ETL pipelines.
[16:34] No expensive infrastructure project. So, this is going to enable faster innovation for you. Number two, there is a clear division of labor. So, Alteryx curates the data.
[16:49] Marketplace takes care of the delta sharing. The technology takes care of the delta sharing. And the consumer consumes the data. And number three, you have governance on both the sides. So, you have Unity Catalog on the Alteryx side and you have Unity Catalog on the consumer side. So,
[17:06] your governance is taken care. So, you are actually meeting your compliance and regulatory requirements right from day one. Now, let's look at the technology that actually enables this.
[17:22] So, traditionally, when it when we talk about data sharing, what does it involve? It involves distributing data copies. For example, today you might be distributing data copies via SFTP server, or maybe you are emailing the emailing
[17:38] or Dropbox or even S3 buckets. So, the problem with that approach is that with the distributing data copies, the data can easily turn stale, or you have security risk, there is no audit trail whatsoever, and it is
[17:55] expensive. So, these are the challenges that Delta Sharing eliminates for you. So, the different between the traditional approach and Delta Sharing is this. So, we are not distributing data copies. What are we doing instead?
[18:11] We are giving you direct access to the data right where it is stored. And that means there is zero data copy. So, think of an analogy. Say it is a similar experience to watching a Netflix movie.
[18:28] So, when you're watching a Netflix movie, do you download the movie? No. You just watch the movie when you need, and then you move about. So, there is no data copy. That's exactly how Delta Sharing works for data. So, there is no data copy. There is no
[18:43] need to build ETL pipelines, and the data is live data that you are having access to. So, if you think about the flow, the provider is having full control and ownership of the data. Delta Sharing,
[18:59] which is part of Unity Catalog, takes care of authentication and authorization. And the consumer, they can receive this data from any platform, any tool, any cloud, any region. So, they don't need to be on Databricks
[19:15] at all. Let's double click and deep dive into how it actually works under the hood. So, here I'm going to focus on Databricks to Databricks sharing. So, firstly each of the organization they're
[19:31] operating in their own environment, in their own network boundaries. So, LSEC is using their own Databricks workspace and the consumer using their own Databricks workspace within their own network boundary. So, first step is to establish a trust between both the Unity Catalogs.
[19:49] And this is a very simple step. It's one-off and it's part of the onboarding process. Once the trust is established, now the consumer can start seeing the metadata on their end. What that means is now they can start this start seeing the schema structure,
[20:07] the table names, the description of the columns, all the metadata. So, it's only metadata at this stage, no data has moved yet. So, step two, this is where the elegant part begins. When the customer is ready to query the
[20:23] data, the request goes to the provider Unity Catalog. Unity Catalog authenticates and it checks what data exactly do you have access to. Depending on that, it generates a temporary token that is scoped controlled
[20:40] and it also generates a short-lived URL and returns back to the consumer. Now, the third part is consumer uses their own cluster compute to access the data directly from the storage using this temporary
[20:55] credentials. That makes it a very secure data sharing. So, if you think about the data, the travel how data traverses, there is no proxies, there is no hop. It's a direct access to the provider's storage. So, that makes
[21:12] that gives you full performance. And also, if you pay attention to the network, because we are talking about regulated industries and banks will ask, so the whole data is getting traversed to through private link, and it is controlled by your firewall controls
[21:30] through network of proxy gateways. So, that makes this whole solution secure by design. Also, remember the auditing? So, all this whole solution is logged on both the sides. So, you have logging
[21:45] LSEG exactly knows who's accessing my data, and the consumer exactly knows who's who's consuming my data. So, it's logged on both the end.
[22:03] Now, I'm going to switch gears and show you what it feels like to be a consumer of this experience. Here, you are receiving data that is AI-ready, which means it's enriched with context. And context is very important
[22:19] for any AI tools to make sense of And because the data is AI-ready, you can right away use this data in Genie. So, if you remember from the keynote, Genie is Databricks' AI conversational
[22:34] tool that allows you to ask a questions in plain language and get answer back. And that works directly on your shared data set. So, use Genie right away and generate value.
[22:53] So, what you're seeing on the screen is the breadth of data that is available for you from LSEG, and customer wants even more. To give you an example of the use case that we are seeing is a quant analyst can use the pricing data from LSEG and they can run exposure analysis and
[23:11] scenario analysis within minutes. So, no ingestion pipeline, no data pipelines, they can generate value right away.
[23:27] Now, let's bring everything together. What makes this whole experience a powerful is Unity Catalog. Because you're going to see all the data right next to your own internal data. And that unlocks discoverability right away. So, your team can discover what
[23:44] you've got access to immediately. Now, they can go about and build AI workflows using all the AI tools and capabilities that the platform offers. Visualize your data and share with your stakeholders using Databricks One tools
[24:01] like Databricks One that you heard about in the keynote. And now you can embed all of these insights in your business applications and your operational workflows. Now, if you think about it, this whole life cycle, right from
[24:18] discovery to em- embed, this whole life cycle is happening in one platform with Unity Catalog governance. So, you have got full lineage, auditability, observability, and the governance through Unity Catalog, which means you
[24:35] are compliant with the all your regulatory and compliance requirements right away. Now, let's talk a little bit about scale at which this can be rolled out.
[24:50] So, LSEG doesn't have just one big monolith data team. So, they have got multiple data teams who have got multiple data products and they are spread across domains. And sometimes these because it's a global company, these data teams, they are spread across
[25:07] the globe. They operate in different geo regions and they sometimes even operate in different clouds. What stitches them together is the very mechanism that we are talking about here, that's Delta Sharing. So, LSEG is using Delta Sharing to share data
[25:24] between teams and to enable collaboration. So, what does enable them is even though they are operating in a federated environment, they still have a centralized governance which is very important for a financial and regulated
[25:39] industry. So, this makes the whole architecture modular and you don't need to re-architecture anything. You don't need heavy investments and the in this results in time to value. So, you build the data products,
[25:54] go live with it, and distribute it within few days. The best part about this architecture is that it doesn't stop at the provider's side. This architecture can propagate to the consumer's side as well. So, on the consumer, the same architecture can be
[26:11] used to distribute the data to their internal teams. So, the consumer's risk team, the trading team, and the asset management team, they can all leverage LSEG data using the same architecture.
[26:27] So, I'll hand over back to Bill. Thank you. I you can see why we're having so much success with this customer. I mean, Akriti really knows her stuff. It's it's fantastic to to learn. I mean, the amount of learning I do every day as
[26:42] well on this. Um the keynote, once again, I'm going to reference the keynote because it was nice that it almost felt like it was our presentation that he was presenting up there. So, I thought, oh, somehow we've got some synchronicity going on here, but it is around the operating model and the
[26:58] governance, and it's really, you know, super important. We actually, um, met Databricks and LSEG met with a customer yesterday. Um, and we were talking about they're they're using LSEG as a reference customer. It was it was quite
[27:13] interesting, and they kept asking us the same question. What are the benefits? What are the benefits? And what are the benefits? And it's not that we struggled to articulate what the benefits are, but we we we talked about it in a more holistic way versus just a we have A, B,
[27:29] C, and D. One of the biggest benefits that we've seen, particularly with this large universal bank, is more, I'm going to use the word unification of thought processes. People are actually buying in. I mean, a lot of times technology is more more of
[27:45] a political thing than a technical thing. People are buying into the platform. We have We have some internal customers coming in from Chicago saying, "We've heard you've got this platform in Europe. Can we connect into it?" We have internal We have customers coming in from Asia saying the same type of thing.
[28:00] They're actually buying into the idea. And in many respects, that's one of the largest benefits. It's much more of a political than actually a technology, but the the the fact that they can leverage the technology. Krity talked about the the maintenance and the freshness, the curation of the data on the LSEG side, which keeps it
[28:17] fresh. The amount of corrections that come in every day is astounding. And LSEG have multiple data teams in the thousands actually looking at the data and correcting it. Because it just happens in the markets like that, particularly during the auctions, if anybody knows about the
[28:33] auctions there. The other big benefit that we've been seeing is the time to prototype. As I mentioned, the the budget holders keep asking, "Well, where's the return on investment? You know, how do we ascertain the ROI?" It's like, well, we've got to we've got to figure out the
[28:49] we've got to prototype it first. With this type of platform, they can actually fail fast, to borrow from Silicon Valley. They can actually say what will work and what won't work on that. And then, of course, you've got the governance and control. The Unity catalog is is
[29:07] so phenomenally important, but it's the permissioning, particularly around AI. AI is consuming so much data. The ingestion of data is phenomenal, but you do need to have the permissioning behind it, entitlements behind it. The regulatory and compliance part of
[29:23] any large bank is huge. Yeah, it takes 20 as they say, 20 years to build a reputation and 1 day to lose it. So, it's all around the governance there. What you notice as well is the immaterial. Because this is a lab
[29:38] environment, we're putting immaterial data in there. The immaterial data means from the markets. The material data is actually their own customer data, which we can't put into a research lab, but of course, you can put it into production where you have the the right the right security around it.
[29:54] But, it's really around these parts here. So, today we read in the newspaper day in and day out about all the data centers being built. All the amount of data centers for the
[30:11] compute. We also hear about if you will, bidding wars for data scientists. We've come to realize that that's not the actual challenge with the banks is the compute or the actual data scientists.
[30:28] It's really around having that platform, as I've mentioned. Having a platform that they can access immediately and actually start to work on their secret sauce. That's really where where where we're trying to get to is to say we have a
[30:44] unified platform here. We have a platform that crosses divisions where your data scientists can plug into. The data is already there. The Unity catalog allows for further permissioning of data, so it's not a free-for-all, but they can actually it's
[30:59] time to prototype, not just necessarily time to market or time to value, it's time to prototype. And that's that's the that's the big thing. So we keep looking at it and we keep saying, "Well, what is the barrier?" The barrier is not the talent. The barrier is not the compute.
[31:15] The barrier is the fact that nothing is actually being built for them and they have to rebuild everything. So this is what we're doing with this large bank. Um it's going really well. And we we we have more and more demand within this bank
[31:31] for the the data scientists to come to connect into it. So we see that very much as a mark of success. Good. We're going to stop there. Thank you, Bill, and thank you for your time. Please fill out your surveys in the app and have a great summit.

Learn more about the Databricks Data and AI platform.

The information provided herein is for general informational purposes only and may not reflect the most current product capabilities or configurations.