Building an Interoperable Enterprise Lakehouse: Goldman Sachs' Iceberg-First Architecture
Summary
- Goldman Sachs built a fully interoperable lakehouse with Apache Iceberg and the Databricks Data and AI platform, establishing a single source of truth across their global financial services operations in 80 countries without data duplication.
- The 'refiners' component orchestrates thousands of data transformation jobs across hybrid on-premises and cloud infrastructure, with approximately 4,000 refiners migrated from on-premises systems to their Iceberg lakehouse.
- Zero-copy Delta Sharing with external clients reduces operational complexity and unlocks new data collaboration capabilities, all governed through Unity Catalog as part of a data strategy that originated from lessons learned in the 2008 financial crisis.
Building an Interoperable Enterprise Lakehouse: Goldman Sachs' Iceberg-First Architecture

Financial services enterprises face the central data problem: data copied everywhere, leading to inconsistency and risk. Goldman Sachs, processing transactions across 80 countries and 175 jurisdictions, required a strategy that unified all data in one place without compromising security, governance, or performance. Their answer: an Iceberg-first lakehouse with semantic data products and zero-copy sharing.
Discover how Goldman Sachs built a fully interoperable lakehouse with Apache Iceberg and Databricks, achieving single source of truth without data duplication. Learn how semantic data models and data products enable enterprise governance, how refiners orchestrate thousands of transformation jobs across hybrid infrastructure, how Iceberg V3 features like variants and deletion vectors unlock new capabilities, and how zero-copy Delta Sharing with clients reduces operational complexity while maintaining security through Unity Catalog.
🤝
Chapters
00:00Goldman Sachs' Iceberg Lakehouse: Foundation and Strategy01:1315 Years of Data Evolution: From 2008 Crisis to Lakehouse02:52Key Tenants: Governance, Experimentation, and Single Source of Truth05:02Semantic Models and Data Products for Enterprise Data08:20Iceberg Adoption at Scale: All Business Units10:14Iceberg V3: Variants, Deletion Vectors, and Enterprise Features11:36Refiners: The Transformation Engine of the Lakehouse14:54Scaling Refiners: Migrating 4,000 from On-Premises15:57Delta Sharing: Zero-Copy Data Collaboration with Clients20:00Demo: Building Derived Data Assets in Iceberg
FAQs
Why did Goldman Sachs choose Apache Iceberg as the foundation for their enterprise lakehouse?
Goldman Sachs chose Apache Iceberg because it provides an open, interoperable format that supports their entire data strategy without vendor lock-in, enabling all business units to use a single table format. They believe Iceberg is the foundation for the future and have built their complete lakehouse and data strategy around it.
What are 'refiners' in Goldman Sachs' lakehouse architecture?
Refiners are the transformation engine of the Goldman Sachs lakehouse, orchestrating thousands of data transformation jobs across hybrid on-premises and cloud infrastructure on Databricks. Approximately 4,000 refiners have been migrated from on-premises systems to run against their Iceberg tables.
How did the 2008 financial crisis influence Goldman Sachs' data architecture strategy?
The 2008 global financial crisis revealed the critical importance of having all data in one place, as calculating counterparty risk and exposure required overnight batch processing across fragmented systems. This experience set Goldman Sachs on a 15-plus-year journey toward a unified data platform, culminating in their Iceberg-first lakehouse on Databricks.
What is Delta Sharing and how does Goldman Sachs use it with external clients?
Delta Sharing is a zero-copy data sharing mechanism that allows Goldman Sachs to share data with external clients without physically copying or moving data to a separate location. Governed through Unity Catalog, it maintains a single source of truth and enforces security controls, unlocking data collaboration that was previously difficult to implement.
Full transcript
[00:07] All right, thanks everyone for joining us here. My name is Ram Naran. I'm a managing director at Goldman Sachs. I head up our data platform and solutions engineering teams. Hi everyone, thanks for joining us today. Good evening. Uh my name is Abishek Naring. I sit in data engineering at GS as well along with Ram. My day remmit is to run any sort of
[00:25] external market data vendor data hypothesis. So bring it in our goan lakehouse ecosystem which we're going to learn today. All right. So today we're going to be talking about our fully interoperable lakehouse with data bracks and unity catalog. Uh we'll start a little bit around like why we got to where we got to. Uh talk a little bit about how we
[00:42] strongly believe iceberg is the foundation for the future and why we have built our entire lakehouse and data strategy around it. We'll talk about how data bricks powers our lakehouse including like a very critical component called the refiners that that we will uh delve into and how it's helping us solve
[00:58] that. Uh some open sharing like how clients share data with us uh which you know we haven't done really in the past that data bricks helped us unlock and then we'll hopefully leave plenty of time for questions. Uh but beginning a little bit with the journey here right like to walk you back in time. uh it was
[01:13] actually around like the global financial crisis of 2008 when we really realized the power of having all our data in one place. By then our businesses had grown complex enough where it was no longer feasible to like you know compute important metrics like
[01:29] counterparty risk, how much margin you owe. Uh and you had to do all of this overnight, right? And calculate your exposure to like entities like Leman. uh and we realized at that time the power of having all our data in one place together, right? Uh and so we've been like on that journey for the last like
[01:46] you know uh 15 years or so or 18 years or so. uh and different iterations of that we built our um on-prem uh data lake in like the 201s uh and that was showing its age and that's when we had needed a new strategy for uh um
[02:04] uh our data going forward and that's that's when we came up with Legend Lakehouse a few years ago right a few tenants none of this should be like you know super surprising uh to anyone in an enterprise um but something that's unique to like financial services Goldman Sachs is a large global
[02:19] financial services firm. So being secured and governed with your data is just table stakes, right? Otherwise, no one would use whatever we had to build. Uh but one of the things we learned from previous iterations of what we did was that we really needed our production
[02:35] grade data to also be available for fast experimentation uh like for for like our analysts and strats and qus who need to be able to experiment, right? And that's something that you know we really didn't get right with our previous iteration and that's one thing that we really wanted to focus
[02:52] on this time around. Uh we also uh wanted to make sure that like you know like part of our challenges was that uh the transparency around cost especially when you're in a large organization with many many silos uh we had a little bit of a socialist system right like where
[03:07] uh like things weren't being accounted for properly. So it was an important tenant for us to get that right. uh and then I'll touch a little bit on single sources of truth throughout the presentation here. But like the single biggest challenge in like you know financial services in a large enterprise is that there is copies of data
[03:23] everywhere right because it is far too easy to copy over data and get things done quickly rather than do the right thing and get it from the source. So we've had to like be very innovative around how we share data in a zero copy fashion and data bricks and iceberg help
[03:39] enable us that. Uh so we'll talk a little bit about that as well. Um anything else you want to mention Abishek? Uh just like you know experimentation and fast boot up that's a that's a cornerstone when you're working with our strategists and and PMs on quantum investment kind of business in Goldman.
[03:55] So that helps us a lot making sure that we keep it as an upfront tenant. All right. Uh so skipping ahead like I'll just briefly touch on our uh reference data architecture. Uh and like how it works is that like we have like things that produce data at Goldman Sachs. call these points of data entry.
[04:11] These could be like trading systems. It could be like a uh like a like a system that uh like traders mark you know end of day positions right or it could be any any number of systems that generate data we call these points of data entry. Uh these then flow down to things we call ads or authoritative data sources.
[04:30] Uh the reason we do that is that bilateral data exchange in an enterprise is very very expensive right like so if everyone has to go to like you know like the producer of the data and get it from there it just ends up being like a very complicated graph of dependencies where
[04:45] uh things don't quite uh end up with the right answer. The the other challenge with that is that producers consumers of the data have to know exactly where they need to go to right and that that can become very complicated and expensive over time. So we landed on this concept called authoritative data sources which
[05:02] uh really capture important concepts that Goldman cares about like trades, risk, margin, um reference data, right? Uh and these are available to the entire enterp entire enterprise via well-defined APIs that we call data
[05:17] models rich data models that capture the full complexity of the data and its association with other pieces of data. Right? Uh and then the next layer is everything that runs on this right like all kinds of applications whether it's our like BI stack or like our um you
[05:34] know like management reporting all document management other other kinds of like you know uh analytics run on top of this like data layer right um and so we've done like a good job of selling all our businesses that they really need to get their data in shape in one place
[05:50] with well- definfined APIs if you want to get true value out of it right now we are on journey there and it's not like everything like is in exactly that spot but we are well along our way to like you know helping uh achieve this stack happen. Uh one thing I wanted to touch about again is like I mentioned uh like having
[06:08] well-defined APIs to access data right so for a long time almost a decade we've been a big buyer in semantic models for our data right because in financial services uh data is not extremely large the way it is in like some other industries but it is very complicated
[06:24] the relationship between different types of data is very context specific and like you need to know how to navigate the graph of data across multiple hops And the cost of getting something wrong can be very high, right? So you need to be very precise about when someone says,
[06:39] "Hey, how did you like you know like price a position?" You need to know did you use a clean price or a dirty price? Right? Like all of these details matter and there's like so many nuances around answering a particular question. So it becomes very important for producers of data to have very well- definfined
[06:56] models that describe their data with all the surrounding information that we thought was useful. a decade ago. So, we've been doing it for a while, but now we realize all that investment might pay off because finally we we get to like teach our AI models how to understand like these semantic models that that
[07:13] that have been built over time. Um and so for the lakehouse we came up with this concept called a data product which is like a evolution of like the data model I mentioned but it has a bunch of useful things like the description of the data sample data points diagrams uh
[07:30] that like how relate like that show how the data connects to other pieces of data obviously ownership what do you need to get access to it within the enterprise and other things that that people might find useful right and you do this for every piece of data that you really care out, you tie ownership to it and you prevent
[07:47] duplication, right? Like you you do not have like two sources of OTC derivatives at the firm, right? Because that would lead to bad outcomes. And so this helps us get towards our goal of like having single sources of truth with well- definfined data products that are easy for humans and for the AI agents to
[08:03] eventually understand. Moving on, uh we want to talk a little bit about like you know what we've done over the last year. We were here last year talking about our journey with data bricks as well. But since then we've managed to uh enable nearly all our business units with separate data bricks
[08:20] uh workspaces that provides like the kind of isolation that they require. Um our lakehouse is built on iceberg right I talked about it earlier. uh we strongly believe in like iceberg as the metadata format that has won the format
[08:35] a wars and it is going to you know like be the foundation on which we base our lakehouse. We strongly believe in interoperability right like and we have like great partners like data bricks who truly believe in that concept as well uh and so like we've given our entire lakehouse
[08:52] uh data which is like you know iceberg on top of S3 um uh data bricks has access to that layer uh and obviously it's done in a controlled way we uh like you know have our standard entitlements mechanism authorization mechanisms applied uh to uh our data bricks uh
[09:09] usage of that data uh and we have all kinds of like control monitoring that help us get comfortable with using this data with highly this highly sensitive data uh in that ecosystem. Um, talking a little bit about iceberg, I don't feel I need to like preach this
[09:25] to this crowd, but literally it is it is like something that's going to help us achieve our dream of like zero copy data and true interoperability across engines, right? Uh, and this is very important to us because we feel we we feel that like different engines are going to be better at different
[09:41] workloads. Uh, and our our businesses like you know like deserve the choice of like you know trying out different engines for different workloads. uh we also like really do not want our data copied from one place to the other when we want to use different engines and so iceberg really helps us get to
[09:57] that to that uh end goal right uh a few things here that we uh have been working on over the last year is uh like iceberg v3 now has support for variants that was like a big gap for us because v2 we could only do so much without the variant data type uh and this unlocks a
[10:14] whole range of possibilities for us uh and so we had a few hiccups like you and making it actually go into production. But uh variant was a game changer in iceberg v3. Uh and also like you know like deletion vectors for faster role level operations, nancond time stamps.
[10:30] There's a whole bunch of really nice uh features coming uh up in V3 that that that we really like and I'm happy to say that entire plant is now iceberg right like all our data on our lakehouse is fully on iceberg wasn't the case when we spoke here last year. uh and we talked
[10:47] about like you know uh how we do not allow direct physical access to the data right we have data models uh which we layer on top of like the iceberg governance layer so like when a consumer accesses a piece of data it all goes through that semantic data model
[11:03] um and I mean obviously like we we we u use unity catalog because it gives us like some really good features like the zero copy interoperability uh we we get to have like a uniform governance layer which we previously did
[11:19] not have um and like the kind of uh optimization features you can unlock right using um unity catalog instead of like you know like the previous mechanisms to do it like really help us get things like lineage access controls these are very important for us um and
[11:36] that's that's how we landed on this I am going to now hand it over to Abhishek to talk about refiners great thank you Ram So um so at the heart of it Goldman Sachs lakehouse ecosystem is a massive ELT machine where we have unlocked a
[11:54] massive use case on that is the transformation part of our data uh and how we are leveraging data bricks for that. So u the whole idea is the data for bunch of data sets has been materialized in an iceberg. All your
[12:10] parsets have been materialized in an iceberg format. They are in our own S3 buckets. Data bricks can understand read that data understand that data. UC has the right permissioning scheme that we have orchestrated on it and then we can express a logic to transform the data
[12:26] and then ingest it back into our ice back into our lakehouse ecosystem as a derived data asset. We can transform through SQL spark like choose your favorite poison. The whole idea there is we focus on back to the lakehouse tenant that we were talking about going as
[12:42] serverless as possible like you you don't want your you know banking and markets you know personals to kind of uh uh look at you know provisioning the clusters during market open and market close when you know the whole world is going frantically crazy on the SpaceX
[12:58] allocation. So that's the idea of like where serless kind of you know really helps you out centralized governance you know you do want if it's your MNPI deal in banking you do want a tighter need to know entitlements control on that data and you do want to make sure that it's segregated to Ram's earlier point every
[13:14] beu every business unit in Goldman is kind of having its own workspace how we uh quickly you know breezing through this how we deploy is we use data bricks asset bundle where we have orchestrated our own inbuilt built homegrown deployment service where a
[13:31] customer of our data let's say an analyst or a portfolio manager or or or a sales engineer they come in they can scribble a very simple data bricks asset bundle yaml in which they'll talk about this is the you know identity I am this is the workspace I want to write to this
[13:48] is the data and this is the software family kit I belong to we call it deployment in in Goldman that's all they just give us that rest of the orchestration is all taken care of by lakehouse infrastructure that GS has created here. Uh yeah, we already talked about serverless you know whether it's like
[14:05] zero infra play like you don't want your strategists and and quants to kind of you know uh chaperon provisioning of the clusters or or you know looking at the spark jobs which they are taking care of fixing all the knobs. So you know that's that's one idea. Autoscaling we talked
[14:21] about whether it's kind of trading spikes end of day you do want the system to take care of it in autoflexing uh reliability and security well you do want all of this u you know data orchestration in your real machine to be well governed well authored well
[14:37] definfined and well owned so the whole idea is you know different workspaces which carry different work workloads from different uh business units in Goldman can have their need to know entitlements enforced Okay, where we are in the journey, we are migrating
[14:54] 4,000 and change refiners from our on-prim system where we were previously onto a datab bricks orchestrated gigantic GS lakehouse ecosystem in that like there are bunch of capabilities that we have used AI like we don't want
[15:09] all our strats and engineers and and data scientists to do the repetitive like you know you know repetitive work so AI takes care of the heavy lifting of that where engineers are just looking at the exceptions, you know, just the brakes. Uh we uh went through a massive
[15:25] uh Spark upgrade from uh you would just laugh at it like initiate 1.x is the version where where we were with very little testing there. But then first thing we did was we used AI to first beef up all the functional tests on our existing codebase. Then used AI to
[15:40] migrate that to the latest and greatest version of Spark. That helped us a lot. uh you know that's such a test generation and validation really like really really we could score using this AI accelerated migration tooling um so Ram did mention like we are
[15:57] a user of delta sharing now called as open sharing so one of the areas of unity that really shines for us is if we are in a certain region in a certain deployment which happens to be same as where our clients are then we can leverage delta sharing approaches to get
[16:15] that data from our client in our compute. The data still remains in our client deployment. We are adding value added enrichment on that data in our global banking and markets business and then we can create a derived data product that our client can get access to. So this kind of a zero copy sharing
[16:32] without any replication without any brittle you know need to write reconciliation framework on like data on both sides if you start copying is really really terrifically useful for us. uh it also improves trust in the shared data because all your permissioning is still orchestrated by a
[16:48] single source of security truth which is sitting in in UC in unity catalog. So we in our world Ram mentioned that we use legend as a data modeling and data orchestration paradigm in Goldman which is also open source to phenos all the entitlement rules all the access rights
[17:04] on that data all bridge from that single source of truth and then we use UC to enforce it in in this kind of a paradigm. Um we have a demo for you today. Uh while I go to the demo, any questions for us so far?
[17:34] Yeah, I think the question is what were uh the trade-offs when we chose iceberg versus like other options that we may have had. Um I think like from the very beginning we saw like the vision in iceberg, right? like you know the the the the concept of being able to query more efficiently
[17:50] than like previous like paradigms that like you know like the hive for example right like it just was like a lot smarter than like it did things a lot smarter than uh like we previously done I think that was one big part of it uh the fact that like there was just like a lot of momentum in the community that like people quickly gravitated towards
[18:05] it just made it a lot easier for us right I think that was like the single biggest point the fact that we were able to then convince our partners like data bricks to add like more add value add-on features uh that helped us use iceberg just like sealed the deal for us. But I think it was really the momentum and
[18:20] like like the gravitation of like the market towards iceberg and it's proven to be like the right choice if you ask me. Yeah. When we were previously on an engine managed table stoages really like there were a bunch of things that were likable right like you like the latency because it'll be fastest because that
[18:36] engine is taking care of their stoages. Uh it'll be like governance etc. the the moment we saw that you know community pull towards in iceberg you can have all these features and latency going almost at par with the engine managed storage with everybody backing it interop is
[18:53] something that we really love right like you don't want to be logged into vendor A or vendor B's their own road maps or features and all the attributes or functionalities that we needed we talked about variant you know that was one of the important features we talked about you know deletion vectors for soft
[19:09] deleting the rows you know to to increase the the velocity how you're doing the DML operations as soon as those things started coming in we were all in right like you know why wouldn't we like an interrupt I mean the single biggest reason was like the our allergic reaction to copying over data right we were copying
[19:24] data all over the place and that leads to bad outcomes right if you're like native format you and you want to use different engines for different like you know like features that they offer you just have to copy the data over and that's just like operationally in the long run it leads to bad outcomes. Yeah.
[19:40] Okay. Let's let's look at the demo. All right. Okay. Uh give me one second. Yeah. Okay. So, what we're going to show you today is uh we're Yes.
[20:00] Yeah. Yeah. We are building a pipeline that would show you that we are ingesting two important data sets. one is the magnificent seven stock prices and the other is the fed fund rates. Now the reason of choosing these two is that there is a inherent relationship between the these two data sets. If I'm playing
[20:17] a role of let's say a Goldman Sachs equity research analyst, you know, I do want to understand what is that relationship. You know, if the Fed run Fed rate hikes, how does it pressure kind of growth or tech tech names? If it goes down, does it lift the valuation?
[20:33] So that is the kind of investment regine analysis that you would try to do to to do your thesis to figure out the price targets. So for the purpose of this demo what we have done is we have created this uh Python based notebook uh infrastructure in which we are uh we
[20:49] have created a I paused it for a second. So we've created a bunch of uh GS data engineering skills to operate on our lakehouse. The operations could be identifying these data structures, reading data from the foreign catalog, ingesting it, ingesting your data as a
[21:06] Pisber format back to your lakehouse, you know, uh, and and and a bunch of other things. So, first things first, what we're going to do is we're going to look at those two data sets, the prices and the fed funds, which have been already materialized into our S3 volumes
[21:23] as a governed iceberg data set by a foreign engine. So here what we'll do is we'll create a federated connection. So that's that's what is being done and we created on that federated connection we created a foreign federated catalog. In that catalog there are two tables which
[21:38] are already existing and we're going to zoom in on the T part of ETL. How we are looking at those data sets transforming them and creating a valuable derived assets. So u when we look at you know this particular one the the one where the cursor is this is the manage catalog
[21:54] in in data bricks in in unity which has u you know u views existing on top of tables which are present in in the foreign catalog which is the one which says 325963 sfdb right so that's how the setup has been done uh that you know
[22:11] your your already ingredient data source are there and you have created a manage catalog on the object objects in that foreign catalog. We're showing you some information schema meta commands. So this guy talks about that in your foreign catalog, the way we have set up our manage catalog is it only has access
[22:27] to that particular S3 volumes and the you know coarse grain control on on the objects there. Then if we look at the schemas in that foreign catalog, it's going to show you a bunch of schemas. The one for the purpose of this demo that we are interested in is this DBX summit trading data. That's the schema
[22:43] we we're talking about today. Coming back to the demo, first things first, we're going to authorize authenticate ourselves. Okay. The second thing we're going to do is we will initialize what we call as a data bricks client in our in our SDK. The initiation of SDK needs a few
[23:00] arguments from the end user. Uh which workspace does the end user want to connect to. In this case, data entpace. The second is which application software kit family does it belong to. So that's the owning dead 325963 and we're going to take this prompt and
[23:16] using our GS data engineering skills for lakehouse we're going to use copilot agent mode to create a code for this. So um again to recap with just two three arguments we are able to instantiate a
[23:32] particular datab bricks client for our this current session. So that means our session is tied to now that client. The second thing we're going to do is we deploy a refiner job. Refiner as in the transformation semantic that we want to bring into those two data sets that we
[23:49] we talked about earlier. So going to the the notebook here. First things we will look at is our code. What exactly is our transformer code going to do? So this is that Python code. It's very simple. For the purpose of the demo, we're going to
[24:04] look at our magnificent 7 tickers. get the data for those tickers from the already ingested pricing data by the other engine. We're also going to look at the second data set which is the fed fund rates which is again it's already
[24:20] present in in in RS3 in as a governed iceberg table and we're going to do some wrangling on top of data. This is the data bricks asset bundle that a user of this ecosystem has to pass in again only a couple of attributes that this will pass in. What is the workspace? what's
[24:36] the software family deployment family that you want to use to package your software as well as that you know you want to use the task type as spark python task and your your python code is that refiner showcase python so pretty much this is all that the customer of
[24:51] lakehouse and goldman has to do internally uh when we deploy this job we use our own written deployment server which will take this data datab bricks asset bundle yaml and fleshes it out to
[25:07] many other facets. See the thing is when we are refining this data, transforming the data, one of the things we have to know is whether the ingredient data sets is right now ready for consumption. It's a batch mode. We're talking about an OOL app system here. So the first things
[25:23] first is whether it's ready for consumption. So there are there are like you know states that we leave on our ingredient data just to depict that. Here we are showing you that that particular fleshed out YAML that we have created which talks about okay where is this ingredient data setting where is
[25:40] the target location for that data uh there is a concept that we have coined internally we call it watermark which simply is the state of the current batch which has not been consumed yet on the ingredient data set okay so the deployment has happened
[25:55] and it's going to return you that job ID that deployment ID next thing is we're going to now run this deploy job again we'll use our lakehouse GS lakehouse skills and there are just two things that this particular prompt needs one is
[26:11] availability URN like where is your source data sets and the digital signature of all the coordinates and the other is target URN which is like you know where do you want to push it the entire signature and the coordinates of this is where I want to push it this is how I want to entitle it so those are
[26:26] just the two things that it's going to use to run this job and and ingest the derived data set back into lakehouse. When uh I run it, it's going to uh you know emit a bunch of logs. So first thing it did was the first thing was the
[26:42] availability URL and the second is it's trying to fetch which what is the lakehouse location where I'm I want to write it. This is the one where availability one the second is the where I want to write it. So that's the right location one. And the third thing is it's going to kick off that refiner job which we go to data bricks notebook and
[26:58] we can see whether that refiner job has come in action. We use serverless compute. So what we're showing here is it takes just couple of seconds to kind of instantiate and we're using the you know performance optimized serverless compute for for for our purpose. So which is which is something that we
[27:15] like. And now the that particular job is running. It's going to take a minute and a half sort of a thing to complete. But while it runs and starts emitting application log, let's come back. We have a polling endpoint at this point which will keep polling what's the current status of that particular job.
[27:32] So based on the job ID, that endpoint will just keep checking the health of that particular data bricks job. It has started emitting these application logs. So at this point if you just you know it's pretty deep but like you know the two data sets that we were talking about the equity prices and the financial uh
[27:48] Fed Fed rates indicator those are the ones that it has read that it has to consume batch one of those. So that's the watermark that was available not yet to be consumed. It has consumed it and it is applying that Python logic to transform that data. So that's the piece
[28:03] that it is is going on. It is continues to pull the job at this point has completed. So now the transformation has happened. It has created that derived data set at this point and we're calling that as EQ fun fed analysis as your derived data set. What it will do is
[28:19] it'll try to find the from that target URN that we passed. Okay, which is the S3 volume location where you have to persist this particular park back. So now we are making the you know the the GS lakehouse ecosystem aware that a new
[28:36] derived data set has been persisted back to that S3 volume the staging bucket that that we call it and from there it gets ingested back into lakehouse as an iceberg table again. So what we have ended up doing here is 5,000 odd records
[28:52] have been uh you know appended to that table. We can go to a data bricks console, go to the go to the notebook and we can try to see that for the batch ID 5 uh you know can we navigate to the
[29:08] correct path and can we look at like you know whether we have indeed inserted 5,28 record here. So it basically what has happened is it's the full circle. We have consumed ingredient data sets on prices and and
[29:24] and and fed rates. We have understood the data. We have orchestrated the right entitlements. We you know set up the spark code to transform that data and the derived data asset has been ingested back into the lakehouse. So that is the step where we have reached. Ram talked
[29:40] about data product as one of the concepts that we have come up with in our lakehouse ecosystem which is a concept that you can cluster contextually related data sets together. It has a well- definfined owner. It has a description. It has you know the relationships encoded in our private
[29:57] semantic layer in such a way that you have to don't have to reverse engineer those relationships looking at the schema. So now what we're doing is we are for this particular data set like what we have inserted we have created a data product already with all this rich metadata we're trying in this Python
[30:12] notebook to import that data product and just look at like you know just the live rows uh you know the latest active rows from that data product why it's important for our some of our you know strategist users data scientists data engineers is they don't want to leave
[30:27] their current playing you know sandboxy environment so if they're on Python they just don't want to leave here and go to let's say a datarix console or some other place. So the entire piece is orchestrated here with like a GS lakehouse a bunch of skills. Apologies
[30:42] for this rendering but anyway the next thing we want to do is doing a bunch of data sciency analysis on this data. So you have that data product. Let's say I want to look at you know some window function based on ticker and date. what has been the count of you know XYZ kind
[30:58] of prices or on on that data that we have in the system let's say I want to look at some overlapping ranges for the fed rates and prices so you can do your you know basic data sciency task on these data sets next thing we want to just quickly showcase here is what is
[31:15] the causality relationship between the prices and the fed funds so the way we want to do it here is the first chart says okay you have these magnificent seven techn names and and and you have the different prices for the last whatever like last year's worth of data
[31:31] then let's just plot them as is. The second thing we want to do here is plot them with a baseline of 100. That way you can compare how they have you know rallied or not you know in tandem with each other. So that's what that that was the whole idea that periods where one stock rallies while the others lag you
[31:48] can easily note that you know the widespread in the price level across tickers can be easily noted. Next chart is about Fed rates. Again, you know, at a baseline, how do you kind of understand it? So, when the Fed rates uh, you know, are raised, the borrowing cost will increase across the across the
[32:04] board and then it'll put the downward pressure on your on your equities. So, here the the Fed cut rates, the cheaper uh capital tends to boost the equities. That will be clear in the next chart. But this is basically your fed rates graph which is going to show you that what's the staircase pattern here. What
[32:20] has happened uh in the recent past here the chart three will bring them together and you will be able to look at the inverse correlation. The places where you did not see the inverse correlation is all the AIdriven you know revenue boost and that rally which is like April
[32:36] May time frame. But otherwise you can clearly see that you know with the the sensitivity of these growth tech names against the fed rates is is is clearly evident here. Yeah, that pretty much kind of wraps up our demo.
Learn more about the Databricks Data and AI platform.
The information provided herein is for general informational purposes only and may not reflect the most current product capabilities or configurations.