Skip to main content

Build Secure AI Foundations: Unity Catalog for Mission Data

Summary

  • Data fragmentation is the single greatest security risk to mission modernization, preventing organizations from building trustworthy AI because critical data remains siloed, inaccessible, or stored in incompatible formats.
  • Databricks Unity Catalog addresses this by unifying governance across data, models, and applications through six elements: discovery, access control, audit, monitoring, lineage, and secure sharing.
  • IDB Invest accelerated project delivery by 60 percent and achieved seven-figure cost savings after modernizing its analytics on the Databricks Data and AI platform, and is now building AI-powered treasury operations and mission-impact applications.

Build Secure AI Foundations: Unity Catalog for Mission Data

Watch: Build Secure AI Foundations: Unity Catalog for Mission Data
Fragmentation is the single greatest security risk to mission modernization. Organizations with siloed data systems cannot build trustworthy AI that serves constituents, delivers on mission, or maintains compliance. this video establishes why consolidating and governing all data in one unified place is the non-negotiable prerequisite for secure, organization-specific AI that actually works.
Learn the six critical elements of enterprise data governance: discovery, access control, audit, monitoring, lineage, and secure sharing. Explore how Databricks Unity Catalog unifies governance across data, models, and applications with fine-grained access control, column-level lineage, and compliance visibility. Through IDB Invest case studies, see how organizations modernize analytics, reduce fragmentation, accelerate project delivery by 60 percent, and achieve seven-figure cost savings while building AI-powered applications for treasury operations and mission impact.
🤝

Chapters

FAQs

What are the six elements of enterprise data governance covered in this video?

The six critical elements are discovery, access control, audit, monitoring, lineage, and secure sharing. Together these elements ensure that organizations know where their data lives, who can access it, how it is being used, and how it flows between systems, forming the foundation for both compliance and trustworthy AI.

How does data fragmentation create security risks for government and enterprise AI?

When data is siloed across multiple source systems, formats, and assets, organizations cannot apply consistent access controls or audit who is using which data. This fragmentation also means AI models are built on incomplete information, producing unreliable outputs that undermine trust in mission-critical applications.

How did IDB Invest use Databricks to modernize its analytics and reduce costs?

IDB Invest consolidated its fragmented data environment onto the Databricks Data and AI platform, which accelerated project delivery by 60 percent and generated seven-figure cost savings. The organization also built AI-powered applications for treasury operations and cash management as a direct result of this unified data foundation.

How does Unity Catalog govern AI models and applications alongside structured data?

Unity Catalog extends its governance framework beyond tables to cover models and AI applications, providing fine-grained access control, column-level lineage, and compliance visibility across the entire AI development lifecycle. This unified approach means organizations can apply the same discovery, audit, and security policies to AI assets that they apply to data.

Full transcript

[00:12] Okay, good uh good afternoon everyone. Um, I'm going to spend time talking about what it means to have a solid foundation of data, which is kind of what we need in order to do all things analytics, all things uh all things AI and that's what we care about in order
[00:27] to deliver to the mission of the public sector agencies that uh many of us uh many of us work for. So, thank you all for thank you all for being here. I've got a question for you. Who is more excited about data than about AI?
[00:43] Okay, this is good. Who is more excited about AI than about data? Okay, some some of you are so excited you raised your hands twice. I love it. I love it. I love it. Maybe it's the free lunch. Okay. So, um the first thing
[01:01] we want to spend some time talking about is this idea that when we have unified data, it actually unleashes value and it unleashes value for for AI. For those of you that have tried to build models, even if you're not trying to do something super advanced like AI, even
[01:17] if you're trying to do something like build a report, answer a question for a boss, the thing that stumps us where we end up spending much more of our time than less is that last piece of data. I can't seem to get the numbers from this region. I can't seem to get the data
[01:33] from this last system. And so what often happens is this is a symptom of the fact that our data estate is fragmented. We've got multiple source systems where our data resides in. We've got multiple
[01:49] different assets, data assets where the data actually is in. And then of course there are multiple formats within which the data exists. And when we have that, our ability to be able to build the value ad services on top to actually
[02:06] extract real value from data which often is dark, often is unusable becomes severely diminished. So what we are trying to do in an ideal world is we want to be able to get to that most sensitive data. We want to be able to
[02:22] get to that last that last corner of where that data exists because it's always that missing piece that we need in order to be able to get to real real value. And so the world that we all yearn for is a world
[02:38] where collaboration is open, where data is nicely organized in domains, where we can get access to data immediately and we don't have to wait. How many of you would like to live in a world like that?
[02:54] Okay. How many of you feel like you live in a world kind of like that? Okay. Okay. This is there's one person. You should find out who uh who she works for. So, there's going to be an influx of job applications uh into uh into that
[03:10] that organization. And you know, the reality is we clean up things, we get things organized, we get things governed, and then what happens?
[03:26] something changes, right? The there's a reorg, there is a new platform, there's a new technology, there's a new boss, there's a new priority, there's a new vision, there's a new strategy, there's a new set of contractors who have slightly different ideas for how things are actually going to going to get done. And the reality is while on the one hand
[03:44] we want to have access to all of this sensitive data, who typically gets in the way? Sorry. Security. Who else gets in the way?
[03:59] Legal. Right. These are our favorite uh our favorite people. I I will say that I spent most of my career in security. I am a recovering siso and so I um some of the things I'm going to talk to you about it has taken me quite some time to
[04:15] get uh get past the world that I lived in and the ethos that I had as a security person to really understand that value from data has to be unleashed. It has to be available and the goal can't simply be security. And so when we look at survey after survey,
[04:32] when I talk to organization after organization, whether it's in the public sector or in the private sector, we see this often. We don't want to do AI. We've got lots and lots of ideas for how to do AI, but we're just not comfortable
[04:49] putting them into production because of security, because of privacy, because of legal. And that's the reality of the world that uh that we live in. And so what does security want to do? They want to secure the cloud. They want to make sure that there is an ATO. They want to
[05:06] make sure that there's strong access controls. They want to make sure that every single thing is compliant. And yes, there's a lot of different compliancies that organizations want. And at Data Bricks, we have almost all of them. And every quarter it seems like
[05:22] we add a new three, four, fiveletter acronym to the list of ways that we are being compliant with some regulation somewhere in the um in the in the world. And when we add all of this up, we find
[05:37] this dichotomy. We find this dichotomy between we want to democratize data. You know, for those of you, how many of you were here early enough for the keynote? Okay. Awesome. Awesome. So then you got the free breakfast too. Good job.
[05:53] Awesome. And so we want to democratize data, but we also want to secure our data. And oftent times this is framed as this false dichotomy of it's either one or the other. This is kind of like this
[06:09] old adage which I never understood and maybe someone can help me understand this. How many of you have heard of having your cake and being able to eat it too? Okay. You know what I don't understand is what is the point of having your cake if you can't eat it?
[06:26] Like is there any value to just saying I have my cake so I get to stare at it all day long but I don't get to eat it? Same thing here. What's the value of having all of this data if you actually can't access it? If you can't actually use it,
[06:41] if you actually can't share it, you can't actually use it to build AI models and to build reports and to extract insights, it feels like an odd thing to be able to do just one of those and not both of those. The way this plays out
[06:56] often in government agencies is this. We've got a long list of services that every agency is tasked with being able being delivering to its constituents, to citizens, to other agencies within the within the government. And we've got a
[07:14] lot a lot of source systems, a very long list of source systems that we this access that this data is stored within or that these systems need access to. However, or these systems need access to the data. However, it doesn't all work
[07:30] as well. And when it doesn't work as well, what do we end up with? We end up with fragmented views of the audience. We end up with legacy technology unable to scale. We end up with a difficulty to make the leap from using descriptive
[07:47] analytics and descriptive insights which are all about the past to being able to have predictive insights which are all about the future. And of course we lack the ability to do it in real time. this morning. In the last couple of hours,
[08:03] I've had conversations with a CIO and a co at two pretty significant federal agencies and both of them complained about exactly this. We have an ocean of data yet we are right in the middle of
[08:19] it and feel like we might be passing out from thirst. And so the reality is we have this data. It just is not very usable in many cases because we feel like we can't use it in a manner that is
[08:34] compliant, in a manner that's trustworthy, in a manner that actually is secure. And when we can't do that, we impair our ability to serve our constituents. we impede our ability to deliver revenue for organizations, agencies that are actually delivering
[08:51] revenue or for others just to be able to deliver on their mission. And in many cases, paradoxically, because of this deluge of data and because of this plethora of murky policy and compliance requirements, we actually make it
[09:07] difficult for ourselves to feel confident of even being compliant with our own rules. and policies. And then the last piece is what we would love to do which in many ways is the promise of data being unleashed and democratized is
[09:24] this notion that we can optimize operations. That data is the key to being able to do more with less. However, if data is not democratized in a secure way, if data is not well-governed, we're not going to be
[09:40] able to do these things. Central to being able to govern our data is going to be of course the catalog. And what we find in currently in many agencies is this idea that there is not
[09:56] just a catalog, there are many cataloges. There's just not a governance. There's many governance structures. And every department and sometimes within a department specific projects have grown up in their own silo because at some point they were super
[10:12] important to some leader. They were given enough budget. They were given a mandate to go do things on their own way. As a result, we've built many many silos of data and AI assets and being built on different platforms over
[10:28] periods of decades. And the reality is every single time one of those silos was built and every single time one of those silos was funded, it was done with the right intentions. It's not like someone sought to go do something in order to
[10:46] make for increased complexity, in order to make for increased ambiguity and increased challenge for the people the people that inherit the organization five or 10 years from now. They actually thought that this was the best approach. So they picked a certain data format.
[11:02] They picked a certain AI model. They picked a certain approach to doing it. Maybe they picked a certain language. Maybe they picked Python or Scala or R. Maybe they picked iceberg or they picked Delta or they picked parquet because
[11:18] there is no shortage of options to pick from. However, as a result of many of those individual decisions being made one at a time in aggregate, what they add up to is that our costs go up. That
[11:34] we have a lot of redundancy in compute and storage. And we often have this inkling to say in order to do this, this is the data that I need and I need all of this data to be brought into my
[11:49] structure and my way of doing it for my project in my department, which feels like the right thing to do at that moment in time because it is going to give us the shortest path between where we are and where we need
[12:05] to be because We are choosing to control our own destiny and we are going to have complete control over all the systems and data that we need to versus being beholden to other parts of the organization.
[12:20] The reality is as we look back over the last 1015 years and the rare view of a lot of those decisions, it becomes more obvious that not taking an enterprise level view and approach to this is
[12:36] actually what got us into this problem to begin with. So it turns out that governance is the key to empowering data users across agency with the right access to sensitive workloads. So we have to figure out how to solve
[12:51] governance. Not to solve governance by saying I am going to do my own governance for my own thing. But how do we do governance in a broader enterprise sense so we get the benefit of what's happening across the um across the
[13:07] agencies in a world where we have strong governance in a world where we can truly feel like we need to have a copy of the data once. So we leverage things like zero copy. We leverage things like we know that there
[13:22] is going to be the ability to get access to the data when there is an appropriate need to know and we're going to be able to do it at scale and we're going to be able to do it at speed. There are six specific things
[13:37] that need to be incorporated into the governance strategy. One is the ability to discover data versus saying I need to suck up all the data and I need to have it here. There needs to be the ability to readily likely AI enabled be able to
[13:55] discover where the data is. Maybe even using constructs like what data bricks is about to announce with data domains to have that data readily organized so it can easily be found regardless of where it is within the agency and perhaps even a broader aperture. across
[14:12] the federal government, which is one of the conversations I had with a CIO this morning is they're looking at how do we make that data sharing better and in this case this particular agency has authority across multiple uh multiple agencies to be able to do that. The
[14:27] second of course is going to be around controlling access. Next is around being able to audit what already happened because we want to make sure there isn't any malfeasants, there isn't any fraud, everyone is following the rules and auditing is a pretty good control to
[14:44] make that possible. And then of course monitoring not just in the past but right now what's happening lineage to be able to extract more insight and more information on where did this data come from? How did this come about? I built this model, but what were all the data
[15:00] sets that made this data made this made up the features that were used to train this model? I would like to know and I would like to know at the point in time that this version of the model was created and the experiment was run. Did it come from some structured data or
[15:16] unstructured data or streaming data? I would like to know it all. And then of course I would like to enable the sharing of data in a governed manner. What organizations are looking for is they want one catalog that has unified
[15:32] governance for data and AI for everything. Remember that first picture I showed regardless of the source, regardless of the asset, regardless of the format, I want to be able to govern it all. It's got to be open. It's got to be able to support multiple engine. It
[15:47] must work well across formats. And it has to understand the mission. It has to have some business context. And this is where Unity catalog and specifically the business semantics capability is all about connecting the data to the context
[16:04] of the mission. And so three things I want to double click on that Unity catalog enables. The first one is around security and compliance. So you want to be able to secure your data AI artifacts
[16:20] with flexible roles and attribute-based access control. So you don't want to have to go in and provision access at that fine grain entitlement level. You'd like to be able to do it at the role level. You'd like to be able to do it at the attribute level. Be able to use tags
[16:37] to dynamically give access to uh data that people need access to versus doing it explicitly one at a time. And you want to be able to automate to the extent possible self-service access with requests and approval flows without tons
[16:52] of integration with thirdparty identity and access management uh technologies. Um and then of course governing the flow of data. This is going to mean having real time column column level lineage for data and for AI and knowing what was
[17:10] done with your data, not just who accessed it, right? A lot of these capabilities are the ones that give us confidence to say we can allow for data to be further democratized because now we have additional eyes and ears. We have additional insights. We have
[17:26] additional controls to make sure that the data is being used in a way to maximize the mission versus eroding it. We want to be able to gain visibility and observability into compliance into cost and uh and and quality because it's
[17:43] not just security that we care about. We also care about practical realities like cost and of course the efficacy of what we're doing is going to be significantly impaired if the quality of the underlying data and the data pipelines is affected and unity catalog will
[18:01] provide uh will provide that. The other thing for those of us that have been around long enough that we've been burned by many a time is we find that there is a platform and it sounds amazing and we go down the path and three or four or five years later the
[18:18] investment sees, the partnership sees and all of a sudden we feel like we are locked in to this one platform. However, it is no longer meeting the needs of the business. So then we pick whatever the next best flavor is for that year and we
[18:34] hope that this is going to be the final solution that is actually going to work for us. So we see our customers experiencing that pain. We want to make sure that we're helping alleviate that pain. And the way we do that is we make sure that a it's got to be unified
[18:50] across your data estate. It's got to be flexible. It's got to work regardless of the data format. We don't want to be in the business of picking winners or losers. We want to be the platform that says we're going to support all of your needs regardless of the platform that
[19:07] you have. And we want your engineers and science data scientists to be empowered to use the tools that they prefer versus us imposing our preference on them. But the way that you make all of this futureproof is it has to be based on
[19:23] actual open-source technology. So, Unity catalog as many of you know uh June of 2024 is when we actually made it open source. So, whether you are using data bricks or not, you can certainly get take advantage of Unity catalog and the
[19:41] rest of the 3.5 billion downloads worth of open- source projects that uh data data bricks very directly uh supports and contributes to. The last piece is AI powered performance optimizations. This
[19:58] is key. If the expectation is that I have to have a very very large army of people doing data governance, it will never get done. This is the conversation I just had 30 minutes ago. How do I convince my leadership that the work to
[20:14] build a data foundation is important? Well, you're not going to get approval for 50 people to do data foundation work. What you need is you need technology to give you that lever to be able to do this at scale with a lot less
[20:30] effort, which is what Unity Catalog does by giving you 20 times faster queries, 50% lower storage costs without you having to handtune every single one of these items. The last thing I wanted to mention as I as I close this out is, you
[20:47] know, many of you are likely thinking the catalog is an important element. Certainly, it addresses many of the risks associated with gaining access to and using sensitive data, but really is that it? Are there other things beyond the catalog that the that data bricks
[21:03] can help with in our ecosystem in managing the broader set of AI risks? It turns out that there are AI risks across four different subsystems that make up any AI platform. It's the world of data, which is what we've really been talking about, but it's also the world of models
[21:20] and applications and unifying governance across them. So, Unity catalog is how we unify governance across them. But if we were to actually ask the question, what are all the risks associated with AI? It turns out that there are 62 risks
[21:35] associated with AI. More of these risks are mitigated by data bricks unity catalog than any other single technology that we have in our portfolio. Hence the focus on unity catalog for today.
[21:51] These 62 risks are the centerpiece to something we call the data bricks AI security framework. What this does is it helps you understand your stakeholders AI use cases. Perhaps you're using one of the many 80 plus you uh solution
[22:08] accelerators that data bicks has published on our website. But then you look at the 62 risks across the 12 components of AI. Each of those are mapped to one or more of the 64 controls that can be used to mitigate mitigate a
[22:24] those particular AI risks. Many of those controls are part of Unity control part of Unity catalog. And then of course the thing that you're going to care about is how can I map the data bricks AI security framework these risks and these
[22:40] controls to NIST or to OASP or to MITER other framework standards industry bodies and we've done all of that work by collaborating with them and at the end you're prioritizing the controls you need versus trying to implement everything. If you're interested in
[22:57] this, you can go to your favorite web browser and uh search engine and just search for DASF and you can get uh you can download the data bricks AI security framework version two. We'll be releasing an updated version in the next uh next few months. And if you are a
[23:14] director or above, we're doing an exec round table. I think we have two or three slots open uh at 3:45 p.m. later today in the um in the continental room. So um I wish I had time for questions but I think I am about a minute and a
[23:29] half over my uh over my time. So um with that what I am going to do is I am really really delighted and excited to introduce our customer speakers to show how we take much of what I talked about
[23:44] which is building that secure foundation for AI and uh Igor and Capil who are both from I the inner American development bank invest IDB invest are going to spend some time talking to us uh about uh about how they went on on
[24:00] their journey. So, Capil and Igor, welcome. Thank you. Thank you, Omar. Let me take that. Thank you very much. Uh, hey everybody. Good afternoon. My name is Igor Valentine. I lead data
[24:18] management analytics and AI for IDBS. Together with Capu here, we're going to be talking a little bit about our journey to get a a scalable and secure AI. And it's clear who is the treasury guy over here, who is the IT guy over here, right? Just by the way that we
[24:34] address it. But anyways, um just getting the ball rolling. Um our journey is started truly uh three years ago where we understood that we had to in order for us to support our business and
[24:52] support the growth of our institution, we understood that we had to empower our business users. looking into solutions that would give them possibilities to truly take the most out of the platform. But without um
[25:09] we have in order for us to really truly talk about what what that journey look for or was looking into uh we have to talk a little bit about our mission. So IDB Invest is the private sector of Interamerican Development Development
[25:24] Bank Group. is IDB group and at IDB group at the IDB invest our main focus is to help fostering developing uh uh the the Latin America and the Caribbean through the private sector supporting investment and sustainable investment
[25:41] through and the private sector growth in the region. And the way that we do that as of now is through mobilizing and capitaliz and using money from the private sectors in the countries that we have and implementing and deploying that um that money in the region is
[25:58] specifically driving innovation and fostering uh new business to grow and to to be executed. Um and and and our journey started by really into looking into four main principles that we had.
[26:14] So in order for us to do a data lad transformation we understood that those four main principles had to be had to be in place execute an analytical ecosystem modernization because we had to start it over there because our analytical ecosystem was
[26:31] truly outdated. As a matter of fact that illustration that we just saw today at the keynote kind of very much resonate with me. our ecosystem was really out uh made us feel like it was maybe 10 years apart from
[26:48] the private sector. Um so after that our conversation started to be we have to truly achieve a golden source that was mandatory for everything that we doing because data fragmentation was a reality from our
[27:04] institutions. I'm pretty sure that's a reality for your institution as well. Excel's everywhere, databases everywhere, SQL's everywhere. So we had to looking into a way to unify everything and truly achieve a single source of the truth. And with one
[27:21] concept in mind, have data governance as an enabler. So data governance had to be something that would help us expedite the usage of data, help us to truly reduce the barriers for those to access
[27:36] our information and most importantly without losing sight of the control and protection of our data, our main assets. And then with everything in place, our concept was let's looking into impact driven use cases. We don't want to
[27:53] execute anything related to data and AI just for the sake of using the technology. We understood that was super important now that our users were empowered. They could be taking care of the daily activities themselves and us from IT perspective would be acting as a
[28:11] center of excellence to truly help them to move the needle to implement use cases that were far advanced into the realities that they had back then. Uh but we cannot do that without choosing
[28:26] the right technology. So we understood that data bricks was the partner for us in choice because they help us to they fulfilled they had many features that fulfilled those needs that we had to start with. I will I will go with the
[28:42] lakehouse. The lakehouse concept the lake house architecture allow us to tap into structure and nonstructure data at the same time supporting multiple languaging processes. The lakehouse helped us to
[28:58] truly uh extract value from the nonstructured data that we had in our institution. To be honest, as an international organization, I'm not so sure if some of you guys know, we sit on tons of reports and analysis executed by
[29:16] our economists and peoples in the field. And most likely back then that knowledge was hidden on several PDF files spread out in our institution on SharePoints and etc. But no one truly had a the opportunity to access that information.
[29:34] Then the Unity catalog remember what I said having data governance as an enabler with the ability to have a fine grain access control. the way to truly spread the information and control in one single console access to all the
[29:52] analytical assets that we were developing. Unity catalog helped us to ensure that we were sharing information across the teams in a safe manner and most importantly applying the rules in one single place to deploy across the
[30:07] board. So lastly, once we achieved those two steps and we understood that data bricks could provide us uh that through the Unity catalog and the lakehouse architecture, we went to look into details and that was we were we were already back then
[30:24] starting to use more and then we tap into clusters policies. Something that was a key enabler to help us to achieve efficiency and collaboration in a very controlled manner. we were able to automatize cluster
[30:40] management. You know, we are in the cloud. Everything that we do costs. Um, so we have to be very mindful. It's a change of mentality for most of those in our analytical ecosystem that were so used to run a query without even question. They were executing queries uh
[30:57] uh without taking care of being concerned about how many data they were pulling in. And because we were in the cloud, everything costs. So we had to do in a way that would be controlled as well. Um also that help us to manage and
[31:12] have a better predictability in our analytical ecosystem. We don't have big pockets and we use government money. So we have to be very mindful about that too. Um with all of those things in place we finally achieve uh the
[31:27] efficiency the collaboration that we were looking for. Now with a concept of single source implemented and the golden source implemented all of our users consume data from a single platform from a single ecosystem leveraged by analytical sandboxes where they cross
[31:44] information with each other. They share analysis that they are doing in a much easier simple manner without sharing excels to everybody that they lose sight of the versions and controls that they had. So all of those things were super important
[32:00] to us because when we started to mature on that route, we started to see the benefits. And as I said, we started three years ago and the benefits now talk by itself.
[32:16] Three years later, we look into our analytical ecosystems. We look into the the the we look into the the the results that we have. And we were s and we are now 60% faster in any data related projects that we executed. That's a huge
[32:33] gain. That's efficiency right there. We increased 85% the amount of users in our analytical ecosystem. Now everybody is consuming from our from our environment and to the point that we change the way that our
[32:49] institution is recruiting. Now we are seeing more and more job descriptions in our institution coming with data bricks exper expertise and knowledge which to us is huge. It's super important. That's the piece of the transformation that
[33:05] we're trying to tr to do to do and to achieve and we reach seven figures on cost mitigation with projects that were all the way from automation of processes into looking into and bringing more efficiency to the
[33:21] institution. I want to bring three cases and we have Capu to kind of do a a deep dive in one in specific. So we are right now not deploying one but three use cases in December all powered by AI technologies
[33:37] or AI features that data bricks provided to us. one and as I said at the beginning all impact driven number one is we are deploying one solution that's it's a virtual assistant embedded into our core lending system that helps our
[33:54] investment officers to pull out information about those hidden knowledge that previously were sitting on on PDF files and they can ask questions to understand better how can they set it up a new project. So that is interaction
[34:12] through natural language processing. They ask questions, we send it to our genius spaces, the genie spaces send it back the responses to it. So that's number one. It's been super useful. Number two is another another example
[34:28] focus on our operations team now into the risk department. As part of the processes that we do to analyze our transactions, we have investment risk investment officers looking into tons of reports coming from the agencies, the rating
[34:45] agencies that we have and also analyzing uh historical information from past projects that they executed. they had to um perform those reading all of those documents and trying to suggest and recommend a risk rating for this
[35:01] particular new project that they are executing with the technology and with data bricks and the the models that we just launched and we are launching we launched as a matter of fact yesterday uh they are capable of data bricks recommends what's the risk rating uh
[35:17] that this new project should have that reduces two weeks in terms of assessment and analyzing one project that they have to do. So, especially looking into the period that we are right now, end of the year where they're trying to run and analyze as
[35:34] much projects as they can, that's a super super u powerful solution delivered for them right there in their hands right now to kind of expedite their process to analyze information. And the last one is the one that we decided to double down right now and I'm
[35:51] going to pass the the the floor to Capil. Thank you. Thank you Eager. So as Eager said, I am a a finance guy. I am not a data scientist. So I will give you the business perspective of what we have uh achieved and before even uh doing that I
[36:07] would like to answer what Omar asked in start of this uh uh session was what is more important to you guys data or AI for me to be honest is data for me AI is like a cherry on the cake the only
[36:22] difference is that cherries increasing in the size and that's the fun part for us so what eager mentioned was my use case. So it was one and a half year back we didn't have any cash system. We
[36:37] didn't have any cash management system in house where we can actually do lot of analysis and actually project our cash cache forecast for coming months. So we had a choice either to buy the off-the-shelf product right or build
[36:53] something in-house and data bricks was picking up in the organization and eager team was doing fantastic job on that and we said okay let's try data bricks and that's where we decided that we're going to go with the data bricks because it was it light it was scalable and it was
[37:12] for us we could see it it it we can easily tailor it for a multicurrency multimarket approach because we are in more than 30 plus countries uh handling all lot of latam markets and the G10 currencies. So we decided to build it
[37:27] internally this uh this system and one of the biggest cha challenge and some of the words you will see repetitive fragmentation was the big challenge for us the data was sitting in different part of the organization. So we have a challenge how to bring all the data in
[37:45] one place from all the internal system. We have a loan system, we have a derivative system, we have a fixed income system, a regular corporate account payable, account receivable, anything which a corporate treasury has, we have all those and then we need to also connect to our banks, our
[38:01] custodians. We have roughly around 100 bank accounts. So we have to connect to all those. So that was a step first step for us to actually work with our IT team to bring data in one place to centralize the data. So that was the key that was I guess the longest time it took us was
[38:19] there centralizing the data and what was saying earlier to make it a single source of truth and institutionalizing it. So that that helped us a lot there. So in order for you to do a cache management the first part is obviously
[38:35] the cache forecast and where it really helps is to have a centralized data. So once we did that we in the datab bricks itself what we did was we created a sophisticated logic which was telling us how to fund the accounts at what point
[38:51] and where was the excess cash how can I generate the alpha of my portfolio and all this analysis we were doing it in the data brick. So it became as for us a basically an analytical tool and these things we were not thinking actually when we started this project but we were
[39:07] gaining and there were a lot of features which were coming there once we did that the idea was okay how I can make it the user friendly how I can actually put the advanced analytics into this so we started with genie worked
[39:25] fine it's a conversational interface like running a SQL query it was great but to me it was a of a limited scope it was giving me the answer but we wanted to do something more then we started partnering with data bricks and I see
[39:41] some of our relationship guys here so what we did was we created a multi- aent tool there the multi- aent tool was doing few things one it was doing data retrieval second it was doing the data analysis of
[39:56] the data which it retrieved D and then it was giving us the business insights or ideas how to invest cash and having this portfolio very efficient in in terms of managing the cash. So all this
[40:12] was done in the playground uh of the uh data bricks and I think the biggest advantage I see this was first time in my life I saw the business was doing lot of stuff and it was a facilitator it as
[40:29] was saying was a enabler for us I have never seen that in last 15 20 years of my career I have seen from the time where we have moved from so-called crystal reports I do not know if you anyone remember the crystal report if anything you need you need to go to your IT and it has to work create a report
[40:45] and it will it the turnaround time was a one month or so but this was something which we can actually do it ourselves so that's what my team is being doing and and what I was mentioning earlier the skill set in the industry is changing today when I actually post a position
[41:01] I'm looking for someone someone who knows about data who has some background there who knows about data bricks who knows about AI so I I think that's that's was uh our journey and what we have been recently doing is taking it to another level. What we are trying to do
[41:18] is we are doing the alert system. What that alert system basically is doing is is we we have created an alert system where based on what I was saying earlier on on retrieving the data using the multi- aent uh features it's actually
[41:36] sending a notifications to the users in the Microsoft teams whenever it sees any kind of anomaly in terms of cash wherever it sees whenever it sees the cash opportunities. Hey, there's a cash opportunity. You can move money from one
[41:51] account to another account. You can do an FX. So, it's doing all those things there. So, that's making it easy that you don't have to go now always as a user. I don't have to go to data bricks always. There are some alerts which I'm getting in the Microsoft team which makes life easier. So, net results is
[42:10] that the being a finance guy, I love the last slide which you was present is seven figure saving. For me that is one way of looking at it respect for me it's actually ability to generate alpha on my portfolio because that's my bread and
[42:25] butter. So I think that's the how I've been using data bricks there uh and it's great to be here. No, fantastic. And and that is a true example of collaboration, right? Something that I said in the early stages and principles that we were looking into once we were we were
[42:42] modernizing our analytical ecosystem. It's of course it's a combination between tools, the right technology, the right people and the right processes in place. And we have to connect all of those things. And uh the reason why we chose data bricks was because it helped
[42:59] us to connect all of those things together and having them being managed by a very small and tiny team very mighty team as well. So um thank you guys very much. It seems that we have you guys we gave you guys three minutes
[43:15] back according to our slides to our timer over here. So we'll be happy to stay here uh down here and uh take any questions of Ormar's back. Um well Capil Igor thank you for uh walking us through how uh how you've been using data bricks. Loved hearing
[43:31] about the cache management system and I love the description of it light. So you decided to build it yourself but without the heavy lift of what typically entails when we think that uh this is a build versus buy decision. So there is sort of that uh that third option which is there's the bill light option or the
[43:48] IT it light option. Um Capil and Eigor will be around so if you have questions for them for me please feel free to come up and I think we have a 15minut break before our next session and I hope you enjoy the rest of the data and AI world
[44:03] tour in DC today. So thank you. Thank you. Thank you.

Learn more about the Databricks Data and AI platform.

The information provided herein is for general informational purposes only and may not reflect the most current product capabilities or configurations.