Southern Company Transforms Utilities Grid with Databricks Lakehouse and NTT DATA
Summary
- Southern Company, one of the largest utilities in the United States, partnered with NTT DATA to replace legacy siloed on-premises systems with a federated Databricks Lakehouse that reduced outage root-cause analysis from 8 days to near real-time.
- A medallion architecture unifies data from meters, power lines, assets, and customer systems, with Oracle GoldenGate providing near real-time ingestion and Unity Catalog governing data access across the organization.
- Five applications—Ramp, Scout, Sphere, Magnus, and Nexus—deliver grid hierarchy visibility, network analysis, and vegetation management, with a multi-cloud strategy integrating Palantir and Google Cloud Platform for flexible AI and analytics workflows.
Southern Company Transforms Utilities Grid with Databricks Lakehouse and NTT DATA

Southern Company, one of the largest utilities in the United States, faced significant complexity managing outage analysis and grid operations across multiple legacy systems. What once required 8 days to identify root causes in utility outages now runs near real-time, powered by a modern Databricks lakehouse architecture partnered with NTT DATA.
Learn how a federated lakehouse design with medallion architecture unified fragmented data from meters, power lines, assets, and customer systems. Discover data product patterns including grid hierarchy, network visibility, and vegetation management that scale across the organization. Explore how Databricks Unity Catalog governs data access, how Oracle GoldenGate enables near real-time ingestion, and how multi-cloud integration with Palantir and Google Cloud Platform enables flexible AI and analytics workflows while maintaining regulatory compliance.
🤝
Chapters
00:00Introduction and Southern Company Challenge: Modernizing Utilities03:09From 8 Days to Near Real-Time Outage Analysis05:49Building the Foundation: Integrated Data for AI Readiness08:13Success Through Organizational Buy-In and Applications10:04Application Portfolio: Ramp, Scout, Sphere, Magnus, and Nexus14:25Grid Analytics Transformation: From Power BI to Web Apps16:55Evolution to Databricks Lakehouse Architecture20:56Medallion Architecture and Federated Lakehouse Design24:27Data Products and Vegetation Management25:46Near Real-Time Ingestion and Oracle GoldenGate28:43Building Reusable Patterns and Long-Term Partnerships31:31AI Agents and Agentic Pipeline Challenges35:09Multi-Cloud Strategy and Ecosystem Integration37:17Conclusion: Building Better Outcomes
FAQs
How did Southern Company reduce outage analysis from 8 days to near real-time?
By migrating from fragmented on-premises SQL servers that could not communicate with each other to a unified Databricks Lakehouse with medallion architecture, Southern Company was able to correlate data from meters, power lines, and customer systems in a single governed platform. Oracle GoldenGate enables near real-time data ingestion, so root-cause analysis that once took 8 days can now run continuously.
What is a federated Lakehouse design in a utility context?
A federated Lakehouse organizes data from multiple operational systems—meters, assets, billing, and federal programs—into a shared medallion architecture while allowing different teams to publish and consume governed data products through Unity Catalog. This replaces the point-to-point integrations between siloed on-premises systems that previously caused delays and fragmented information at Southern Company.
What applications did Southern Company build on the Databricks Lakehouse?
Southern Company's application portfolio includes five tools: Ramp, Scout, Sphere, Magnus, and Nexus, covering use cases such as grid hierarchy visualization, network visibility, and vegetation management. The team evolved from Power BI reports to full web applications, enabling richer analytical experiences for grid operations teams managing power delivery across Georgia, Alabama, and Mississippi.
How does Southern Company integrate Databricks with Palantir and Google Cloud Platform?
Southern Company uses a multi-cloud integration strategy that connects Databricks with Palantir for specific analytics workflows and Google Cloud Platform for flexible AI workloads, allowing teams to use the best tool for each problem while maintaining regulatory compliance. Unity Catalog provides the governance layer that enforces data access controls consistently across all integrated systems.
Full transcript
[00:08] Um so we have with us here Kaiden Whitbeck from Southern Company and Gregory Stollman from NTT Data talking about powering the grid. Very thank you.
[00:27] All right. So as she mentioned, I'm Greg Stollman with NTT Data. We're a gold partner with uh Databricks. I lead our strategy and solutions group, so working a lot with organizations kind of taking the next journey, modernizing their platform, and really kind of enabling analytics and AI and kind of the next generation of use cases.
[00:43] Um with Kaiden, um long-time partner with Southern Company, they're one of the largest utilities in the US. I think they're the number two. They service Georgia, Alabama, and Mississippi from power, gas, and a variety of other analytics uh or utilities um services.
[00:59] Yeah, thanks, Greg. Um appreciate the opportunity to come out here and talk to everybody. Uh my name is Kaiden Whitbeck. I'm a manager of data engineering and architecture in our power delivery um data analytics organization. So we are in the business. I've spent a lot of my
[01:16] career in IT, and I'm going to try not to say TO because at Southern Company we call IT TO for technology organization. But it's uh the IT. So I have uh spent a lot of time with NTT over the years. I first started working with them about 10 years ago at our
[01:32] subsidiary up in Chicago. Um and since that time, found them a real a real good partner to work with. So today we'll just go through some of our journey at Southern Company of Databricks, and um hopefully you guys can get something out of it, and also
[01:47] appreciate you cutting lunch short. It's right in the middle of lunch, so um glad to be here. So with most organizations, uh utilities aren't very aren't any different, right? Their applications, where the data's coming from, is very complex. On the
[02:03] screen is their current state of applications from meters to power lines to customer payments to billing to federal programs helping them either you give backs or discounts to different different customer segments. So, they have the same challenges a lot of
[02:19] different organizations do, where they need to bring data in, they need to organize it, cleanse it, and be able to have a platform to be able to deliver based off of their complex organizational data. Being able to do this in a silo, what where we were maybe six or seven years
[02:35] ago, on-prem systems, was very difficult. It caused delays, fragmented information. There was, you know, on-prem SQL servers that didn't talk to each other, so you had to make hops, duplicate data, all those different, you know, legacy cumbersome technologies. So, really be
[02:52] able to move forward past this and being able to get onto a platform like Databricks, being able to organize the data, really helped them enable a variety of different analytic use cases. Yeah, we do come from an industry that has a lot of legacy debt, a lot of
[03:09] systems out there, a lot of complexity that interrupt our data. So, that results in things like this. We have a real outage request from our COO that comes to us. 5-6 years ago, that would come to us, and I would be like, "Give me 8 days and I'll get you the information." Even though a storm just
[03:25] happened, we have customers that were out, things were impacted, but he's still asking the questions, "What happened? Why did it happen? What's the impact?" And we have to go to all these systems here to get that information. And so, we're we know these systems. This is our vegetation management
[03:41] system, our AMI and metering system. We have outage systems, and asset systems. The system we have, as as Brett mentioned, we're three states, so we have different op codes, we call them, across these states. They quite often have different systems. So,
[03:57] you're talking different source systems for outage management. Um some of these are consolidated, different solutions for vegetation management. None of this is talking, right? We We know We know individually what's happening. We know our outage information. We know when we did the
[04:12] trimming on the trees. You know, we know the age and condition of our assets. But, you know, 5 years ago we didn't really have this platform or this place where we could easily connect it all together and integrate it into a single view. You know, so we can compare one time, you know, from the last time the
[04:27] trees were trimmed to this time. Think about that with the age of the assets, the condition of the assets. So, none of that was possible. And we didn't have, of course, a unified data model. So, without that, there was a lot of manual work, 8 days of it.
[04:47] And bringing that 8 days to something near real time is kind of that iceberg, right? They need to We need to lay the foundation. We need to build the platform. Um you know, you can do a lot of this in silos. I kind of look at the way people are enabling AI and advanced analytics today as people were within Power BI or
[05:03] Tableau a decade ago, where they would store a lot of the business logic inside of the semantic models, where it wasn't shareable, it was siloed. So, if you got a customer base that had certain aggregations or joins, you couldn't get it out, right? So, being able to take that back into the
[05:18] platform like we're doing with the visualization to data engineering, the same thing for AI, right? You can do a lot of data engineering cleansing within notebooks to produce an ML model, but then it's locked into that single source, right? The It's not shareable,
[05:33] it's not consumable. So, being able to take all of that, shift it left into a platform to where we can take the data, create data products that are reusable, that create data products that can service multiple different use cases across the organization. It helps speed
[05:49] up um the time to decision, be able to go from proactive going from reactive to proactive looking at the data. So, this really the AI depends on all these integrated data sources. Really enabling that foundation under the water stuff
[06:05] you can't see produces the cool stuff that you can do now on top of that data once it's cleansed, once it's aggregated, once it's joined into a common data model. All right, so I guess that the AI's important, right?
[06:22] AI takes data to run. I think the uh the thing I like to point out here is back to the days of Terminator, right? So, we had the movie Terminator, this guy comes from the future, he's an AI, he knows everything. If we would have just fed him bad data, the movie would be over, right? Um so, from from siloed systems to
[06:39] integrated data pipelines, uh we need to to get our data pipelines in place. And before this, we had systems that were specific schemas, right? So, there was a schema for every system, every opco has its own schema. Um we still have some of that today because
[06:55] the raw layer is low loading that data at that level, but it's uh in the past, there was no shared data model, we had manual joins and reconciliation as we're going through those eight days to manually figure out what happened, uh one-off analytics um workflows. Let's say you're just making requests, you're
[07:10] putting filling out forms to different groups to like say I need to download this data from this skater device. Um it's taking a long time. After we're centralized, we have pipelines that are running, right? So, they're they're they're known, they're standardized schemas across domains, and you have
[07:26] curated reusable data sets. So, this is repeatable analytic workflows, and this is where we want to get. So, if I talk to the business about these things, they're just like, "What are you talking about? Schemas? You want repeatable?" They don't get They don't get this. They don't get the work. They just hear it
[07:42] takes me 8 days to get the data. And what the business cares about is these applications. They want to go to one application. The data's behind the scenes, right? It works. It's all integrated. So, it takes a while to get there, but this
[07:57] is basically we're in a place now where we have like I'll call flagship applications for different areas. And how we got there was was a lot of work, a lot of buy-in from the business. Um some things that were really big successes for us were uh
[08:13] getting a communication person on our team. So, we have somebody on our team that like is really good at graphic design, at communicating, rolling out um you know, emails to people, releases of applications. Uh they're telling the new features that are coming up. And uh so,
[08:28] that's that's really just something that a lot of us technical people aren't as good at, and they can make it look really good. Um looking good is important, but also the buy-in is important. So, we have big these uh circles we call coins that represent our different applications.
[08:45] And uh they were key in helping design these and working with our business partners to design what these look like. And you think an image should just be Let me tell you, engineers and developers have different ideas of what images should look like from the business. And you have to like get everybody in the same
[09:00] room, and there's a lot of arguing about this. And like, I like that color better. I don't like that doesn't represent my business. We have three opcos that are all like trying to come together in one place for this application. I would say the benefit is is getting in the room and getting buy-in from everybody. So, they're there, and then when they see this this
[09:16] image, they're like, "Hey, I was the one that helped to I I was the one that fought for it to be blue, fought for the arrow over the the next thing." So, uh and and we've seen evidence of that where we'll be on Teams meetings on video calls, and they'll have, you know, their bookshelf behind them. And we
[09:31] actually print these coin we call coins into metal physical artifacts that we give to everybody. They put them on their bookshelves. They display them on the video calls. They're proud to be part of the team, part of the program, right? So, then they're like bought in because these things are iterative. We're changing them all the time.
[09:47] There's new releases coming out. So, that's been a real success for us. Um as far as the uh the coins and what they do. So, the first three here, we have this blue one, we call it ramp. It's kind of the uh the past, present, and future of power delivery data. So,
[10:04] ramp is all about um what's happened. It's a reporting. It's on regulatory and compliance, right? So, different different states, they have different regulatory bodies, but um you can go to one place and say, "How are we doing, you know, on our metrics?"
[10:19] The next one is um scout. So, that's our outages. So, current outages, what's going on. You can go in there and say uh for this premise, you know, what's the history of outages for the last 5 years? Um so, that's that's very helpful. The next one is called sphere. It has this
[10:36] little icon that looks like a torna- a hurricane in it, right? Cuz it's very weather focused. So, um this is the future. So, this is the storms are rolling into the southeast where a lot of our customers live. We need to figure out where our crews should adjust, you know, where should they should position,
[10:52] where the outages are going to happen most likely based on models. Um weather is super unpredictable, right? So, we're always trying to tweak what those models look like. We've actually hired a meteorologist on our team who helps with that, gathers data. Um we're building more and more mature data products to
[11:07] help feed this application. So, we even have um situation where we're buying um devices that we put on our poles, right? That are little weather stations. And uh we can we think that some of those might source the data better than even the
[11:22] stuff we get from NOAA or from weather services we buy data from. So, we're always measuring um and seeing like, you know, is there a better better answer to this weather data, but um that's Spear. So, and and also these are accessible outside the company. That's been a a battle with security, but you
[11:40] can go to your phone, right, and be out in the field and and connect to these. So, uh the next one, the red icon here, is called Magnus. Um that's our our network underground. So, these are These are our customers that are the most sensitive for outage. Think under cities, under stadiums. These um devices
[11:57] are in vaults, you know, submersible, um so that water can hit them and they don't get impacted. They're highly redundant. And uh but we need to monitor them for to see what the condition is of them. And and so we are doing a lot of condition maintenance, condition-based maintenance on these assets um in this
[12:13] application to try to help people uh the business understand what what needs to be replaced sooner. The green icon here is um our customer, so it has the house on it to signify behind the meter. Um so, it's all of the the things like what's the customer's
[12:30] profile here, what kind of power usage are they using, are they somebody who's charging electric vehicles, do they have multiple um AC units? So, we can tell, you know, of course, by watching the power there and um profile the customer. So, that's that. This last one's interesting. Um we call it Nexus and uh
[12:48] it started out we we used to call it our sandbox environment. And it was a lot of Power BI applications that we're trying to figure out like, is this a good metric, you know, we're we're we're innovating here. It's kind of like an incubator. Well, the business didn't really They were like, "It's a sandbox.
[13:03] We don't really want to go look at that." That's all saying, "That's interesting." So, we said we probably need to rebrand this into something that's more of a product. So, Nexus sounds cool and it's like, you know, not Lots of times the business might not even realize it's Power BI. And it's not always Power BI. I mean, the rest of these are all like web apps running on
[13:20] Azure, right? So, they're more the traditional deployment of a custom app um that points back to Data Bricks. But this one could be whatever. It could be whatever thing we're innovating on. We put it on one place and then Yeah, Nexus. And just to kind of give an example of
[13:36] working with uh um the you know, product ownership, working with uh the customers, there was one thing that we developed recently which was the grid analytics and grid insights. So, having data coming from multiple different sources, being able to aggregate or being able to aggregate it in one spot brings challenges in and
[13:52] of itself because now you have a lot of data, right? Um, you can think about all the different power lines, the different poles, the transformers, the meters, the feeders, all the ways that you can deliver power out to your different customers. That's now all documented and tracked with GIS um and different
[14:08] waypoints, different hierarchies from those different pieces of how it's put out. Um, that was originally developed once we put it in into Power BI, but it wasn't performing. It was 10 to 15-second load times once they changed filters just because of the size of the data wasn't really, you know, cumbers-
[14:25] it was too cumbersome for Power BI to handle in memory. We tried direct query, mix of direct query plus stuff in memory. Um, and then, you know, one thing we did is um and he mentioned that a lot of things are applications now. So, we took it and this was right when
[14:40] they started releasing lake base. It wasn't approved yet, so it's something that is on the road map as it being able to make it faster, but we were able to we were able to vibe code and create a application um on top of a uh serverless SQL warehouse hitting the Databricks
[14:57] tables and we reduced the we basically recreated all the Power BI reports over a few months um with the business, with the different integrations, and uh we had the refresh time went from 15 seconds down to about a second or two. Um, we're looking to improve that once
[15:12] lake base and then also Databricks applications become available internally. I think we're working on getting those things approved internally by Southern Company now that they are GA, but it's another way that we're utilizing a system that is being able to scale large data to be able to deliver it faster and quicker to the end users.
[15:35] Yeah, so here's the slide of the journey of where the industry has gone. And you know, we had data warehouses back in 1990s running off relational databases. These things work great. We still have them now. We still use them now on premise. We have We are transitioning to Databricks, but the thing about these
[15:51] these warehouses is they weren't always the best for adapting to changes. So, some change happens to the source system, you have to ripple that change through the extraction transformation and loading layer and then change the the star schemas and the dimensions to adjust for that. All that took time.
[16:07] But So, we and also they didn't track really any kind of semi-structured data, right? So, enter the the lakehouses in sort of the 2005 to 2015 time frame. We still have this as well, retiring this year though,
[16:22] running on Hadoop. And but we got a lot of mileage out of this. It was one of the first places where we could load unstructured data, image files, all of our meter data or detailed meter data is here, right? So, we could run ML models on these and that was that was great, but you've probably heard them
[16:39] referred to as data swamps at times during that time. The data would go in, you wouldn't get it back out. So, enter Databricks. We've been using Databricks since about 2019 and it's kind of the best of both worlds. We have the structured data warehouse, we have the lakehouse, we
[16:55] have the compute in the cloud. We have things like Unity Catalog, metadata, we have a scalable platform with compute that can scale, right? So, all these things are amazing and sort of our journey to the lakehouse. And so we have the federated lakehouse concept
[17:12] at Southern Company. And we started with our customer data, like I said, back in 2019, that was a lot of data um focused on obfuscating customer data and sharing that data with Delta sharing out to customers um that needed to see we we wouldn't share the actual customer
[17:28] information, but we'd share like some usage and they'd use that and we could share that with some of our partners. Um so, that's an example of a data product that we're building off of this data customer domain. Every one of these domains, these white bubbles, have our IT teams and stewards behind them.
[17:44] They're supporting them. Um we have about we're coming up on about 15 of these now, but um so, they're growing. They're not all displayed here, but the next one we did was the Southern Gas. I came from the gas organization for most of my career and we built a lake house there also off the customer data. It had
[18:00] um some products, the gray bubbles are are products. So, we have like safety and damage products off that data. Um that kind of followed up with I'm going to tell you about our shared lake house. So, we have a domain that's that's owned by IT. I had sort of flashes of of um
[18:16] Ollie talking about Panther this morning. So, Panther was that acquisition they made for security, if you guys heard that in the keynote. Um since this is an IT-centric domain here, that's a lot of some of the data they're putting in that. So, I feel like we've been ahead on like I I don't know what Panther's going to do, but it sounds
[18:32] like it's going to pull a lot of the security data together into a domain, right? Sort of like we do this here, but we saw other uses for this cuz we have a lot of other data at our company that spans across the opcos, not just IT security data. So, like we're standing up a fleet lake house on that. That's like fleet like vehicles. So, power
[18:49] delivery organization happens to own most of the vehicles, but there's also some in gas, some in nuclear. They're all spread around of our around our holding company. So, it fit in that lake house and the really cool thing about it is is we can't scale up each of these data domains for a team every time. We
[19:04] need to have like teams that can reproduce like like my team in the business is going to be able to own this fleet lake house more directly, right? And not have to TO and our the IT they'll still be involved. They'll administrate the cluster. They'll do some governance. But at the end of the day, we're the ones
[19:20] that going to be totally loading the data from a software as a service the business went and bought. Um interesting, it's lot easier for my team to load data that comes outside our company than from on premise. So, the on premise data is very controlled by a lot of our our TO. And and there's it's
[19:36] sensitive data. There's a good reason for that. But um Luckily, we have Unity Catalog there in the middle that lets us hook all these together. So, these these all look defined, but they're really it's really about who owns these different domains and who knows the data and then who can
[19:52] approve access to it in Unity Catalog at the end of the day. So, this is evolving for us. Yeah, and and one of the areas we've mentioned repeatable products is we recently did a project with driver safety for the fleets. So, taking all the IoT sensors from the cars, the
[20:08] drivers badges once they badge into the car to see kind of what they're doing. That was one different which is one organization within Southern Company. But that was able to be scalable now since it's repeatable, we have a pattern. We're going to get other data from gas, other different opcos, and
[20:24] bring it into the same model. So, it becomes a product that is reusable. It increases time to value and time to market for those different other areas versus you know, we have the data and now redoing that project somewhere else and then in another silo. So, it's in this
[20:40] kind of modular format where data is available in the cloud through Delta Sharing and through different other kind of you know, mini lakehouses and we're able to bring it in to create one product across the organization.
[20:56] This is the slide for the medallion architecture and what I'll hit here, you guys know this, but what I'll hit is who loads what part of this medallion architecture. At Southern Company, I find it's very important to include our TO or IT counterparts on the raw layer. So, their job is to get us a one-to-one copy of that source system.
[21:12] They're better equipped to maintain that than we are in the business. They know about the upgrades, the outages that are happening in in the IT side. So, let them manage that raw on-premise data. It gets very, very complicated. You saw your Greg's first slide there.
[21:27] It a lot of secure layers they have to go through. So, they take it there. My team and the business teams can take it over in the curated and aggregated layer. They give us a one-to-one copy. We can watch that, build history. We may want to, you know, every time something change, we may want to track that in our star schema in the
[21:43] curated layer, and then aggregate it further, change the data around for a final data product at the end. And And the key is who does it. So, my team can work further over in the in the higher layers here.
[22:01] So, this is example of a data product we partner with NTT with. And And Greg mentioned it earlier. It's I'll hit real quick what a data product is. It's a defined schema that has business meaning. It's governed access, quality controls, reusable transformation logic. You do stuff that's built once and used multiple times. So, these are the legal blocks of
[22:17] our of our data. We have the block we can build other things on top of it. Data becomes, you know, valuable because it's in the curated and aggregated layers at that point, standardized, validated. These are all the things we want with data products. Repeatable, so we don't have one-off analysis. So, the
[22:33] grid hierarchy product that NTT built for us is a good example of this. It's pretty much getting clean copies of our data off the grid. When you have the transformer substation, transformer banks in the substation, meeting those up to the feeders and the secondary lines, and tracking that. And that can
[22:50] all change, right? Depending on what's going on with the self-healing grid. So, real important to track the history of that over time so we can get the system-level domain insights, check the products on top of that. So, now we're building load aggregation and network visibility products on top of this base
[23:06] data. So, we can figure out things like intermittent outages, right? Those are really tough to get. The grid changes and it's configured different from one day to the next as storms roll in as things change. You know, one day a device might be overloaded, the next day it's not. So, you need to see the history of that over time and these
[23:22] products allow us to sort of correlate with load aggregation with outage information all in one place. So, that gives us products that give us things like operational analytics for outage and reliability. My team can build products beside that that do observability because the data's no good
[23:38] if it's not good quality. So, we have products that sit right there that we monitor ourselves. IT might be doing something with their one-to-one copy over in the in the raw, but when it hits our layers, we're watching it and we're saying that loaded last night or that didn't load, right? So, that's an observability product that we we build
[23:55] and monitor and we're evolving that too cuz there's lots of solutions for data quality and observability. We kind of have a custom one built right now. But, um business reporting, we have a lot of engineers out in the field that are pretty smart. They have Power BI. They can go get this same data product and
[24:11] build their own dashboards. If we like it, maybe we'll pull it into Nexus and show it to a bunch of other people, right? So, we're constantly working with people in the business. And then, we have AI and predictive models and NTT's been great. They have some great data scientists and and engineers that are helping us build
[24:27] those products like this. Yeah, just one of the things we did was their vegetation management a handful of years ago and being in Atlanta, if anybody's from there, been there, it's coined city in a forest. We're very proud of trees in Atlanta keeping our
[24:43] old growth trees and a lot of neighborhoods we live in were built in the 40s and 50s. So, and we saw above ground power lines. So, when we're trimming, you know, they come by trim trees to make sure everything's more reliable. There's no you know, overhanging trees on the power lines so the storms don't affect power
[24:59] outages. You know, being an Atlanta-based company when we were working with them, you know, we did have a lot of people jokingly come to us like, "Can you guys, you know, auto program into your AI predictive models that our feeder gets trimmed an extra time or two?" Just joking around, but um you know, it is kind of neat to
[25:14] be able to see, you know, what how trimming affects um you know, different outages, where different growth is, and I think we're trying to take it the next step and start to use satellite images in the future to be able to look at, you know, where the uh growth is, you know, from a
[25:29] overhead perspective versus just a reliability history pattern based off of different, you know, events, different outage patterns, and how the, you know, different trees and vegetation grow. Yeah, totally right. The area We have a a lakehouse coming up for aerial service, another one of those shared lakehouses, right? So, we fly a lot of
[25:46] lines. It's not always coordinated when people are flying. We'd like to put that all in one place, optimize it a lot more. So, yeah. So, that takes us from 8 days to near real-time, right? So, we're doing all these steps that I've already talked
[26:01] about. Um don't need to rehash that, but we get to near real-time. All our data isn't near real-time, of course. Uh most the majority of our data still comes in on a nightly batch, and that does great for most everything we need. But, there are absolutely things we need need need real near real-time, and we need to work
[26:18] with our IT counterparts very closely with those. It's very complicated. We have a number of different solutions. We use Oracle GoldenGate to get key tables over over to our lakehouse um near real-time to get to those sub-second type responses.
[26:33] Um and it doesn't always it doesn't always go sub-second. There's like a lot of complexity involved with near real-time. But, we are at a place where we can offer that now, and um you know, allows faster real-time analysis and accelerated decisions, and
[26:49] um some reduced manual engineering effort. So, how they kind of we enabled this all together? It wasn't something that, you know, NTT did on their own. It wasn't something some Southern company did on their own. It's really kind of a deep partnership between us, between the business, between the end users. Uh
[27:05] there are also other SIs uh within the organization that we need to work with. Um you know, it's a large organization, so we're all doing a part. So, having that kind of relationship with between ourselves, between other different entities, and Databricks, you know, they're very heavy in uh with solution architects coming in, working with the
[27:22] team, letting us know where, you know, there's new things that are coming about that can help enable Unity Catalog when it was uh brought up uh and how we can migrate to Unity Catalog a few years ago. Uh the cloud infrastructure, obviously, the scalable platform that we've talked
[27:37] about. So, you know, it's not it's no longer siloed in a data warehouse on some uh server that is not accessible. The infrastructure gives us the ability to scale up as needed to give us the compute so we can access um all the different large data sets that we have.
[27:53] Um the one team mindset really kind of goes with that partnership relationship. We work together. We have a lot of employees inside of Southern Company that have a Southern Company ID to log in. You know, they join meetings, and I think one of our client partners jokes
[28:10] often that they don't really even know that they're a part of NTT. They just think they're a Southern Company employees. So, and sometimes that might not be the best for our visibility as an organization about what we're doing for them, but it just kind of shows the relationship that we work together. We're not really trying
[28:27] to do anything on our own without without the partnership of the other parts of the organization, and we want to deliver for that final end goal. And the last thing is the agile development. We have to be agile in these environments with multiple different partnerships. We're all working on different parts of in an
[28:43] overall engagement to deliver that final end goal. So, the agile gives us the ability to focus on certain areas, certain tasks to deliver a product. Yeah, NTT does a really good job of sticking with us, making sure that we're in the room when they're in the room. A lot of our business wants to make sure
[28:59] there's employees with contractors at the same time. They do a good job of getting right in that line, but we're investing in NTT and the people on the team, too, because they're learning our data along with us, right? They become very knowledgeable of the data, so it's definitely a partnership between our companies.
[29:20] This kind of really kind of goes back to re-reusable, um, you know, assets that we're building. So, we're building patterns, not one-offs. This gives us the ability to scale, right? So, we can be able to scale, um, either ingestion frameworks, we can scale products that we're building out that can be shared across the organization. Um, we don't
[29:37] want to use new unproven vendors, uh, because it doesn't really, you know, there's hiccups, there's security, there's different reasons. So, we're using existing partners like Databricks. They're founded in what we're trying to accomplish and really get help accelerate in and deliver the products that we need to.
[29:53] Um, so, we're using existing contracts, partners like ourselves, but there's a few other different SIs that are within the organization. Um, the solutions are built with subject matter experts in the room. Kaiden mentioned bringing in meteorologists, but also
[30:08] organizations within the business, right? We need to bring in, you know, it can't be an IT-led engagement because it's never going to hit the mark. You're kind of blindly throwing at what the target is. So, we need to understand what we're delivering, what they need, the different data elements, the critical data elements to make decisions
[30:26] and the key metrics they need to continue to drive their business forward. Yeah, and my takeaway here is the leverage of existing, um, contractors and partners. Like, you just don't want a whole bunch of these. There's You want the ones that know your data, that are there for the long haul. And if you start getting new
[30:42] people in all the time, it just takes too much especially in data engineering, it takes too too much time to ramp up the people. So, you get a lot of synergies by having long-term relationships like ours, you know. So, this is more of a kind of a future vision slide. We're not doing everything
[30:58] today. So, there is you know, that top layer is building our our agents, building our orchestration layer all the way down to the bottom where you have self-healing. We're not at self-healing at Southern Company yet today. It's in the future, but this is just kind of you know, from start to
[31:14] beginning how we're going to build everything together to have the foundation and be ready for agentic AI operations. Yeah, this is definitely a strategy slide, right? So, you're thinking for us it is anyway. So, strategy is what you're trying to do right now today to get to that strategy, but long-term like
[31:31] we're not self-healing yet, but we want to get there. Um I was very interested in Ali's comments this morning about the problems with AI agents and I'm like, man, he's still on my slide, right? This is kind of the same thing that we have and I can tell you a story about bumps in the road of
[31:47] code of vibe coding, right? So, we're working with our forward deployed Databricks engineers and we have this workshop planned about 2 weeks ago. We've you know, NTT developed a power app for us that's you know, it's running running against SQL Server on premise. We want to that
[32:04] we want to use this workshop to vibe code it. We're going to take the power app, download the YAML out of it, has all the configuration. We're going to feed all this to the planning agent, right? So, we use Copilot for GitHub. So, we have Copilot for GitHub. We're feeding the the YAML files to it. We're feeding like a transcript where somebody
[32:21] demoed the application tell sort of what how that how the application works. We have a guide, a user guide we're feeding to the planning agent and we're going to say, all right, let's build a Databricks app. So, it's going to all run on Databricks now, Databricks app native And um, it works great. The night before
[32:37] for prepping for the workshop, we get to the workshop and we try to run it. It's running off the expensive um, Opus model, right? So, that's all of a sudden, it's all shut down. Like, and out of coincidence, the only models that are turned off are the expensive ones. We're like, "What's
[32:54] going on here?" Uh, so we call the help desk. They don't know what's going on. Nobody's communicated this. Uh, they tell me post to this webs this website where the where the the where they're uh, watching this, the group that manages that. And they did respond pretty quickly later that day, but blew up our workshop. Um,
[33:11] we tried using uh, Sonnet, which is a lesser model, right? And it just wouldn't do the job. It didn't do it like it did the night before. So, um, I'm not going to say it all goes smoothly, but what had happened is IT had basically put in governance of the cost, right? And they had just said,
[33:27] "You know what? That's the expensive model. We don't want one person to just spend hundreds of dollars on that model." It was a thing that made sense to IT, but blew up our our So, this is the rocky road you're living in with this this uh, agentic agentic pipelines and building things.
[33:43] But, in the future, uh, we did solve that, you know, it's it's amazing what the vibe coding can do. It um, can create the applications very quickly. Uh, and we're getting to the point where we want to use that more and more. I still want to get to the place where we deploy it because if vibe codes it and then
[34:00] you're like, "There's a change. Is it really going to adapt the way you think it is?" I'm really curious how genie code is going to help that. I mean, we had to do that all on a machine. A lot of prerequisites go into like putting all the components on that machine to do it. So, I'm really encouraged that, "Okay, Databricks now
[34:17] has, you know, the genie code out there." They announced that it's going to be like all hopefully in one place, like Databricks does. So, um, should streamline that effort so that we can do things like this slide says, you know, "I want to build a new pipeline. It should connect from here to here, build
[34:33] that out, pass it down to the specialized agents. They're going to do your data quality checks, your lineage checks, um push it to the to the um platform layer and the application layer, and then at the end um you should have these agents that are monitoring and doing self-healing. As you said, we still keep the human in
[34:49] the loop, make sure our guidelines are followed at the end. Um if you want to deploy what it recommends, but I think when you have things like schema drift in the future and some of those things that break your pipelines, I I feel encouraged that AI is going to be able to handle that pretty pretty quickly.
[35:09] Okay, so this is the uh Databricks has all our data slide, right? So, and and that's great for us. It's one platform, one source of truth, all the things that that gets us governed data, eliminates duplications, simplifies downstream models.
[35:24] Um you have control without being locked in. Uh this is highlighted by our business going out. They have budgets for things. Um our transmission organization has a lot of needs right now. If you're following the power Do you know the power industry right now? There's all
[35:39] these data centers out there that are requiring transmission lines. So, we have billions of dollars in capital construction projects. And guess what's blocking them? The planning of them, the deployment of them. A lot of times it's the data, right? We don't have the data from the supply chain, we don't have the data um to to do the planning we want,
[35:56] and so they've been able to go out and take budget and justify bringing in other partners like Palantir, right? So, Palantir has a pretty good um foothold on that in our organization. They do They've done a great job unblocking that data for us. And the good thing about
[36:11] that is is we hooked it up right to Databricks. We have, you know, a connection that is real time. They can do uh compute pushdown, so they can write their models in um AIP in Palantir, and push that down for execution on Databricks, and then return
[36:28] the results, right? So, that open architecture helps us We have Google Cloud Platform out there. We have a a fairly new CIO who is like, we want to do multi-cloud. Um we want It's the AI Center of Excellence that is brought in this Google, and they do they create
[36:45] agents, they call them Atoms, they can go do units of work just like agents do, but they can also consume the data from Databricks, and we have been here, we have those data products. That's not going away, right? These are the blocks that everybody else can use, so it's been a a good story that we have
[37:01] Databricks in front of that and and can maintain some control of our data. And so, in our final story or final slide, you know, Databricks is a great platform, but I think it's just part of the equation. So, it's not you know, we
[37:17] get Databricks and then all of a sudden everything's solved, right? So, you know, there's a lot of planning, there's a lot of migrations. It's really starting to build better pipelines, um having faster trusted decisions from those pipelines because you know your data's accurate, working with their internal TO organization to ensure that
[37:34] data's replicated out, it's cleansed, it's processed, and it's delivered to the business so they can make much quicker decisions instead of in 8 days, either next day or near real time. And then this gives us better outcomes, right? So, we have data faster, we have data in our hands to make those
[37:49] decisions, we can understand where the outages are. You know, everybody here has power going to their house when there's a storm, the last thing you want is your power to be out for a long period of time, but we can get that quicker, we can they can spin up crews and be able to understand where the power's out and get it restored to the
[38:06] different customers. All right, well, we thank you for your time. Appreciate it. Yeah. Any
Learn more about the Databricks Data and AI platform.
The information provided herein is for general informational purposes only and may not reflect the most current product capabilities or configurations.