Skip to main content

Baker Hughes Data Transformation: From Excel to AI-Driven Insights

Summary

  • Baker Hughes transformed from a legacy report factory—where data was downloaded from systems, scrubbed in Excel, passed through multiple people, and posted to SharePoint—to a trusted analytics platform by building a unified data lakehouse using Databricks medallion architecture.
  • The transformation delivered zero-touch data sourced directly from high-quality real source systems, paired with ThoughtSpot's agentic analytics Spotter capability for natural language access, giving business users a direct relationship with their data.
  • An open contribution model and single-path-to-production automation democratized data engineering at scale, while new data visibility uncovered operational issues that had previously been hidden by pre-filtered, manually assembled reports across Finance and Manufacturing.

Baker Hughes Data Transformation: From Excel to AI-Driven Insights

Watch: Baker Hughes Data Transformation: From Excel to AI-Driven Insights
Baker Hughes transformed from a legacy report factory into a strategic platform powerhouse in two years by consolidating on a unified data lakehouse. The challenge: data scattered across systems, scrubbed through spreadsheets, and disconnected from the business. The solution: building trust through transparency, quality data, and putting analytics directly in the hands of users.
Learn how Databricks medallion architecture combined with ThoughtSpot's agentic analytics enables business users to have direct relationships with their data. Discover how automation, single-path-to-production, and open contribution models democratize data engineering at scale. See real examples of how data visibility uncovered operational issues and transformed decision-making across Finance and Manufacturing.
🤝

Chapters

FAQs

How did Baker Hughes move from Excel-based reporting to AI-driven analytics?

Baker Hughes rebuilt their data foundation using Databricks medallion architecture to create zero-touch pipelines sourced directly from high-quality real source systems, eliminating the cycle of manually downloading, scrubbing in Excel, and uploading to SharePoint. This video outlines a two-year journey including team reorganization, automation, and a single-path-to-production model.

What is ThoughtSpot Spotter and how does Baker Hughes use it?

ThoughtSpot Spotter is an enhanced agentic analytics capability that allows business users to interact with data through natural language queries, going beyond traditional dashboards. Baker Hughes uses it alongside Databricks Genie as part of their analytics strategy to give business users a direct relationship with their data without requiring SQL skills.

What is an open contribution model and why did Baker Hughes adopt it?

An open contribution model allows engineers and analysts beyond a central data team to contribute data assets to the platform under standardized governance, democratizing data engineering at scale. Baker Hughes adopted this approach to expand their platform's coverage while reducing the bottleneck of a single centralized team, accelerating analytics delivery across Finance and Manufacturing.

What did Baker Hughes discover when they first showed executives data from their new platform?

When Baker Hughes delivered zero-touch data from their medallion architecture, business executives initially said 'That can't be right' because it differed from the pre-filtered, manually assembled data they had previously seen. This reaction revealed previously hidden operational issues and misalignments between what executives saw in dashboards and what operational units reported.

Full transcript

[00:07] So, here we have the usual forward-looking statement that you've seen 10 times already today. Um It's important for us to get feedback, so please do complete your survey at the end of the session or at some point today. That'd be great. So, we're going to talk a little bit
[00:22] about how Baker Hughes moved from disconnected data to um AI driven insight. That's it. Okay. So, my name is Paul Thompson. I am I
[00:38] lead the implementation of the Baker Hughes data and AI strategy. I'm based out of Houston, Texas. I'm Jennifer Lowell. I'm one of the data architects in the group uh who has helped to actually implement the solutions.
[00:53] A little bit about us, Baker Hughes is an energy technology company rewriting the energy equation by leveraging our unique position between the industrial outcomes and the energy sources. So, I want to walk you through a little
[01:09] bit of a high-level roadmap of our journey from disconnected data to trusted AI-driven insights from the over the past 2 years. Um we'll talk about some of the problems we faced and we'll also talk about how we reorganized our group
[01:25] to face those challenges and then we'll highlight some of the detail of how we accomplished this, too. So, that can't be right. This is the first thing everybody tells says every time we show them some um up-to-date data that's based on quality
[01:42] data from real source systems because the business executives are so used to dealing with this data was downloaded from this system, scrubbed in Excel by this person, passed on to this next person, given to a third person to put in a SharePoint and then somebody
[01:57] points a a dashboard at that SharePoint and then they may update that once a month or they may not. Um but when you actually give them zero touch data which is sourced from qual- high quality um medallion architecture that we've built,
[02:13] first thing they always say is, "That can't be right. That isn't what I saw last month." So, that's where we start. So, just to give you a little bit more about that starting state, we really
[02:28] have this difference between what our executives would see and what our operational units would be providing and they rarely aligned in the middle. Data would be pre-filtered. It would be missing values. They would take out
[02:44] things that they thought were irrelevant to their views but really had business value. It was really difficult to understand what the source of truth was for the data. One of the other challenges was reference data. Each group seemed to have a different mapping of what regions
[03:01] and product companies and business units they were displaying. It would be shown in different ways and have slightly different groupings depending on how they wanted to look at it. One of the best examples of this is when
[03:16] they would go with these presentations and on the slide you would see an image of a snapshot of an Excel spreadsheet that somebody had scrubbed and made to look the way that they wanted but there was zero visibility into the data
[03:33] underlying that and you couldn't click into it and show them what exactly that data was sourced from. So, we had a senior executive come to us and he says, "I feel like I'm driving a car down the highway with a blacked-out windshield.
[03:49] I've got no gauges. I've got no fuel gauge. I've got no speedometer. I've got a rearview mirror. All I can do is see where I've been. And he asked us, "Give me Take the Take the cover off the windshield. Give me some fuel gauges. I need to know how
[04:05] much gas I've got. Am I going to reach the destination? How fast am I going? Am I running out of oil?" So, it was a really good way of for him to sort of describe to us the challenges that the business faced when trying to make data-driven decisions.
[04:23] So, how did we develop from that starting stage? Well, I was brought into the group um to lead the delivery of our largest data group. I'm not taking credit for everything. I just I came in as a software guy. I often say to people, "I'm not a data guy, really. I'm a
[04:39] software guy. I just got abducted by the data people." But, actually that turned out to be a strength because organizing the data group to run like a software organization really, really helped. So, we were more productive and we're more responsive to business needs. But,
[04:56] so what did we really change? Well, the big thing was implementing an Agile methodology. So, we have prioritized backlog of everything we need, various levels of detail. It's constantly evolving to meet the business requirements. We have
[05:12] really vetted requirements and acceptance criteria. We um adjust our sprint lengths according to the needs of the business unit. Not every business unit wants to have a a bi-weekly sprint review. They might want a monthly sprint review. That's fine. We
[05:27] can We can adjust to that and we can adjust our workload and the the the work that we commit to in each sprint at at each um at those points at a at a pace that the business needs. Um Acceptance I mentioned acceptance
[05:43] criteria. We make sure we meet them. We self-organize to meet those things. Somebody might be working on a dashboard one one one sprint and they might be doing some data engineering another sprint. We use all the skills we have on the team. Um we
[05:59] use lineage to determine downstream impacts of changing a data model. What we have used to have in the in the bad old days is someone would add a column or remove a column or change the meaning of a column without really any caring about anything that's downstream of it. And then the next thing you know is
[06:15] someone says, "My dashboard doesn't work. The revenue report is broken." Things like that. So, we actually uh taking a much more software development approach to develop to doing data engineering. Um we also implement a lot of automation
[06:33] so that our data engineers can focus on writing SQL cuz everyone understands the language of data and the automation take care takes care of all the rest. Yeah. So, one of the ways that we really enforce the automation and how that's
[06:49] the key to how we industrialized our group was to enforce that single path to production. All code that all of our engineers write must be committed through get branches passing through our CI CD path pipelines.
[07:05] Every new table that they create has to have automated test cases associated with it that run at deployment and making sure that we are enforcing things like uniqueness. Change data capture is there by default.
[07:22] It's automated into the system so that our engineers can just focus on creating the data model and they don't need to work on making sure that we're capturing the changes. This is also handled as part of schema evolution. So, if we have a new
[07:37] requirement to add another attribute, we can easily do so without having to truncate a table and losing all of that data. It's there. It's automated, so we don't have to worry about it.
[07:55] So, building the analytical views was a heavy lift, but it was a necessary one. We have all these um strategic ERPs and other data sources that we're pulling together to deliver data for the organization. Um we had stale, poor-quality data in
[08:11] non-production environments, which made it really difficult to develop things. Um so well, we also had deprecated data sources, but they were still maintained. No one ever shut off the pipelines that was feeding essentially nothing. So, we you know, we we we decided to leverage a
[08:28] data catalog so that we could actually explore the data and see what we actually had. Yeah, and that was actually a really important part of our journey because when we were handed Databricks, we had no idea what we even had available to us. So, that data catalog was really important for us to be able to
[08:44] understand what we had at our fingertips, what we were lacking, so that we could have that added in and ingested for our engineers to be able to leverage. Um the next problem that we faced was not having adequate buy-in from the business
[09:00] for having the subject matter experts to help us understand the context and logic for the business requirements themselves. They would say, "We need this report, and it needs to look like this one that we've been producing for the past 10 years, but I have no idea
[09:16] how they get those numbers because when we'd roll it up, it didn't look the same. So, we had to do some reverse engineering to try and get to those same data points. And then once we figured out those numbers, then we'd go, 'Okay.' Sit down with the business. Is this
[09:32] really what you want to be doing? Is this really what you want to be showing? These are the transformations that we had to perform to be able to get to this data. Um the the matching of that data was obviously a a huge challenge for us.
[09:48] Another piece of it was the reference data. So, because we as I said before, each of the groups represented their business in slightly different ways. Everybody had all of these different mapping tables. So, standardizing that reference and helping people understand
[10:05] this is what is the definition of this region and this product company and this business unit was really important to how we could proceed.
[10:20] So, once we had this structure in place, we also had all of this visualization. When we started, we had no great way of connecting directly to the data sources for our visualization platforms. Where in different segments, we were using different visualization tools. We have
[10:36] Power BI, we have ThoughtSpot, and we have Tableau. Um our users wanted canned reports for the most part. And we would see that sometimes they wouldn't necessarily know
[10:52] what they wanted until they had it in front of them. One of the great things about ThoughtSpot was the ability for our our customers, our internal business users, to be able to create their own answers and their own live boards so
[11:09] that they could slice and dice the data the way that they needed to be able to see it for their reporting and for their business insights. Um we found that this is really important, especially for our finance groups. They love using ThoughtSpot
[11:26] because it gives them that ability to really delve into the data the way they want to see it. So, we have here a quick little quote. Uh ThoughtSpot offers that the data can be adjusted with just a few clicks. There's no heavy lift required,
[11:42] and they can react to what the business needs at that point in time. So, talking about changing business priorities and requirements. Our data needs vary wildly. Some people want head count
[11:58] reports, some people want invoices, people want cash, people want purchase orders, whatever it may be. Once we had that medallion architecture in place and we built these robust, repeatable, automated solutions, providing the business with custom-tailored, specialized views no
[12:13] longer required the significant effort that they did in the past. So, um we were at this stage of maturity where we have a lot of gold schemas, a lot of gold views that our business users are are looking at through various visualizations. Some of them are looking at it in Databricks.
[12:30] Now we can start switching on the more advanced smart features of things because we've got data that that can actually feed into those systems without having to pre-scrub it. So, when we started, we started with Genie because that is what we had
[12:46] accessible to us within our organization. The refinement of the instructions and the example queries that we had to load into the system took time. And we found that one of the major challenges was user enablement. That is a critical component because often what we found
[13:04] was our business users were asking questions outside the scope of that intended Genie space. Even though it was broadcast to them exactly what it was for, they would still ask questions outside of it. So, this became a challenge
[13:20] to make sure that they understood and that we could react to what they were actually asking. Um, the embedding of business-friendly metadata and the clear definitions for these business users is still a work in progress, but it is critical so that we
[13:37] have that accurate AI result to their questions. Obviously, advancing technologies, as we've seen this week, are enabling so many new capabilities, and the success of those really depends on the data
[13:52] quality, the semantic design, and that user enablement piece. So, here are some example questions that our users have asked. You know, they want to
[14:07] know about revenues in certain locations. Um, what assets are available in certain regions? Uh, what is our operational health? Where are we in this market?
[14:25] The understanding of what data is available so that you don't have somebody looking at a revenue report and asking about asset management has been kind of the challenge. Part of that is monitoring. So, making sure that you look at what your users are doing with your spaces,
[14:43] with your models, is really important to understanding what the business actually needs so that you can build out the capabilities that they need. Um, things that you should also be considering are the benchmarks. So, making sure that you have set business
[15:00] logic queries that you can test against to make sure that your model is not drifting or giving you inaccurate results because even though you have the tagline on these, make sure that you verify these responses before you make
[15:16] business decisions with this. A lot of times they won't always necessarily do that double-check. So, it's also important to make sure that you do those checks to make sure that you're giving quality data.
[15:33] This has also led to some really interesting discoveries and outcomes for our business units. So, as they get to play with the data and ask questions of it, they have found things in the data that they weren't aware of because they didn't have the visibility because of
[15:49] that scrub data. So, an example of this is we had a case where we were demoing the dashboard that was brand new to one of the businesses and they said, "That that that can't be right. We had payment and advanced term customers
[16:07] who had open shipments." And you're like, "Th- There's no way. There's no way that they have open shipments." So, th- this is what the data says. They went back and they looked and yes, the shipments had been there.
[16:25] There was no payment recorded in the system. What we found was that there was a business process issue. The payments had been made, but that information hadn't actually made it into the systems of record. So, people were going around the
[16:40] process to complete and and send these shipments without having the the proper pieces in place. Other things that we have found are with the shipments not properly logging that they have been completed. So, it was skewing our KPI
[16:57] dashboards and making things look like they were really not in a good state when they were. It was again a process issue where people were not completing the steps that the business expected them to complete to have that accurate data.
[17:19] So, now on our journey, we're starting with Spotter. If you're not familiar with Spotter, this is ThoughtSpot's um uh answer, I will say, enhancement of the Genie capabilities. It's been built from the ground up with business users in mind.
[17:34] It's easy for them to understand and build their model because it's a conversation. On the surface, Genie and Spotter look very similar. We find that Genie is better suited to our engineering-type
[17:50] users because it's more SQL-focused and they are going to be able to code and specify, really see under the hood, what is required. Whereas Spotter, it's the conversation. It's really, really powerful in how it allows you to
[18:07] actually train the model. So, one of the really cool things is with the the training is you you have the conversation. You say, "Hey, I want to give you some business context so that you can answer better."
[18:22] And you can define, this region is also named this region. And give it all of those things. And another powerful piece is that you can ask it what it is unclear about in your model and it'll tell you, "Hey, I need a little more
[18:38] context in this area. Can you give me some more definitions?" And then it'll, you know, cycle with you and say, you know, "Is this right?" And you can give it a little bit more information so that it can give you exactly what your users
[18:53] need. And again, they don't need to worry about the SQL at all. So, really, this context is the differentiator. Uh Genie needs you to provide it the instructions or code snippets, whereas Spotter allows you to have that conversation to define your business
[19:09] logic. Okay. So, we've now got data in the hands of users. We've got people looking at data and they're having conversations with their data. But we had far more demand for data
[19:25] products than we would ever have the capacity to supply. I could double my team size, which is already an enormous team, and I'd never meet the requirements of the business in the time frame that they would want. So, last late last year we devised a way for other groups to develop on our tech
[19:41] stack. So, earlier this year we launched what we call the open contribution model. So, we give training to our business users. If they if they're tech-savvy enough, they usually get either uh contractor in or they have somebody who's really good with data.
[19:56] So, we give them training on our standards, our coding practices, our tool set, and we let them at it. And then when they want to when they're ready to deploy, members of my team, Jennifer or others, would then um
[20:12] approve the pull request, basically, to get it into the system. So, they can't put junk into our system. We are controlling what they put into the system, but we didn't have to build it all. And we've found that this has really helped some of our business users. We've kind of gone from a data lakehouse
[20:28] architecture that was only used by the data and AI group and nobody else could use it, and we've now gone to democratizing, really democratizing data across the company. So, we have a little quote here from one of our uh the leader of one of our business units, um and his team of three built a whole
[20:46] set of custom reports just for them, deployed them to production, they can maintain them, and all we had to do was give them some training and approve some of their code. That's all we had to do. So, where do we go next?
[21:02] Well, we've built the foundation, we've got people putting parts of the house on the foundation. We just need to widen the foundation now. So, we're going to be mastering data across many more business domains to provide more and more context so we can
[21:17] connect the dots together to enable AI to answer more and more types of question for the business. So, we're implementing reference 360 and Informatica product so that we can really get some control over our reference data because at the moment, although we've done made great efforts
[21:33] to standardize our reference data, changes to that reference data continue to happen. Um and now we've got the business caring about data because they're talking to their data. We've got them to accept real data rather than data that's been washed through 100 Excel spreadsheets
[21:49] and two two dashboards along the way. They're looking at data that's sourced directly from systems. They trust that data. They actually care about the quality of that data. So, it's really revitalized some of our data governance initiatives which were We all know that data governance is very, very difficult to get the business
[22:05] to accept. So, that's where we're going next. Um that's all we have for you today. Thank you.

Learn more about the Databricks Data and AI platform.

The information provided herein is for general informational purposes only and may not reflect the most current product capabilities or configurations.