Skip to main content

Building Trusted Generative BI: Lessons from Virgin Atlantic

Summary

  • Virgin Atlantic's Head of BI Tom Barber explains that their first generative BI project failed because speed without trust eroded decision-maker confidence, revealing that business meaning must be codified in governed metrics before AI can reliably answer questions.
  • Unity Catalog Metrics Views anchor Databricks Genie's answers to authoritative, stakeholder-approved metric definitions, ensuring that AI-generated results are explainable, consistent, and defensible across the organization.
  • Virgin Atlantic designs separate decision workflows for analysts and business leaders, treats metrics as products with formal owners and managed lifecycles, and embeds governance by design to make trusted generative BI scalable across the organization.

Building Trusted Generative BI: Lessons from Virgin Atlantic

Watch: Building Trusted Generative BI: Lessons from Virgin Atlantic
Generative BI promised faster answers, but Virgin Atlantic discovered that speed without trust erodes confidence. When decision-makers cannot defend AI-generated answers, adoption stalls. This case study reveals how trusted generative BI requires more than governance: it requires business meaning to be consumable by AI, starting with governed metrics signed off by stakeholders.
Learn how Databricks Unity Catalog Metrics Views anchor AI answers to authoritative definitions, how Databricks Genie enables explainability and lineage visibility, and how to design decision workflows tailored to analysts and leaders separately. Tom Barber, Head of BI at Virgin Atlantic, shares practical lessons on treating metrics as products, embedding governance by design, and scaling trust across the organization.
🤝

Chapters

FAQs

Why did Virgin Atlantic's first generative BI project fail?

The initial generative BI attempt failed because it delivered speed without trust — decision-makers could not defend AI-generated answers to stakeholders, causing adoption to stall. Tom Barber explains that the root cause was business meaning not being codified in a form that AI could reliably use, rather than a technology failure.

What are Unity Catalog Metrics Views and how do they support trusted AI?

Unity Catalog Metrics Views are governed metric definitions that have been signed off by stakeholders and stored as authoritative sources of truth for KPI calculations. Databricks Genie uses these metric views to anchor its answers, ensuring that AI-generated results are consistent, explainable, and aligned with the definitions used in official financial and operational reporting.

How does Virgin Atlantic design analytics for different user personas?

Virgin Atlantic designs separate decision workflows for analysts and business leaders, recognizing that each group has different needs for depth, explainability, and speed when consuming insights. This persona-based approach ensures that Genie responses are appropriately scoped so that both groups can trust and act on the answers they receive.

What does it mean to treat metrics as products at Virgin Atlantic?

Treating metrics as products means defining each KPI with a formal owner, a governed definition, and a managed lifecycle — similar to how software products are built, maintained, and retired. Virgin Atlantic applies this approach to ensure one metric has one definition across the organization, eliminating the conflicting numbers that previously emerged between different teams' dashboards and reports.

Full transcript

[00:08] Good afternoon everyone. Welcome to the session. 3:00 p.m. on a Wednesday, so I hope you enjoyed your lunch and are ready to learn about Virgin Atlantic's journey. So, before we get started just a couple of little housekeeping bits. So, as you've probably heard 4,000 times already, uh complete surveys
[00:25] that'll be in the app for after this. And probably the main bit of housekeeping here is an and I'm really sorry for this, but being someone who travels around a lot, I do have to rush very quickly after this session for a 5:30 flight from SFO. So,
[00:41] I would love to take loads of questions, but I'll pop my details up on the last slide. So, if you want to follow up, that is absolutely fine. Just give me a message, but I do just need to escape very quickly. If anyone can give me a lift, happy to chat on the way. Cool. So, um what we're covering today.
[00:58] I'm basically going to talk about what worked and what didn't in our journey from traditional um BI, as I like to call it, into generative um trusted BI at Virgin Atlantic. Now, has anyone in the room flown with Virgin Atlantic? Oh, okay, great.
[01:16] Sometimes you do US conferences and go, "No idea who you are." Um yeah, really quickly, started in 1984, 40 aircraft now, which often surprises people. Massive partners with Delta. Um and yeah, it's a great experience. So, if you haven't been on board, jump
[01:31] on. Um so for those that have flown, you must know about that feeling of that pressure that you get when you descend or take off and your ears just suddenly pop. Now, the airline industry is covered in pressure.
[01:48] It's operational, commercial, especially at the moment with fuel prices. Um and that also applies to the analytics world as as So, I also know pressure well cuz not only do I lead the Beyond Advanced Analytics
[02:04] team for Virgin Atlantic, which can be very pressurized sometimes, but a really great experience. I also like to fly planes, throw myself out of planes, and somehow manage to DJ in clubs just on the side just for just a little bit about me. Living in the UK, clearly by
[02:20] the accent. Um and yeah, so when we talk about pressure, the pressure in the cabin, it wasn't for us about technology cuz pressure just came from the changing expectations of our business, of our
[02:36] leaders, of the analysts in the business. There are really three pressures that I kind of think about when I think about traditional BI. One of those is speed. Clearly, people want answers quicker.
[02:51] A lot of people speak about reporting cycles, and you're producing a deck on a Friday. It's then got to be ready for the Monday. You then edit the deck, go back, and keep going around this horrible cycle. But, the pressures of the business just make everyone want those answers quicker. The complexity, increasingly, and for those in the room,
[03:08] I'm sure in your businesses, you're finding that it's not just about one domain anymore. It's about multiple domains, multiple things cross-cutting commercial, a bit of customer, a bit of you know, engineering, all that kind of stuff. And traditional BI just doesn't stack up to that.
[03:24] And the expectation of our leaders, especially with ChatGPT and other agents like that, you just end up in a position where everyone expects answers that are accurate and quick and right at their fingertips. They don't want to be opening dashboards. They don't want to
[03:41] be asking for a new tweet to a dashboard or a new table to be a feed or anything like that. And that creates you a problem because in the traditional BI world, you really are in that kind of report factory, as I like to call it. Um which we need to get out of.
[03:59] So, the business wanted answers, but BI was still delivering reports. My team was just churning out reports, and this needed to change. So, Wilbur and Orville, the first people to take flight. Wilbur and Orville, right?
[04:17] Um now, Wilbur and Orville, for Virgin Atlantic and those that have flown with us, you may know that they are actually the names of our little salt and pepper shakers that you get with your in-flight meal. So, a little bit trivia, just to make sure everyone's awake. But, um maybe pop your hand up, and
[04:34] where was Wilbur and Orville's first flight? For North Carolina. North Carolina, correct. Kitty Hawk. So, for that fabulous answer, I'll give you the most niche conference swag ever. And there's a little Wilbur and Orville.
[04:50] There was. And it does say pinch from Virgin Atlantic on the bottom, but that's fine. Um good marketing. So, unlike Wilbur and Orville, our first attempt didn't go so well. Um it we, yeah, we introduced these
[05:06] little guys. It didn't go too well. So, I like the Wilbur and Orville story here because everybody remembers that flight. Um but nobody really talks or remembers about the prototype. No one knows how long Wilbur and Orville, well, they they do, but we don't all go, "Oh, they took
[05:22] like 4 days to do this, 5 days to do that." Every organization has that prototype story, and that was our first attempt. So, our first use of AI for BI, it looked amazing. Um we we saw that um
[05:39] we're in this this place where we could um it was seconds, not hours. There was natural language questions going on, huge engagement cuz everyone was excited, and we thought we'd cracked this. But the problem was that business meaning was missing.
[05:55] Our speed um was just just started to really not equal the trust. Our analysts had all the context in their heads um or written down or is in kind of tribal knowledge, and that was kind of doing a lot of hidden work that we didn't really
[06:11] appreciate. And but an analyst work differently to leaders, and that was quite a it's always been a thing, but I think it was a real realization for me especially in the team that says, "No, we really need to think about the different behavior of analysts and leaders or
[06:27] decision makers." And really that summarizes it. It wasn't a failure. It just revealed that the actual problem here was that we weren't scaling trust. So, our first use of AI revealed that it was we we needed to work out how to scale that trust.
[06:48] So, hopefully that sounds a little bit familiar um because the scenario here is that what was happening was that people would be asking um an AI agent of these questions. They'd then be going, "Oh, it looks right." But it sounds right. But then the first
[07:04] thing they'd tend to do is go to a dashboard and then maybe another dashboard and then maybe ask an analyst to go, "Is it actually right?" Which is not where you really want to be. So, we kind of coined that term and I was thinking about this and I was like, "Actually, there's a gap here and it's
[07:20] the gap in trust." So, as AI adoption gets better over time uh of quicker over time, sorry. And you want to make sure your decision confidence is high because that's what you need in a trusted BI organization,
[07:36] you get this kind of effect where actually, as the tr- as the the increases, the confidence starts to decrease because answers are coming quicker and it just starts to the business context disappears, the definitions aren't trusted, and you keep
[07:52] asking why that that answer is generated. So, ultimately, we kind of coin this a little bit the the trust gap. The governance was there, but the trust experience wasn't. So, we did a load of great stuff with Unity
[08:07] Catalog. We were one of the first um customers in in Europe to um to be completely Unity Catalog. Uh our governance systems are really great, but that wasn't enough to solve this challenge. So, ultimately,
[08:22] the answers arrived faster than the confidence. So, we have to reframe that problem. We thought we had a problem with AI and the technology. So, we thought Genie was struggling. We thought other things were struggling. We thought ChatGPT just
[08:38] wasn't going to cut it, which is mad to say now. Um and clearly we were wrong with that. So, it's not about governance, as I say, because we had a really great data foundation with governance. The team had done a great job in getting frameworks
[08:53] in place for our data quality. And it's not about lineage because clearly Databricks gives you some really great tooling to understand the lineage all the way uh from report to source. AI ended up just exposing um a business
[09:09] meaning problem. Like as I said before, it lived in analysts' heads. Meaning wasn't reusable and trust wasn't consumable by AI. So, we didn't rebuild trust. We thought of this as trust just needs
[09:27] to be consumable by AI, which is was the realization here. So, how did we think about doing that? So, we look at trusted generative AI. And this is our Cruise L student looking
[09:43] at our North Star. So, our North Star, you've got to be thinking about governed metrics. And starting with a shared sign-off of the definitions. I know lots of people in the room or lots of organizations will have massive spreadsheets, jargon busters, all these kinds of
[09:59] technologies, but it's all about having those governed metrics. And we found that Unity Catalog uh metrics views were a real hero here because we could treat those as the authoritative assets for those signed-off definitions. Now, the important bit here in the signed-off
[10:15] definitions is that we spent a lot of time thinking, "How do we get these metrics definitions really clear, all worked out, and importantly, signed off and agreed by multiple different stakeholders?" So, for this, we made
[10:31] sure we had a senior-level stakeholder, a business area uh contextual owner, like an analytics lead, or someone who leads an analytics team in commercial. And it was countersigned, for want of a better term, by my team as well. So, we had that joint understanding of the
[10:46] data. And this isn't a massive long task of writing really long documents because clearly tools like ChatGPT and others can generate this information really quickly to get to a first pass. And then you just have to get that signed off and tweaked. Um and then we, you know, store
[11:03] it as a markdown file, pop it in a Git repository, and then you've got your trusted definition. And importantly, at the on each of those metrics definition, there is the one single Unity Catalog metrics view that is the authoritative source for that metric.
[11:18] Uh and we'll come on to that a little bit later. So, AI itself needs to be explainable. So, the answers need to be grounded in that governed data. It's a bit of a common phrase now, but the govern data, make sure your data
[11:34] foundation is correct is impera- like it's really important. Got to have visible lineage and explainability. And then the conversational access can come with AI BI Genie. Or as Databricks like to keep renaming it. Um
[11:49] This slide's already out of date after yesterday. Um so then you've got the decision workflows. So this is quite important because when the team we never really thought really deeply about decision workflows. This is all about how you have those experiences that are
[12:04] guided to each persona. Your analyst and leader workflows do differ. And we'll go into that in a second. And governance is embedded in the design of both of those workflows. So I really like this phrase actually because a lot of the time people go,
[12:20] "Well, metrics definitions." People will just still argue and debate them. But it's not about is which one is right or is this right. It's about No, this is the right metric, the authoritative metric. And therefore let's have a conversation
[12:35] about why this one that you're presenting to me is not correct or may not be repre- representative the right thing. And that could be the detail of um filters, you know, slight changes in context, that kind of stuff. But the point is you've always got that authoritative view to go back to, powered by Unity Catalog metric views.
[12:56] So a little bit more detail on what that one metric's one metric really means. So metrics aren't outputs. Um they are products. And that was another key bit because I think a lot of people will think, "Okay, let's just write a metric into a document." But
[13:12] actually we we treat these in a full product life cycle from drafting into um into iterative uh updates. If you need more data quality rules, if you need more a slight tweak to the calculation, new context for different business areas. And that's was a really important
[13:28] difference. So then you get one metric, one definition, and the right ownership on there. And then that gives you that single entry point for every single decision to be made on net promoter score, on time performance.
[13:44] So then your metrics we use become your authoritative layer. And the important bit here is that generative BI can only answer questions if it can ground to trusted metrics. Sounds obvious, but not always something that we didn't focus on it at start.
[14:08] Cool. So yeah, talking about the shared persona, the different personas, but importantly, the different personas need that shared trust. So we have our analysts. Now analysts are all about exploring, validating, hopefully, and refining
[14:23] those, you know, the decisions and the the insights that they're doing. And we see that as a really powerful feature in Databricks Discover, and that's the experience that we push towards our analysts. Then our decision makers, leaders, they're all about asking, they want to
[14:38] explain easily, and they act on decisions. Again, this slide is now two days out of date. So Databricks one, now Genie one, is the experience that we push our decision makers down. And then the important bit there is we've discussed bits of this, is that
[14:53] there's a shared foundation in the middle where you've got your governed metrics, one metric, one definition. The explainability is there, so every answer is explainable and defensible, most importantly. And then the lineage is really clear if there are challenges.
[15:11] So that's where trust starts to scale, when every persona starts from the same governed foundation, and they have a clear starting point to go to. So one little trick we have here, which I think actually the technology is catching up with this is that we have a start here
[15:27] concept. So our analysts have environment in Databricks, they can go into their workspace and they will have what we call a start here notebook. Sounds simple, but it is literally almost like a a 101 textbook that they can access that will explain
[15:43] by each domain what are the certified metrics, where can I get the answers, how do I join net promoter score to on-time performance with you know, examples and really clear ways of doing it. So they can really easy work that through. And then the
[15:58] leaders they just need to know that that here's your authoritative certified genie spaces or genie agents now. And yeah, I can go in there and know that that NPS or OTP metric is correct. And then you've got a situation there with
[16:14] that shared foundation where the analyst and the decision maker just know exactly that they're looking at the same numbers. And the analyst can dive in, the decision maker can be confident and the whole world is a happier place.
[16:32] So it's all about designing this for explainability. So as I said before, faster answers are not enough. Users must be able to defend those answers. And there's a really good parallel here with how you know, strict you need to be in the world of aviation engineering. And cuz that explainability needs to be
[16:47] visible, the lineage visible and the context visible. So an answer nobody can defend is not decision grade BI. It's just not going to not going to cut the mustard. It's a nice little UK term.
[17:03] So in the in the airline industry, our pilots might you know, convert from one aircraft type to another. And that takes a lot of lessons and looking back and that's what we do different what we do differently. So I think here what we do differently
[17:20] really importantly, start with those decisions. It's not about dashboards. It's not about like drawing out charts on a page. I used to have a stakeholder that would always um literally draw on a piece of paper and give it to me and say, "I want this."
[17:35] Um which is not starting with a decision. Should be designed from those decisions, design the metrics that design uh to that drive those decisions first. Um don't dive straight into jumping on models and semantics and designing tables or even going into Genie code and
[17:52] saying, "Let design me this data model." You need to have that decision and the metric first. And then spend like a lot less time on the interfaces, as I say, and more time on thinking about analytics and the
[18:09] assets you're creating for analytics as products. Cuz I I think that was the key moment where my team essentially changed from a report factory to an analytics products um development house or I can't figure the right term, but yeah, we're we're now
[18:26] firmly everything is an analytics product, and that includes the metrics, includes the governance, includes the data quality checks, includes all of that great stuff. And ultimately, assume that trust will be questioned by analysts, by leaders, by anyone.
[18:41] And that means you just need to think about how you can design for that trusted experience. And the simple answer is make the explainability visible. And if Genie can tell you how it's got to an answer, even better. So, trusted Gen BI, generative BI,
[18:59] starts with trusted analytics products. You just can't do it with a mesh and a you know, a myriad of different um assets and dashboards.
[19:18] Cool. So, on a landing into LAX like this 350, um the takeaways So, trusted Genosys BI is an evolution of BI, of course. It's not a replacement. So, shared metrics have to become have
[19:34] to come before shared AI. As we said, uh that's one definition, one source of truth. The explainability is non-negotiable. If a user cannot defend it, it just won't happen. The leaders will not act on it,
[19:49] and you start to get in a bit of a mess. And then, trust must scale by design. And I think the really important bit there, even though it's tiny on the slide, is by design. Thinking about how you actually build assets to have that
[20:04] trust like integrated in by design is so important, rather than just assuming it is going to be there. And, you know, as I said earlier, a lot of our pieces there are just having, you know, that those firm metrics definitions documented, version controlled,
[20:22] and making sure we've got those start experiences aligned to each persona. That means that your governance can then scale from day one because we've got that pattern. And what's what we're really finding now is because we've packaged all of that up,
[20:38] and we're almost version controlling our what we call our governance framework and what a good analytics product looks like, it means we can scale that a lot quicker. The example being, we can now and basically almost cookie-cutter across
[20:53] the business, and use using tools like Genie code, Open AI Codex, um and a few others, we compare our analysts with some agentic processes working purely on, you know, on Git repositories um and all this this
[21:09] knowledge and information in the background which we're really excited about Genie and ontology as well that could really help this. And basically rinse and repeat and say I've got this new decision flow, I've got this new metric. You know the decision made you know the personas that we're dealing with, you know the kind of
[21:25] business context. Now go off and create me 80% of the trusted assets that will create this new analysis product but brand it as a product and talk about it less as tables and more as a new on-time performance punctuality
[21:40] and product set which would include a Genie, includes a metrics view, includes a couple of tables with drill down and that just means that you're in a much stronger position. So ultimately trusted generative BI isn't
[21:55] created by connecting AI to data simply. Um it's created by connecting it to trusted analytical products. It's been a pleasure to talk to you all. Thank you very much.

Learn more about the Databricks Data and AI platform.

The information provided herein is for general informational purposes only and may not reflect the most current product capabilities or configurations.