Skip to main content

From Cost Mystery to Cost Mastery: COPA Airlines' Lakehouse Efficiency

Summary

  • COPA Airlines migrated from Teradata to Databricks, multiplying compute capacity by five while maintaining cost levels and reducing a critical analytics query from 12 hours to 1 hour.
  • The company implemented SQL warehouses with governance-enforced consumption caps, real-time Delta Live Tables for boarding feeds, and a federated cost ownership model that distributes accountability between technology and business units.
  • By implementing a medallion architecture with Unity Catalog integration, COPA met GDPR and personal identifier compliance requirements while achieving 17 times faster competitive analysis and real-time operational insights across 360 daily flights.

From Cost Mystery to Cost Mastery: COPA Airlines' Lakehouse Efficiency

Watch: From Cost Mystery to Cost Mastery: COPA Airlines' Lakehouse Efficiency
Airlines operate under razor-thin profit margins with complex, interconnected global operations. COPA Airlines, achieving the world's third-strongest operating margin and Latin America's best on-time performance, faced a critical challenge: how to scale analytics and real-time capabilities without exponential data costs. Legacy Teradata infrastructure bottlenecked performance, limited cost visibility, and restricted agility. The company needed to migrate to a modern platform supporting native big data processing, real-time streaming, and predictable cost management.
Discover how COPA Airlines migrated from Teradata to Databricks, multiplying compute capacity by five while maintaining costs. By implementing SQL warehouses with governance-enforced consumption caps, real-time Delta Live Tables for boarding feeds, and federated cost ownership between technology and business units, COPA reduced critical analytics queries from 12 hours to 1 hour. The medallion architecture with Unity Catalog integration enabled compliance with GDPR and personal identifier regulations. Results: predictable cost management, 17 times faster competitive analysis, real-time operational insights, and velocity to deliver new products on live data.
🤝

Chapters

FAQs

Why did COPA Airlines migrate from Teradata to Databricks?

COPA Airlines faced significant challenges with Teradata including costly capacity upgrades, inability to separate compute from storage, and resource contention among analytics teams competing for shared CPUs and memory. The legacy platform also suffered availability incidents when poorly optimized queries compromised the shared infrastructure.

How did COPA Airlines control costs after migrating to Databricks?

COPA implemented SQL warehouses with governance-enforced consumption caps that prevent teams from exceeding budgeted compute spend without approval. They also adopted a federated cost ownership model that distributes accountability between the technology team and business units, giving each stakeholder visibility and responsibility for their own consumption.

What performance improvements did COPA Airlines achieve after migrating to Databricks?

COPA Airlines achieved five times the compute capacity at equivalent cost compared to Teradata, reduced a critical analytics query from 12 hours to 1 hour, and delivered 17 times faster competitive analysis. Real-time Delta Live Tables also enabled live boarding and operational feeds that were not possible on the legacy platform.

How does COPA Airlines handle regulatory compliance in their Databricks data platform?

COPA Airlines implemented a medallion architecture integrated with Unity Catalog to meet compliance requirements for GDPR and personal identifier regulations. Unity Catalog provides lineage, access control, and auditability across all data assets, enabling the airline to demonstrate compliance while running analytics on sensitive operational data.

Full transcript

[00:08] Hi everyone. Thank you for coming. My name is David Garcia. I'm director of software and data engineering for Copa Airlines. I have been in Copa for the last 14 years. So I have like 30 years of experience in technology. So
[00:23] I was working when some of you was burning. I mean in this case, in this opportunity, I will be sharing the journey of Copa Airlines. How we migrate from data data to data bricks
[00:39] in efficient way. How we discover the mystery of the cost and how we reach the the mastery of the cost management. Copa Airlines is just to give you a background. It's
[00:54] airline that is based on Panama City in Central America. Our fleet is a lot of Boeing 737. We We have the network of 85 cities. Right now, as a matter of fact, it's 87 because we have a new two additional
[01:12] destination. We fly to 32 countries. And for 11 year consecutive year, we have been the you know the the the more punctual airline in Latin America and to and to time in in the world.
[01:28] So do you imagine So the the model is a hub. So that's mean that we bring passenger from south to Panama and then from Panama to the north and vice versa. So the complexity in Panama is really really high. And the data that is coming
[01:46] because we have 360 flight going and coming. So it's a really complex operation. The challenge was imagine that we are we have 360 flights and Teradata
[02:03] was, you know, in high demand and we have multiple incidents because the Teradata was down because, you know, for example, a rookie analyst was perform a really hard queries which not optimized and and compromise this the the the availability
[02:19] of the of the services. So, that was one of our our challenge. Other was, how we can split the compute between analytics team. In Teradata, everything was together. So, we have CPUs, memory, and storage
[02:34] together and we could not uh split the storage from the from the CPU and memory. So, if we have to upgrade Teradata, we have to spend like a $10,000 $10,000 per month additional
[02:49] to move to the bracket that we have to the next bracket. Just forget we need additional gigabytes to store more data. The other thing was um uh the analytics team is starting to compete to use us the CPU and and
[03:05] memory. So, that's was a you know, a really really problem for for us. The scope of our project was migrate all of the data that was in Teradata to Databricks in in this case to Databricks and uh re-engineering all of the business logic.
[03:22] The when we implemented Teradata 10 years ago that was, you know, the king of the warehouse at the time uh all of the business logic was inside Teradata. So, that's was a really good business for Teradata, not for Copa
[03:37] because when we decide to move out of Teradata, we have to invest more money to to extract all of the logic, okay? Um the other thing is migrate all of the Tableau dashboard and also the the Alteryx workflow that we already
[03:52] have it to repointing to Databricks. And implement a tool that help us to to implement a data governance because in in our current in in that time, 2 year and a half ago, Teradata doesn't have any tool to to to make the
[04:09] the governance. So, we define three phases. The first the the phase one was exploring. Our strategy was we want to be uh cloud agnostic vision. So, that's
[04:25] mean if we have to if if our vision was cloud agnostic, we cannot use, for example, Redshift. We cannot use Synapse. We cannot use Oracle and the other one. That's mean that the the list of tool that we have to choose at that time was or Snowflake or
[04:42] Databricks. That were the option, okay? So, we we made that um um that selection and then we start to go uh data we did a a deep dive. How we want to make the selection between Snowflake and Databricks?
[04:59] And the reason and the reason was this. We want to implement a big data platform. A native big data platform. At that time, 2 years and a half ago, Snowflake was a really good data warehouse. Traditional data warehouse that are
[05:16] evolving to become in a big data. That's mean is not was not native big data platform at that time. That was the reason why we did we decide to go to the second phase that was proof of concept
[05:31] and implement a that proof of concept with Teradata. So, what we did? We select two tables of 65 million records to make a join in Teradata and then in the same compute
[05:47] condition in Databricks. You know what? Teradata, the response was 5 minutes. Every time that we made that, 5 minutes. But in Teradata, the first run was 38 seconds.
[06:04] Amazing. And the second run was 11 seconds. So, when the when we see that, we say, "You know what? This is a no-brainer." We have nothing to think about it because maybe when the when you are talking, for example, with Snowflake,
[06:19] Snowflake say, "Hey, a database cannot manage uh data warehousing because it's a it's it's for just for manage a huge amount of data." So, that's what's the the the second phase. The other thing was we tried to to to teach uh machine
[06:36] learning model in Teradata and the process never ending. When we make the same process to run in Databricks, uh the learning process was completing in 8 hours. So, that's what's a no-brainer
[06:52] for us. The third phase was we have to select who will be the vendors that will be helping us in um in extracting business logic from Teradata to in in this case, we select
[07:08] PySpark for as a technology stack. So, the conversion of the code to the the logic. And uh we were working with the XE, AWS, Infosys, BCT, uh Impetus. And uh that time, um
[07:25] all of those um was uh doing a very deep dive analysis in the in the code. And uh the the vendor that make, you know, very deep dive analysis because it implement a agent in Teradata, was
[07:40] analyzing all of the components, was Impetus. We select Impetus. And the second one we we start to to to to request some resource from Visity that was helping us in the in the last phase of the project.
[07:56] So, how we want to the the the data governance? So, we have in the in the in the left side we have the ingestion layer that basically we have um the a batch
[08:12] process where we can ingest data in the in the medallion architecture of Databricks. Or and also we have Data IQ that is another tool that allow you to design the process inside Data IQ Data IQ but you can run inside the
[08:27] Databricks. So, that's was an advantage of the flow that we already have it. And then the streaming injection using PySpark for the streaming. So, in top of it we we use the the Unity
[08:43] Catalog. We integrate the Unity Catalog with the Active Directory so we can uh implement the data governance and decide which people have access to the to a specific data. We are the airline industry is a industry that is really regulated so we
[09:01] have to we have to be compliance with PII, GDPR, and all of the data protection frameworks. So, with the Unity Catalog allow us to implement a really good governance on top of it because everybody that want to
[09:16] access a specific data go through the Unity Catalogs. So, we can decide who have access to which data and in that in that moment. So, the final architecture that we define was uh
[09:31] the ingestion, Data IQ, PySpark, and then I'm I'm working for the technology team. But in Copa we have federated the organization. So the analytics team is in the line of business. The technology is in our hand. So we
[09:48] manage the single source of truth. That mean that the data that should should be sharing for everybody. And the business has their own S3. We are using in Amazon. Of course. Uh bucket. Customer spending has
[10:03] S3 operation. has the S3. The revenue management have S3. And they are responsible for those specific uh uh S um S3 buckets. In the consumption layer, that is we
[10:18] will start that to discover the mystery. We we design to implement uh SQL uh warehouse of Databricks in order to put a cap of the consumption of the DBUs. Okay? Because if you bring the
[10:36] serverless, for example, to a a rookie analyst, can consume all of the DBUs that you already contracted because it doesn't have that ex- enough expertise to use in in a right way the resources that we have. So.
[10:56] What were the benefits? Imagine that is something that I that is a secret only for you, okay? We have a process that we gather all of the all of the first of our competitors. Okay? We gather all of all of that data and we compare with our first in order to make adjustment to our first. Okay?
[11:11] Please don't share with anybody. Okay? So that's process uh spend two hours in Teradata. Two Sorry. 12 hours. Running to process all of the the first
[11:27] of our competitors and I'm going to the Teradata. In Teradata bricks, that process is taking only 1 hour. Okay? They're also The The other thing was
[11:42] real-time analytics. Teradata doesn't have that functionality. So, Teradata is a data warehouse that manage a batch processing that you have to to feed maybe in 5 minutes, 10 minutes, but you cannot manage that in milliseconds, for example.
[11:57] And also, a predictable growth of of the couple of the data of the data growth. For example, we assign uh SQL warehouse for revenue management team. We assign a SQL warehouse to the customer. We have We assign a um SQL
[12:13] warehouse for operation. So, each team has a different compute. Our use case Our business case was not reduce our cost. Was increase the, you know, the the compute capacity. As a matter of fact,
[12:28] each uh SQL warehouse that we assign to each team has the same capacity that the entire Teradata that we had in the past. So, we are multiplied by five the capacity of of compute. What's uh
[12:44] Talking about the the first example, when we gather all of the the first of our competitors, that's mean in 12 hours in Teradata, we have to to wait until 12 hour to design to make adjustment to our fare.
[13:00] Right now, we can do it one 1 hour later. So, the the the impact in the business was huge. On another hand, the on-time performance that is one of the key thing of Copa or key goals,
[13:15] uh we implement a real-time uh delta live table with a streaming um injection of boarding. So, if If are flying one of you are in flight flying in Copa, when you are scanning your passport
[13:31] in the airport, we are receiving that feed in the real time in a second. So, with this, we can know each we have a problem with one specific airport in one specific country, and we can, you know, uh trigger any action in order to avoid
[13:47] that the flight can departure later, you know? This is all the other thing that was uh uh a really really good resource. The other thing was of course, the decision speed. If we have the data in the right time,
[14:05] the business can make decision right away. So, that's was one other other other the business impact for our implementation. So, myths. The five myths.
[14:20] One, Lakehouse cost is unpredictable. If you implement the right governance, that's mean if you implement the SQL warehouse, you everything has a cap of compute that can that can has can
[14:36] have access. You don't have the possibility that anybody can make a for example, a query that are affecting uh the entire ecosystem. You just will be affecting the people that is using this specific SQL warehouse.
[14:55] That's the reason why we say the governance is really important. And also, how you manage the cost. I'm going to I'm going to talk later regarding that. The myth number two, serverless is always more expensive. Not really. It's a powerful tool. For example, we use only for those
[15:13] specific process that has been going through a uh development life cycle. I mean that you are testing it. You want know how you know it's a you know how is the behavior of that job. You know that and you know that the
[15:28] source of the of the data is not changing many time. Doesn't mean that that you can receive a corrupt file and you're processing something that never ending. This is the thing that you have to avoid to move to serverless. So if you know specifically what is the the behavior of
[15:45] that specific job, that is a good candidate to manage in serverless. Okay? Also people that have enough expertise developing queries with Optimus way. That is that is other other other option
[16:02] for assign to one person or the person that have that kind of expertise. You can assign serverless access but not for everybody. Okay? The number three is the you need serverless from the day one? No. You can access just using the SQL
[16:17] warehouse. You don't you assign serverless just for the thing that you are ready the thing that I already mentioned regarding what is a good usage of serverless. Okay? And the last one. The you can cap the consumption. You can you
[16:34] can do it if you create a SQL warehouse. The SQL warehouse for example in our case you use the size you know SQL warehouse managed by size like a teacher size that you have XS and and whatever.
[16:50] And you can assign the pen of the of of the specific person the consumption. You assign exactly the the size that they specific teams need. Okay?
[17:06] The cost. How we do it? So if you know you have a users that are accessing through a SQL warehouse and going directly to to the data that is in in parquet format or Iceberg, whatever option do you use. But you need,
[17:22] for example, multiple and the the group is very big, you can make auto scaling to that SQL warehouse and say, "You know what? You can expand all up to three machine." With this, you have a cap
[17:37] of usage of the DBUs in Databricks. So, that mean that nobody can make a not optimized queries and are consuming all of the DBU in Databricks. You have a cap. The cap is is is managed by the SQL warehouse. So, that is a good good way
[17:54] to to manage that kind of um of challenge, for example. Copa Airlines, for example, is a is a airline that is really cost focused. Why? Because the airline industry is really complex. You are managing in an airport that you
[18:10] are not the owner of that airport. You are mana- you are flying for over a over a country, and the company the country can have a volcano, and the volcano start, you know, and you you you you your your network is affected for that. So, many things can happen from 1
[18:26] minute to another. The only good thing that you could you could do is have the control of the your cost, okay? IT governance. If you see in the light blue, in the JSON,
[18:42] we we assign that cost specifically for the technologies cost center. Okay? The single source of truth we manage for the technology cost center. But the the S3 bucket that is assigned to
[18:58] the to the business user, we charge that cost directly to the business. Why? Because in the in the S3 bucket, we don't have it limit. So, the people can put every data that they wanted.
[19:14] But, the only way that one body, one person can, you know, be convinced to be to be, you know, conscientious regarding the cost is that they have to pay for that. That's the reason why we assign the S3
[19:30] bucket to pay directly by the business. In the consumption ledger, we have the data the SQL warehouse assigned for each people. And we define a service level agreement. Okay, I'm bringing you a capacity,
[19:47] a limited capacity, because you have to escalate you can escalate up to three machine, nothing else. If you need more, you have to sit down with technology again and discuss what is the the extra consumption that you need. And we
[20:02] rewrite the service level agreement again. Okay? And I we avoid surprise, so you know that we have the cost in a spiking the cost in in in only one day.
[20:18] For the for the the the the companies that have citizen ID, uh this is the way that we manage. So, we have people that are in the business that develop their own their own components. The injection was directly charged to
[20:34] the their consumer. Of of course, the the S3 bucket as charging to the line of business. And the the consumption in the in the SQL warehouse are managed by a service level agreement as well. So, so that's the way that we manage and
[20:50] we put cap in the consumption of the of of of our internal customer. That's what the way when the we started 2 year and a half ago analyzing the the mystery of the cost, and that's the way that we reached it
[21:06] the mastery of the cost. Because right now our uh our business case, as I mentioned at the beginning of my presentation, my our business case was not reduce the cost, was keep the same cost but multiply the capacity.
[21:23] With this, right now, we have the same the same cost, but we have five times the capacity. So, that's the way that we have a really happy customer right now because they they queries run the knowing minutes in hours versus
[21:40] 12 hour, 24 hours. Right now, they have in minutes. So, that's the way that Copa discovered how a start from the cost mystery to the cost mastery. Okay? Thank you for coming.

Learn more about the Databricks Data and AI platform.

The information provided herein is for general informational purposes only and may not reflect the most current product capabilities or configurations.