Skip to main content

From Snowflake to Databricks: Consolidating Enterprise Data Platforms at Česká Spořitelna

Summary

  • Česká Spořitelna, the largest bank in the Czech Republic and part of Erste Group, reversed a prior decision to standardize on Snowflake and selected Databricks as their strategic data platform after a rigorous TCO and performance evaluation.
  • The bank validated its decision using TPCDS benchmarks at scale, testing open data format performance including Iceberg and Delta with Liquid Clustering to confirm that consolidation to Databricks required no performance or cost compromises.
  • Databricks' openness — support for open data formats like Iceberg, exclusive governance groups for regulatory compliance, and a unified AI and data platform — were key factors in consolidating from a fragmented Oracle DWH, Cloudera, and Snowflake stack.

From Snowflake to Databricks: Consolidating Enterprise Data Platforms at Česká Spořitelna

Watch: From Snowflake to Databricks: Consolidating Enterprise Data Platforms at Česká Spořitelna
Financial institutions face fragmented data infrastructure across on-premise systems, multiple cloud platforms, and legacy databases. This creates data governance challenges, duplicated costs, and barriers to implementing unified AI strategies. Česká Spořitelna, the largest bank in the Czech Republic and part of the Erste Group, recently evaluated this problem and made a strategic pivot.
Learn how Česká Spořitelna is consolidating Oracle DWH, Cloudera, and Snowflake workloads onto Databricks as their unified data platform. Diego Fanesi and Karol Zemko present the technical evaluation behind this decision, including rigorous TCO and performance analysis using TPCDS benchmarks at scale. Discover how open data formats like Iceberg, Liquid Clustering, and exclusive governance groups enable enterprise consolidation without performance or cost compromises.
🤝

Chapters

FAQs

Why did Česká Spořitelna switch from Snowflake to Databricks?

Česká Spořitelna initially pursued a consolidation to Snowflake but ultimately selected Databricks as their strategic platform after evaluating openness, performance, and total cost of ownership. Key factors included Databricks' support for open data formats like Iceberg, unified AI and data capabilities, and exclusive governance groups that meet regulatory requirements for a financial institution.

What are TPCDS benchmarks and how did Česká Spořitelna use them?

TPCDS is an industry-standard benchmark suite for evaluating data warehouse query performance at scale across complex analytical workloads. Česká Spořitelna used TPCDS benchmarks as part of an open-source benchmarking methodology to objectively compare Databricks against alternatives, including testing open data format performance across Delta and Iceberg with Liquid Clustering.

What is Liquid Clustering in Databricks and why does it matter for performance?

Liquid Clustering is a Databricks optimization technique for Delta tables that incrementally re-clusters data based on access patterns, improving query performance without requiring full table rewrites. This video covers it as part of the performance and TCO evaluation that helped support Česká Spořitelna's decision to consolidate onto Databricks.

What challenges does a fragmented data stack create for a bank?

A fragmented stack across Oracle DWH, Cloudera, and Snowflake creates duplicated costs, inconsistent governance, and barriers to implementing a unified AI strategy. Česká Spořitelna identified consolidating to a single platform as the prerequisite for reaching AI readiness, allowing them to avoid the migration complexity and governance gaps that come with maintaining multiple independent data platforms.

Full transcript

[00:08] Hello everyone. I'm Diego Fenzi, lead specialist solutions architect of uh for data bricks and uh today I have here with me Carl Zamco from Chesca Airst group. Uh today in this session we are going to talk about um AI readiness and
[00:24] how consolidation to a single data platform is key for reaching AI readiness. Um before we start I wanted to do a uh a
[00:42] little intro. Um we see here probably you have seen those pictures um numerous times. Um we have here Gartner uh quadrants and also the forester wave in three different
[00:57] areas of the the big data kind of landscape. um data science and machine learning, the um database cloud uh solutions and
[01:13] the data lakeouses. This is kind of this represents kind of the perception of the market of the datab bricks platform. It can clearly show that our customers uh believe that datab bricks is
[01:28] providing u uh market leader performances on all those areas of uh uh use cases. And this is key because when you're looking to unifying to a single platform, you really don't want to
[01:44] migrate to a platform where there will be compromises or a lot of compromises on the TCO aspect and in on other aspects. So uh Chesca is part of the group. Airst group is one of our strategic customers
[02:01] in the financial sector and Chesca is their biggest division. So two years ago, Chesca was uh unifying um around uh unifying on Snowflake. So migrating all their data platforms on
[02:16] Snowflake. Well, today they selected data bricks are the uh as their main uh strategic data platform and they are now migrating all their use cases, all their data platforms to data bricks.
[02:34] So uh Carol's team actually deployed the first workspace into group three years ago and uh now today Carl is leading the data platforms engineering team at Chesca and he's responsible of all the
[02:50] cloud data platforms. Um so we are going to hear from him why they made this choice, what was what were their evaluations and uh how they ended up into making this decision. Please welcome Carol Zamco.
[03:13] Thank you very much. So let me give you a bit of a context about Chescasponia so you know what we're dealing with. uh CHK Paronia the biggest bank in the Czech Republic. We serve individuals uhmemes, municipalities, other state entities and
[03:29] also large corporation. uh and we also provide additional financial market services and as was previously mentioned we are part of uh ERS group which is one of the leading banking groups uh in central and eastern
[03:48] Europe and this is really to showcase that we as a banking entity have quite a lot of data and at the same time we have a varied let's say portfolio of uh workloads we do this
[04:04] includes collaborative projects across uh across the group. So about our journey, seven years ago, we adopted datab bricks as a tool mainly to
[04:21] really help us with classifying uh our our customers and this was in the form of models, propensity to buy models. uh so we are talking about classical machine learning predating AI and at the time it
[04:39] was very useful for this but it didn't saw broader adoption in the bank then four years ago business has showed keen interest in datab bricks and they started using it not just for classical
[04:55] models but also preparing data for marketing journeys to Salesforce and Here again datab bricks proves itself as a platform not just for uh machine learning and data science but also to do data engineering.
[05:13] Then three years ago as was previously mentioned after some showcases to to the rest of the group a PC was organized where we deployed the first uh datab bricks workspace through our deployment standardized pipelines in the tenant and
[05:31] this was really for to try it out a bit themselves. Then one year ago, we productionalized our first AI use case. Uh the use case was for preparing personalized uh tips about
[05:47] clients for banking advisors and from then many followed and now we are getting to the more interesting part. 10 months ago adopts data bricks as the groupwide platform solution. Afterwards, six months ago,
[06:02] Chescasperonia adopts its uh adopts it as its main data platform. This was from let's say the previous snowflake. From this, you can see a clear trend of acceleration.
[06:21] Now, where are we today? So, we are standardizing our entire stacks. This includes not only let's say migration but also uh integrating our existing in-house tooling and processes into the platform. We are also migrating basically
[06:39] everything we have cloud, snowflake, sas and other platforms and either in part or fully and we are doing more and more collaborative projects across the the group and all this is to achieve this
[06:55] idea of a unified platform and a unified data strategy. And basically all this uh context was to really ask the question Why did we decide to do this pivot and
[07:10] what were the different reasons for the accelerated adoption? So maybe let's start by looking at some of the challenges of having such a large technical stack. First of all, the data is fragmented. you have operating data
[07:29] quality uh governance model and you cannot really maintain an standard across all your stack. This is especially relevant since we have a strategy of AI first adoption and the creation of a unify AI semantic
[07:46] layer is either very difficult or basically impossible. Next, having platforms with equivalent capabilities drives it cost. You obviously pay extra on maintenance since you need to have
[08:02] multiple qualified personnel to maintain the platforms, but even more so often times they do it independently of one another. Uh so this really drives basically just maintaining the existence of having them around and you have the
[08:19] licenses on top. So if you deal with licenses obviously you have the cost but there is also another aspect of added complexity where don't get me wrong a license model can work but optimal
[08:34] results are only if you precisely estimate your needs for a specific time period. So this ideally doesn't fluctuate in the next year and if you need to buy hardware because you are running on
[08:50] premise solutions this is just a cherry on top. So this can really throttle growth then another issue that we have is complication with groupwide collaboration between multiple entities because we didn't had a unified stack.
[09:06] Obviously, this doesn't need necessarily to be your problem, but on a let's say just company level, this also limits sharing of information and knowledge because everybody is working with a different stack.
[09:23] And the last hurdle that we needed to overcome basically with having such a wide portfolio is imposing some sort of standardized AI governance. Now this is very difficult in the best cases to have across the platform but you have
[09:41] the added complexity of having both cloud platforms and onremise platforms and for European companies this is especially I would say tricky since besides the local regulations and compliances that you need to follow
[09:58] there is also the overarching one of the EU European EU act on top of it. So all this what I said was some let's say
[10:15] challenges from having a large stack. Okay. But let's look at what basically push us into this migration into this pivot and do it so quickly as we are doing it round. So we did quite a few uh migration viability
[10:33] assessments for different platforms and across all of them we basically find some similar uh information. The first one it will probably not be a surprise but they can be quite costly and time uh
[10:49] intensive especially if you are doing multiple of them at the same time. So this is the obvious but beside it we find that lift and shift methodology of simply moving code bases from one platform to another
[11:04] doesn't really work as well as you would expect and it's more beneficial to put in the effort and to refactor the code.
[11:21] So why is this? Uh well first of all you are basically if you are migrating in the lift and shift method you are paying resources to translate your technical debt into a new platform. This is extra poignant if you are let's say have a 20-year-old Oracle uh DBH on onrem.
[11:40] Other aspect of this is well if you are just lifting the code and putting it into a new platform you wrote the code in the original uh place and it's not surprised that it's probably running best there. So you are
[11:57] effectively not utilizing the the new paradigms and the new shifts and tools that you have uh at your disposal if you are just lift and shifting it. So all of this basically defeats the purpose somewhat of migrating.
[12:14] Also there is no better time to basically uh future proofing than when you are migrating. And what we found is that really the previous tools such as lakebridge without the AI uh
[12:31] integration wasn't working very well for us. You would really even with simple datab bricks assistant we had much better results. So now with tools such as legbridge being integrated with genie code this really opens the possibility
[12:47] of new new adoption. Then there is the aspect of openness. You have iceberg that basically makes the platform more open. But the most important thing is
[13:02] you don't have to migrate really the the data from one data type to other. It's basically just a flag. There is native support the performance also doesn't suffer. Other aspect is that the decoupled storage and data. So for
[13:18] example with snowflake or with fabric you have this aspect that you need the compute to access the data. Here it is in your native basically storage or S3 bucket. Then there is the aspect of integration and connectors where we
[13:34] have now with the expansion more connectors to use and at the same time we also can integrate our in-houseuilt tools into data bricks and now with as was said in the keynote open sharing is
[13:51] another thing that will be expanding this integration capability. Another thing why we are basically bullish on datab bricks right now is that we as a banking entity need to also
[14:06] handle regulatory workloads uh be it to the Czech national bank or big to the rest of our group as a parent entity and not to oversimplify but with these types of workloads
[14:27] costs are not really your primary interest. There are things such as stability, reliability, auditing and security which take precedence in this. Based on our initial analysis, newly implemented features such as disaster
[14:42] recovery, failover across regions, exclusive groups, all this really opens the door to now be able to do these kinds of workloads in data bricks. So in detail disaster recovery with failover to other regions means you have a better
[14:58] stability the the workloads will not fail you are able to basically continue uh exclusive access for groups that's another very interesting feature that we are quite extensively testing right now if you are not uh aware of this it
[15:14] basically changes the governance model from a cumulative one when all your accesses are basically uh added together to an exclusive where you switch to this group and you only retain the privileges that you have and we are using this to
[15:29] basically create a form of enforcement of data contracts on the platform level. Other thing which are quite specific for us but we need for example uh hardware security model encryption which we are
[15:45] able to do on datab bricks not always possible on other platforms and last but not least some sort of verbos auditing where we can really track every action that is being done and with all these tools we are basically investing Q3 and
[16:01] Q4 to really tackle this not just from the platform level but also from the operational level of our bank to be able to basically migrate the entirety of Oracle and those regulatory workloads into data bricks.
[16:21] So final thoughts for my part, these were all the reasons that we currently are investing in this strategy and pivoting from Snowflake. But maybe something that you noticed is I didn't really talk about performance much or basically not at all. And the reason for
[16:38] this is because Diego Fantasy did an excellent analysis on our behalf and I would like to basically let him tell it about himself. Thank you very much.
[17:03] Um yeah uh so essentially we have some few open questions still right can data bricks perform on SQL use cases right can is the uh SQL performance on datab bricks at the same level as other platforms and um how data
[17:21] bricks and snowflake perform on open data formats and how TCO is impacted when we bring in something like iceberg or other data open data formats um this is very important because sometimes companies and customers will
[17:38] consider architectures they will not really test it they will not really troutly um evaluate it from all the aspects especially from the query performance and the TCO aspect then they start
[17:53] migrating into it and then they find that it's not feasible for them. So um when we talk about performance and TCO evaluation, this is not an easy topic.
[18:09] Many customer will go through um uh PC's but getting also the PC right is not trivial. The focus u here are some best practices. um that I've learned across
[18:26] customers and across PC's. So you should rather focus on the biggest consumption drivers. So biggest data sets, biggest queries, longunning jobs because you don't want to spend your time into those
[18:42] queries that are going to save a few cents of a dollar and then actually have something that is lower on the biggest query that are going to compose 60 70 80% of your over overall consumption.
[18:58] Um, another aspect is that you are not executing this exercise as an a data engineering team. Well, not data engineering team, but as as an engineering team of a product, but
[19:14] rather you should do this exercise in a way that should really reflects your point of view, which is cost. You want to see the the the end cost and evaluate the end cost of
[19:30] your platform of your architecture. So for that reason you should com do a comparison that is on idle cost parity. So that means you don't look at instances that are being provided or
[19:45] other aspect or cores or stuff like that. When we are looking at serverless offerings, those are technical details that doesn't really make a difference for you, what it really matters is cost. Cost at
[20:02] idle time. Okay, cost at idle time because the the the warehouse is not going to be fully utilized at every second. you're going to have uh large utilizations and that's where query cost is important. And the second
[20:20] aspect is when you don't have a full utilization of the warehouse, that's where uh idle cost also becomes important. If more use cases are part of the same PC's, always sum them up all together
[20:37] and compare the total overall TCO, not use case by use case. Why? because you want to see really what is the platform that is overall able to give you the highest level of savings.
[20:57] When we look at benchmarks like TPCDS, you probably are familiar with this. Um we publish that on data sets of 10 terabytes, 20 terabytes, 30 terabytes, you can save um 20%, 26% and 41%
[21:14] on costs between snowflake and data bricks. However, this test is only the sequential query test. Not really doesn't really bring a concurrency into the picture. And the second objection that we get from customers when we show
[21:30] this is well 10 terabytes data set my data sets are hardly that large. Well let's see if this is true. So um on the right hand side you can see
[21:46] uh the structure of the of the test TPCDS and on the left hand side you can see the um um the sides of the tables. So first of all we are talking about 10 terabytes split over 24 tables.
[22:04] Second 10 terabytes is the raw data. So it's uncompressed CSV format. after ingestion is much smaller. We are talking about three terabytes. This is not because we filter data but because
[22:21] delta tables but all the uh also other formats also iceberg is particularly efficient in compressing data because there is no filter filter data. So the results we have seen in the
[22:36] previous slides are really executed on a data set that looks like this. And look there is only one table that is above one terabyte size and we have three facts tableable and then we have a long tail of very small tables. Why? Because this is clearly
[22:53] um star schema. So we have dimension tables and fact tables. I'm sure that if you now go and look at your workspaces, you will be able to find more than one data set that is
[23:08] either bigger or at least of the same size of this. However, there are some challenges especially when we look at TPCDS implementations across all vendors and
[23:26] those are most implementations of uh TPCDS only focus on sequential query execution. So it doesn't really bring concurrency into the picture. And you know if if we look at other implementation most implementations only
[23:42] use the data generator and the queries but disregard most of the specs uh of TPCDS because TPCDS is a third-party industry standard benchmark that is very strict and very rigid.
[23:57] Okay. Some vendors will go to the extent of uh using pre-ingested data set. Did they went through and uh and fully optimize those to the at a level that is probably not feasible
[24:13] for production use cases? Well, nobody knows, right? And in some cases they will also use proprietary tools which you know the transparency of all of this is not really uh it's not easy to gain transparency over what they have done and how they
[24:30] obtain those numbers and ETL and concurrency often remain uncovered. So this uh kind of uh was the motivating factor for me to start my work and redevelop the TPCDS starting from the
[24:48] official specifications. I did it with JMeter. Why? Because JMeter is open-source tool uh that uh many um that you know generally in the market um is used to benchmark.
[25:04] I metered ingestion and concurrency tests not only power tests. Today we are going to see just ingestion and concurrency for time reasons and uh I tested it on multiple data sets. So 100 gigabytes, 1 tabyte, 10 terabytes and I tested it on iceberg
[25:20] delta and native tables on snowflake. I compared at idle cost parity. Of course, there is not um there is never a perfect um correspondence between um
[25:38] uh data bricks and snowflake but they go very close when we are on on the same sides and then I fix the number of clusters to one on both platforms. Why? because I truly want to see how concurrency is going to uh push to the limit that
[25:55] single cluster to the point where we are forced to scale and add more um infrastructure. Uh the tiers I consider is uh datab bricks enterprise and snowflake enterprise. So those are the results for the
[26:12] injection costs and um I was able to measure 66% lower cost on ETL for the one terabyte scale 78% cost on the 10 terabyte scale
[26:30] and uh we still have some savings on the smallest data sets which is about 10% and of course those are just 30 cents okay because the the full test uh was about like a few dollars. So of course
[26:46] the savings in this case is very little and that's the reason why you should focus on the what is going to drive most of the cost in your platform. So the biggest data set and the biggest sizes.
[27:04] This is the results of concurrency test. So we are talking about 20 parallel query streams. Uh the order of the query is also decided by the specifications. I follow the same order. And uh this is
[27:19] datab bricks delta using liquid clustering versus snowflake native tables using their uh clustering functionality. And um oops uh 25% cost savings on BI and analytics
[27:35] for one terabyte, 29% uh savings on BI and analytics for the 10 terabyte scale and 2.4% cost savings on the smallest one, the 100 gigabytes. Again, we are
[27:50] talking about a few dollars. So uh it's just 10 cents, right? which is pretty much the same thing we have seen in ingestion.
[28:06] Now what happens when we bring iceberg into this picture? Well, one thing to notice is that we already achieved open data format on data bricks while we did not achieve this on snowflake, right? But if we truly want to standardize on on iceberg, how does TCO and query
[28:23] performance change on both platforms? So in this overview, this is just Snowflake. So is Snowflake versus Snowflake. Snowflake managed iceberg
[28:39] um gen one and gen two and snowflake native tables gen one and gen two. And you notice how the already you have an increase in price, right? The TCO is 15 to 26% higher,
[28:57] 5 to 19% higher for the 1 TBTE and 23 to 27% higher for 100 GB. Now there is one aspect. So the two percentages, one is gen one to gen one.
[29:12] The second percentage, the second percentage is gen two to gen two comparison. There is no price difference on snowflake whether you are quering an iceberg or an uh snowflake native table.
[29:28] So all this difference is coming from lower query performance. This means you're paying more to get less to effectively uh degrade the experience of your users.
[29:46] Um what happens if we we look at data bricks right delta versus iceberg? Well the savings uh you can see it for yourself from 21 38 to 45% savings.
[30:03] uh but the most important part is that delta versus iceberg is 0% difference across all data sets. This is very important because as datab bricks you have seen it this morning in the keynotes um we are effectively
[30:21] publishing uh releasing new features in the context of iceberg but we really don't care if you want to use iceberg or delta our implementation of iceberg also has a liquid clustering in our platform that is not available on the other platforms
[30:38] and this is also one of the uh main reasons why the costs on our platform are so This is um a little bit about the query performance of concurrency test. So this
[30:54] is percentages. Um, you notice how essentially data bricks on iceberg is able to outperform even uh snowflake native tables on gen 2 on pretty much
[31:11] all the sizes I tested. So to me this what means for a platform to be open, right? um a platform to be truly open it means
[31:27] there is no performance or TCO gap not even feature gap between their own proprietary format if any this case data bricks doesn't have it and an open- source data format right if you are not
[31:42] able to show this you are not truly open because customers will start implementing on top of iceberg and then they will be forced to migrate straight back to native tables because of costing reasons, because users need faster
[31:58] queries or um simply because there are feature gaps between the two formats. What about adaptive? So, Snowflake has this new warehouse type which is called
[32:15] adaptive and it it essentially means autoscaling. So does autoscaling is effectively uh built to handle high concurrency use cases. So what happens if we start using adaptive and why adaptive was not part
[32:32] of my results? Well, there is actually a good reason. I tested adaptive. I did two runs with adaptive. So first thing, first approach was okay, it's autoscaling. So I will just um configure it to completely be uncapped, right? You
[32:52] can have as much as much hardware as you want. You figure out how much you need for those queries. And the r the tests ran very fast, but costs were to the roof. So apparently the autoscaling of adaptive
[33:08] overprovisioned hardware pretty much all across the the the queries and this uh resulted in quite a big waste. Then I decided okay then I will kind of cap the size to the same size I used for
[33:25] the other tests and I will also reduce the throughput ratio to four and see how it behaves. Right? and costs are much more reasonable, right? But still higher that compared to all the other tests. So,
[33:41] you know, this led me to think it's extremely hard to size it correctly and um you know, I I felt it was almost unfair to include it into the results. Um so to
[33:58] simplify, I decided to exclude it. So um don't trust me don't trust what I presented it tested it for yourself I am open sourcing the code of my benchmark
[34:15] you can take it you can run it uh by default it's uh simplifi what happened ah okay um it's simplified to run on a datab bricks notebook but you don't have to you can even export the jmx files and
[34:31] the query and run it on your laptop. And um this may be useful also to test your own architectures especially if you're implementing unified data layer. Those tests only covered uh managed iceberg
[34:47] for snowflake and managed iceberg for data bricks. when you are using external tables on iceberg on both platforms things can change right
[35:07] so um key takeaways so we see we saw how uh consolidating to a single platform is a key aspect for to reach AI readiness this is because especially for the financial sector you want something you want a unified uh governance layer across all your
[35:22] assets across ML, AI, functions, data, federated sources and so on. That will also give you a full coverage in audit audit logs who who has accessed
[35:38] what and uh and so on. So um also when you when we talk about AI readiness not all your AI applications are going to run inside datab bricks or your uh data platform
[35:54] that you chose. So open data formats also become important. The ability to access your data from another platform also become very important. and you don't want that this choice of open data
[36:11] formats will bring any compromises in terms of query performance in terms of uh TCO and so on. So that was it for for me. Uh if you have any questions I will be
[36:27] around for quite some more time and um yeah happy to answer it later. Thank you.

Learn more about the Databricks Data and AI platform.

The information provided herein is for general informational purposes only and may not reflect the most current product capabilities or configurations.