Skip to main content

Foundation First: Building Tax Data Strategy That Scales with AI and Databricks

Summary

  • Starbucks' tax function built intentional data infrastructure using Alteryx for flexible transformation logic and Databricks as the governed data backbone, replacing fragmented Excel-based and email-based workflows that created data wrangling bottlenecks and compliance risks.
  • Concrete use cases enabled by the governed foundation include automated apportionment calculations, functional interview planning, and self-service audit work paper generation—all requiring trusted data before AI can participate reliably.
  • This video argues that without a trusted data foundation, AI in regulated functions like tax introduces new compliance risks because fast, polished-looking outputs can be based on unreliable data, and organizational leaders—not AI—own the outcome when something goes wrong.

Foundation First: Building Tax Data Strategy That Scales with AI and Databricks

Watch: Foundation First: Building Tax Data Strategy That Scales with AI and Databricks
Tax operations are drowning in data wrangling. Spreadsheets, PDFs, emails, and fragmented sources across ERP systems create a bottleneck that prevents tax teams from doing strategic work. Without a trusted data foundation, automation and AI introduce new risks: fast hallucinations, compliance nightmares, and decisions based on unreliable outputs. Starbucks reimagined this by building intentional data infrastructure.
Discover how Alteryx provides flexible transformation logic and Databricks supplies governed data backbone to transform tax operations. Learn the evolution from reactive data cleaning to proactive foundation building, including concrete examples: automated apportionment calculations, functional interview planning, and self-service audit work paper generation. See how tax teams move from Excel-based analysis to governed Databricks datasets that enable compliance, analytics, and AI-powered tax insights at scale.
🤝

Chapters

FAQs

How is Starbucks using Databricks in its tax function?

Starbucks built a governed data layer on the Databricks Data and AI platform as the backbone for tax analytics and AI automation, replacing fragmented Excel-based and email-based workflows. The architecture uses Alteryx for flexible transformation logic alongside Databricks for governed data storage, enabling use cases like automated apportionment calculations and self-service audit work paper generation.

What data challenges do tax teams face when adopting AI?

Tax teams work with voluminous and complex data from multiple sources including ERP systems, Excel files, email bodies, PDFs, and images, all of which require significant wrangling before AI can process them reliably. This video explains that without a trusted data foundation, AI automation creates compliance risks because inaccurate outputs can look polished and convincing, influencing decisions that carry regulatory consequences.

What are the four foundational pillars for tax automation and AI described in this video?

This video outlines four foundational pillars for tax automation and AI, with data identified as the most critical because it underpins every other capability. The speaker Shreya Coram from Starbucks emphasizes that each pillar is necessary and that poor-quality or ungoverned data makes automation unreliable regardless of the sophistication of the AI technology used.

Why is a governed data foundation critical before deploying AI in tax operations?

Starbucks' Shreya Coram explains that when AI generates outputs at speed without a trusted data foundation, errors can appear polished and convincing, making them harder to catch and more dangerous in regulated functions like tax where incorrect outputs can create compliance nightmares. The philosophy in this video is that trust and reliability—not execution speed—become the binding constraint on AI adoption when data foundations are missing.

Full transcript

[00:08] Garbage at the speed of light is a saying that we're all too familiar with when it comes to automation. What happens when GenAI joins the party? That garbage at the speed of light might show up with some surprises.
[00:23] And the output might look very convincing, very polished. But the AI doesn't own the outcome. Whose butt's in the hot seat when something goes wrong? It's us. It's the leaders across the organization
[00:41] who took an action based on an output that looks reliable. Based on a report that maybe had some citations that weren't quite valid.
[00:58] When the speed of execution is no longer the constraint, trust and reliability is what will block scale and also become the compliance nightmare when something goes wrong and foundations fail.
[01:18] Hi everyone, I'm Shreya Coram. I lead tech and data enablement for the tax function at Starbucks. And we've been reimagining our function there with tech and data empowerment, uh refocusing what our teams are able to do and deliver, unlock new growth,
[01:35] and strategic value to partner with the business. Thank you so much for being here. Um I don't know if it was the tax data angle of it that brought you here or just filling in a session before the happy hour start.
[01:50] Either way, thank you so much for being here. There's different versions of a slide like this. The foundational pillars for
[02:06] automation and AI. I tend to ground into these four. Data is what we're going to dive in deeper today, but I do want everyone to take away that each of these pillars is critical and serves a crucial role in unlocking
[02:23] the power of possible with your functions. So, data. Data is hot topic here, hot topic at Days, and that has been a huge focus of my work.
[02:39] Setting not just the strategy with data, but truly building data-driven operations for the tax function. With that, design an architecture that can serve us now as well as in the future. Data's at the core of much of what we do
[02:56] in tax. We have a variety of data challenges from voluminous data, complex data coming from different sources. Uh not just ERP, um and some companies have more than one ERP, but really
[03:12] different sources and formats. It could be that Excel file that someone in accounting does to support their account. And tax needs it to get to the underlying detail for the determination and calculation. Could be within the body of email, text
[03:28] or images. We have lots of PDFs and images from contracts and reports as well as notices. So, with all of this work just to get at the data, historically tax would spend a lot of
[03:43] time data wrangling. Um a lot of time just to get the data in a format where it was readable and relevant to their work. You know, so they could do the tax part of tax work as a part of the data work
[03:58] for it. Tax also has the challenge of frequent changes. This becomes really relevant for both our data as well as our design and how we approach our builds.
[04:15] We require flexible design and quick adaptability. Let's imagine maybe in response to tariffs, business reroutes supply chains for how coffee gets around the world and with that VAT
[04:32] and GST determinations come up anew and you've created new intercompany flows that you have to report on and track. Maybe there's tax legislation that's coming in right at quarter end and not only does your team need to
[04:48] quickly model out everything to the pennies of EPS impact, but one of the items has a significant cash impact and you need to know what it means for that next estimated payment you have going out the door.
[05:04] Or maybe the business makes a change to your customer loyalty plan and all of a sudden you have a new book tax adjustment that you need data for and also to be able to calculate and track.
[05:20] So I purposely said all that so you aren't reading the slide ahead of me, but on the left data challenges probably a lot that you can relate to. What I really want to highlight is the last one. The requiring of lightning fast
[05:36] adaptability. We need to have the rules logic relevant to tax process and then use as close to the process owner as we can to those who own the outcomes.
[05:56] This is about where we've been focused. So, how do we go from challenges to data empowerment? On the right side has been um our main themes, but I'll just zero in on getting the right data in the right place in time,
[06:12] ready to consume, ready for those tax processes, and flexible to adapt to those needs. This journey we've been on had growing data challenges intersect
[06:29] with the rise of low code. And that emergence positioned us for the data empowerment that we've been on the course for the past several years. Being able to do that at the people and process level.
[06:53] Enter Alteryx. Uh my world was opened to Alteryx about eight years ago. And it was a true game-changer for tax. We were one of the first to adopt it within all of finance. I think we were second only to supply chain. And it remains a core tool that we use
[07:10] today. Just within tax, we have dozens of workflows deployed. You can see some examples at the bottom there. These are delivering higher data quality, accelerated turn time, and the running total over this eight-ish years
[07:26] has redeployed thousands of hours for the tax team to focus on more strategic work. And as well as it re-reinvesting and innovating the function. Some of these use cases are more
[07:41] straightforward reconciliations of just states or staging data. A couple highlights of that I'll just mention to give you a feel for them. Apportionment is a fun one. Uh so that's
[07:56] how you figure out how much income's going to be taxed in all the different states. So it's bringing property data, payroll data, revenue data. For many that comes from different systems and sources. And then you layer on top of that all of
[08:13] the states and several cities have their own set of rules of what's within or out of the base, how the calculations determine it. So that's a lot of data to bring together and rule logic to manage on top of it.
[08:28] Another example I like to mention, um mainly because it doesn't involve summarizing financial data, we have a workflow that just helps the team plan for functional interviews. So I don't know if anyone's participated with that with their tax team, but R&D and engineer teams, we tend to have to have
[08:45] conversations to understand the nature of the work and then to figure out what that means from a tax reporting standpoint. The common audit requests that's for tax authorities, we've been really focused on the US ones, but typically there's a
[09:02] common set of work papers and data that they ask for at the start of an exam. And this was one of our first Alteryx deployments was for sales tax audits. And we actually created it so that could
[09:18] be self-service for the team. So just enter your parameters of the audit and just get it straight from gallery. Where we want to take this evolution next is to connect the job trigger actually into SharePoint cuz that's
[09:34] where the tax teams are otherwise doing their work, logging the other um information about the audit. So, in the future, our vision's to have those work papers just show up in the same folder that they're working in.
[10:02] The key or one of the key things I'd like you to take away is that's not enough for the data to exist. That data needs to be ready, relatable, and trusted for the processes consuming them and the end users who are taking action on it.
[10:17] So, left are some of the qualifiers that we look for, but for time, I'm going to switch gears to the right cuz this is the more that we're looking that we've been unlocking with a data layer. So, we did some amazing things with data
[10:32] transformation solving for some of those challenges. But, here's what's next for the unlock. So, having data flows and date managed data connections for those on-demand pulls and scheduled job runs,
[10:48] so that the data can get where it's needed. Streamlining transformed data to serve multiple processes. I think of this is like a highway system where there's different off-ramps to deliver the data tailored to the each process.
[11:09] And storing the tax process data to enable more. This is what I'm really excited about with Databricks and I'm going to go into a quick use case about that.
[11:26] Bringing it together, Alteryx is our managed logic layer. So, our definitions, our rules, our policies, everything that is taxes owning and managing, that's where we're primarily doing that through Alteryx.
[11:42] And it allows us to have that flexible design and quick adaptability to meet the needs of the team when they need it. And then Databricks is that governed as going to be that backbone to govern and scale that data. And Genie is going to let us do cool
[11:58] things once we have everything in place. So, um we're still exploring that, but allowing users to interact with dashboards, even create create them instead of us having to guess what they might want to see, uh getting narrative insights on
[12:16] trends and outliers and blind spots. All right. I chose a use case that's more accounting data than tax data to hopefully make it a little bit more relatable for the room.
[12:38] For intercompany data, I don't know how many of you dabble in that space, but you might think that if company A and company B are held by the same parent company, that maybe the data challenges wouldn't be the same. Um
[12:53] yeah, they uh they still have their own challenges. So, just think for this example, tax needs transaction data, balance data, needs to be able to match between party A and party B, also needs to know what it was for, agree it back to contracts,
[13:11] check the amounts within um agreed-upon ranges and benchmarking. And then of course the tax authorities want to have a bunch of forms that we have to fill out to to say everything. So,
[13:29] this was our this has been our process before adding Databricks. So, I just rattled off most of the data sources on the left. Nothing too crazy. A lot of it's structured. Um mostly the financial statements as well as the tax managed list. Um
[13:46] maybe less structured. Our transfer layer, we have some managed data connections, but still working on that area. Um Alteryx has um done a lot, so we have solved for a lot of our um transformation needs moving
[14:02] from what was done in Excel to Alteryx. But, look what's happening on the right. What are my tax people doing? It's all Excel land, so wherein tax is you know, needing to look at the output for their review, their calculation to
[14:19] get it on a return, it's going into Excel. And it's living in different places. So, what happens when someone needs to do analysis for planning or scenario modeling? They're opening up Excel.
[14:39] So, where are we going? Adding Databricks is not only going to aid us in compliance. Um I know I have the like a lovely picture here, but just imagine Databricks is both bringing the data together in the upstream, so we have a cleaner pick up point. And then it's
[14:56] also housing that tax process data to be able to in that changes the consumption interface of where tax users. And I'm going to cap them on how many Excel sheets, but I still gave them at least one.
[15:13] So, with that Databricks is going to give us that governed data lake with trusted outputs that drive outcomes.
[15:30] So going from data ready to AI ready, I've been focused on data readiness for most of this talk. To be AI ready is to have meaning beyond the numbers. To enable teams to focus on that strategic work with adaptable design at the speed of tax or insert your business partner name
[15:47] here. The foundational investment is key to unlocking that power of possible for AI and whatever comes next.
[16:06] So as you probably gathered, this is sponsored by Alteryx and I will be hanging out with them at their booth tomorrow. So I definitely go you know, there is still at their booth for a little bit today, but they'll be around the next 2 days. So stop by and um hear about the cool things that
[16:22] they're doing with Databricks. I hope that you learned something. Um maybe got inspired to focus on foundation as well as the magic of AI. And if you want to chat with me, I'll be
[16:38] at the Alteryx booth tomorrow like mid-morning 10:30-ish to about noon. You can also find me on LinkedIn. Um I love to talk all things data, automation, governance, and empowerment. And um with that, thanks again so much
[16:56] for tuning in and I don't think I could do questions with this uh voice arrangement, but if anyone does want to ask me a question, I'll hang out for a few minutes at the end too.

Learn more about the Databricks Data and AI platform.

The information provided herein is for general informational purposes only and may not reflect the most current product capabilities or configurations.