Skip to main content

7-Eleven: Accelerating ML Delivery with Agentic AI and Databricks Asset Bundles

Summary

  • 7-Eleven and LTM delivered an end-to-end ML anomaly detection pipeline for fuel equipment across 85,000 stores in under two weeks—compared to a traditional six-week cycle requiring five-plus engineers—using Windsurf as an agentic AI force multiplier for one principal engineer.
  • Databricks Asset Bundles, GitLab, and Delta Lake provided the deployment, governance, and ingestion infrastructure, with automated ServiceNow ticket generation triggered by the anomaly detection model to initiate maintenance workflows without manual intervention.
  • Success depends on skilled engineers who can define context, establish guardrails, and govern AI-generated output, with modular progressive development being the key approach that makes agentic code generation practical at enterprise scale.

7-Eleven: Accelerating ML Delivery with Agentic AI and Databricks Asset Bundles

Watch: 7-Eleven: Accelerating ML Delivery with Agentic AI and Databricks Asset Bundles
Operational resilience at 7-Eleven depends on predictive maintenance across 85,000 stores generating continuous sensor telemetry. When critical fuel equipment fails, millions in revenue are at risk. Instead of the traditional six-week development cycle requiring five plus engineers, 7-Eleven partnered with LTM to deliver an end-to-end ML solution in under two weeks using agentic AI.
Discover how Windsurf, GitLab, and Databricks Asset Bundles accelerated code generation and deployment through modular, progressive development. Learn how one principal engineer, paired with agentic AI as a force multiplier, built a complete anomaly detection pipeline with automated ServiceNow ticket generation, governed Delta Lake ingestion, and continuous monitoring at enterprise scale.

Chapters

FAQs

What business problem did 7-Eleven solve with this Databricks ML pipeline?

7-Eleven needed to detect anomalies in fuel equipment at thousands of gas station locations before failures occur, since cumulative downtime across many locations represents significant potential revenue loss. The solution uses anomaly detection models that automatically generate ServiceNow maintenance tickets when equipment signals a potential failure, enabling proactive intervention without manual monitoring.

How did agentic AI tools like Windsurf accelerate the development timeline?

Windsurf served as a force multiplier allowing one principal engineer to accomplish what would traditionally require five-plus engineers over six weeks, delivering the complete solution in under two weeks. The approach relied on modular, progressive code generation—establishing context first, then generating and validating components incrementally—to keep AI-generated output focused and testable.

What role do Databricks Asset Bundles play in the 7-Eleven solution?

Databricks Asset Bundles provide the deployment infrastructure that packages and deploys the ML pipeline components—including Delta Lake ingestion, anomaly detection jobs, and monitoring workflows—in a governed, reproducible way across environments. This video explains that Asset Bundles enable the modular architecture that makes agentic code generation practical by providing clear boundaries for each component.

What are the key lessons from 7-Eleven and LTM about using agentic AI for ML delivery?

The key takeaways from this video are that agentic AI tools dramatically reduce delivery time but require skilled engineers to define context, establish guardrails, and validate AI-generated output. Success depends on a solid Delta Lake foundation for data governance, integration with enterprise systems like ServiceNow, and a modular development approach that keeps each generated component focused and testable.

Full transcript

[00:08] Good afternoon everybody. It is uh 110 right now. So it's time to start our presentation. Um welcome to um the presentation organized by LTM and 7-Eleven on agendicai part of the data bricks data summit.
[00:29] This is the legal statement that they want us to display. So it's very standard. Have to display it for a few seconds. Thank you very much. As a reminder, do not forget please to complete your surveys. Surveys always help us receive feedback and we improve
[00:45] our presentations and make for a better future event here at Datab Bricks. Thank you for that. And with that we're ready to move to the main part of this presentation. This is
[01:00] called accelerating operational resilience at scale and it is a collaboration between 7-Eleven and LTM. Uh the particularly we're looking at how 7-Eleven's integration of data bricks and aentic AI
[01:18] um help us accomplish this operational resilience at scale. We're going to go through some uh challenges, present your our solution to us, discuss the benefits, the lessons learned, and then some final thoughts.
[01:33] At the very end, uh this is a lightning presentation, so it's going to be only 20 minutes total. But at the very end, we will try to have three four minutes for questions if you have any. So, let me start then a closer look at
[01:48] what we are talking about. Now 7-Eleven is a company that has many many thousands of stores worldwide about 85,000 right now internationally. Uh we are housed at the headquarters in the Dallas Fort Worth area where we are
[02:04] taking care of the North America which is United States and Canada and that is about 13,000 stores. Out of those many thousands of stores have gas stations. The particular project is about fuel equipment. In other words, gas stations
[02:19] and fuel pumps and it is equipment down down time making sure that uh we call for maintenance when the equipment requires that. When critical equipment fails, sales are intermediately impacted. As you can understand, uh the
[02:36] cumulative effect of downtime across thousands of locations represent represents a massive potential revenue loss and we want to avoid that. Um so let's see now what is the traditional approach to developing such a solution
[02:53] that will protect the revenue stream loss due to equipment downtime. Traditionally you will have to do to acquire your talent right so identifying assembled teams uh takes time to make sure that you have all the required
[03:08] skills the cost of development so your highly skilled team members would have to uh work on extended timelines a project like this typically require a team of five plus resources and the time to market there's a cycle
[03:25] until you actually have this solution in production going live to all those thousands of stores that we have. That's a traditional approach and now we are doing a more modern approach using Aentic AI
[03:41] and let's talk about what this looks like. Okay, so what you see at the top is the broader environment where we operate. Here we have the store equipment on one end and the delta lake which is a datab bricks uh actually all
[03:57] plat platform at 7-Eleven is on the Azure cloud. So it is Azure datab bricks cloud and we have the delta lake. Then we have the machine learning modules that implement the solution that detects when the equipment is down and produces
[04:15] some alert. And then finally we have uh fully automated generation of service tickets uh and work orders. Um in our environment this happens as a service now uh work incident ticket and work
[04:30] orders and there's full integration end to end across all that. Uh the agentic workflow specifically um includes code generation and that in this example was done using wind surf as an ID the integrated
[04:47] development environment uh integrated with gitlab and then finally uh receiving um the datab bricks going to data bricks environment where it is running as a production job. Uh
[05:03] Rajes, thank you Nick. Um so um what we do is we first ingest the data of uh all the equipments the which includes a telemetry and uh sensors into our data lake. Um and then we use machine
[05:19] learning modules built on data bricks. Um that would do this anomaly detection. Um and then what it would do is it would create all the alerts and then we integrate that with the alerting module which uh we have as a service now. So we
[05:34] create all our uh alerts the anomaly uh detection alerts into service now and then there is a help desk team that uh is monitoring the these incidents uh and then they would action on these incidents. uh either they would uh create the work orders. So for the for
[05:51] the technicians to go and fix uh or they would um um instruct the store managers to kind of do the self um applied u solutions to kind of get get it back uh online. Um
[06:07] next slide please. Um and this is how we uh have automated our uh or accelerated our delivery. So we use Vincerf which is uh the the agentic uh code developer tool uh that is approved within our
[06:24] organization. Um and uh this is uh the current situation as of now uh that we are using windsurf. Um and then we use that IDE to generate the code. Um and then we have the gitlab. So essentially um windsurf is outside of data bricks.
[06:42] Uh we need to have a way to have them into datab bricks. So we use the GitLab and then automated CI/CD module pipelines through databicks asset bundle and then we then deploy them into databicks.
[07:00] So this kind is a kind of a snapshot of how we use winerf uh to u to to accelerate the whole thing. So the the most important thing that we did is to spend time in establishing the context. It's not u that like we start right away with writing the code. We uh spend more
[07:17] most of the time in giving a very lengthy context uh explaining how we would do it to a uh to a developers to or to a set of people who are going to do the development. um and make sure that they ask and let them ask us
[07:33] questions after what we have explained and then clarifying all the answers just to make sure that there is a a proper understanding between uh the developers before they start writing the code. And then what we do then we modularize the
[07:49] whole uh development into uh and then we let the windsf develop each module review each module just to make sure that it is um it it's good enough uh to be moved forward with
[08:04] um and then we go one module at a time and then after all of it we we must make sure that they are integrated into one uh pipeline and we did that uh so it is uh um a sequential not sequential it is kind of
[08:20] a progressive um um gener code generation and uh review and then make sure that we um we know exactly what's been generated before we move into the next thing and then we integrate all of that with the automated CI/CD module uh
[08:37] on GitLab and then we use datab bricks asset bundle which would uh do all the deployments into datab bricks So then to return back to the big picture of all this, uh the traditional approach as I mentioned earlier would
[08:54] require typically a team of two or three engineers for this particular project and uh most likely five plus weeks for the full development effort until we put it to production and that is a higher cost and slower time to market. With the
[09:10] agentic approach that we followed in this project, we used only one principal engineer and that was actually the LTM resource that we had uh as a contractor um highly skilled principal engineer though and wind surf as a force
[09:26] multiplier. So what we managed to do here we deliver this in under two weeks end to end the whole thing and it was the same quality but delivered at a fraction of time. And here I want to clarify that when I say two weeks I mean
[09:41] until we were ready for deployment and that is nothing sort of amazing. Uh we did spend about one week slowly developing the context and uh uh all the details about the environment about the project uh multiple back and forth uh
[09:58] prompts. The prompts were typically two pages long. The response is sometimes five pages long each prompt right and there were many many rounds of that. Um, and that's how we managed to do it. And the other thing I want to clarify is that
[10:14] it's not that it was our fault. We forgot about the deadline or we were doing other things and then we rushed to the end. Not at all. Actually, there was an external crisis with some other team and we volunteered to help them and they didn't even know it was possible to do
[10:30] it that fast, but we did with the help of this agent AI approach. uh the impact now for the organization is that we managed to protect millions of dollars in revenue stream and uh here now a technical clarification why do I say we
[10:48] protected the revenue stream right and I do not say we generated sales dollars that's because uh when you have a gas station think about it it it may have eight pumps 12 pumps 16 20 pumps sometimes all right so if one of
[11:04] them is down that does not necessarily mean that the average revenue stream you get from one this one pump is going to be completely lost right because the pumps are not generating revenue independent of each other. When one of them is down other pumps will pick up
[11:21] some part of the um of of the traffic that usually is picked up by this. But this is not always the case. So the smaller the gas station also some locations are more convenient and sometimes the customers do not have the time. They get frustrated with their
[11:37] experience at one pump and then they leave before they try the others or maybe the entire station is too busy and loss of one point of sale also is makes a difference and uh therefore there was this one uh benefit and we were able to
[11:54] quantify it in terms of dollar amounts and also the other was that we're able to deliver this project with minimal resources. uh speed to market was uh reduced down to two weeks instead of the five plus week that we have taken
[12:09] underway otherwise. Key takeaways uh 7-Eleven collaborated with LTM to implement this agent AI approach. Uh we operated within a unified Delta Lake environment. So it was very important
[12:25] that we had uh uh datab bricks and also wind surf also gitlab uh all of them um seamlessly uh even even the documentation of this final project
[12:41] which was done in confluence was generated by wind surf in this case and we had some very nice um summary of uh the solution posted over there for the rest of the team Um then um what I wanted to mention also is
[12:58] that agentic AI dramatically reduces engineering friction but uh strong CI/CD and government governance are still essential. Um also um small highly skilled teams can
[13:15] deliver enterprise scale ML solutions when they are empowered with the right tools. future considerations now what we could have done differently if we could start all over uh the truth is that at the
[13:31] time we were at a more preliminary phase of our agentic AI journey at 7-Eleven because this is a project we're referring back now to October 2025 right and things are already very different so one thing that is of primary importance
[13:46] is the semantic layer I'm sure you're hearing about this uh continuously right now um at the time we did not have any right now we have efforts underway to build the context layer at 7-Eleven and
[14:02] by context layer now I'm referring to closer to the data ontologies but I'm also referring to uh if you think in in terms of a aentic ID environment the skills and the rules that you are to follow and just on a higher level uh
[14:19] speaking to from the human point of view is we have uh compliance, we have the legal environment where you operate. We have the industry best practices and you have your own company policies. But then also specific to every project, you have
[14:36] the project goals and you have the micro environment in which your own project operates. All those are part of the context. So um rather than trying to build very long prompts so that you can establish
[14:52] that context which is what we did in this particular implementation that we're talking about. One could create a more organized and more structured context layer so that the agents connect to that and they already have the context and you don't have to repeat it
[15:09] and then your prompts can be much shorter and much more project specific. Uh would you would like to add anything um Rajes? Oh go ahead I'm good Nick. Okay. So, um, another thing that we're doing currently
[15:25] at 7-Eleven is we try to integrate this with Jira, right? To give you some, uh, broad reference here, uh, the development effort we had in this project, if I were to summarize it was
[15:40] use 18 prompts and each one of them is about one page long. The way we actually did it back in October was not exactly like that. It was similar to that. It was more like we offer a prompt that was about two pages. We get a response that was about five pages long and then we go
[15:56] and read it paragraph by paragraph and then we say can you please lay out the architectural approach for this solution and then we say step one step two step three and then the next prompt would be on your step three I have a comment I have a question why did you offer this approach is wouldn't it be better to do
[16:12] it this other way and then the agent would say no actually you're right there are pros and cons and then there would be an analysis of all the pros and cons so going back and forth now if we could do this once again I would say that the optimal size would be 18 prompts of about one page each. Those in theory
[16:29] could have been zer story descriptions and then instead of writing them as prompts to the ID, you could just write them as Jira stories and then give the ID the Jira ticket number and then say refer to this Jira ticket and please implement this for me and then that's
[16:44] what we are doing right now. We have started this integration with Jira. Um another thing to mention is is windserve the only ID no obviously not we have genie at the day the bricks that we are considering right now uh there is codex
[16:59] there is cloud so there are other approaches uh I'm not trying to say now that w wind surf is the best one but at the time it was the one that was approved for 7-Eleven and that's why we used it um and finally there is the command line
[17:15] interface the datab brick CLI integration that one could use from within windsurf to run the code in datab bricks. Um, and this is the end of our presentation.

Learn more about the Databricks Data and AI platform.

The information provided herein is for general informational purposes only and may not reflect the most current product capabilities or configurations.