Skip to main content

Real-Time IoT Pipeline with Databricks: From Raw Signals to Analytics

Summary

  • Milwaukee Tool built DataLink, a real-time IoT platform that collects high-fidelity sensor data at 20 kHz from power tools in development and streams it through MQTT brokers and ZeroMQ into Databricks, reducing data availability latency from 24 minutes to seconds.
  • Spark Declarative Pipelines parse binary sensor payloads into structured time-series signals, enabling ML engineers to run experiments and root cause analyses in near-real time rather than waiting for the next batch data cycle.
  • The platform cut problem-to-solution time for engineering investigations from 5–6 weeks to hours, and Milwaukee Tool plans to scale the architecture globally to 100+ hubs to support product development and innovation across the trades.

Real-Time IoT Pipeline with Databricks: From Raw Signals to Analytics

Watch: Real-Time IoT Pipeline with Databricks: From Raw Signals to Analytics
Real-time IoT data collection at scale is critical for manufacturing, but infrastructure challenges limit speed and agility. Milwaukee Tool generates terabytes of high-fidelity sensor data from tools in development, collected at 20 kHz with 50 simultaneous variables. Their previous approach required 24 minutes from data generation to cloud availability, creating bottlenecks that delayed engineering iteration and root cause analysis.
DataLink is their real-time IoT platform using Wi-Fi collection, MQTT brokers, and ZeroBus streaming into Databricks. Spark Declarative Pipelines parse binary sensor payloads into structured signals in seconds, enabling instant ML experimentation and real-time alerts. Learn how this solution reduced problem-to-solution time from 5-6 weeks to hours, and Milwaukee Tool's path to deploying 100+ hubs globally.
🤝

Chapters

FAQs

What is DataLink and what problem does it solve for Milwaukee Tool?

DataLink is Milwaukee Tool's real-time IoT platform for collecting high-fidelity sensor data from power tools under development. Before DataLink, it took 24 minutes from data generation to cloud availability, delaying ML experiments and root cause analysis; DataLink streams data in seconds, enabling engineers to iterate much faster.

How does Milwaukee Tool collect and stream IoT sensor data into Databricks?

DataLink collects sensor data at 20 kHz with up to 50 simultaneous variables from Wi-Fi-connected tools in development labs. The data flows through MQTT brokers and ZeroMQ into Databricks, where Spark Declarative Pipelines parse the binary sensor payloads into structured time-series signals ready for ML models and real-time alerting.

What impact did DataLink have on Milwaukee Tool's engineering process?

DataLink reduced problem-to-solution time for engineering investigations from 5–6 weeks to hours. With data available in near-real time, ML engineers can identify issues, test hypotheses, and validate fixes within the same day rather than waiting for the next batch data cycle.

What is Milwaukee Tool's plan for scaling DataLink globally?

Milwaukee Tool plans to deploy 100+ DataLink hubs globally to support tool development operations across their worldwide manufacturing and engineering network. The architecture is designed to scale with Databricks handling increasing data volumes without requiring significant changes to the ingestion or processing pipelines.

Full transcript

[00:07] All right, good afternoon everyone. I appreciate you coming to the our presentation. Just a reminder about the forward-looking statement and then also there'll be the survey at the end for you to fill out.
[00:26] But uh we're Milwaukee Tool and today we're here to talk about a real-time streaming a data streaming platform that we built using ZeroMQ. So my name is Max Merget, uh senior manager of ML engineering and with me is Karat Mocha, one of our talented ML
[00:44] engineers. Little background on Milwaukee Tool. So we're over 100 years old as a company. Uh we were founded and are still headquartered in southeastern Wisconsin and currently we have over 25,000
[01:01] employees across the globe supporting all of our operations. For those not familiar, to sum it up uh concisely, I'd say that we are a solutions provider for professional users in the trades. Think uh plumbers, carpenters, electricians that work on
[01:19] things from residential all the way up to the largest commercial projects like the buildings that we've been in all week for the conference. And now some of you are probably asking like why is a power tool company presenting at a data and AI conference and that is 100% a valid question.
[01:36] Uh and the answer is because over the last few decades we've had to become more of a technology company uh to continue delivering uh the performance and productivity uh and safety that our users demand from us.
[01:54] And if you take a look at our history, that that really shows. So AI is getting a lot of attention right now, but for Milwaukee Tool, that underlying mindset is is very familiar. Um we've always led through very purposeful innovation. All the way back in 2008 when we were putting embedded
[02:09] systems into power tools to when we were pioneers with lithium-ion technology and brushless motors, and now AI and ML, uh our approach has main remained uh consistent. We don't adopt technology for the sake of technology. We're always looking to
[02:26] apply it when it helps solve real-world problems for our users. And that's why AI is not a shift in who we are, it's just a continuation of how we've always innovated. And compared to the rest of the industry, our approach is it's different. Uh first and foremost, we
[02:43] start with the user problem. Again, AI is not the goal, improving productivity, safety, and performance is, and looking to apply the right technology to solve real job site challenges. Second, we we build it and we prove it
[02:58] in the field. Um our solutions are developed and tested on real job site locations, taking advantage of those thousands of people we have working across the globe. So, we're getting out of the lab and ensuring that they perform where it matters most.
[03:13] And third, we move faster through integration. Uh we're looking to connect data across design, testing, manufacturing, uh to create faster iterations and accelerate innovation. And then most importantly, our culture is really what drives it all. Uh we're constantly
[03:30] investing in our teams and embracing rapid iteration. And again, you know, AI is important part of the story, but it's it's really only one part. What matters most is that it's never a single piece of technology, it's applying the right technologies to
[03:46] solve that user problem. Uh so, depending on the challenge, it might involve hardware, software, AI, or or a combination of all of those. Uh and again, that that focus remains consistent. It's it's delivering that solution to help the user uh in their jobs that they do each and every day.
[04:08] So, let's take a look at an example then that really highlights this. Uh so, just earlier this year at World of Concrete, which is one of the the biggest trade shows that we go to, we announced our newest M18 FUEL circular saw, and it's a it's a strong example of how that innovation model has
[04:23] continued to evolve. Um and again, you can see that this example is bringing together multiple of those technologies all working together. And in this case, the the problem that we were trying to solve was both performance and safety during demanding
[04:39] cutting applications. Um this is a tool that our users are going to go and grab when they have, you know, big jobs to complete. They're not just doing a quick cut, they're doing large applications all day long. Um and so, we're bringing together battery, motor, embedded systems, and and machine learning to to
[04:55] really achieve all of that. And what we're doing here with with AI and ML is helping the tool continuously interpret what's happening in real time, and then recognizing conditions that may indicate a kickback, uh which if you're not familiar is when the blade of the
[05:11] tool can bind up and actually kick back at the user. Um and so, we're we're trying to detect that, react to it, and then respond almost instantaneously. Uh the result being a safer and smarter tool.
[05:30] Now, unlike other use cases for for AI and ML, where you're working with source data such as, you know, transaction information or things where the data is generated organically, um we we don't have that liberty. We need to go and create this data that's used to develop features such as auto stop. And this gets back to the root of
[05:46] understanding our user and the factors influencing the problem we're trying to solve. So, staying with auto stop, uh this feature was actually originally designed for our flagship M18 fuel drill, and it required understanding all types of information such as
[06:01] what accessory is on the tool, what material you're drilling into, the state of charge of the battery, the ambient environment, is it an old tool or a new tool, and you know, and then how's the user interacting with it? Are they holding it one hand, two hand? Are they using the side handle? Are they drilling
[06:18] overhead, right? There's there's lots of different things that can go into this. And um all of that allowed us to create a solution that prevents kickback, but also avoids shutting down too early and being a nuisance. So, we're we're using all of that information to dial in
[06:34] something that truly is focused on productivity and safety. The problem is every time we add a factor to our experimentation and data collection, this scales exponentially. So, that can result in terabytes of data being generated just to develop a single
[06:50] feature on a single tool. And then if you step back and take a look at the full NPD cycle, uh there's lots of activities beyond ML features that can take advantage of this data. From setting day-in-the-life targets to
[07:06] life testing products, all of these activities are improved with more and richer information. And of course that improvement allowed us to move from prototype to mass production knowing that our solution is going to meet uh the needs of our users that we're constantly focused on.
[07:23] So, with that I'm going to hand it over to Karat now, who's going to dive into the details about what that data is and how we collect it. Can you all hear me? Cool. Thanks, Max, for all the context there. Um and thank you all for being here again. I know it's pretty late in
[07:39] the conference and one of the last presentations, so I you guys signing up and uh hearing hearing us out. Uh let's take a deeper dive on what our data actually looks like. So, often times when we're developing features or
[07:54] doing any type of analysis on the tool as engineers, we are collecting a lot of real-time time series data. Uh this data is variables within our firmware that is spinning the tool. This data is sensor data coming from the
[08:09] sensors that our tools equipped with, statuses, flags, errors, all of that kind of stuff. For some of our more advanced use cases, we can sample this data up to 20 kHz in terms of sampling frequency and record
[08:26] up to 50 variables simultaneously. Uh so, you know, think of workloads like machine learning. That's where we really push the bounds of how many variables we're recording and how fast we have to record them.
[08:43] But, this presents us currently, at least in our business it did, uh with several challenges. So, our engineers, when we surveyed them, we found uh they wanted more and more bandwidth constantly. Our current systems weren't able to handle that type of developmental data collection from
[08:59] the tools. Uh we started to find that the time to access from data generation to data consumption by our engineers or what we call our end users in the engineering side, uh was really long. They wanted it
[09:14] faster. And lastly, it the process complexity of setting up the data collection process on our tool, uh you know, being able to extract it, analyze it, query it in the cloud was a little bit cumbersome.
[09:32] So, diving deep a little bit further into uh those three aspects. If we look at our current system, we are reaching its peak and we'll talk about what this current system looks like uh, a few slides later. But, right
[09:48] now we can sample about maybe 75 tools streaming data into our system um, all simultaneously, pushing about 16 GB an hour and we are seeing a whole bunch of issues. Our data is extracted in big batches. So, if we
[10:03] have any type of uh, failures, crashes, or the service goes down, we are losing thousands of applications in that batch. We are also seeing the system crash more and more often as more and more tools
[10:19] are starting to record this type of high fidelity data. And then, what it results for me personally is a lot of late-night calls with our global teams to try and figure out and bring the systems online or explain somebody the process of resetting a device remotely.
[10:36] But, this demand isn't stopping here. By the end of this year uh, or early into next year, we anticipate a lot more tools hopping on this type of data collection in our organization, which can lead to over 4,000 tools streaming data,
[10:53] 900 GB being uploaded every hour. So, how are we going to handle this workload globally? The second problem was time to access. So, with our solution currently, we are
[11:10] doing um, uploads by connecting a tool using a hard wire into uh, into our laptops using a client app. And these uploads go into AWS. AWS will do some type of parsing on this data, insert it into some tables where we can consume
[11:26] it. And then, the users have to query the data and then extract it. Now, if you notice some of those times here uh, that we face currently, from data generation to consumption, we're looking at about 24 minutes. What this means or another way to look
[11:43] at it is when I spin a drill and I do one application, I have to wait 24 minutes before I can query that data in the cloud and actually use it for any type of analytics. That's pretty long when it comes to engineering or I like to say engineers in our organization
[12:00] absolutely do not have the patience for that. And lastly, the process of setting up this data collection and actually doing it, especially being a global manufacturer and having operations globally, we have to
[12:17] communicate this across the oceans, right? And you look at the amount of steps an operator has to follow to gather this data from a tool that's undergoing testing. They have to remove the tool from testing, connect it to some laptop, extract the data, wait for
[12:34] the data to download, erase the memory, put the tool back on testing. Without our data collection, all of these steps get it eliminated. So, um engineers in the organization, we need to incentivize them to collect more and
[12:49] more high fidelity data, not deter them because the process is cumbersome, right? So, all these three big challenges, um we we had to solve them by coming up with a solution.
[13:04] Here's how we do this currently, and we've tried collecting this data over Bluetooth in the past, so our tools are equipped with One-Key technology which can use Bluetooth to collect this data. But, Bluetooth simply cannot handle the type of bandwidth that we are looking
[13:21] at. With 50 variables streaming at, you know, over 1 kHz at times, there there's no way Bluetooth can can reach those uh that type of workloads. But, the nice part about Bluetooth that we saw was the time to access was
[13:37] really nice because we could constantly extract this data from the tool while skipping the process of connecting it to some client app. This also sped up the test efficiency a little bit because the operators did not
[13:52] have to spend time to take tools off of test. The tool was collecting the data while the test was ongoing. But, Bluetooth didn't give us what we were looking for in terms of the data itself. Hence, why we took the onboard storage
[14:07] approach. So, developmentally on our tools, we will put some type of onboard storage device on them to be able to collect all these variables. This gave us exactly what we were looking for in terms of recording capability, but the time to access is what I was explaining earlier with 23 minutes, it
[14:25] was way too long before we could start to consume the data. The operator efficiency was really low because they had to follow so many steps to try and gather the data.
[14:41] Okay. So, now we're arriving at how do we solve this problem? This is where we developed our solution called DataLink. So, DataLink is uh an IoT-based device that uses neither Bluetooth nor a hardwired connection to extract data from our tools, and then we
[14:57] leverage the power of ZeroBus to land that data in two DataBricks right away, uh and use Spark declarative pipelines to parse it and get it into consumer-ready tables. There's three key aspects to the system.
[15:13] How the tool sends the data to the hub. How the hub uploads the data to the cloud, which is using ZeroBus. And then, how does the cloud parse and get the data ready for consumption? So we'll take a deeper dive into those three details.
[15:35] So instead of using Bluetooth or any other wireless technology, since we were doing this project developmentally, we had a little bit more leeway on what we could use within our organization. And we decided to go the Wi-Fi route. We were really inspired by the capabilities of AirDrop or other
[15:54] uh CarPlay or other levers that the industry uses to stream large amounts of data really fast. So the hub projects a Wi-Fi access point. The the part of this access point is not to
[16:09] be connected to the internet, but really just to use Wi-Fi as a protocol to transfer data. So the tools connect to that Wi-Fi access point. We have a MQTT broker running on the hub, and the tools are publishing their raw payloads as on the topics on the MQTT
[16:27] broker. So this is how we were able to get past that first step of collecting high-fidelity data in real time um from our tools.
[16:43] So now that we have landed a bunch of zeros and ones on the hub from the tool, how do we get it onto the cloud? And this is a little bit of a multi-step process. So the hub is going to take that raw payload and then partially parse it. This allows us to extract some metadata
[17:01] from each payload that comes onto the hub. So now we know, okay, we have a lot of signal or sensor data that came in. This is the tool it came from. This is the hub it came from. This is the dictionary ID that it relates to. Then the hub that's running a
[17:18] microservice based application, is going to store this into a local database. Our upload service within that application is going to establish a stream with ZeroBus. We use a GRPC ZeroBus stream, and it
[17:34] will send this data onto our bronze level table using ZeroBus right away. Great. Now we have the data that was generated on a tool, landed in the cloud within maybe one hop.
[17:54] Now that the data is in the cloud, we need to take the bunch of zeros and ones and parse it. This is where the Spark declarative pipeline really shines. So, each time an engineer wants to collect data on a tool, we have them set up a dictionary of what they're collecting, which is stored in Databricks.
[18:10] The Spark declarative pipeline is able to look up this dictionary, and now it knows how to parse the binary payload. So, we can use things like UDF functions to parse that payload into signal level data. So, you can see I've converted, you know, a bunch of zeros and ones into
[18:26] something like current and voltage, we which we can actually analyze. So, now this data gets inserted into our final uh silver level tables, where engineers can query it, we can do analytics on it, we can analyze it further.
[18:43] And if my luck's not too bad today, I might actually show you a live demo that might work here. So, we'll see how that works out. All right. So, I have my hub here. I have it paired with a tool. Uh in case in this case, I have a drill. And on my
[19:01] laptop, I have a app running, which is basically querying the final tables and plotting that data for me once it comes in. So, it's monitoring the data that's coming in from this tool. It'll give you some latency numbers of the difference
[19:16] in time between the data landing from the tool to the hub, and then the data when it parses and lands into our silver level tables. So, that 6 seconds is the whole data flow. So, we'll run it and let's see if it shows up. Touch
[19:33] wood. Is there anybody here? Um it takes a few seconds. Uh because the Spark pipeline will churn through all the binary payloads, and you can see a new application has showed up there with all our secret.
[19:48] That, too. So, there you go. We have a new application. I'll take that applause. I was really nervous about this one. But, this is not the only data that our pipeline is parsing.
[20:05] So, within the pipeline, we are sending the tool data. Alongside the tool data, we are also sending hub heartbeats, what tools are connected, what is their status, signal strength, what are the issues, logs, have we failed a payload, is there
[20:22] something that came in wrong, did we miss a sequence ID. All of these All of this information can be parsed by the same pipeline. So, once it lands, the pipeline looks at the event information, what kind of an event it is, and parses the data accordingly, and
[20:38] pushes it into all different tables for us to consume from. So, now we've taken the data from the tool all the way to the cloud. Uh what happens after?
[20:57] Our users in this case are consuming the data for multiple different purposes. So, we talked about a machine learning use case. On our side, what this can look like is an engineer having the capability to try multiple machine learning models virtually within Data Bricks in real time as the data is
[21:14] coming in from our test labs. This way they don't have to go through the effort of translating a model, trying to put it on an embedded device, and then recording some test results. They can do everything within Data Bricks almost in real time.
[21:29] And the second use case is where the engineer has put 15, 20, 30 tools on test thing for their project, and now they're being spawned with some alerts that the tools are disconnecting, or they see a value out of threshold, um
[21:44] and we can actually generate those alerts within Data Bricks and send to all different personnel to take action. So, with that said, I'll hand it back to Max to show you some more insights on what we are going to do in the future with
[22:00] the solution. Thank you, Karat. And let's give him another round of applause for that demo, for getting it to work. Not the the easiest thing to do in front of a live audience. Um so, getting back to some of the the business purpose here, you know, as
[22:16] Karat was uh explaining, one of our use cases is when we are doing developmental testing, um and especially something like life testing, where we are putting the tool through an accelerated life cycle to make sure that we're going to
[22:32] hit the the demands of the users, and that when we give this tool to the user, it's going to last 3, 5, 10, you know, however many years we're we're guaranteeing that tool to last for. And so, if we take a look at this process of what that would look like without Data Link, uh you can see there's there's a number
[22:48] of steps here, such as, you know, first we're discovering that issue, uh but then, you know, this is likely happening in one of our global global environments. So, we have to first get those tools out on the water and and and start that that transport process. Uh once we get the tools in hand, we then
[23:04] have been having some back and forth with the global team and we're trying to understand what exactly they did to cause that issue and then we're trying to recreate that issue in whichever lab these tools were sent to. Once we're able to to recreate that issue, awesome. Now we're going to do it
[23:20] a bunch of times and start collecting data, try to figure out that root root cause and then once we have an idea of what root cause is, we can start to analyze the data and start to come up with solutions. Well, across that time we just spent five or six weeks of development time
[23:37] from hey, we had an issue to hey, we have a solution. Um, which doesn't sound like a lot of time, but again we're looking at project cycles that happen within a year to 18 months, so that that six weeks is a is a lot of time. Now that we have data link,
[23:53] that process looks a lot different. Again, we discover that issue, but now that we've discovered that issue and we have a timestamp of when that issue occurred, we can automatically start going and querying the data and trying to isolate those events and understand
[24:08] what happened. And now um, instead of spending all that time sending tools back and forth, trying to recreate the failure, we're already jumping ahead and and looking at that data. So ultimately this is reducing the time between identifying a solution or
[24:24] identifying a problem and then deploying a solution. So if you take these numbers and start to scale them, these weeks that we save per issue, you know, you scale that on the number of issues that can occur across the number of builds in each project, across the number of active NPD
[24:39] projects, we're looking at thousands of hours of engineering time that we're saving because we've shifted from reactively fighting to now proactively problem solving.
[24:54] And again, Milwaukee Tool, we're a global company, we have thousands of people working all over the globe. And so, currently we have data link deployed at our engineering and test facilities around our headquarters in Southeastern Wisconsin. And the project teams that have been using it, uh who have been beta testing
[25:11] it for us, they they absolutely love it. They think it's phenomenal and are already planning and putting in requests so that they can put it on most every project going forward. Uh so, now we've started working with our larger testing and operation teams to start to roll this ability to our
[25:27] global enterprise. Uh and you can see where some of those locations are across the globe there. Uh and so, once we do that, the result will be over 100 hubs in all of these locations communicating with those thousands of tools that Kraut was
[25:42] talking about uh by the end of the year, providing our development teams with the the insights that honestly up until now were basically just a a pipe dream. Um and of course, all of that to say that the end result are better products faster in the hands of our users.
[26:01] So, with that I want to thank you for your time.

Learn more about the Databricks Data and AI platform.

The information provided herein is for general informational purposes only and may not reflect the most current product capabilities or configurations.