Skip to main content

Databricks for Enterprise AI: Autonomous Satellite Operations at Scale

Summary

  • Blue Origin built a complete data and AI infrastructure on the Databricks Data and AI platform, processing over 6.5 petabytes monthly across 4,300 users to support critical decisions about vehicle certification, lunar landing, and human spaceflight.
  • The Mint platform automatically generates zero-human-intervention data pipelines, while BRAIN is a multi-agent system that manages satellite anomalies using Delta Lake, Unity Catalog, and the Model Context Protocol.
  • A data flywheel with medallion architecture enables continuous model improvement through operational feedback, and the lakehouse approach was chosen over traditional time-series databases for autonomous satellite operations.

Databricks for Enterprise AI: Autonomous Satellite Operations at Scale

Watch: Databricks for Enterprise AI: Autonomous Satellite Operations at Scale
Blue Origin processes over 6.5 petabytes of data monthly across 4,300 users, managing critical decisions about vehicle certification, lunar lander deployment, and human spaceflight. Scaling this data ecosystem to support thousands of autonomous pipelines and AI agents requires rethinking data architecture, metadata governance, and how data is prepared for intelligent systems.
Discover how Blue Origin built a complete data and AI infrastructure on Databricks, from the Mint platform that automatically generates zero-human-intervention pipelines to BRAIN, a multi-agent system managing satellite anomalies using Delta Lake, Unity Catalog, and the Model Context Protocol. Learn how a data flywheel with medallion architecture enables continuous model improvement through operational feedback, and why a lakehouse beats traditional time-series databases for autonomous satellite operations.
🤝

Chapters

FAQs

What is Blue Origin's BRAIN system and what does it do?

BRAIN is a multi-agent system built on the Databricks Data and AI platform that manages satellite anomalies, performing root cause analysis using Delta Lake, Unity Catalog, and the Model Context Protocol. It enables mission operators to respond to anomalies without manually querying multiple separate systems.

Why did Blue Origin choose Databricks over traditional time-series databases for satellite operations?

Blue Origin found that the lakehouse architecture on Databricks better served their need to combine structured system data with multimodal data including documents, drawings, and telemetry streams. The ability to converge analytics, AI, machine learning, enterprise search, knowledge graphs, and RAG in one platform was a decisive factor.

What is Blue Origin's Mint platform?

Mint is a platform Blue Origin built on Databricks that automatically generates zero-human-intervention data pipelines, scaling data ingestion and transformation without manual pipeline authoring. It is a core part of their data flywheel strategy that feeds operational feedback back into AI models for continuous improvement.

How does Blue Origin use the Model Context Protocol (MCP)?

Blue Origin uses the Model Context Protocol to expose lakehouse data to AI agents in a governed, secure way, allowing systems like BRAIN to query relevant data without requiring direct database access. This approach maintains the security and governance boundaries required for critical aerospace operations.

Full transcript

[00:09] Hello everyone. How are YOU DOING TODAY? THIS IS GREAT. I know that's a strange thing to say standing here at a conference in front of a slide with our rocket on it, but
[00:25] hear me out. My name is Clark Stevens. I'm the director of data and AI for Blue Origin. And today I want to tell you why I have one of the coolest jobs in the world. Think about what we actually do. We take millions of data points
[00:42] from design, manufacturing, test, operations, and we use those data to make decisions about whether a vehicle carrying a national security payload makes it to orbit, about whether a lunar lander lands on
[00:59] the moon, about whether a human comes home. The rocket is the output. Data is the process. We are a data company that builds rockets.
[01:15] And today I'm going to show you what that means in in practice. Because here's the thing. Rockets are built from atoms, but they're designed, manufactured, and flown with bits.
[01:35] Let me give you a sense of scale. At Blue Origin, data is everywhere. Every month on our data platform, we have more than 4,300 distinct users submitting over 20 million queries,
[01:51] scanning over 6 and 1/2 petabytes of data every month. That's not a data warehouse anymore. That is a central nervous system. Engineers, manufacturing technicians,
[02:06] supply chain planners, mission operators. They're all reaching into that same data fabric to make decisions about whether a rocket flies.
[02:23] That fabric is built on Databricks. And I want to be specific about what Databricks does for us. Because it's more than a lakehouse. Databricks is the connective tissue between our structured system data and our multimodal data. The documents, drawings, telemetry
[02:41] streams, time series that describe how our rockets are built and how they fly. It's where analytics, AI, machine learning, enterprise search, knowledge graphs, rag all converge.
[03:00] When I talk about any of the use cases today, assume Databricks is underneath it. Let me walk you through how we actually use it. When our engineers author a new design, they're not starting from scratch.
[03:16] They're building on top of decades of part history, test results, and supplier data. All landed in Databricks from hundreds of systems all across the company. We've connected engineering drawings to machine readable metadata.
[03:33] So designs flow cleanly to manufacturing. Engineers can query the full life history of a part. Every heat treat, every design change, every inspection, every non-conformance
[03:49] without opening five different applications cuz we've brought them together all in one place. Along with the unstructured documents that describe them. On the factory floor, we use Databricks to power work order analytics we never could before.
[04:06] We analyze historical execution data across hundreds of thousands of work orders to surface time estimates, identify unclear instructions that correlate with downstream non-conformances, and generate quality scores that help manufacturing engineers
[04:22] improve work packages before they ever hit the floor. We've also built a global priority number on top of Databricks that provides a unified rank across every work order, manufacturing order, purchase order. So, our MRP can finally give a single
[04:39] coherent answer to the question, "What part should I build next?" This only works because Databricks lets us join data across supply chain, manufacturing, engineering, all in one place. For New Glenn certification on the Space Force's National Security Space Launch
[04:54] Program, we built an integrated development environment that allows our external US government partners to log in and have secure role-based access control to the specific data they need for vehicle certification and mission assurance activities.
[05:12] And for mission operations, we built a mission operations AI co-pilot for New Glenn that helps mission controllers find procedures, past anomaly reports, and flight data.
[05:28] Pulling structured telemetry from Databricks alongside the procedure documents and anomaly reports that describe them. That multimodal join is the whole game. Then, GenAI burst onto the scene.
[05:44] As As may guess, our founder is a major AI optimist. Using GenAI Blue as a leadership mandate and a requirement for us to accomplish our goals in the necessary time frames.
[05:59] But all of this specialized knowledge I've been talking about doesn't exist in a large language model. These models are genuinely extraordinary. I mean that. Their breadth of knowledge and reasoning. I use them every day. And then you ask one to tell you about
[06:16] the pedigree of an engine rolling down the manufacturing line. It'll just invent a supply chain. It'll do so with confidence. With citations.
[06:32] It's like hiring the smartest person you've ever met and then realizing that person knows nothing about your industry, your company, or your products. Great at trivia night. Not so much on the factory floor or the mission control room. So the question became
[06:49] how do we plug Blue Origin's data into AI? Before I get to the plumbing, I have to talk about the data itself. Because if your data isn't accurate, annotated, connected, and trustworthy,
[07:05] your agents will fail. And they'll fail quietly and at scale. That's why we invested heavily in making our data AI ready. We built a metadata management service on top of Data Bricks
[07:23] that ingests, stores, and synchronizes comprehensive metadata. Table and column descriptions, system ownership, domain classification, lineage across our entire data corpus. We developed an AI data readiness
[07:39] algorithm that scores every data object on four dimensions. Does it have clear ownership? Is it unique in the catalog? Is it well described? And does it have descriptions?
[07:55] Tables scoring above 70% are flagged high AI readiness and surface to agents first. Tables below 40% trigger agentic workloads that help system owners improve their metadata. With GenAI doing the drudge work and system owners validating.
[08:11] The results. Agents don't waste tokens searching across thousands of undocumented tables. They go straight to the data that's been vetted. The answers are faster, cheaper, and more trustworthy.
[08:31] But accurate and annotated data is not enough. It has to be fresh. Think about what a manufacturing planner needs when a part fails inspection 2 hours before it's supposed to roll down the factory line. Or an AI agent needs when a manufacturing leader asks you to analyze
[08:47] work orders open right now. If your data is 3 hours old, your answers are 3 hours old. And at a space company, that's the difference between a confident answer and a hallucinated guess.
[09:03] So, we looked at our legacy ingestion platform, and we didn't like what we saw. We had thousands of legacy pipelines. Each one hand built by a central team of data engineers. New data onboarding could take weeks as
[09:20] they cleared up behind a large backlog and waited for the next sprint to start. Data structure changes could take days and required a lot of coordination with the business. When late detected schema changes broke pipelines,
[09:35] they interrupted business operations and pulled engineers off their real jobs to firefight. Large data sets could only ingest on schedules of 3 to 24 hours and only a very small group of people could publish data. That means every new data need had to
[09:51] queue up behind a centralized team. So, we built Mint. Metadata driven ingestion. Mint is our homegrown, event-driven, near real-time auto-generation platform.
[10:07] And I want to describe how it actually works because it's one of the most elegant things that we built and I personally am very proud of it. When a system owner wants to publish their data, they don't file a ticket with the central team. They don't wait for a data engineer or
[10:23] an expert to start. They onboard their system of record to our eventing platform, register an event schema, and start publishing. Mint is listening.
[10:39] The moment a new event schema uh starts in our registry, Mint automatically creates the subscription, spins up the processor, lands the data in S3 as parquet, and makes it available in Databricks. No human in the loop. No central team required.
[10:56] The pipeline creates itself. And because it's event-driven and incremental, we're driving average source to lakehouse latency from 98 and 1/2 minutes to 10 minutes or less.
[11:12] Mint does three things that matter for AI. First, it gives our agents access to fresh data. So, when an AI agent answers the question, "What work orders are blocked right now?" the answer is minutes old, not hours.
[11:27] Second, it's a single ingestion, multi-target framework. Data and metadata are loaded once and then made available simultaneously to Databricks, to our knowledge graphs, to our rag vector stores, and our
[11:42] metadata catalogs. One pipeline, every consumer. Third, it democratizes publishing. Any system owner at Blue can make their data available to the entire company and
[11:57] every AI agent without going through a central data team. Near real-time data for AI is not a project at Blue anymore. It's just how Blue Origin works.
[12:17] The Model Context Protocol, MCP, is how we expose our data platform to AI in a standardized, secure way. At Blue Origin, the Databricks MCP server is by far the most used MCP at the company. Let me give you some numbers. Every month, more than 3,000 distinct
[12:33] users generate more than a million invocations. And the Databricks MCP has 5.5 times the usage of the number two MCP at the company.
[12:48] Year-to-date, monthly users is growing 26% and invoch- invocations are growing 180% month over month. That divergence is important. User growth is strong. 26% is excellent. But invocation growth running seven
[13:03] times faster than user growth means something specific. Users aren't just trying it once. They're coming back and they're coming back consistently. Workflows are getting deeper. Agents that were built six months ago to
[13:18] answer one question are now answering 10. The Databricks MCP isn't a novelty at Blue. It's core data and AI infrastructure.
[13:34] So, that's the foundation. Fresh data. AI-ready metadata. A multimodal knowledge graph. And an MCP layer that agents can actually build on. When a manufacturing leader at Blue Origin asks our agents, "Do I have enough labor to complete the
[13:50] work orders in my work center that I need to this week?" They're not getting a hallucinated guess. That agent is going to the Databricks MCP, querying high-readiness data objects with validated metadata, and returning a result in seconds.
[14:06] That is enterprise AI. That is agentic AI, grounded in real enterprise data. We're moving from being data-rich to intelligence-driven. And we're doing it because the mission demands it.
[14:21] Increased launch cadence. Lunar landers. Satellite constellations. None of that happens if our employees are spending significant amounts of their week hunting for siloed information all over the place.
[14:40] That's why, frankly, I think I have one of the coolest jobs in the world. We're building the data and AI engine that builds the road to space. And that road is getting faster, cheaper, and more reliable every single day. Now that I've laid out the data ecosystem for you,
[14:55] the natural question becomes, "What do you build on top of it?" To answer that question, I'm going to hand it over to my colleague, Theo. Theo is going to walk you through a multi-agent system built on the foundation to manage satellite
[15:10] telemetry. Theo, over to you. Thank you, Clark. Hi, everyone. My name is Theo Talvacchia. I'm the senior for ground software uh for the Blue Ring program under the Blue National Security.
[15:27] If you're not familiar with Blue Ring, it is a highly maneuverable, multi-destination, multi-mission spacecraft. And it's one of our key components of providing spacecraft to national security space. The thing that I want to first start off with is that there are two fun facts
[15:44] about Well, I'm here essentially to introduce you to BRAIN. BRAIN is the Blue Ring Blue Resilient Autonomous Network Environment. It is purely a The goal of Blue Ring is very much like Clark said, is a central nervous system that's meant to leverage
[16:00] AI for the ground system for satellite operations. Two fun facts about BRAIN. First and foremost, it's a If you're familiar with quantum mechanics, it's an it's a dynamical object that moves through space-time. BRAINs are. Second fun fact. Last year in I used the exact
[16:18] spelling in a AWS re:Invent post. Um and what's funny about it is I got a lot of DMs about that I should use AI to check spelling. So, while I did not use AI to check spelling, we are using AI to fly
[16:33] satellites. Let's pause there for a little bit and kind of give the state of the industry in terms of what satellite operations is today. First, satellite operations has a scale problem. I would say probably a very massive one. Say you have four vehicles, you have 10
[16:51] contacts a day where you're trying to troubleshoot and make sure the vehicle is doing the activities that you planned as well as its health is in the right configuration. Say you have six subsystems. All of a sudden, you've essentially gotten to 240 decision windows daily where something
[17:07] could potentially go wrong. More importantly than that, somebody has to be watching. Usually that's sometimes a system, but most likely it has to be an operator. So, that brings the question of cost in terms of operator. And the satellite industry are very you might be familiar that you cannot scale your operators
[17:23] linearly with your vehicles. So that's just a fact. You can no no management is going to approve that head count. So the re-economic reality is solved for the industry by robotic automation. However, robotic automation brings its
[17:38] own problems. First and foremost, that introduces additional tools, additional software, COTS products, scripting. And what that creates, it creates a larger cognitive load on the operators. So while you may think you're switching context between multiple tools and
[17:54] software, you're forcing operators to essentially learn more tools, learn more context, and constantly be switching between various items to do their work. What does that cognitive load look like? Let's talk about it. So an operator workflow, let's say you have an anomaly
[18:11] of sorts, and all of a sudden you have to operator has to pull telemetry from system A, cross-reference with potential shift notes from the prior prior shift, look at commands that were sent to the vehicle, and tons of various context switching
[18:27] between all of these. The problem is is that that point all of this cross-correlation occurs in somebody's head, which is the operator, has to be relayed up, and more often than not the output is also the problem. It's a PDF of a root cause analysis of an anomaly
[18:43] that no one reads 6 months later. So how do we plan to solve that? And the solution to that is Brain. We designed Brain to be a multi-tiered, multi-agent system that ingests everything from the ground system.
[18:59] Every mission plan, every command sent, every anomaly triaged, every console decision, every operator action, everything. The differentiator though, last year when we started this project was agents were novel. They're no longer
[19:15] novel. Agents are commodity. Nobody is stopping from doing this anywhere. And what we realized as we started production is that the differentiator is the memory stack. We worked with experts at both AWS and Databricks to design the memory
[19:32] architecture. Multi-tiered memory architecture is goes beyond just session or episodic memory. Okay. We don't want just that. We don't want an agent to just simply remember what an operator said. We want entity and event tracking and traceability across the entire
[19:48] satellite's lifetime. Procedural memory that simply captures all heuristics, okay, from thousands of operator decisions. And memory is great, but if you're not closing the feedback loop, you're wasting your time. So that's as
[20:03] important as memory. The feedback loop is that the resolution of any anomaly, any action, eventually makes its way into the institutional knowledge of the entire system. So that means that if you have an anomaly, that resolution is available for
[20:19] knowledge for every agent, every vehicle, every future shift. The system has to get smarter every one somebody touches it. If you don't, you're failing. The data backbone for all of that is Databricks.
[20:35] And the most important part is that the learning compounds from that data. How do we see this going into the future, right? So our goal is to use the same architecture, but a broader scope. Not just maneuver plans. We're talking
[20:52] potentially contact margins, fleet optimization, flight dynamics, everything. And the reason why is because we want to get to an end state where nominal fleet operations for vehicles, whether that's Blue Ring, Terra Wave, run autonomously
[21:09] with minimal human intervention only in novel scenarios. And the reason why we want that is because we want the compounding effect. We essentially want every investigation, every operator action, every validation to make everything smarter. At Blue, and
[21:26] specifically with Brain, we're focused on the data flywheel. We want iteration n plus one to be smarter than iteration prior n. And the most important part of all of this, this has to scale with your constellation.
[21:43] Blue Rings are vehicles that maybe won't be as large of a constellation, say, as TerraWave, but it needs to be able to handle the scalability. Flying four vehicles should be the same load on your system as flying 40 vehicles. And more importantly, the
[21:59] cognitive load on the operators needs to stay flat. What we want now is to go through a demo of Brain. And I want to talk to you a little bit about this specific use case that we're showing, which is a root cause analysis.
[22:17] At the core of this is a root cause analysis orchestrator engine. And we've broken this down into triage, investigation, and learning. The orchestrator orchestrator agent receives the anomaly trigger. Doesn't try to solve the problem there. It reasons first.
[22:34] Which specialists to call in what order? First, in this case, it calls the telemetry analyst. The telemetry analyst has very specific goals. One is to query the Delta Lake telemetry table through a Databricks SQL warehouse.
[22:49] Billions of rows of decommutated telemetry have to be be able to be pulled, and they're all partitioned by vehicle and time. Anomaly classifier. This calls a trained model on Databricks model serving.
[23:05] This is registered in your Unity catalog using MLflow. And it returns a probability score as long as it as well as a feature attribution. Knowledge specialist. This does a semantic retrieval over any runbook data. Looks at prior incidents, assesses
[23:22] how common they are relative to the current one, anomaly reports, hardware and software documentation, ICDs, ops procedures, all of it. And finally, your resolution advisor. No calls, just absolute reasoning.
[23:39] This takes all of the evens evidence that was provided by it all the others and recommends recovery actions. At the end, the system writes the completed investigation back to Delta Lake. This is the closing of that feedback loop, right? The Delta Lake as
[23:55] an episode, but that episode becomes training for the next investigation. So, with that said, let's switch over to the demo. Okay. All right.
[24:11] So, this is Brain, and this is the home page of what we deem to for the operators to utilize. When we go and when operators log in for Blue Ring operations, this is one of the first pages they will see. There are additional tools and apps that will be
[24:27] built into it, but for the focus of this demo, we'll showing this part. So, the fleet overview, if you're familiar with satellites, essentially is to show all of your vehicles in the fleet. And not only just show them, but also provide immediate status into what
[24:43] each subsystem is doing for each vehicle, as well as data for uh ground contacts, so any sort of performance issues that your antennas might be having, as well as the next plan for each of the vehicles, for example, and subsequently all of your subsequent
[24:59] contacts, their expected purpose, and their tie to all of their maneuver plans. So, Blue Ring Flight 1 has entered essentially a state of safe mode. What
[25:15] is safe mode? For satellites, most satellites essentially have a vehicle mode. We have assumed in this scenario that the vehicle has gone was out of contact, reentered contact, we downloaded telemetry, and the state of the vehicle
[25:31] is shown as safe mode. Immediately, Brain pops up an investigation at the very top. BR1 has entered safe mode. ADCS limit has exceeded.
[25:47] So, we kick off the investigation. Okay. This is going to run in real time, and I'm going to not only walk you through the presentation through the demo, but also show you what is happening in the background within Databricks. All of the items that you see on here are directly tied to Databricks
[26:02] functionality, and Brain is tying all of that information together in a final root cause analysis. First and foremost, I will show you what is on the left-hand side. So, these are past investigations. This isn't a mock demo. This is actually
[26:19] data. It is using mock data, but it is a data that is sitting in Databricks with actual calls and actual agents that are in AWS Agent Core production. So, you can see all the times I've run this demo. Every single time it generates a
[26:37] RCA episode. Let's talk about what it's doing. So, you saw in the you've seen the architecture, you've seen kind of the flow. First, it calls the telemetry analyst. The telemetry analyst immediately goes to the warehouse, pulls SQL queries for all of the items that were deemed related to
[26:54] the ADCS momentum limit. And as you can see, reaction wheel speeds as well as the actual cause the actual potential cause of this, which is the total momentum of the vehicle. This is a very common thing in satellites, by the way. Essentially, your reaction wheels
[27:10] cannot dump all of the momentum, and eventually they saturate. So, they have to reset and recompute the um recompute themselves for the 6° of freedom. So, in this case, the total momentum exceeded set limit of 15 Newton meter seconds.
[27:26] Then, the anomaly classifier kicks in. Talked about the anomaly classifier. It looks at the uh brain anomaly detection, which is a model served within Databricks. Now, for the everyone essentially can develop their own anomaly models. In our
[27:42] case, use the very simplistic isolation forest. In real-life operations, you're more likely to use something like an autoencoder uh for better classifications of anomalies. So, while this is running, I'm going to just walk through the Databricks side.
[27:58] One of the first items I mean know if you need to make this bigger. Okay. One of the first items that we use within Databricks and plan on using for operations. We're not planning to just use brain for operations, but also for subsystem engineers. Subsystem engineers if is a very key component, and the tools
[28:14] you provide them essentially means that operations can do things faster, smarter, and respond to uh anomalies quicker. So, the issue is that in this case, alerts. Oh, command control command and control systems have alerts. These alerts
[28:31] typically are linked to limits. What we are trying to do here is pull those limits into Databricks. So, all of the data is not only in the for the operator of the command control system, but also if you see a bigger picture of the limits that are getting uh defined. In this case, we've defined a limit of
[28:47] 14.75 m seconds, and as you can see, this was crossed at a specific time. Utilizing the alert system, we can also define we can also provide direct emails to subsystem leads based on subsystems that
[29:04] they are working for with their specific alerts. So, they immediately know potential potential issues. Let's talk about our data. Okay. We worked with Databricks to define a specific schema that works for
[29:19] satellites within within the warehouse. And the goal of this is that satellite telemetry comes in raw. Our ground system decommutates that data, and it has to be parsed and stored in a specific schema. Most common, if you've worked with satellites, you'll recognize
[29:36] a lot of this information. Spacecraft ID, contact ID, and so on. We partition this data on two things. Spacecraft spacecraft ID and and time. The key thing here is one of we have the data, and you can see here sample data.
[29:53] But the other item that we utilize and we really like in Databricks is the lineage. So, lineage essentially allows us to immediately see the traces of where all of this data is utilized. And the reason for that is for my main point, traceability, right? We want operators
[30:09] to have visibility into everywhere where this data is being utilized. Let's talk about queries. So, the SQL editor will pretty provide it very much will be a useful tool to subsystem engineers.
[30:25] Very often in satellite operations, and I find myself repeating, is you've had to deal with CSV files. Yes, Excel files, cuz that's how telemetry is passed. Somebody somewhere will sit in front of a computer, ask to be download the telemetry, pull it, send it off.
[30:41] And that becomes very consuming. When you run anomalies, you then have to cross-correlate to Excel spreadsheets. Not fun. So, in here, we plan on providing this tool to subsystem engineers to essentially
[30:57] do their own queries and provide access to the data. Subsequently, you can actually use things like jobs and tasks at a nightly basis to deliver the telemetry directly to the subsystem engineer, so they stop bothering operations. Telemetry dashboard. Once we switch to
[31:13] the agents that are bought by the way, still running, I think. Uh yeah. Past the anomaly, but I'll go back there in a second. Telemetry dashboard. One of the other tools that we plan on providing to subsystems is the ability to create their own dashboards. We want them to have access to the data and
[31:30] it's near real-time. The goal of this is that they can troubleshoot their own issues, see seasonalities, trends that are occurring within their subsystems. But also in operations, provide the capability to do things and view things like torque health, wheel speeds,
[31:45] and the total momentum. If you notice, this is essentially the same exact plots that Brain generated over the same period. I do have two additional things on here. One, the utilization percentage. This is how you immediately see that spacecraft one has exceeded the max momentum.
[32:01] Utilization period utilization percentage at 101. Also, you'll see the RCA episodes. All of these are on here. These are all the recorded, everything from episode ID to trigger event to root cause, and finally, your confidence score at the
[32:17] right. Let's continue. For anomaly detector, this essentially runs all of the capabilities for the isolation forest. This gets kicked off. We're on version seven currently in this case. And when we don't do the anomaly detection, the
[32:33] goal there is to run this at an expected period of time backwards. Since I think that finished, let's go see how it runs. So, as you can see here, the anomaly classifier has defined a moderate degradation in the period before the actual momentum
[32:50] started to increase, right? So, what you're seeing here is that we were flagging this potential separate this potential saturation of the reaction wheels prior to momentum increasing. This is key aspect of
[33:05] satellite operations. If somebody is watching this, as soon as the degradation finishes, they would essentially be able to be alerted. This is also would be running at a much faster pace. I think this is running like 6 hours. So, ideally anomalies you're going to run every hour look back. Knowledge
[33:22] specialist. Knowledge specialist is grabbing all of the data and from the ops procedures, ICDs, documents, everything that is tied to potential ADCS issues. We also generate
[33:37] and the stick is a little bit slow. We also generate a root cause analysis fishbone diagram. This is taking all of the prior events that all of the other agents have resolved and putting them, stacking them in order relative to time leading up to the root to the actual issue, which is we were out of contact
[33:53] for a period of time, we got in contact and immediately entered safe mode. And finally, when this is all said and done, it goes and generates the reasoning agent. Pulls everything from the total system
[34:08] momentum and generates a full RCA root cause analysis report. This is very detailed and at the end, funny enough, you do get a PDF file that you can download. But, because this is stored in Databricks, that file will not will not be lost and it can always be referenced in the
[34:24] future. Here you can see all the anomalies that all of the prior versions that we've run. And then finally, if you did want to use Spark, we'll have this capability.
[34:39] Uh Clark may not enjoy the larger Spark spend budget, but if you do have this capability, we will want to let subsystem engineers run their own analysis within Databricks. Typically, those are sitting on somebody's laptop, somewhere hidden in a
[34:55] cube, and we want all that analysis to be able to be pulled here and shared with the rest of the team so that everyone is aware, for say like GNC or thermal, all the analysis that is done that lets us know we're in a good state.
[35:13] Let's pop over to the demo. Finally, the the last part of the demo is that specialist contribution. We now, as you can see at the bottom, we are capturing that investigation as a structured episode for future learning. This episode is recorded, it's stored in Databricks, it can be referenced later
[35:29] by the knowledge base and knowledge specialist.
[35:46] So, that's Brain. There's a lot more functionality in Brain that we are planning on integrating. The first thing that I want to clarify is this is a production environment. This is not, outside of the mock data, we're using Databricks. We're designing
[36:01] everything with subsystem engineers, and subsequently, our actual agents, our real agents, deployed in the production environment in AWS Agent Core. You're probably wondering at this point, why Databricks? For satellites, why not
[36:16] a time series DB? Most common in the industry, you'll hear various names, which I won't say, for time series as a solution to this. Do we have a time series? Yes, we do in our ground system. However, that time series is for our local short-term usage. If
[36:33] you want fast queries, sure, operators can query items very fast within that. What Databricks is in this case is our long-term storage, where all of the analysis, all of the seasonality trends, everything that feeds potentially back into AI and T to per- potentially
[36:50] provide solutions comes in. So, why not a time series? We evaluated them. They're great at one thing, fast queries over timestamp data. But satellite operations isn't just a time series problem, okay? We need telemetry.
[37:07] We need runbooks. We need anomaly episodes. We need anomaly models, episode history, fine-tuning data sets, provide capabilities for operators, provide capabilities for subsystem engineers, and SMEs running pipelines. All governed, and this is the key point,
[37:22] all governed, all image, all correlatable from the same platform. We don't want people running things off of their notebooks. We don't want people run- running things off on the side. Okay. And here's the more important long-term play.
[37:38] Medallion architecture. All of the data that you saw is getting fed into bronze, our bronze layer. The problem is is that's great for now, but as we expand brain, as we start getting more data, we're looking at fine-tuning. We look to layer- layering those to bronze, silver,
[37:55] and then subsequently gold. The goal there is that bronze is your decommutated, silver is your cleared and calibrated, and gold is where the learning happens. You're going to create a gold that is curated for training, labeled anomalies, and fine-tuning.
[38:11] Every time brain investigates an anomaly and an operator validates that conclusion, that episode flows back into gold for your model retraining. The anomaly detector you just saw version seven okay, was trained yes on prior versions one through six.
[38:27] The goal at Blue and specifically brain is to keep the data flywheel turning. More investigations, more labels, better labels, better models leads to better investigations the next time around. And because it's all Delta Lake, every
[38:43] version of that table is time traveled and auditable. You cannot build that loop on a time series database. You need a lakehouse. And that for us is Databricks. Thank you.

Learn more about the Databricks Data and AI platform.

The information provided herein is for general informational purposes only and may not reflect the most current product capabilities or configurations.