AI

NASA's Code That Kept Astronauts Alive for 9 Months

Butch Wilmore and Suni Williams weren't supposed to stay on the ISS for 9 months. A busted Boeing Starliner forced NASA to improvise a rescue with SpaceX. Here's the actual software and ML keeping them alive — and why it's nothing like the code you write.

NASA's Code That Kept Astronauts Alive for 9 Months

Two astronauts boarded a spacecraft for an 8-day mission. They came home 9 months later on a completely different vehicle. That's not a bug report. That's a full systems failure cascade that NASA had to solve in real time.

I've shipped software that broke production. In 2021, a misconfigured Nginx rule on my SaaS took down billing for 6 hours and cost me roughly $4,200 in churned trials. I thought that was bad. These folks were orbiting Earth at 17,500 mph with no ride home.

So I went deep on what actually runs up there. And it's fascinating.

The Starliner Problem Was a Software Problem First

Before Butch and Suni were stranded, Boeing's Starliner had helium leaks and thruster failures. But the part nobody's talking about? The fault detection software kept flagging issues and then... not flagging them. The onboard fault management system is supposed to catch anomalies and hand off to ground control. It did. Sometimes. Inconsistently.

That's the scariest kind of bug. Not the one that always fails. The one that fails maybe.

NASA uses a system called FDIR — Fault Detection, Isolation, and Recovery. It's running on radiation-hardened processors that'd embarrass a 2004 laptop. We're talking IBM RAD750 chips. Single-core. 400MHz. They run this on purpose because cosmic radiation flips bits in regular silicon. Your M3 MacBook would go insane up there.

The ML Actually Running on ISS

Here's where it gets weird. NASA isn't running PyTorch on the ISS. They're not doing anything you'd recognize as modern ML. The "machine learning" up there is closer to what we'd call statistical process control — models trained on years of sensor data that flag when something's drifting outside its normal band.

Think: oxygen partial pressure has been 3.2 sigma above baseline for 11 minutes. That's not a neural net. That's math your stats professor taught you and you forgot.

The real ML lives on the ground. NASA's Johnson Space Center runs predictive maintenance models in Python — actual sklearn pipelines — trained on telemetry from every ISS mission since 2000. When Starliner's thrusters started misbehaving, those ground models were already surfacing anomaly scores before the human flight controllers caught it visually.

The astronauts aren't flying the ISS. Software is flying the ISS. The astronauts are there to fix the software when it breaks.

How SpaceX's Dragon Actually Saved Them

SpaceX Crew Dragon runs on a completely different stack. Triple-redundant flight computers, touchscreen interfaces running on Linux, and guidance software built with Model-Based Design in MATLAB Simulink. Yeah. MATLAB. Don't @ me, it works.

The autonomy stack on Dragon uses sensor fusion to dock automatically with the ISS. It's running an extended Kalman filter — again, not sexy, not a transformer model, just solid 1960s math implemented really well. That's the lesson. Boring reliable beats clever fragile every single time.

NASA had to certify Dragon for a crew return mission it wasn't originally manifested for. That certification process ran parallel risk models comparing Dragon's failure probability against keeping Butch and Suni on ISS indefinitely. Real expected value calculations. Real tradeoffs. This is literally how every infra decision should work and almost never does.

What Indie Hackers Can Actually Steal From This

Stop chasing the fancy model. The ISS doesn't run GPT-4. It runs Kalman filters and fault trees that engineers have refined for 40 years. Your users don't need the newest thing. They need the thing that doesn't go down.

Redundancy isn't glamorous. Triple-redundant systems mean you built the same thing three times and shipped none of the cool features. NASA does it anyway. Your single-instance Postgres probably shouldn't be your only bet either.

And the biggest one: when your system fails, the question isn't "how do we fix this fast." It's "how do we keep people safe while we fix this." Butch and Suni sheltered on ISS for 9 months. Sometimes the right answer is slow down, don't make it worse.

OPEN IN REEDL_ FEED →← Back to feed