Dev

H5N1 Hit Australia. The Data Systems That Caught It

Australia just confirmed its first H5N1 case. The surveillance infrastructure that detected it runs on satellite telemetry, genomic pipelines, and distributed sensor networks most engineers have never heard of. Here's how it actually works, and where it almost failed.

H5N1 Hit Australia. The Data Systems That Caught It

The bird is dead before anyone knows it's sick. That's the problem.

H5N1 doesn't announce itself. By the time a farmer notices dead poultry, the virus has already moved. Australia confirmed its first case this year, and the response looked clean from the outside. It wasn't. The detection system held, but barely, and the reasons it almost didn't are exactly the kind of distributed systems failures I've been watching kill production for 12 years.

The surveillance stack nobody talks about

Australia runs a wildlife disease surveillance network called AUSVETPLAN layered on top of the National Notifiable Disease System. Sounds boring. It's not. Under the hood you've got passive sensor nodes at wetland sites transmitting GPS-tagged mortality event data via Iridium satellite constellation. Not cell towers. Satellite. Because the birds dying in the Macquarie Marshes aren't doing it near a Telstra tower.

The Iridium Short Burst Data protocol gives you 340 bytes per transmission. That's it. So the edge devices run compressed binary schemas, not JSON. Every engineer who has ever watched a team decide to send JSON over a constrained telemetry channel because it's easier to debug has cost their organization real money. We had an outage in 2021 at a 60-person climate tech company I was advising. Their field sensors were transmitting 4KB JSON payloads over SBD. Backpressure killed the ingestion pipeline in under 6 hours during a mass mortality event in the Sacramento Delta. We lost 11 hours of sensor data. The investigation cost $340k when you factor in regulatory response and resampling flights.

Don't do this.

Genomic sequencing as a real-time system

When a sample tests presumptive positive for influenza A, it goes to the CSIRO Australian Centre for Disease Preparedness in Geelong. They're running Oxford Nanopore MinION devices for rapid field sequencing alongside Illumina NovaSeq for confirmation. The MinION gives you a clade identification in under 6 hours. The NovaSeq gives you the full genome in 24. That gap matters.

Here's the thing: the real-time genomic data pipeline feeds into a GISAID submission queue and simultaneously into Australia's national biosecurity dashboard. That dashboard is pulling from at least four upstream systems with different SLAs. I've seen this architecture before. It looks fine until one upstream source goes stale and nobody notices because the dashboard keeps rendering with cached data. Silent staleness. It's the worst kind of failure because the system looks healthy.

Satellite imagery and the movement problem

Post-detection, the question becomes spread. Where did infected birds go? Australia uses Planet Labs Dove satellite imagery at 3-5 meter resolution combined with eBird citizen science data ingested through an AWS pipeline. The correlation model runs in SageMaker. It's trying to predict waterfowl movement corridors in near-realtime.

The model's biggest enemy isn't bad data. It's latency mismatch. Citizen eBird observations have a 0-72 hour submission lag. Satellite passes happen on a fixed cadence. You're joining two streams with wildly different temporal guarantees and calling the output realtime. It's not. It's a best-effort approximation that looks authoritative on a map.

I've seen this kill production dashboards at two different companies. The map looks confident. The underlying join is lying to you.

What actually held

The detection held because Australia invested in redundant confirmatory pathways. Multiple labs. Multiple sequencing technologies. Cross-validated against FAO reference databases before the public announcement. That's not glamorous. That's just not cutting corners on the confirmation pipeline, which half the systems I've reviewed absolutely do.

The glamorous part is the satellite tech. The unglamorous part is why it worked. Redundancy. Schema discipline. Not shipping JSON over 340-byte satellite links.

OPEN IN REEDL_ FEED →← Back to feed