Everyone benchmarks Redis on a laptop and thinks they're ready for production. They're not. Here's what actually happens when you push Redis past 500k ops/sec and why your single-threaded assumptions will destroy you at scale.
Redis processed 847,000 operations per second for us. Then it stopped. Not crashed. Just... stopped responding. Latency went from 0.3ms to 47 seconds in about 90 seconds flat. This was Black Friday 2022. We had 12 million concurrent sessions and a $2.3M/hour revenue run rate. It took us four hours to recover.
Here's the thing: Redis wasn't the problem. We were.
Redis is single-threaded for command processing. You've heard this. What nobody tells you is what that actually means under load.
One slow command blocks everything. Everything. That KEYS * someone ran on a 4GB keyspace? That's not just bad practice, that's a production kill switch. I've seen this kill production at three separate companies. We literally put KEYS in our Datadog monitors as a banned-command alert after 2022. If it fires, PagerDuty wakes someone up immediately.
At 1M ops/sec, your event loop has roughly 1 microsecond per operation to stay current. You blow that budget on one bad LRANGE with 50,000 elements and you've now queued 200 operations behind it. They queue faster than they drain. That's your latency cliff right there.
Redis 7.x on a c6g.2xlarge can do around 1.1M simple GET/SET ops/sec. Measured. Real. But that number is basically fiction in production.
Your actual throughput ceiling is set by three things, in this order: network bandwidth, connection count, and command complexity. Not CPU. Not memory. Network and connections first.
A single Redis connection maxes out around 100k ops/sec because of round-trip latency. You want 1M ops/sec? You need pipelining, connection pooling, or both. We run 50 persistent connections per app pod with pipeline batches of 100 commands. That's how you actually get there.
Pipelining isn't optional at scale. It's the difference between 100k ops/sec and 1M ops/sec on identical hardware.
We were using Redis Cluster with 6 shards. Looked great on paper. The problem was our key distribution was garbage. Hot keys.
One shard was handling 73% of our traffic because every active session was hitting a leaderboard sorted set keyed by a single global tournament ID. One key. One shard. Completely saturated while five others sat at 8% CPU. Redis Cluster doesn't rebalance hot keys. It can't. The key lives where it lives.
The fix was ugly. We added a read replica per shard specifically for that sorted set, used client-side routing to spread reads across primary and replicas, and added a local in-process cache in front using Caffeine with a 500ms TTL. Latency dropped from 47s back to 0.4ms. The in-process cache absorbed 60% of the read volume before it ever touched Redis.
You need redis-cli --latency-history running in production. Always. Not Datadog's Redis integration alone, the actual latency history tool. It shows you microsecond jitter that CloudWatch will happily average into oblivion.
Set maxmemory-policy explicitly. If you haven't set it, it defaults to noeviction. At 1M ops/sec with a memory leak or traffic spike, noeviction means Redis starts returning errors instead of evicting old keys. I watched this take down a checkout flow for 22 minutes because nobody had read the defaults.
Don't run Redis on the same host as anything else at this scale. I don't care how good your cgroups configuration is. Noisy neighbor on CPU scheduling alone will add 2-3ms of tail latency. That's an eternity.
One million ops/sec is achievable. It just looks nothing like the benchmarks.