System Design

URL Shorteners Aren't Simple. Here's What Breaks.

Everyone thinks a URL shortener is a toy problem. It's not. Here's what actually happens when you hit 10 billion redirects a day and why the 'obvious' database choice will destroy you at scale.

URL Shorteners Aren't Simple. Here's What Breaks.

Bit.ly processes roughly 10 billion clicks per month. That's not a fun fact. That's a constraint that invalidates every 'just use Postgres' answer you've ever given in a system design interview.

The Write Path Is Actually Easy. Stop Worrying About It.

Most people spend 80% of their design time on URL generation. That's backwards. Writes are rare. Reads are everything.

Here's the math: assume 100M new URLs created per day (generous) vs 10B redirects. That's a 100:1 read-to-write ratio. Your entire architecture should reflect this asymmetry and most people's don't.

For generation, use base-62 encoding over a counter. Not a hash. Hashes give you collision headaches and wasted retry logic. A distributed counter from something like Redis with INCR gives you monotonically increasing IDs you convert to short strings: 0-9, a-z, A-Z, 62 characters, 7 characters gets you 62^7 = 3.5 trillion unique URLs. You're fine for a decade.

Key insight: the short code is just a base-62 representation of an integer. There's no magic.

The Part Nobody Talks About: The Read Path Will Humiliate You

In 2021, a fintech I was consulting for had 8M daily active users hitting their internal redirect service. They'd built it on RDS MySQL, single primary, one read replica. Seemed fine. Then a campaign went viral on Twitter, 40x normal traffic in six minutes, and they spent $340k over three days in emergency RDS scaling, engineer overtime, and a third-party incident response firm because their p99 redirect latency hit 14 seconds. Fourteen seconds. For a redirect.

The fix wasn't a bigger database. It was accepting that hot URLs are hot forever.

You need a two-tier cache. First tier: in-process cache in your application servers using something like Caffeine (Java) or cachetools (Python), LRU eviction, max 10k entries per pod. Zero network hop. Second tier: Redis cluster with consistent hashing across nodes. The top 0.1% of URLs absorb 80%+ of your traffic. Cache them aggressively with a TTL of 24 hours and you've just eliminated 80% of your database load before the request even reaches Redis.

Cache invalidation on URL deletion? You accept eventual consistency. A deleted URL might still redirect for up to 24 hours. That's a product decision, not an engineering failure. Make it explicitly.

The Database Decision Everyone Gets Wrong

People reach for Cassandra because 'it scales' and then discover they need secondary indexes on custom domains and expiry times and suddenly they've built a Cassandra anti-pattern in production.

Use Postgres with a single table: (short_code TEXT PRIMARY KEY, long_url TEXT, created_at TIMESTAMPTZ, expires_at TIMESTAMPTZ, user_id UUID). Shard by the first character of short_code, that's 62 shards max, each handling roughly 56B rows at the trillion-URL scale. Add PlanetScale or Citus if you want managed horizontal sharding. The relational model gives you joins for analytics, foreign keys for user data, and partial indexes for expiry cleanup that Cassandra simply can't match cleanly.

Key insight: choose boring technology for as long as possible. Cassandra's operational complexity costs more than its scaling benefits at anything under 50TB.

Redirect Type Is a Product Decision Disguised as a Technical One

301 vs 302. This matters more than people think. 301 is permanent, browsers cache it, your analytics stop working because the request never hits your servers again. 302 is temporary, every redirect hits your servers, you get click data but you pay the infrastructure cost. Most URL shorteners default to 302 and charge for analytics. That's not accidental. It's the business model.

OPEN IN REEDL_ FEED →← Back to feed