System Design

Design internet-scale systems, step by step

No server-room tours — just the decisions interviewers actually probe: capacity, data flow, trade-offs, and the blueprints behind the machines you interact with daily.

23 chapters3 tracks4 free+19 Plus~23 min read

Start here

01
Start here

System Design in 30 Minutes

Core

A repeatable talking skeleton for design rounds: requirements → estimation → components → trade-offs.

walkthroughskeletonestimation
  • Clarify functional and non-functional requirements first.
  • Estimate scale: QPS, storage, cache hit ratio.
  • Sketch the data flow with boxes for client, API, services, storage.
  • Always close with a bottleneck and a trade-off discussion.
1 min · Chapter 1 of 23Read the chapter

High-level design

02
High-level design

Scale from Zero to Millions

PlusMedium

Walk a single-server app up to millions of users: web tier, data tier, cache, CDN, stateless services and sharding.

scalingload balancerdatabase replicationcdnstateless
  • Start with one server, then split the web tier and data tier so each scales independently.
  • A load balancer plus master/slave replication removes the single points of failure.
  • A cache tier, CDN, and a stateless web tier keep latency flat as traffic grows.
  • When the data tier tops out, shard by a well-chosen key — the hash is the hot spot.
1 min · Chapter 2 of 23Read the chapter
03
High-level design

Back-of-the-Envelope Estimation

PlusEasy

Turn vague requirements into QPS, storage and availability numbers using powers of two, latency tables and SLA nines.

estimationqpslatencyavailabilitycapacity
  • Know the units: 2^10 ~ 1 thousand, 2^20 ~ 1 million, 2^30 ~ 1 billion.
  • Anchor on latency: memory ~100 ns, disk seek ~10 ms, network round trip ~150 ms.
  • Nines map to downtime: 99.9% = ~8.8 h/year, 99.99% = ~52.6 min, 99.999% = ~5.3 min.
  • Worked example: Twitter — ~3 500 QPS and ~55 PB of media over 5 years.
1 min · Chapter 3 of 23Read the chapter
04
High-level design

Rate Limiter

Easy

Design a rate limiter used at the API gateway.

limiterapitoken bucket
  • Token bucket vs sliding window — when each fits.
  • In-memory vs Redis-backed counters.
  • Distributed limiter coordination and fallbacks.
1 min · Chapter 4 of 23Read the chapter
05
High-level design

Consistent Hashing

PlusMedium

Distribute keys across nodes with minimal migration when the cluster grows or shrinks.

consistent hashingshardingdistributedring
  • Hash keys and nodes onto one ring; assign each key to the next node clockwise.
  • Adding a node only moves keys in its own arc — about 1/K of the cluster.
  • Powers Memcached, Dynamo, Cassandra and CDN sharding.
  • Use virtual nodes to average out imbalance from few physical nodes.
1 min · Chapter 5 of 23Read the chapter
06
High-level design

Key-Value Store

PlusHard

A distributed Dynamo/Cassandra-style store: partition on a hash ring, replicate 3x, tune consistency with a quorum.

kv storeconsistent hashingquorumvector clocksstable
  • Single-server tricks (compression, append-only log, caching) only take you so far.
  • Partition with consistent hashing, then replicate to N=3 with sloppy quorum + hinted handoff.
  • Vector clocks detect conflicting versions; Merkle trees reconcile replicas with minimal transfer.
  • Writes go commit log -> memtable -> SSTable; reads use a Bloom filter to skip tables.
1 min · Chapter 6 of 23Read the chapter
07
High-level design

Unique ID Generator

PlusMedium

Generate unique, numeric, 64-bit, time-ordered IDs at 10 000+ per second without a central auto-increment.

idssnowflake64-bitdistributedsequences
  • Requirements: unique, numeric, 64-bit, ordered by creation time, 10 000+ IDs/sec.
  • Multi-master (+k steps), UUID (128-bit) and a ticket server each fail a requirement.
  • Snowflake: 1 sign + 41 timestamp + 5 datacenter + 5 machine + 12 sequence bits.
  • 41-bit millisecond timestamp earns ~69 years from a custom epoch; 4096 IDs/ms per machine.
1 min · Chapter 7 of 23Read the chapter
08
High-level design

URL Shortener

Easy

Design a service that turns long URLs into short keys at scale.

shortenerhashingstorage
  • Base62 encoding and collision handling.
  • Read-heavy workload: cache-first reads.
  • Redirects as HTTP 301 vs 302 trade-offs.
1 min · Chapter 8 of 23Read the chapter
10
High-level design

Notification System

PlusMedium

Push, SMS and email at 10M+ sends a day — decoupled per channel, durable, retried and respectful of user settings.

pushsmsemailapnsfcmqueue
  • iOS push goes through APNs, Android through FCM; SMS and email use third-party providers.
  • One message queue per channel: an outage in one provider never blocks the others.
  • Persist every event, dedupe by event ID, and retry failures — at-least-once beats data loss.
  • Check opt-in settings, use templates, rate-limit per user and watch queue depth.
1 min · Chapter 10 of 23Read the chapter
11
High-level design

News Feed

PlusMedium

Design the timeline behind a social product.

feedfypsocial
  • Pull vs push (fan-out) delivery models.
  • Ranking pipeline and pre-computation.
  • Sharding user data and timeline caches.
1 min · Chapter 11 of 23Read the chapter
12
High-level design

Online Chat

PlusMedium

Design a messaging system with presence, delivery and read receipts.

chatwebsocketpresence
  • WebSocket vs long-polling for realtime transport.
  • Message ordering and idempotent delivery.
  • Presence service and last-seen semantics.
1 min · Chapter 12 of 23Read the chapter
13
High-level design

Search Autocomplete

PlusHard

Return the top-5 most searched queries for a prefix in under 100 ms using a cached trie rebuilt offline.

typeaheadtrietop-kqueriessuggestions
  • Match only at the start of a query; return top-5 by historical frequency; no spell check.
  • A trie with cached top-k per node plus a ~50-char prefix cap answers in O(1).
  • Rebuild the trie weekly from aggregated logs — never update it per keystroke.
  • Browser-cache results (~1 h), sample logs, and shard the trie by leading characters.
1 min · Chapter 13 of 23Read the chapter
14
High-level design

Design YouTube

PlusHard

Upload and stream video at YouTube scale — 2 B users, 150 TB/day of uploads and a $150k/day CDN bill.

videostreamingcdntranscodinghls
  • 2020 reality: 2 billion MAU and 5 billion videos watched every day.
  • At 5 M DAU, uploads hit ~150 TB/day and CDN egress ~$150 000/day.
  • Upload: blob storage -> transcoding -> transcoded storage + completion queue -> CDN.
  • Stream from the CDN (MPEG-DASH/HLS); long-tail videos serve from cheaper origin storage.
1 min · Chapter 14 of 23Read the chapter
15
High-level design

Design Google Drive

PlusHard

File storage and sync for 10 M DAU: 4 MB blocks, delta sync, deduplication, and long-poll notifications.

cloud storagefile syncblocksdeduplicationlong polling
  • 50 M signed-up users x 10 GB free = 500 PB of capacity; ~240 upload QPS peaking at 480.
  • Files split into <= 4 MB blocks, hashed, compressed, encrypted, then chunk-synced to S3.
  • Delta sync + block deduplication slash bandwidth; long polling keeps clients in sync cheaply.
  • Strong consistency via ACID metadata DB — caches are invalidated on every write.
1 min · Chapter 15 of 23Read the chapter
16
High-level design

Caching & CDN

PlusEasy

Place caches from browser to database and survive stampedes without serving stale data.

cachingrediscdnstampede
  • Cache the hot read paths; stale-by-a-little is usually acceptable.
  • Layers: browser headers, CDN, Redis/MemoryCache, then the database.
  • Cache-aside reads with invalidate-before-write keep stale entries out.
  • Beware stampedes — use single-flight in .NET to collapse concurrent misses.
1 min · Chapter 16 of 23Read the chapter
17
High-level design

Load Balancer

PlusEasy

Spread traffic over a server pool at every tier — DNS, CDN, L4/L7 — while staying healthy.

load balancerdnsl4l7health check
  • DNS, anycast and CDN route at the edge; L4/L7 split inside the DC.
  • Algorithms: round robin, weighted, least connections, IP/consistent hash.
  • Health checks plus passive failure detection keep traffic off bad nodes.
  • Load balancing forces stateless backends — sessions move to Redis.
1 min · Chapter 17 of 23Read the chapter
18
High-level design

Message Queue & Event-Driven

PlusMedium

Asynchronous messaging to decouple services, absorb traffic spikes and keep flows consistent.

kafkapub-subevent-drivenqueuesaga
  • Queues decouple producer from consumer and buffer traffic spikes.
  • Point-to-point delivers once; pub/sub fans out to every subscriber.
  • Kafka is a partitioned, replayable log; RabbitMQ excels at routing.
  • Sagas plus idempotent consumers keep cross-service flows correct.
1 min · Chapter 18 of 23Read the chapter
19
High-level design

Microservices & API Gateway

PlusMedium

Split a monolith at domain boundaries, front the services with a gateway, and stay resilient.

microservicesapi gatewaycircuit breakerdecomposition
  • Cut at domain boundaries; each service owns its data store.
  • The API gateway centralises auth, routing, rate limiting.
  • Timeout + circuit breaker wrappers prevent failure cascades.
  • Observability with correlation IDs and tracing is mandatory.
1 min · Chapter 19 of 23Read the chapter

Low-level design

20
Low-level design

Parking Lot

Easy

Model vehicles, spots, entrance ticketing and payments.

classesenumsingle responsibility
  • Enum spot types and vehicle compatibility rules.
  • Ticket ↔ spot lifecycle and fare calculation.
  • Where the strategy pattern removes if/else sprawl.
1 min · Chapter 20 of 23Read the chapter
21
Low-level design

Library Management

PlusEasy

Domain model for catalogue, borrow/return and reservations.

bookingstatesdomain
  • BookCopy vs BookTitle: why copies matter.
  • State machine for borrow → overdue → returned.
  • Hold queue management and notifications.
1 min · Chapter 21 of 23Read the chapter
22
Low-level design

Expense Splitter (Splitwise)

PlusMedium

Model groups, expenses, shares and simplifying settlements.

groupsbalancessettlement
  • Expense with per-member share amounts.
  • Compute net balances; simplify transfers greedily.
  • Handle rounding so balances always sum to zero.
1 min · Chapter 22 of 23Read the chapter
23
Low-level design

Vending Machine

PlusMedium

State-driven machine with coins, inventory and change.

statefsmmoney
  • State pattern: idle, selecting, dispensing, out-of-stock.
  • Coin validation, change computation and refunds.
  • Inventory guards against over-dispensing.
1 min · Chapter 23 of 23Read the chapter

Practice the decisions, not the diagrams

Pair these blueprints with our track problems to turn theory into talking points.