7 System Design Concepts Explained in 10 Minutes

ByteByteGoAbout 6 min readJun 26, 2025Watch original
THE SUMMARYAI-generated

Key Concepts

CAP Theorem, Consistency, Availability, Partition Tolerance, Eventual Consistency, Conflict Resolution (Last Write Wins, CRDTs, Merge Functions), Load Balancing (Layer 4, Layer 7, Least Connections, Least Time), Consistent Hashing, Circuit Breakers (Closed, Open, Half-Open), Rate Limiting (Token Bucket, Leaky Bucket, Fixed Window, Sliding Window), Monitoring (Metrics, Logs, Traces, Events), Alerting (Static Thresholds, Anomaly Detection, Composite Alerts), Service Level Objectives (SLOs).

CAP Theorem and Consistency vs. Availability

The CAP Theorem states that a distributed system can only guarantee two out of three properties: Consistency, Availability, and Partition Tolerance. Since network partitions are inevitable, the real choice is between consistency and availability during partition events.

  • Consistency: All nodes see the same data at the same time.
  • Availability: Every request to a non-failing node receives a response.
  • Partition Tolerance: The system continues to operate despite network failures.

Examples:

  • Google Spanner (Consistency): Uses atomic clocks and synchronized time across global data centers to maintain linearizable transactions. During network partitions, the majority partition remains available for reads and writes, while minority partitions become read-only. Trade-off: Minority partition loses write availability.
  • Amazon DynamoDB (Availability): Continues accepting writes during network partitions using eventual consistency and resolves conflicts using "last write wins" based on timestamps. Trade-off: Users might occasionally see stale data.

Key Argument: The choice between consistency and availability depends on the application's requirements. Banking systems typically need consistency, while social media feeds can tolerate eventual consistency.

Eventual Consistency and Conflict Resolution

Eventual consistency promises that if updates stop, all replicas will eventually converge to the same state.

Benefits:

  • Writes can complete immediately without waiting for confirmation from all replicas, improving performance and availability.

Conflict Resolution Strategies:

  • Last Write Wins: Uses timestamps to pick the most recent update. Simple but can lose data.
  • Conflict-Free Replicated Data Types (CRDTs): Uses mathematical properties to guarantee that all replicas converge to the same state regardless of update order.
  • Application-Defined Merge Functions: Allows developers to write custom logic for resolving conflicts based on business rules.

Example: Amazon's shopping cart uses eventual consistency. You can add items even if some servers are temporarily unreachable.

Modern Systems: Systems like DynamoDB often achieve consistency within milliseconds under normal conditions, making eventual consistency barely noticeable.

Load Balancing

Load balancing distributes incoming requests across multiple servers.

Types:

  • Layer 4 Load Balancers: Operate at the transport layer (IP addresses and TCP/UDP ports). Fast but limited in routing intelligence.
  • Layer 7 Load Balancers: Operate at the application layer (HTTP headers, URLs, request content). More powerful but computationally expensive.

Load Balancing Algorithms:

  • Least Connections: Routes new requests to the server handling the fewest active connections.
  • Least Time: Factors in how quickly each server responds, avoiding slower servers.

High Availability: Load balancers are deployed in primary-secondary configurations with heartbeat protocols. If the primary fails, the secondary takes over within milliseconds.

Consistent Hashing: Ensures the same client consistently hits the same server, which is critical for maintaining session state.

Consistent Hashing

Consistent hashing solves the problem of data redistribution when adding or removing nodes in a horizontally scaled system.

Process:

  1. Hash each node to determine its position on a circular hash ring.
  2. Hash the key and walk clockwise around the ring until you hit the first node.
  3. Replicate the data to the next n-1 nodes clockwise on the ring.

Benefits:

  • Adding or removing nodes only requires moving k/n keys (where k is the total number of keys and n is the number of nodes) instead of nearly all keys.

Examples: Amazon DynamoDB and Apache Cassandra use consistent hashing for horizontal scaling.

Circuit Breakers

Circuit breakers prevent cascading failures in distributed systems.

States:

  • Closed: Requests flow normally.
  • Open: Requests are blocked immediately, returning fast failures.
  • Half-Open: A few test requests are allowed through to check if the service has recovered.

Process:

  1. The circuit breaker tracks successful and failed requests to a service.
  2. When the failure rate exceeds a threshold, the circuit breaker trips to the open state.
  3. In the open state, requests fail immediately, preventing resource exhaustion.
  4. After a timeout period, the circuit breaker moves to half-open and allows a few test requests through.
  5. If the test requests succeed, the circuit breaker closes. If they fail, it opens again.

Example: Netflix pioneered this pattern with their Hystrix library.

Key Insight: Fast failures are better than slow failures that consume resources.

Rate Limiting

Rate limiting controls how many requests a client can make within a given time window.

Algorithms:

  • Token Bucket: Accumulates tokens at a fixed rate up to a maximum capacity. Each request consumes a token.
  • Leaky Bucket: Processes requests at a constant rate, smoothing out traffic spikes. Excess requests are queued or dropped.
  • Fixed Window: Counts requests in discrete time intervals. Prone to boundary effects.
  • Sliding Window: Uses a rolling average that smooths out the boundary problems of fixed windows.

Distributed Rate Limiting: Requires coordination across multiple instances.

  • Redis-based Rate Limiters: Use Lua scripts to atomically increment counters and set TTLs, ensuring consistent rate limiting across the service cluster.

Tiered Rate Limiting: Implements different limits for authenticated versus anonymous users and progressively stricter limits as suspicious patterns are detected.

Monitoring and Alerting

Monitoring provides visibility into system behavior and performance. Modern observability focuses on four key signal types:

  • Metrics: Time-series numerical data (CPU usage, request rate, error counts).
  • Logs: Structured records of discrete events with contextual information.
  • Traces: End-to-end request flows showing how a single request moves through the distributed system.
  • Events: Significant occurrences like deployments or configuration changes.

Alerting:

  • Static Thresholds: Work for stable metrics but fail when traffic patterns change.
  • Statistical Anomaly Detection: Learns normal patterns and alerts on deviations.
  • Composite Alerts: Combine multiple signals to reduce noise (e.g., high CPU, increasing error rates, and slow response times).

Goal: Service Level Objectives (SLOs) that measure user experience.

Synthesis/Conclusion

The video outlines key concepts for building reliable and scalable distributed systems. It emphasizes that reliability isn't about preventing all failures, but about handling them gracefully. The choice between consistency and availability, the use of eventual consistency with conflict resolution, load balancing techniques, consistent hashing for horizontal scaling, circuit breakers for preventing cascading failures, rate limiting for protecting against overload, and comprehensive monitoring and alerting are all crucial components. The video concludes by advising to start simple, measure everything, and add complexity only when there is clear evidence that it is needed. Not every system needs all these patterns, and the concepts covered provide the tools to make informed trade-offs.

AI summaries can miss context or contain errors. Check important details against the original video.

Go a little deeper.

Have a question about this video? Load its transcript to open the video chat.