Consistency Levels
⚖️ Master tunable consistency! Learn ONE, QUORUM, ALL and balance consistency vs availability vs latency!
📖 Instagram: The QUORUM Success Story
Instagram stores 2+ billion photos and 500+ million daily active users' data in Cassandra. Challenge: How do you ensure photo uploads appear immediately for the user, while maintaining data durability? Instagram's solution: QUORUM consistency for both reads and writes. Configuration: RF=3 across 3 datacenters. Write path: When user uploads photo, Cassandra writes to 3 replicas but waits for only 2 acknowledgments (QUORUM = 2 of 3). Latency: ~15ms write confirmation (vs ~5ms for ONE, ~45ms for ALL). Read path: When viewing photo, read from 2 replicas (QUORUM), use newest timestamp. Result: Strong consistency (users always see latest photo), High availability (survives 1 node failure), Acceptable latency (15ms feels instant). Trade-off: Slightly slower than ONE, but guarantees data safety. 10 years of operation: Zero major data consistency incidents. Instagram engineers: "QUORUM is the sweet spot - strong enough for consistency, fast enough for user experience!"
💻 Interactive Consistency Level Console
Practice different consistency levels! See how they work in real-time with latency, availability, and consistency trade-offs.
🎯 Quick Scenarios - Click to Explore:
🎯 Click scenarios above to see how each level works
✓ Real-time latency and behavior simulation
⚠️ Critical: Default Consistency Level is ONE!
Cassandra's default consistency level is ONE - which means fastest performance but eventual consistency only! This is fine for caches or analytics, but for critical production data (user accounts, financial transactions, orders), you MUST explicitly set QUORUM or higher. Many beginners don't realize this and wonder why they occasionally see stale data!
⚖️ What are Consistency Levels? (Beginner's Guide)
Imagine you have 3 copies of your address book: one at home, one at work, and one in your car. When you update your friend's phone number, how many copies need to be updated before you're "done"?
📝 Real-World Analogy
ONE: "I'll update just the copy at home, that's good enough!" ← Fast, but risky - if you look at the work copy later, it might have the old number.
QUORUM (Majority): "I'll update 2 out of 3 copies before considering myself done." ← Balanced approach! Now you're guaranteed that if you check 2 copies, at least one will have the new number.
ALL: "I must update all 3 copies right now or I'm not done!" ← Most secure, but problematic - if your car is in the shop and you can't access that copy, you can't make ANY updates!
In Cassandra, consistency levels work exactly like this! They control how many replica nodes must respond before a read or write is considered successful. This is called "tunable consistency" - you can adjust the dial per query!
ONE - Speed Demon
How it works: Contact 1 replica only
Write: Wait for 1 acknowledgment (~5ms)
Read: Return from 1 replica (~5ms)
Guarantee: Eventual consistency
Risk: May see stale data
Best for: Caches, analytics, logs, counters
Example: Page view counters, session data
QUORUM - Sweet Spot ⭐
How it works: Contact majority of replicas
Write: Wait for (RF/2)+1 acks (~15ms)
Read: Read from (RF/2)+1 replicas (~15ms)
Guarantee: Strong consistency
Survives: 1 node failure (RF=3)
Best for: 99% of production!
Example: User profiles, orders, transactions
ALL - Maximum Safety
How it works: Contact ALL replicas
Write: Wait for ALL acks (~50ms)
Read: Read from ALL replicas (~50ms)
Guarantee: Strongest possible
Risk: Unavailable if ANY node down!
Best for: Rarely used!
Example: Critical compliance, audit logs
💡 Key Insight: Tunable Consistency
Unlike traditional databases where you get one consistency model for everything (usually strong ACID), Cassandra lets you choose per query! This means:
- Same table, different levels: Use QUORUM for writes, ONE for analytics reads
- Per-query control: Set level in your query, not globally
- Trade-off control: You decide: Fast + Eventually Consistent OR Slower + Strongly Consistent
- Flexibility: Critical data gets QUORUM, non-critical gets ONE
🔢 QUORUM Mechanics: Step-by-Step Animation
🎯 Key Takeaway
QUORUM means "majority" - with RF=3, you need 2 acknowledgments. The beauty is that the client doesn't wait for all 3 replicas! As soon as 2 replicas respond (at ~15ms), the write is considered successful. Replica 3 will eventually complete the write asynchronously, but the client already got confirmation. This is the perfect balance between speed (faster than ALL) and consistency (stronger than ONE).
📝 Write Path: ONE vs QUORUM vs ALL
👁️ Read Path: The Stale Data Problem with ONE
🚨 Why This Happens
With ONE consistency, Cassandra picks ONE random replica to contact. When you write with ONE, only that one replica gets updated immediately - the other replicas are updated asynchronously in the background. If you then read with ONE, Cassandra picks another random replica, and there's a 66% chance (2 out of 3) it will be one that hasn't been updated yet! This is called "stale read" or "eventual consistency" - the data will eventually be consistent, but not immediately.
🧮 Strong Consistency Formula: R + W > RF
🎓 Master This Concept!
The formula R + W > RF is your golden rule for strong consistency! Here's why it works: If you write to W replicas and read from R replicas, and R + W > RF, then mathematically there MUST be at least 1 replica in both sets (pigeonhole principle). That overlapping replica guarantees you'll read the latest data. QUORUM + QUORUM always satisfies this (2 + 2 = 4 > 3), which is why it's the perfect choice for strong consistency!
✅ Best Practices
Use QUORUM for Production
Default is ONE - change it!
QUORUM = strong consistency
Survives node failures
~15ms acceptable latency
Match Read/Write
R + W > RF = strong consistency
QUORUM + QUORUM = guaranteed
ONE + ONE = eventual only
Choose based on needs
Avoid ALL
ALL fails if 1 node down
Poor availability
Slow (~50ms)
Use QUORUM instead
💼 Top 5 Interview Questions
Answer: QUORUM = (RF / 2) + 1 = majority of replicas.
With RF=3: QUORUM = (3/2) + 1 = 2 replicas
With RF=5: QUORUM = (5/2) + 1 = 3 replicas
Guarantees: Can tolerate (RF-QUORUM) failures = 1 failure with RF=3
Consistency: R + W > RF gives strong consistency (QUORUM + QUORUM = guaranteed)
ONE: Fastest (~5ms), eventual consistency, may read stale data
QUORUM: Balanced (~15ms), strong consistency, survives 1 node failure ⭐
ALL: Slowest (~50ms), strongest guarantee, fails if ANY node down ❌
Recommendation: Use QUORUM for 99% of production workloads
Formula: R + W > RF
Where R = read consistency level, W = write consistency level
Example with RF=3:
- QUORUM write (2) + QUORUM read (2) = 4 > 3 ✓ Strong!
- ONE write (1) + ONE read (1) = 2 < 3 ❌ Eventual only
- ONE write (1) + ALL read (3) = 4 > 3 ✓ Strong!
Best practice: QUORUM/QUORUM = strong + available
LOCAL_QUORUM: Majority of replicas in LOCAL datacenter only
Example: RF=3 per DC → LOCAL_QUORUM = 2 in local DC
Latency: ~10-15ms (doesn't wait for remote DCs)
Use case: Multi-DC deployments - 99% of writes should use LOCAL_QUORUM
vs QUORUM: QUORUM counts ALL replicas globally (slower, ~100ms cross-DC)
Problem: ALL requires ALL replicas to respond - if ANY node is down, operation fails!
Availability impact:
- RF=3: If 1 of 3 nodes down → ALL fails completely
- Normal maintenance? ALL fails
- Network blip? ALL fails
Latency: Slowest replica determines latency (~50ms+)
Better alternative: Use QUORUM - nearly as strong, but survives failures
Only use ALL for: Critical compliance where you absolutely must guarantee all copies written (rare!)
Responsive Ad