Consistency Level: ONE

ONE - Speed & Eventual Consistency

โšก Master the fastest consistency level! Learn when ONE is perfect, when it's dangerous, and how to use it safely!

๐Ÿ“– Netflix: When ONE is Perfect

Netflix serves 260+ million subscribers globally with billions of viewing records per day. For viewing history and watch progress, Netflix uses consistency level ONE. Why? These operations happen every second as users watch content. Challenge: Writing to QUORUM would add ~10ms per write = billions of extra milliseconds daily! Netflix's solution: Write viewing progress with ONE (~3-5ms latency). Trade-off accepted: If user switches devices immediately, might see slightly old progress (e.g., "You're at 10:30" instead of "10:45"). Impact: Acceptable! Users rarely notice 15-second progress lag. After a few seconds, data converges (eventual consistency). Performance gain: 3x faster writes, 60% lower CPU usage, billions saved in infrastructure costs. Netflix engineers: "For non-critical, high-volume data where eventual consistency is acceptable, ONE is a massive performance win!"

๐Ÿ’ป Interactive ONE Simulator

See how ONE works in real-time! Explore different scenarios and understand when it's safe vs dangerous.

cqlsh@speed-cluster

โšก Quick Scenarios - Click to Explore:

โšก ONE Consistency Simulator Ready!
๐ŸŽฏ Click scenarios above to see ONE in action
โœ“ Learn when ONE is safe vs dangerous

โš ๏ธ Default Trap: ONE is Cassandra's Default!

CRITICAL: If you don't explicitly set a consistency level, Cassandra defaults to ONE! This catches many beginners off-guard. They write code like INSERT INTO users... thinking they have strong consistency, but they're actually using ONE! For production data, ALWAYS explicitly set QUORUM or higher. ONE is only safe for specific use cases like caches, analytics, and temporary data.

โšก What is Consistency Level ONE?

Imagine you're sending a group text to 3 friends about dinner plans. With ONE, you hit send and as soon as just one friend receives it, your phone says "Message Delivered!" You don't wait for all 3 friends to get it - you assume they'll all get it eventually. This is "eventual consistency".

๐ŸŽฏ Real-World Analogy: The Speedy Mailman

You have 3 mailboxes (one at home, one at work, one at vacation house). You want to update your address with the post office.

With ONE (Speed Priority):
The mailman says "I'll update your home mailbox RIGHT NOW and come back for the other two later." You get confirmation in 5 seconds! Fast! But if you immediately check your work mailbox, it still has the old address. After a few hours, all 3 mailboxes have the new address (eventual consistency).

Perfect for: Address updates for non-urgent mail (magazines, ads)
Dangerous for: Bank statements, legal documents, medical records

This is exactly how ONE works in Cassandra!

In technical terms: Consistency Level ONE means Cassandra only waits for 1 replica to acknowledge a write or provide data for a read. It's the fastest possible consistency level (~3-5ms) but provides only eventual consistency - meaning your data will eventually be consistent across all replicas, but not immediately.

โšก

Write with ONE

Behavior: Send write to all replicas, wait for 1 ACK
Latency: ~3-5ms (fastest!)
What happens: 1 replica confirms immediately
Others: Updated asynchronously (background)
Use when: Speed > Immediate consistency

๐Ÿ‘€

Read with ONE

Behavior: Contact 1 random replica
Latency: ~3-5ms (fastest!)
What you get: Whatever that 1 replica has
Risk: May be stale if replica behind
Use when: Okay with slightly old data

๐Ÿ”„

Eventual Consistency

Guarantee: All replicas will eventually match
Timeline: Usually within milliseconds
Mechanism: Background sync (read repair, anti-entropy)
Trade-off: Speed now, consistency later
Perfect for: High-volume, non-critical data

๐Ÿ’ก Key Insight: Speed vs Consistency Trade-off

ONE is Cassandra's "speed mode". You're essentially saying: "I care more about low latency than immediate consistency." This is perfect for:

  • Session data: User logged in? Don't need all replicas to agree immediately
  • Caches: Slightly stale cache data is fine - it's just a cache!
  • Analytics/metrics: Pageview count off by 1-2? Acceptable!
  • Temporary data: Shopping cart before checkout? Can be eventually consistent

But it's dangerous for critical data like user profiles, orders, financial transactions, or anything where seeing old data would confuse users or cause business problems.

๐Ÿ”ง How ONE Works: The Complete Picture

Write with ONE: Only Wait for 1 Acknowledgment (RF=3) โšก Client Writes with CL=ONE ๐Ÿ“ก Coordinator sends write to ALL 3 replicas (parallel) but only waits for 1 ACK Replica 1 โšก Node: 10.1.0.1 โœ“ ACK T = 3ms FASTEST RESPONSE! Replica 2 Node: 10.1.0.2 โณ Writing T = 8ms Client doesn't wait! Replica 3 Node: 10.1.0.3 โณ Writing T = 15ms Async background Success at 3ms! โฑ๏ธ Complete Timeline T=0ms: Client sends write, coordinator forwards to all 3 replicas T=3ms: โœ… Replica 1 ACKs โ†’ CLIENT RECEIVES SUCCESS! T=8ms: Replica 2 ACKs (client already got success) T=15ms: Replica 3 ACKs (background, client doesn't care) T=15ms+: All 3 replicas have data โ†’ Eventually consistent! ๐ŸŽฏ Key Point: Client latency = 3ms (fastest replica), not 15ms (slowest)! That's the power of ONE!

โšก Performance Magic Explained

Notice how the client only waited 3ms even though Replica 3 took 15ms? This is the magic of ONE! The coordinator sends the write to all replicas in parallel, but the client only waits for the fastest one to respond. The slower replicas complete in the background asynchronously. This gives you 3-5x faster latency compared to QUORUM, which would wait for 2 replicas (~8-12ms). For high-volume, non-critical data, this performance boost is massive!

๐Ÿ‘ป The Stale Data Problem: Why ONE Can Be Dangerous

Stale Read Scenario: The User Profile Update Nightmare Step 1 (T=0): Initial State All 3 replicas have: email = "[email protected]" R1: old R2: old R3: old Step 2 (T=5ms): User updates email with CL=ONE UPDATE users SET email='[email protected]' WHERE id=123; R1 NEW โœ“ R2 OLD โŒ R3 OLD โŒ Step 3 (T=10ms): User immediately reads with CL=ONE SELECT email FROM users WHERE id=123; Cassandra picks ONE random replica to read from... Lucky! (33%) R1 โœ“ Gets NEW email Unlucky! (66% chance!) ๐Ÿ’ฅ R2 OR R3 โŒ Gets OLD email! ๐Ÿ’ฅ The User Experience Nightmare Action: User updates email to "[email protected]" Result: System says "โœ“ Email updated successfully!" Problem: User refreshes profile page immediately... 66% chance: Sees "[email protected]" โŒ (Read hit R2 or R3!) User thinks: "My update didn't work! The system is broken!" User refreshes again: Now hits R1, sees "[email protected]" โœ“ User confusion: "Why does it keep changing back and forth?!" โœ… SOLUTION: Use QUORUM for user-facing data to guarantee they always see their latest changes!

๐Ÿšจ Why This Is So Common

This exact scenario happens thousands of times per day in production systems using ONE! Here's why:

  • Probability math: With RF=3 and ONE, there's a 66% chance (2 out of 3) your read hits a replica that wasn't updated yet
  • Human behavior: Users often update data and immediately refresh to verify the change
  • Replication lag: Even though background replication is fast (~10-50ms), it's not instant
  • Default trap: Many developers don't realize ONE is the default and wonder why they see stale data
  • Production impact: Users think the system is broken, file support tickets, lose trust

This is why ONE should NEVER be used for user profiles, account settings, orders, or anything users directly interact with!

โœ… When ONE is SAFE (Perfect Use Cases)

ONE is perfect when you need maximum performance and can tolerate eventual consistency. These are scenarios where slightly stale data won't cause user confusion or business problems.

๐Ÿ’พ

Caches & Session Stores

Why SAFE: Stale cache = minor, re-fetch from source
Example: Redis replacement, API response cache, CDN cache
Impact: Cache miss โ†’ regenerate (acceptable)
Benefit: 3-5x faster writes, 5x higher throughput
Real: Netflix viewing cache, Twitter timeline cache
Pattern: If data has TTL, ONE is usually safe

๐Ÿ“Š

Analytics & Metrics

Why SAFE: Approximate counts perfectly fine
Example: Pageviews, clicks, impressions, heat maps
Impact: Count off by 1-2? Completely acceptable!
Benefit: Handle millions of events/second
Real: LinkedIn profile views, Google Analytics style
Pattern: If it's a counter, ONE works great

๐Ÿ“

Logs & Events

Why SAFE: Append-only, order doesn't matter
Example: App logs, activity streams, audit trails
Impact: Logs eventually consistent = totally fine
Benefit: Never block application on logging
Real: Most companies for non-compliance logging
Pattern: Write-heavy, read-light? Use ONE

๐Ÿ”„

Session Data

Why SAFE: Temporary, short-lived, non-critical
Example: Login sessions, shopping cart (pre-checkout)
Impact: Worst case = user logs in again (minor)
Benefit: Ultra-fast session creation/updates
Real: E-commerce sites for browsing sessions
Note: Switch to QUORUM when cart โ†’ order!

๐Ÿ“บ

Viewing History

Why SAFE: Slight progress lag acceptable
Example: Netflix watch position, Spotify plays
Impact: "At 10:30" vs "10:45" = user doesn't care
Benefit: Billions of writes/day possible
Real: Netflix confirmed this exact use case!
Pattern: High-frequency updates? ONE shines

๐Ÿ”

Regenerable Data

Why SAFE: Can be recomputed from source
Example: Search indexes, thumbnails, materialized views
Impact: Data loss? Just regenerate (no big deal)
Benefit: Don't block on derived data writes
Real: Image processing pipelines, search engines
Pattern: If it's derived, ONE is often perfect

๐ŸŽฏ The ONE Safe Use Pattern

Notice the pattern? ONE is safe when data is: (1) High-volume, (2) Non-critical, (3) Temporary or regenerable, and (4) Not directly user-facing in a way where staleness would be confusing. If your use case fits this pattern, ONE can give you massive performance gains with minimal risk!

โŒ When ONE is DANGEROUS (Never Use Here!)

ONE is dangerous when users directly interact with the data and would be confused or upset by seeing stale information. These are critical business scenarios where eventual consistency causes real problems.

๐Ÿ‘ค

User Profiles

Why DANGEROUS: Users see their own stale changes!
Example: Name, email, phone, profile photo, settings
Impact: "I just updated this, why is it wrong?!" ๐Ÿ’ฅ
Probability: 66% chance user sees stale data
Result: Support tickets, user frustration, lost trust
Use instead: QUORUM or LOCAL_QUORUM

๐Ÿ’ฐ

Financial Transactions

Why DANGEROUS: Money MUST be accurate!
Example: Bank balance, payments, invoices, credits
Impact: Legal liability, regulatory violations ๐Ÿ’ฅ
Horror: Balance shows $100 โ†’ stale read shows $0
Result: Lawsuits, compliance fines, audits fail
Use instead: ALL or QUORUM + lightweight transactions

๐Ÿ›’

Orders & Inventory

Why DANGEROUS: Business logic breaks down
Example: Order status, stock counts, reservations
Impact: Overselling, double-bookings, angry customers
Horror: Item out of stock but shows available
Result: Revenue loss, refunds, bad reviews
Use instead: QUORUM writes, LOCAL_QUORUM reads

๐Ÿ”

Authentication & Security

Why DANGEROUS: Security holes!
Example: Passwords, API keys, permissions, tokens
Impact: Unauthorized access, data breaches ๐Ÿ’ฅ
Horror: User changes password, old still works!
Result: Security incident, compliance violation
Use instead: QUORUM or ALL for auth changes

๐Ÿ“‹

Critical Settings

Why DANGEROUS: System behavior inconsistent
Example: Feature flags, quotas, rate limits, configs
Impact: Features randomly on/off, limits not enforced
Horror: Rate limit disabled on some nodes only
Result: DoS attacks, resource exhaustion
Use instead: QUORUM for all config changes

๐Ÿ“ง

User-Generated Content

Why DANGEROUS: Creator can't see their work!
Example: Posts, comments, messages, reviews
Impact: "Did my post work?" confusion
Horror: User posts comment, doesn't see it, posts again
Result: Duplicate content, user frustration
Use instead: QUORUM for all UGC writes

๐Ÿšซ The Golden Rule for ONE

Simple decision framework: Ask yourself: "If a user saw data from 100 milliseconds ago instead of right now, would they be confused, upset, or would business logic break?" If YES โ†’ Don't use ONE! If NO โ†’ ONE is probably safe. When in doubt, use QUORUM - it's only ~10ms slower but guarantees strong consistency.

โšก Performance Benefits: Why ONE is So Fast

Latency Comparison: ONE vs QUORUM vs ALL (RF=3) Write Latency (Lower is Better) ONE: 3-5ms โšก โœ“ Wait for 1 replica only - Always picks fastest! QUORUM: 12-15ms Wait for 2 of 3 replicas - Slower but strong consistency ALL: 40-50ms (or timeout!) โŒ Wait for all 3 replicas - Slowest, fails if any node down Throughput (Higher is Better) ONE: 50,000 writes/sec/node 1x QUORUM: 15,000 writes/sec/node 3.3x ALL: 5,000 writes/sec/node 10x
โšก

3-10x Faster Writes

ONE: 3-5ms average latency
QUORUM: 12-15ms (3x slower)
ALL: 40-50ms (10x slower)
Why: Only wait for fastest replica!
Impact: Sub-10ms API responses possible

๐Ÿ“ˆ

3-10x Higher Throughput

ONE: 50k writes/sec/node
QUORUM: 15k writes/sec (3.3x less)
ALL: 5k writes/sec (10x less)
Why: Less coordinator overhead
Impact: Handle massive event streams

๐Ÿ’ป

40-60% Lower CPU

ONE: 40% CPU usage
QUORUM: 65% CPU (1.6x more)
ALL: 80% CPU (2x more)
Why: Less waiting, less coordination
Impact: Lower infrastructure costs

๐Ÿข Real Companies Using ONE (Production Success Stories)

๐Ÿ“บ Netflix: Viewing History at Massive Scale

Scale: 260+ million subscribers, billions of viewing events per day
Use Case: Watch progress tracking ("You're at 12:34 in Episode 3")
Why ONE: Users watch content continuously - every second generates an update. QUORUM would add 10ms per write = billions of extra milliseconds daily!
Configuration: Write progress with CL=ONE, RF=3
Trade-off Accepted: If user switches devices immediately, might see progress from 10-15 seconds ago (e.g., "10:30" instead of "10:45")
Impact: Users rarely notice or care - data converges in ~50ms
Performance Gain: 3x faster writes, 60% lower CPU, billions saved in infrastructure
Netflix Engineering: "For high-volume, non-critical data where eventual consistency is acceptable, ONE is a massive win!"

๐Ÿฆ Twitter: Timeline Caching

Scale: 500+ million users, billions of tweets
Use Case: Cached timeline data for faster loads
Why ONE: Timeline cache = derivative data, can be regenerated from source
Configuration: Write cache entries with CL=ONE, TTL=300s
Trade-off: Stale cache entry? Re-fetch from source (acceptable overhead)
Impact: 5x more cache writes possible
Note: Original tweets use QUORUM - only cache uses ONE

๐Ÿ’ผ LinkedIn: Profile View Counters

Scale: 900+ million members, billions of profile views
Use Case: "Who viewed your profile" counts and analytics
Why ONE: Approximate counts totally acceptable for analytics
Configuration: INCREMENT counter with CL=ONE
Trade-off: Count might be off by 1-2 views temporarily
Impact: Can handle millions of view events per second
User Experience: "127 views this week" vs "128 views" = user doesn't care!

๐ŸŽฏ Common Production Pattern

Notice the pattern? All these companies use ONE for high-volume, non-critical data and QUORUM for critical data. This hybrid approach is the production best practice!

โœ… Best Practices for Using ONE Safely

๐ŸŽฏ

1. Never Use ONE by Default

Problem: ONE is Cassandra's default!
Trap: Many developers don't realize this
Solution: Explicitly set CL in code
Pattern: QUORUM as default, ONE for specific cases

โ“

2. Ask the Staleness Question

Question: "Would 100ms old data confuse users?"
Yes: Use QUORUM
No: ONE might be okay
Remember: 10ms slower >> user confusion

๐Ÿ”„

3. Use Hybrid Approach

Pattern: Different CL per table
Critical: QUORUM (users, orders)
High-volume: ONE (logs, metrics)
Benefit: Performance where safe, consistency where needed

โฑ๏ธ

4. Consider TTL Data

Pattern: If data expires, ONE is safer
Examples: Session cache (TTL=1h)
Logic: Will be deleted anyway
Benefit: Max performance for temp data

๐Ÿ“Š

5. Monitor Stale Reads

Action: Add staleness metrics
Alert: If >100ms consistently
Tool: Read repair helps
Pattern: Measure before assume

๐Ÿงช

6. Test with ONE First

Pattern: Start QUORUM, test ONE later
Process: Launch with QUORUM
Then: Identify high-volume, safe tables
Benefit: Safe migration, proven safety

โš ๏ธ Common Mistakes to Avoid

  • โŒ Using ONE for user profiles: Users see stale data after updates
  • โŒ Not explicitly setting CL: You'll get ONE by default!
  • โŒ Assuming "fast enough": QUORUM only ~10ms slower
  • โŒ Using ONE for auth: Security nightmare
  • โŒ Not testing staleness: Always measure actual behavior

๐Ÿ’ผ Interview Questions & Answers

1
What is consistency level ONE and when should you use it?

Complete Answer:

Consistency level ONE means Cassandra only waits for one replica to acknowledge a write or provide data for a read, out of the total RF (replication factor) replicas. It's the fastest consistency level (~3-5ms latency) but provides only eventual consistency.

How it works:

  • Write with ONE: Coordinator sends write to all replicas but only waits for 1 ACK before returning success to client. Other replicas update asynchronously in background.
  • Read with ONE: Coordinator contacts 1 random replica and returns whatever data that replica has.

When to use ONE (SAFE cases):

  • Caches: Session cache, API response cache - stale cache = minor, just re-fetch
  • Analytics/metrics: Pageview counts, click tracking - approximate counts acceptable
  • Logs: Application logs, event streams - append-only, eventual consistency fine
  • Viewing history: Netflix watch progress - slight lag acceptable to users
  • Temporary data: Shopping cart before checkout, session data with TTL

When NOT to use ONE (DANGEROUS):

  • User profiles: Email, name, settings - users will see their own stale changes (66% probability!)
  • Financial data: Balances, transactions - must be accurate, legal requirements
  • Orders/inventory: Stock counts, order status - business logic breaks with stale data
  • Authentication: Passwords, API keys - security hole if old credentials still work

Performance benefits: ONE gives 3-5x faster writes and 3-10x higher throughput compared to QUORUM/ALL. Netflix uses ONE for billions of viewing events per day, saving massive infrastructure costs.

Key insight: ONE is perfect when you need maximum performance and can tolerate eventual consistency (data converges in 10-50ms). Always ask: "Would 100ms old data confuse users?" If yes โ†’ QUORUM. If no โ†’ ONE is safe.

2
Why can ONE consistency level return stale data? Explain with an example.

Complete Answer:

ONE can return stale data because of replication lag - the time between when one replica is updated and when all replicas are updated. With ONE, you only wait for 1 replica, so other replicas might still have old data.

Detailed Example - User Profile Update:

Setup: RF=3 (3 replicas), all using CL=ONE

Timeline:

  • T=0ms: All 3 replicas have email = "old@example.com"
  • T=5ms: User updates email with CL=ONE
    • Write reaches Replica 1 first
    • Replica 1 ACKs โ†’ Client gets SUCCESS โœ“
    • Replica 2 still has "old@example.com" โŒ
    • Replica 3 still has "old@example.com" โŒ
  • T=10ms: User immediately refreshes profile page (READ with CL=ONE)
    • Lucky (33% chance): Read hits Replica 1 โ†’ Returns "new@example.com" โœ“
    • Unlucky (66% chance): Read hits Replica 2 or 3 โ†’ Returns "old@example.com" โŒ STALE!
  • T=50ms: Background replication completes, all replicas now have "new@example.com" โœ“

User Experience:

  • User updates email
  • System says "โœ“ Email updated successfully!"
  • User refreshes page โ†’ sees OLD email (if unlucky)
  • User thinks: "My update didn't work! The system is broken!"
  • User refreshes again โ†’ now sees NEW email (if lucky)
  • User is confused: "Why does it keep changing?"

Why 66% probability of stale read?

With RF=3 and only 1 replica updated, there's a 2/3 chance your read will hit one of the 2 replicas that don't have the latest data yet. This is basic probability: 2 stale replicas out of 3 total = 66%.

Solution: Use QUORUM for both reads and writes. With QUORUM, 2 replicas are updated during write, so when you read from 2 replicas, at least 1 MUST have the latest data (R + W > RF = 2 + 2 = 4 > 3).

3
Netflix uses consistency level ONE for viewing history. Explain why this is safe and what trade-offs they accept.

Complete Answer:

Netflix uses CL=ONE for viewing history because it's a high-volume, non-critical use case where the performance benefits massively outweigh the minor risk of stale data.

Why ONE is SAFE for Netflix viewing history:

  • High volume: Billions of viewing events per day (users watching content continuously)
  • Performance critical: Every second of watching generates an update - QUORUM would add ~10ms per write = billions of extra milliseconds daily
  • Non-critical data: Watch progress isn't business-critical like subscription data or billing
  • User tolerance: If user switches devices and sees "You're at 10:30" instead of "10:45" (15 seconds behind), they don't care - they just click play and it resumes
  • Quick convergence: Data becomes consistent in ~50ms - by the time user actually switches devices, data is already consistent

Trade-offs Netflix accepts:

  • Slight progress lag: If you watch on TV, then immediately switch to phone, you might see progress from 10-15 seconds ago
  • Eventual consistency: All replicas eventually have correct progress, but not immediately
  • Acceptable user experience: Users rarely switch devices so quickly that they notice, and when they do, they just seek to where they left off

Performance gains Netflix gets:

  • 3x faster writes: ~3-5ms with ONE vs ~15ms with QUORUM
  • 3-5x higher throughput: Can handle billions more events with same infrastructure
  • 60% lower CPU usage: Less coordination overhead
  • Billions saved: Massive infrastructure cost savings over QUORUM

Important contrast - What Netflix uses QUORUM for:

  • User profiles (email, name, settings)
  • Subscription data
  • Billing information
  • Payment methods

Key lesson: Netflix uses a hybrid approach - ONE for high-volume non-critical data (viewing history), QUORUM for critical data (user accounts, billing). This is the production best practice: use the right consistency level for each use case!

Netflix engineers quote: "For non-critical, high-volume data where eventual consistency is acceptable, ONE is a massive performance win. The benefits far outweigh the tiny chance of stale reads."

4
What's the default consistency level in Cassandra and why is this dangerous?

Complete Answer:

The default consistency level in Cassandra is ONE, and this is dangerous because most developers don't realize this and unknowingly get eventual consistency when they expect strong consistency.

Why ONE as default is dangerous:

  • Silent failure: Developers write code like INSERT INTO users... thinking they have ACID-like strong consistency
  • Production surprises: Everything works fine in testing (single node or low load), but in production users start seeing stale data
  • User confusion: Users update their profile, refresh, and see old data โ†’ "The system is broken!"
  • Support tickets: Flood of complaints about "updates not working" when actually it's eventual consistency
  • Wrong assumptions: Developers coming from MySQL/PostgreSQL assume strong consistency by default

Common scenario where default ONE causes problems:

// Developer writes this code (no CL specified)
INSERT INTO users (id, email) VALUES (123, 'new@email.com');

// Cassandra uses default CL=ONE
// Only 1 replica updated immediately
// User refreshes page...
SELECT email FROM users WHERE id=123;
// 66% chance of reading from stale replica!
// User sees old email โŒ

Real production horror story:

  • E-commerce site launches with Cassandra
  • Developers never explicitly set consistency level
  • Gets ONE by default
  • Users update their shipping address
  • System confirms "Address updated!"
  • User places order 5 seconds later
  • Order ships to OLD address (read hit stale replica)
  • Customer complaints, refunds, lost revenue

How to protect against this:

  • Always explicitly set CL: Never rely on defaults
    // ALWAYS do this
    session.execute("CONSISTENCY QUORUM");
    session.execute("INSERT INTO users...");
  • Set driver defaults: Configure your driver to use QUORUM as default
  • Code review checks: Verify CL is explicitly set in all queries
  • Testing: Test with RF=3 in staging to catch staleness issues
  • Monitoring: Alert if seeing lots of read repair (indicates staleness)

Why Cassandra chose ONE as default:

  • Historical reasons (early Cassandra optimized for availability over consistency)
  • ONE gives best performance and availability
  • Assumes developers will read docs and set CL appropriately (dangerous assumption!)

Best practice: Treat QUORUM as your default, use ONE only for specific high-volume, non-critical use cases where you've consciously decided eventual consistency is acceptable. Always explicitly set consistency level - never rely on defaults!

5
Compare the performance characteristics of ONE vs QUORUM vs ALL. When would you choose each?

Complete Answer:

Performance Comparison (RF=3 cluster):

ONE (Fastest):

  • Latency: 3-5ms (wait for 1 replica)
  • Throughput: ~50,000 writes/sec per node
  • CPU usage: 40% under load
  • Consistency: Eventual only
  • Availability: Works even if 2 nodes down
  • Why fast: Only waits for fastest replica, minimal coordination

QUORUM (Balanced) โญ:

  • Latency: 12-15ms (wait for 2 of 3 replicas)
  • Throughput: ~15,000 writes/sec per node (3.3x less than ONE)
  • CPU usage: 65% under load
  • Consistency: Strong (if R+W>RF)
  • Availability: Works with 1 node down (tolerates RF-QUORUM failures)
  • Why slower: Must wait for 2 replicas, more coordination overhead

ALL (Slowest):

  • Latency: 40-50ms or timeout (wait for all 3 replicas)
  • Throughput: ~5,000 writes/sec per node (10x less than ONE)
  • CPU usage: 80% under load
  • Consistency: Strongest possible
  • Availability: โŒ FAILS if ANY node down!
  • Why slowest: Wait for ALL replicas, slowest one determines latency

When to choose each:

Choose ONE when:

  • โœ… High-volume writes (billions/day)
  • โœ… Performance is critical (< 10ms required)
  • โœ… Data is non-critical (caches, metrics, logs)
  • โœ… Eventual consistency acceptable
  • โœ… Data is temporary or regenerable
  • Examples: Netflix viewing history, Twitter timeline cache, LinkedIn profile view counters, session data, application logs

Choose QUORUM when:

  • โœ… 99% of production use cases!
  • โœ… User-facing data (profiles, accounts)
  • โœ… Business-critical data (orders, inventory)
  • โœ… Need strong consistency
  • โœ… Can tolerate ~15ms latency (almost everything can)
  • โœ… Need fault tolerance (survives 1 node failure)
  • Examples: User profiles, e-commerce orders, authentication, account settings, payment data, anything users directly interact with

Choose ALL when:

  • โœ… Critical compliance requirements
  • โœ… Audit logs that must be on every replica
  • โœ… VERY rarely used in production!
  • โŒ Avoid: Poor availability (fails if any node down)
  • โŒ Avoid: Slow performance (10x slower than ONE)
  • Better alternative: Use QUORUM + verification for most "ALL" use cases
  • Examples: Financial audit logs for regulatory compliance, critical security events

Production recommendation:

  • Default: QUORUM (safe, proven, good enough for 99% of cases)
  • Optimize: Switch high-volume, non-critical tables to ONE
  • Avoid: ALL (use QUORUM instead unless you have specific regulatory requirements)

Real numbers from production benchmark:

  • ONE: 50k writes/sec, 3ms P50, 8ms P99
  • QUORUM: 15k writes/sec, 15ms P50, 45ms P99
  • ALL: 5k writes/sec, 50ms P50, 200ms P99

Key insight: QUORUM is only 3x slower than ONE but gives you strong consistency and fault tolerance. For most use cases, that trade-off is worth it. Use ONE only when you've consciously decided performance is more important than immediate consistency.

Advertisement

Responsive Ad