ONE - Speed & Eventual Consistency
โก Master the fastest consistency level! Learn when ONE is perfect, when it's dangerous, and how to use it safely!
๐ Netflix: When ONE is Perfect
Netflix serves 260+ million subscribers globally with billions of viewing records per day. For viewing history and watch progress, Netflix uses consistency level ONE. Why? These operations happen every second as users watch content. Challenge: Writing to QUORUM would add ~10ms per write = billions of extra milliseconds daily! Netflix's solution: Write viewing progress with ONE (~3-5ms latency). Trade-off accepted: If user switches devices immediately, might see slightly old progress (e.g., "You're at 10:30" instead of "10:45"). Impact: Acceptable! Users rarely notice 15-second progress lag. After a few seconds, data converges (eventual consistency). Performance gain: 3x faster writes, 60% lower CPU usage, billions saved in infrastructure costs. Netflix engineers: "For non-critical, high-volume data where eventual consistency is acceptable, ONE is a massive performance win!"
๐ป Interactive ONE Simulator
See how ONE works in real-time! Explore different scenarios and understand when it's safe vs dangerous.
โก Quick Scenarios - Click to Explore:
๐ฏ Click scenarios above to see ONE in action
โ Learn when ONE is safe vs dangerous
โ ๏ธ Default Trap: ONE is Cassandra's Default!
CRITICAL: If you don't explicitly set a consistency level, Cassandra defaults to ONE! This catches many beginners off-guard. They write code like INSERT INTO users... thinking they have strong consistency, but they're actually using ONE! For production data, ALWAYS explicitly set QUORUM or higher. ONE is only safe for specific use cases like caches, analytics, and temporary data.
โก What is Consistency Level ONE?
Imagine you're sending a group text to 3 friends about dinner plans. With ONE, you hit send and as soon as just one friend receives it, your phone says "Message Delivered!" You don't wait for all 3 friends to get it - you assume they'll all get it eventually. This is "eventual consistency".
๐ฏ Real-World Analogy: The Speedy Mailman
You have 3 mailboxes (one at home, one at work, one at vacation house). You want to update your address with the post office.
With ONE (Speed Priority):
The mailman says "I'll update your home mailbox RIGHT NOW and come back for the other two later." You get confirmation in 5 seconds! Fast! But if you immediately check your work mailbox, it still has the old address. After a few hours, all 3 mailboxes have the new address (eventual consistency).
Perfect for: Address updates for non-urgent mail (magazines, ads)
Dangerous for: Bank statements, legal documents, medical records
This is exactly how ONE works in Cassandra!
In technical terms: Consistency Level ONE means Cassandra only waits for 1 replica to acknowledge a write or provide data for a read. It's the fastest possible consistency level (~3-5ms) but provides only eventual consistency - meaning your data will eventually be consistent across all replicas, but not immediately.
Write with ONE
Behavior: Send write to all replicas, wait for 1 ACK
Latency: ~3-5ms (fastest!)
What happens: 1 replica confirms immediately
Others: Updated asynchronously (background)
Use when: Speed > Immediate consistency
Read with ONE
Behavior: Contact 1 random replica
Latency: ~3-5ms (fastest!)
What you get: Whatever that 1 replica has
Risk: May be stale if replica behind
Use when: Okay with slightly old data
Eventual Consistency
Guarantee: All replicas will eventually match
Timeline: Usually within milliseconds
Mechanism: Background sync (read repair, anti-entropy)
Trade-off: Speed now, consistency later
Perfect for: High-volume, non-critical data
๐ก Key Insight: Speed vs Consistency Trade-off
ONE is Cassandra's "speed mode". You're essentially saying: "I care more about low latency than immediate consistency." This is perfect for:
- Session data: User logged in? Don't need all replicas to agree immediately
- Caches: Slightly stale cache data is fine - it's just a cache!
- Analytics/metrics: Pageview count off by 1-2? Acceptable!
- Temporary data: Shopping cart before checkout? Can be eventually consistent
But it's dangerous for critical data like user profiles, orders, financial transactions, or anything where seeing old data would confuse users or cause business problems.
๐ง How ONE Works: The Complete Picture
โก Performance Magic Explained
Notice how the client only waited 3ms even though Replica 3 took 15ms? This is the magic of ONE! The coordinator sends the write to all replicas in parallel, but the client only waits for the fastest one to respond. The slower replicas complete in the background asynchronously. This gives you 3-5x faster latency compared to QUORUM, which would wait for 2 replicas (~8-12ms). For high-volume, non-critical data, this performance boost is massive!
๐ป The Stale Data Problem: Why ONE Can Be Dangerous
๐จ Why This Is So Common
This exact scenario happens thousands of times per day in production systems using ONE! Here's why:
- Probability math: With RF=3 and ONE, there's a 66% chance (2 out of 3) your read hits a replica that wasn't updated yet
- Human behavior: Users often update data and immediately refresh to verify the change
- Replication lag: Even though background replication is fast (~10-50ms), it's not instant
- Default trap: Many developers don't realize ONE is the default and wonder why they see stale data
- Production impact: Users think the system is broken, file support tickets, lose trust
This is why ONE should NEVER be used for user profiles, account settings, orders, or anything users directly interact with!
โ When ONE is SAFE (Perfect Use Cases)
ONE is perfect when you need maximum performance and can tolerate eventual consistency. These are scenarios where slightly stale data won't cause user confusion or business problems.
Caches & Session Stores
Why SAFE: Stale cache = minor, re-fetch from source
Example: Redis replacement, API response cache, CDN cache
Impact: Cache miss โ regenerate (acceptable)
Benefit: 3-5x faster writes, 5x higher throughput
Real: Netflix viewing cache, Twitter timeline cache
Pattern: If data has TTL, ONE is usually safe
Analytics & Metrics
Why SAFE: Approximate counts perfectly fine
Example: Pageviews, clicks, impressions, heat maps
Impact: Count off by 1-2? Completely acceptable!
Benefit: Handle millions of events/second
Real: LinkedIn profile views, Google Analytics style
Pattern: If it's a counter, ONE works great
Logs & Events
Why SAFE: Append-only, order doesn't matter
Example: App logs, activity streams, audit trails
Impact: Logs eventually consistent = totally fine
Benefit: Never block application on logging
Real: Most companies for non-compliance logging
Pattern: Write-heavy, read-light? Use ONE
Session Data
Why SAFE: Temporary, short-lived, non-critical
Example: Login sessions, shopping cart (pre-checkout)
Impact: Worst case = user logs in again (minor)
Benefit: Ultra-fast session creation/updates
Real: E-commerce sites for browsing sessions
Note: Switch to QUORUM when cart โ order!
Viewing History
Why SAFE: Slight progress lag acceptable
Example: Netflix watch position, Spotify plays
Impact: "At 10:30" vs "10:45" = user doesn't care
Benefit: Billions of writes/day possible
Real: Netflix confirmed this exact use case!
Pattern: High-frequency updates? ONE shines
Regenerable Data
Why SAFE: Can be recomputed from source
Example: Search indexes, thumbnails, materialized views
Impact: Data loss? Just regenerate (no big deal)
Benefit: Don't block on derived data writes
Real: Image processing pipelines, search engines
Pattern: If it's derived, ONE is often perfect
๐ฏ The ONE Safe Use Pattern
Notice the pattern? ONE is safe when data is: (1) High-volume, (2) Non-critical, (3) Temporary or regenerable, and (4) Not directly user-facing in a way where staleness would be confusing. If your use case fits this pattern, ONE can give you massive performance gains with minimal risk!
โ When ONE is DANGEROUS (Never Use Here!)
ONE is dangerous when users directly interact with the data and would be confused or upset by seeing stale information. These are critical business scenarios where eventual consistency causes real problems.
User Profiles
Why DANGEROUS: Users see their own stale changes!
Example: Name, email, phone, profile photo, settings
Impact: "I just updated this, why is it wrong?!" ๐ฅ
Probability: 66% chance user sees stale data
Result: Support tickets, user frustration, lost trust
Use instead: QUORUM or LOCAL_QUORUM
Financial Transactions
Why DANGEROUS: Money MUST be accurate!
Example: Bank balance, payments, invoices, credits
Impact: Legal liability, regulatory violations ๐ฅ
Horror: Balance shows $100 โ stale read shows $0
Result: Lawsuits, compliance fines, audits fail
Use instead: ALL or QUORUM + lightweight transactions
Orders & Inventory
Why DANGEROUS: Business logic breaks down
Example: Order status, stock counts, reservations
Impact: Overselling, double-bookings, angry customers
Horror: Item out of stock but shows available
Result: Revenue loss, refunds, bad reviews
Use instead: QUORUM writes, LOCAL_QUORUM reads
Authentication & Security
Why DANGEROUS: Security holes!
Example: Passwords, API keys, permissions, tokens
Impact: Unauthorized access, data breaches ๐ฅ
Horror: User changes password, old still works!
Result: Security incident, compliance violation
Use instead: QUORUM or ALL for auth changes
Critical Settings
Why DANGEROUS: System behavior inconsistent
Example: Feature flags, quotas, rate limits, configs
Impact: Features randomly on/off, limits not enforced
Horror: Rate limit disabled on some nodes only
Result: DoS attacks, resource exhaustion
Use instead: QUORUM for all config changes
User-Generated Content
Why DANGEROUS: Creator can't see their work!
Example: Posts, comments, messages, reviews
Impact: "Did my post work?" confusion
Horror: User posts comment, doesn't see it, posts again
Result: Duplicate content, user frustration
Use instead: QUORUM for all UGC writes
๐ซ The Golden Rule for ONE
Simple decision framework: Ask yourself: "If a user saw data from 100 milliseconds ago instead of right now, would they be confused, upset, or would business logic break?" If YES โ Don't use ONE! If NO โ ONE is probably safe. When in doubt, use QUORUM - it's only ~10ms slower but guarantees strong consistency.
โก Performance Benefits: Why ONE is So Fast
3-10x Faster Writes
ONE: 3-5ms average latency
QUORUM: 12-15ms (3x slower)
ALL: 40-50ms (10x slower)
Why: Only wait for fastest replica!
Impact: Sub-10ms API responses possible
3-10x Higher Throughput
ONE: 50k writes/sec/node
QUORUM: 15k writes/sec (3.3x less)
ALL: 5k writes/sec (10x less)
Why: Less coordinator overhead
Impact: Handle massive event streams
40-60% Lower CPU
ONE: 40% CPU usage
QUORUM: 65% CPU (1.6x more)
ALL: 80% CPU (2x more)
Why: Less waiting, less coordination
Impact: Lower infrastructure costs
๐ข Real Companies Using ONE (Production Success Stories)
๐บ Netflix: Viewing History at Massive Scale
Scale: 260+ million subscribers, billions of viewing events per day
Use Case: Watch progress tracking ("You're at 12:34 in Episode 3")
Why ONE: Users watch content continuously - every second generates an update. QUORUM would add 10ms per write = billions of extra milliseconds daily!
Configuration: Write progress with CL=ONE, RF=3
Trade-off Accepted: If user switches devices immediately, might see progress from 10-15 seconds ago (e.g., "10:30" instead of "10:45")
Impact: Users rarely notice or care - data converges in ~50ms
Performance Gain: 3x faster writes, 60% lower CPU, billions saved in infrastructure
Netflix Engineering: "For high-volume, non-critical data where eventual consistency is acceptable, ONE is a massive win!"
๐ฆ Twitter: Timeline Caching
Scale: 500+ million users, billions of tweets
Use Case: Cached timeline data for faster loads
Why ONE: Timeline cache = derivative data, can be regenerated from source
Configuration: Write cache entries with CL=ONE, TTL=300s
Trade-off: Stale cache entry? Re-fetch from source (acceptable overhead)
Impact: 5x more cache writes possible
Note: Original tweets use QUORUM - only cache uses ONE
๐ผ LinkedIn: Profile View Counters
Scale: 900+ million members, billions of profile views
Use Case: "Who viewed your profile" counts and analytics
Why ONE: Approximate counts totally acceptable for analytics
Configuration: INCREMENT counter with CL=ONE
Trade-off: Count might be off by 1-2 views temporarily
Impact: Can handle millions of view events per second
User Experience: "127 views this week" vs "128 views" = user doesn't care!
๐ฏ Common Production Pattern
Notice the pattern? All these companies use ONE for high-volume, non-critical data and QUORUM for critical data. This hybrid approach is the production best practice!
โ Best Practices for Using ONE Safely
1. Never Use ONE by Default
Problem: ONE is Cassandra's default!
Trap: Many developers don't realize this
Solution: Explicitly set CL in code
Pattern: QUORUM as default, ONE for specific cases
2. Ask the Staleness Question
Question: "Would 100ms old data confuse users?"
Yes: Use QUORUM
No: ONE might be okay
Remember: 10ms slower >> user confusion
3. Use Hybrid Approach
Pattern: Different CL per table
Critical: QUORUM (users, orders)
High-volume: ONE (logs, metrics)
Benefit: Performance where safe, consistency where needed
4. Consider TTL Data
Pattern: If data expires, ONE is safer
Examples: Session cache (TTL=1h)
Logic: Will be deleted anyway
Benefit: Max performance for temp data
5. Monitor Stale Reads
Action: Add staleness metrics
Alert: If >100ms consistently
Tool: Read repair helps
Pattern: Measure before assume
6. Test with ONE First
Pattern: Start QUORUM, test ONE later
Process: Launch with QUORUM
Then: Identify high-volume, safe tables
Benefit: Safe migration, proven safety
โ ๏ธ Common Mistakes to Avoid
- โ Using ONE for user profiles: Users see stale data after updates
- โ Not explicitly setting CL: You'll get ONE by default!
- โ Assuming "fast enough": QUORUM only ~10ms slower
- โ Using ONE for auth: Security nightmare
- โ Not testing staleness: Always measure actual behavior
๐ผ Interview Questions & Answers
Complete Answer:
Consistency level ONE means Cassandra only waits for one replica to acknowledge a write or provide data for a read, out of the total RF (replication factor) replicas. It's the fastest consistency level (~3-5ms latency) but provides only eventual consistency.
How it works:
- Write with ONE: Coordinator sends write to all replicas but only waits for 1 ACK before returning success to client. Other replicas update asynchronously in background.
- Read with ONE: Coordinator contacts 1 random replica and returns whatever data that replica has.
When to use ONE (SAFE cases):
- Caches: Session cache, API response cache - stale cache = minor, just re-fetch
- Analytics/metrics: Pageview counts, click tracking - approximate counts acceptable
- Logs: Application logs, event streams - append-only, eventual consistency fine
- Viewing history: Netflix watch progress - slight lag acceptable to users
- Temporary data: Shopping cart before checkout, session data with TTL
When NOT to use ONE (DANGEROUS):
- User profiles: Email, name, settings - users will see their own stale changes (66% probability!)
- Financial data: Balances, transactions - must be accurate, legal requirements
- Orders/inventory: Stock counts, order status - business logic breaks with stale data
- Authentication: Passwords, API keys - security hole if old credentials still work
Performance benefits: ONE gives 3-5x faster writes and 3-10x higher throughput compared to QUORUM/ALL. Netflix uses ONE for billions of viewing events per day, saving massive infrastructure costs.
Key insight: ONE is perfect when you need maximum performance and can tolerate eventual consistency (data converges in 10-50ms). Always ask: "Would 100ms old data confuse users?" If yes โ QUORUM. If no โ ONE is safe.
Complete Answer:
ONE can return stale data because of replication lag - the time between when one replica is updated and when all replicas are updated. With ONE, you only wait for 1 replica, so other replicas might still have old data.
Detailed Example - User Profile Update:
Setup: RF=3 (3 replicas), all using CL=ONE
Timeline:
- T=0ms: All 3 replicas have email = "old@example.com"
- T=5ms: User updates email with CL=ONE
- Write reaches Replica 1 first
- Replica 1 ACKs โ Client gets SUCCESS โ
- Replica 2 still has "old@example.com" โ
- Replica 3 still has "old@example.com" โ
- T=10ms: User immediately refreshes profile page (READ with CL=ONE)
- Lucky (33% chance): Read hits Replica 1 โ Returns "new@example.com" โ
- Unlucky (66% chance): Read hits Replica 2 or 3 โ Returns "old@example.com" โ STALE!
- T=50ms: Background replication completes, all replicas now have "new@example.com" โ
User Experience:
- User updates email
- System says "โ Email updated successfully!"
- User refreshes page โ sees OLD email (if unlucky)
- User thinks: "My update didn't work! The system is broken!"
- User refreshes again โ now sees NEW email (if lucky)
- User is confused: "Why does it keep changing?"
Why 66% probability of stale read?
With RF=3 and only 1 replica updated, there's a 2/3 chance your read will hit one of the 2 replicas that don't have the latest data yet. This is basic probability: 2 stale replicas out of 3 total = 66%.
Solution: Use QUORUM for both reads and writes. With QUORUM, 2 replicas are updated during write, so when you read from 2 replicas, at least 1 MUST have the latest data (R + W > RF = 2 + 2 = 4 > 3).
Complete Answer:
Netflix uses CL=ONE for viewing history because it's a high-volume, non-critical use case where the performance benefits massively outweigh the minor risk of stale data.
Why ONE is SAFE for Netflix viewing history:
- High volume: Billions of viewing events per day (users watching content continuously)
- Performance critical: Every second of watching generates an update - QUORUM would add ~10ms per write = billions of extra milliseconds daily
- Non-critical data: Watch progress isn't business-critical like subscription data or billing
- User tolerance: If user switches devices and sees "You're at 10:30" instead of "10:45" (15 seconds behind), they don't care - they just click play and it resumes
- Quick convergence: Data becomes consistent in ~50ms - by the time user actually switches devices, data is already consistent
Trade-offs Netflix accepts:
- Slight progress lag: If you watch on TV, then immediately switch to phone, you might see progress from 10-15 seconds ago
- Eventual consistency: All replicas eventually have correct progress, but not immediately
- Acceptable user experience: Users rarely switch devices so quickly that they notice, and when they do, they just seek to where they left off
Performance gains Netflix gets:
- 3x faster writes: ~3-5ms with ONE vs ~15ms with QUORUM
- 3-5x higher throughput: Can handle billions more events with same infrastructure
- 60% lower CPU usage: Less coordination overhead
- Billions saved: Massive infrastructure cost savings over QUORUM
Important contrast - What Netflix uses QUORUM for:
- User profiles (email, name, settings)
- Subscription data
- Billing information
- Payment methods
Key lesson: Netflix uses a hybrid approach - ONE for high-volume non-critical data (viewing history), QUORUM for critical data (user accounts, billing). This is the production best practice: use the right consistency level for each use case!
Netflix engineers quote: "For non-critical, high-volume data where eventual consistency is acceptable, ONE is a massive performance win. The benefits far outweigh the tiny chance of stale reads."
Complete Answer:
The default consistency level in Cassandra is ONE, and this is dangerous because most developers don't realize this and unknowingly get eventual consistency when they expect strong consistency.
Why ONE as default is dangerous:
- Silent failure: Developers write code like
INSERT INTO users...thinking they have ACID-like strong consistency - Production surprises: Everything works fine in testing (single node or low load), but in production users start seeing stale data
- User confusion: Users update their profile, refresh, and see old data โ "The system is broken!"
- Support tickets: Flood of complaints about "updates not working" when actually it's eventual consistency
- Wrong assumptions: Developers coming from MySQL/PostgreSQL assume strong consistency by default
Common scenario where default ONE causes problems:
// Developer writes this code (no CL specified) INSERT INTO users (id, email) VALUES (123, 'new@email.com'); // Cassandra uses default CL=ONE // Only 1 replica updated immediately // User refreshes page... SELECT email FROM users WHERE id=123; // 66% chance of reading from stale replica! // User sees old email โ
Real production horror story:
- E-commerce site launches with Cassandra
- Developers never explicitly set consistency level
- Gets ONE by default
- Users update their shipping address
- System confirms "Address updated!"
- User places order 5 seconds later
- Order ships to OLD address (read hit stale replica)
- Customer complaints, refunds, lost revenue
How to protect against this:
- Always explicitly set CL: Never rely on defaults
// ALWAYS do this session.execute("CONSISTENCY QUORUM"); session.execute("INSERT INTO users..."); - Set driver defaults: Configure your driver to use QUORUM as default
- Code review checks: Verify CL is explicitly set in all queries
- Testing: Test with RF=3 in staging to catch staleness issues
- Monitoring: Alert if seeing lots of read repair (indicates staleness)
Why Cassandra chose ONE as default:
- Historical reasons (early Cassandra optimized for availability over consistency)
- ONE gives best performance and availability
- Assumes developers will read docs and set CL appropriately (dangerous assumption!)
Best practice: Treat QUORUM as your default, use ONE only for specific high-volume, non-critical use cases where you've consciously decided eventual consistency is acceptable. Always explicitly set consistency level - never rely on defaults!
Complete Answer:
Performance Comparison (RF=3 cluster):
ONE (Fastest):
- Latency: 3-5ms (wait for 1 replica)
- Throughput: ~50,000 writes/sec per node
- CPU usage: 40% under load
- Consistency: Eventual only
- Availability: Works even if 2 nodes down
- Why fast: Only waits for fastest replica, minimal coordination
QUORUM (Balanced) โญ:
- Latency: 12-15ms (wait for 2 of 3 replicas)
- Throughput: ~15,000 writes/sec per node (3.3x less than ONE)
- CPU usage: 65% under load
- Consistency: Strong (if R+W>RF)
- Availability: Works with 1 node down (tolerates RF-QUORUM failures)
- Why slower: Must wait for 2 replicas, more coordination overhead
ALL (Slowest):
- Latency: 40-50ms or timeout (wait for all 3 replicas)
- Throughput: ~5,000 writes/sec per node (10x less than ONE)
- CPU usage: 80% under load
- Consistency: Strongest possible
- Availability: โ FAILS if ANY node down!
- Why slowest: Wait for ALL replicas, slowest one determines latency
When to choose each:
Choose ONE when:
- โ High-volume writes (billions/day)
- โ Performance is critical (< 10ms required)
- โ Data is non-critical (caches, metrics, logs)
- โ Eventual consistency acceptable
- โ Data is temporary or regenerable
- Examples: Netflix viewing history, Twitter timeline cache, LinkedIn profile view counters, session data, application logs
Choose QUORUM when:
- โ 99% of production use cases!
- โ User-facing data (profiles, accounts)
- โ Business-critical data (orders, inventory)
- โ Need strong consistency
- โ Can tolerate ~15ms latency (almost everything can)
- โ Need fault tolerance (survives 1 node failure)
- Examples: User profiles, e-commerce orders, authentication, account settings, payment data, anything users directly interact with
Choose ALL when:
- โ Critical compliance requirements
- โ Audit logs that must be on every replica
- โ VERY rarely used in production!
- โ Avoid: Poor availability (fails if any node down)
- โ Avoid: Slow performance (10x slower than ONE)
- Better alternative: Use QUORUM + verification for most "ALL" use cases
- Examples: Financial audit logs for regulatory compliance, critical security events
Production recommendation:
- Default: QUORUM (safe, proven, good enough for 99% of cases)
- Optimize: Switch high-volume, non-critical tables to ONE
- Avoid: ALL (use QUORUM instead unless you have specific regulatory requirements)
Real numbers from production benchmark:
- ONE: 50k writes/sec, 3ms P50, 8ms P99
- QUORUM: 15k writes/sec, 15ms P50, 45ms P99
- ALL: 5k writes/sec, 50ms P50, 200ms P99
Key insight: QUORUM is only 3x slower than ONE but gives you strong consistency and fault tolerance. For most use cases, that trade-off is worth it. Use ONE only when you've consciously decided performance is more important than immediate consistency.
Responsive Ad