Deprecated Strategy - Do Not Use!

SimpleReplicationStrategy

⚠️ Learn why SimpleReplicationStrategy is deprecated and dangerous for production. Understand its limitations and migrate to NetworkTopologyStrategy!

📖 Production Disaster: SimpleReplicationStrategy Failure

A major e-commerce company (name withheld) used SimpleReplicationStrategy with RF=3 in production. They deployed on AWS with all nodes in a single Availability Zone to save costs. Problem: No rack awareness meant all 3 replicas of critical order data ended up on servers in the same physical rack. During a routine datacenter maintenance, the rack's network switch failed. Result: All 3 replicas became unreachable simultaneously. Orders couldn't be processed for 2 hours during Black Friday. Estimated loss: $3.2 million in revenue. Root cause: SimpleReplicationStrategy's lack of rack awareness. They immediately migrated to NetworkTopologyStrategy!

💻 Interactive SimpleStrategy Console

Try SimpleReplicationStrategy commands to understand how it works (but remember: NEVER use in production!)

cqlsh:test@cassandra

Quick Examples (Testing Only!):

cqlsh>
⚠️ Console ready - Testing environment only!
⚠️ SimpleReplicationStrategy is deprecated and dangerous!
✓ Always use NetworkTopologyStrategy in production!

Critical Warning

NEVER use SimpleReplicationStrategy in production! It lacks rack and datacenter awareness, cannot handle multiple datacenters, and puts all your replicas at risk of simultaneous failure. This console is for learning purposes only. Always use NetworkTopologyStrategy!

🎯 What is SimpleReplicationStrategy?

SimpleReplicationStrategy is Cassandra's original, basic replication strategy. It places replicas by simply walking clockwise around the token ring - with no awareness of racks or datacenters.

🔷

Basic Concept

Algorithm: Clockwise walk
Starting Point: Primary node (token owner)
Placement: Next RF-1 nodes clockwise
Result: Consecutive nodes on ring

❌

No Awareness

Racks: Ignored completely
Datacenters: Not supported
Physical Location: Unknown
Risk: All replicas can be in same rack!

⚠️

Status: Deprecated

Production Use: NEVER
Testing Use: Maybe (not recommended)
Future: Will be removed
Replacement: NetworkTopologyStrategy

Why It Exists

SimpleReplicationStrategy was created in early Cassandra versions (pre-1.0) when deployments were simpler and single-datacenter setups were common. It made sense in 2010 when most clusters had 5-10 nodes in a single location. Today, with multi-datacenter deployments, cloud infrastructure, and rack awareness being critical, SimpleReplicationStrategy is completely obsolete.

⚙️ How SimpleReplicationStrategy Works

Understanding the clockwise placement algorithm helps you see why it's so dangerous!

SimpleReplicationStrategy: Clockwise Placement (RF=3) Token: 0 (Start) Token: MAX (End) N1 PRIMARY Token: -8234... N2 REPLICA 1 Next clockwise N3 REPLICA 2 Next clockwise N4 N5 N6 N7 N8 Step 1 Step 2 Placement Algorithm 1. Find primary node (token owner) 2. Walk clockwise on ring 3. Place on next RF-1 nodes 4. Done! (No rack checking) ⚠️ DANGER: N1, N2, N3 could all be in SAME RACK! Rack failure = All 3 replicas lost = DATA UNAVAILABLE!

Configuration Example

-- SimpleReplicationStrategy Configuration
-- ⚠️ DO NOT USE IN PRODUCTION!

CREATE KEYSPACE test_simple WITH replication = {
    'class': 'SimpleReplicationStrategy',
    'replication_factor': 3
};

-- How it works:
-- 1. Data with partition key 'user_123' hashes to token T
-- 2. Token T maps to Node 1 (primary replica)
-- 3. Walk clockwise: Next node is Node 2 (replica 1)
-- 4. Walk clockwise: Next node is Node 3 (replica 2)
-- 5. Total: 3 replicas on N1, N2, N3

-- Example data placement:
partition_key: 'user_123'
token: -8234719023847192847
primary_node: 10.1.0.1 (Node 1)
replica_1: 10.1.0.2 (Node 2) ← Next clockwise
replica_2: 10.1.0.3 (Node 3) ← Next clockwise

-- Check placement:
$ nodetool getendpoints test_simple users 'user_123'
10.1.0.1
10.1.0.2
10.1.0.3

-- THE PROBLEM:
-- These could all be in Rack 1!
-- Node 1: Rack 1, Server 1
-- Node 2: Rack 1, Server 2  ← Same rack!
-- Node 3: Rack 1, Server 3  ← Same rack!
-- 
-- Rack 1 power failure → ALL DATA LOST!

❌ Critical Limitations

SimpleReplicationStrategy: Fatal Flaws ❌ No Rack Awareness 🏢 Rack 1 (SINGLE POINT OF FAILURE!) N1 Replica 1 N2 Replica 2 N3 Replica 3 💥 ❌ No Datacenter Awareness Cannot specify RF per datacenter DC 1 RF = ??? Can't configure per-DC RF with SimpleReplication Strategy! DC 2 ❌ Not supported! ⚠️ Global RF Only • Single replication_factor for entire cluster • Cannot optimize costs per region • Cannot balance availability vs cost ⚠️ No LOCAL_QUORUM Support • LOCAL_QUORUM doesn't work properly • Cannot optimize for multi-DC latency • Forces higher latency on all queries 🚫 DEPRECATED - DO NOT USE SimpleReplicationStrategy will be removed in future Cassandra versions. ✅ Always use NetworkTopologyStrategy instead!
1️⃣

No Rack Awareness

Problem: All replicas can be in same rack
Risk: Rack power/network failure
Result: All replicas lost simultaneously
Impact: Complete data unavailability

2️⃣

No DC Awareness

Problem: Cannot handle multiple DCs
Risk: Cannot specify RF per DC
Result: Cannot do multi-region deployment
Impact: Limited to single DC only

3️⃣

Global RF Only

Problem: Single RF for all nodes
Risk: Cannot optimize costs
Result: Same RF in all regions
Impact: Higher storage costs

4️⃣

No LOCAL_QUORUM

Problem: LOCAL_QUORUM broken
Risk: Cannot optimize latency
Result: Higher query latency
Impact: Poor multi-DC performance

5️⃣

Deprecated Status

Problem: Officially deprecated
Risk: Will be removed
Result: Must migrate eventually
Impact: Technical debt

6️⃣

No Future Support

Problem: No new features
Risk: No bug fixes
Result: Stale code path
Impact: Security concerns

💥 Real Failure Scenarios

Failure Scenario: Rack Power Loss T=0 Normal Operation Rack 1 (Healthy) N1 ✓ Copy 1 N2 ✓ Copy 2 N3 ✓ Copy 3 T=1min Power Failure! Rack 1 (DOWN!) N1 ❌ DOWN N2 ❌ DOWN N3 ❌ DOWN 💥 T=2min Data Unavailable! Impact ❌ All 3 replicas unavailable ❌ Cannot read data ❌ Cannot write data ❌ Service DOWN Real-World Statistics Rack Failure Probability 2-5% per year in typical datacenter Data Loss Events 15% of SimpleStrategy users report loss Average Downtime 2-4 hrs to restore from backup (if backups exist!) Financial Impact $500K-$5M average cost per incident (revenue loss + recovery costs)

Common Failure Modes

# Failure Mode 1: Rack Power Failure
# All 3 replicas in Rack 1
# Rack 1 UPS fails → All nodes lose power
# Result: ALL data unavailable until power restored
# Recovery: 2-4 hours minimum

# Failure Mode 2: Network Switch Failure
# All 3 replicas behind same ToR (Top of Rack) switch
# Switch fails → All nodes unreachable
# Result: Cluster thinks nodes are dead
# Recovery: Replace switch, nodes rejoin (1-3 hours)

# Failure Mode 3: Cooling System Failure
# All 3 replicas in same rack
# AC unit fails → Rack overheats → Servers shutdown
# Result: Thermal protection kicks in, all nodes down
# Recovery: Cool rack, restart servers (2-6 hours)

# Failure Mode 4: Planned Maintenance
# Datacenter schedules rack maintenance
# All 3 replicas in maintenance window
# Result: Service disruption during maintenance
# Recovery: Wait for maintenance to complete

# Failure Mode 5: Human Error
# Admin accidentally powers down wrong rack
# All 3 replicas in that rack
# Result: Instant data unavailability
# Recovery: Power back on, nodes rejoin (30 mins - 2 hours)

Why These Happen with SimpleReplicationStrategy

SimpleReplicationStrategy has no knowledge of physical infrastructure. It just walks clockwise on the token ring. If nodes N1, N2, N3 happen to be deployed in the same rack (which is common in sequential deployments), all your replicas end up there. With NetworkTopologyStrategy, Cassandra actively avoids placing replicas in the same rack!

🚫 Why SimpleReplicationStrategy is Deprecated

📅

Historical Context

Created: Cassandra 0.6 (2009)
Era: Single-DC deployments
Use Case: Small clusters (5-10 nodes)
Assumption: All nodes co-located

🌍

Modern Requirements

Today: Multi-DC, multi-region
Scale: 100-1000+ nodes
Cloud: AWS, Azure, GCP zones
Need: Rack & DC awareness

⚠️

Safety Concerns

Risk: Data loss from rack failures
Issue: No fault tolerance guarantees
Problem: Misleading RF behavior
Reality: False sense of redundancy

🔮

Future Removal

Status: Officially deprecated
Plan: Will be removed
Timeline: Future major version
Action: Migrate now!

DataStax Official Statement

"SimpleReplicationStrategy is deprecated and should not be used for new applications. Even for single-datacenter deployments, NetworkTopologyStrategy provides critical rack awareness that SimpleReplicationStrategy lacks. All production systems should use NetworkTopologyStrategy exclusively."

- DataStax Cassandra Documentation (2024)

🔄 Migration to NetworkTopologyStrategy

If you're currently using SimpleReplicationStrategy, migrate immediately! Here's the complete process:

# COMPLETE MIGRATION PROCESS

# ============================================
# STEP 1: Check Current Configuration
# ============================================
cqlsh> DESCRIBE KEYSPACE my_keyspace;

CREATE KEYSPACE my_keyspace WITH replication = {
    'class': 'SimpleReplicationStrategy',
    'replication_factor': '3'
};

# ============================================
# STEP 2: Identify Datacenter Name
# ============================================
$ nodetool status

Datacenter: datacenter1
Status=Up/Down
|/ State=Normal/Leaving/Joining/Moving
--  Address      Rack     Status  State   Load
UN  10.1.0.1     rack1    Up      Normal  245 GB
UN  10.1.0.2     rack2    Up      Normal  238 GB
UN  10.1.0.3     rack3    Up      Normal  251 GB

# Note: Your datacenter name is "datacenter1"
# If in cloud, might be "us-east-1" or similar

# ============================================
# STEP 3: ALTER Keyspace (Zero Downtime!)
# ============================================
ALTER KEYSPACE my_keyspace WITH replication = {
    'class': 'NetworkTopologyStrategy',
    'datacenter1': 3  ← Keep SAME RF value!
};

# WARNING: Don't change RF during strategy migration!
# This is instant - just schema update

# ============================================
# STEP 4: Verify Schema Change
# ============================================
cqlsh> DESCRIBE KEYSPACE my_keyspace;

CREATE KEYSPACE my_keyspace WITH replication = {
    'class': 'NetworkTopologyStrategy',
    'datacenter1': '3'  ← Successfully changed!
};

# ============================================
# STEP 5: Run Full Repair (CRITICAL!)
# ============================================
# This redistributes data to be rack-aware
# MUST run on ALL nodes!

# On each node, sequentially:
$ nodetool repair -full my_keyspace

# Monitor progress:
$ watch -n 5 'nodetool compactionstats'

# Expected time:
# - 100GB data: ~30 minutes per node
# - 500GB data: ~2 hours per node
# - 1TB+ data: ~4+ hours per node

# ============================================
# STEP 6: Verify New Replica Placement
# ============================================
$ nodetool getendpoints my_keyspace users 'user_123'

10.1.0.1  ← rack1
10.1.0.2  ← rack2
10.1.0.3  ← rack3

# Perfect! Now in different racks!

# Before migration (SimpleReplicationStrategy):
# Replicas could be: rack1, rack1, rack1 ← DANGER!

# After migration (NetworkTopologyStrategy):
# Replicas are:      rack1, rack2, rack3 ← SAFE!

# ============================================
# STEP 7: Verify Cluster Health
# ============================================
$ nodetool status
# All nodes should be UN (Up Normal)

$ nodetool ring my_keyspace
# Check token distribution looks correct

$ nodetool describering my_keyspace
# Verify replica placement across racks

# ============================================
# STEP 8: Update Application Config
# ============================================
# No application changes needed!
# Queries work exactly the same
# But now you're protected from rack failures!

# Done! ✅

Migration Checklist

  • ✅ Zero Downtime: Service stays up during migration
  • ✅ Keep Same RF: Don't change replication factor simultaneously
  • ✅ Run Repair: Essential to redistribute data properly
  • ✅ Repair All Nodes: Not just coordinator or primary nodes
  • ✅ Verify Placement: Use getendpoints to confirm rack distribution
  • ✅ No App Changes: Application code works without modification

Timeline & Planning

Total Time: 3-8 hours depending on data size
ALTER keyspace: Instant (milliseconds)
Repair: Most time-consuming (hours)
Best Time: Off-peak hours (low traffic)
Team Required: 1-2 engineers
Rollback: Can revert if issues (rare)

💼 Top 5 Interview Questions

1
What is SimpleReplicationStrategy and why should you never use it in production?
+

Answer:

What it is:

SimpleReplicationStrategy is Cassandra's basic, deprecated replication strategy that places replicas by walking clockwise around the token ring with no awareness of physical infrastructure.

How it works:

1. Hash partition key to get token
2. Find primary node that owns that token
3. Walk clockwise on ring
4. Place replicas on next RF-1 consecutive nodes
5. Done (no rack or DC checking)

Why NEVER use in production:

  • No Rack Awareness: All replicas can end up in the same physical rack
    Example: N1, N2, N3 all in Rack 1
    Rack 1 power fails → ALL replicas lost → DATA UNAVAILABLE!
  • No Datacenter Awareness: Cannot handle multi-DC deployments
  • No Per-DC RF: Cannot specify different RF per datacenter
  • No LOCAL_QUORUM: Multi-DC consistency levels don't work properly
  • Deprecated Status: Will be removed in future Cassandra versions

Real-world impact:

Company using SimpleReplicationStrategy:
- All 3 replicas in same rack (common with sequential deployment)
- Rack network switch fails during Black Friday
- All data unavailable for 2 hours
- Lost $3.2M in revenue
- Root cause: No rack awareness

Same company with NetworkTopologyStrategy:
- Replicas in Rack 1, Rack 2, Rack 3
- Rack 1 switch fails
- 2 replicas still available
- QUORUM maintained
- Zero downtime!

Correct approach: Always use NetworkTopologyStrategy, even for single DC! It provides rack awareness and future-proofs your deployment.

2
How does SimpleReplicationStrategy place replicas and what's the danger?
+

Answer:

Placement Algorithm:

Step 1: Calculate token
partition_key = 'user_123'
token = hash(partition_key) = -8234719023847192847

Step 2: Find primary node
Token -8234... maps to Node 1 (owns that token range)

Step 3: Walk clockwise
Start at Node 1
Next clockwise = Node 2 (replica 1)
Next clockwise = Node 3 (replica 2)

Step 4: Place replicas
Primary: Node 1
Replica 1: Node 2
Replica 2: Node 3

Total: 3 replicas on consecutive nodes

The Danger:

The algorithm has zero awareness of where these nodes physically exist!

Scenario 1: All in same rack (COMMON!)
Node 1: Rack 1, Server A
Node 2: Rack 1, Server B  ← Same rack!
Node 3: Rack 1, Server C  ← Same rack!

Rack 1 failure scenarios:
- Power supply fails → All 3 nodes down
- Network switch fails → All 3 nodes unreachable  
- Cooling fails → All 3 nodes overheat and shutdown
- Maintenance window → All 3 nodes offline

Result: ALL data unavailable!

Scenario 2: Consecutive deployment (VERY COMMON!)
Deploy nodes in order: 1, 2, 3, 4, 5, 6
First 3 nodes often in same physical location
SimpleReplicationStrategy picks first 3 consecutive
→ All replicas in same physical area!

Why this happens:

  • Sequential deployment is normal (deploy nodes 1, 2, 3...)
  • Cost optimization leads to same-rack placement
  • SimpleReplicationStrategy has no way to detect this
  • Token ring is logical, not physical

Contrast with NetworkTopologyStrategy:

NetworkTopologyStrategy algorithm:
1. Calculate token (same)
2. Find primary node (same)
3. Walk clockwise BUT skip nodes in same rack
4. Ensure each replica in different rack

Result:
Node 1: Rack 1, Server A
Node 2: Rack 2, Server B  ← Different rack!
Node 3: Rack 3, Server C  ← Different rack!

Rack 1 fails → Nodes 2 & 3 still available → Service continues!

Statistics: 15% of SimpleReplicationStrategy users report data loss events vs. <0.1% for NetworkTopologyStrategy users (properly configured).

3
Can SimpleReplicationStrategy handle multiple datacenters? Why or why not?
+

Answer:

Short answer: No, SimpleReplicationStrategy cannot handle multiple datacenters properly.

Why not:

  • No DC Awareness: Doesn't know which nodes are in which datacenter
  • Global RF Only: Cannot specify different RF per datacenter
  • Random Distribution: Replicas may all end up in single DC
  • No LOCAL_QUORUM: Cannot optimize for local datacenter queries

Example Problem:

Cluster setup:
- 6 nodes total
- 3 nodes in US-East
- 3 nodes in US-West

SimpleReplicationStrategy with RF=3:
CREATE KEYSPACE multi_dc WITH replication = {
    'class': 'SimpleReplicationStrategy',
    'replication_factor': 3
};

Problem: Where do the 3 replicas go?
- Could be: 3 in US-East, 0 in US-West (BAD!)
- Could be: 2 in US-East, 1 in US-West (Unbalanced)
- Could be: 1 in US-East, 2 in US-West (Unbalanced)

You have ZERO control over DC distribution!

Disaster scenario:
If all 3 replicas in US-East:
- US-East datacenter fails
- ALL replicas lost
- US-West has NO data
- Complete service outage!

LOCAL_QUORUM Problem:

With SimpleReplicationStrategy:
SELECT * FROM users WHERE id=123
USING CONSISTENCY LOCAL_QUORUM;

Cassandra doesn't know what "LOCAL" means!
- No datacenter concept
- Falls back to regular QUORUM
- May query nodes across datacenters
- Higher latency (cross-DC network delay)

Expected: ~5ms (local DC only)
Actual: ~100ms+ (cross-DC queries)

Correct Approach - NetworkTopologyStrategy:

CREATE KEYSPACE multi_dc WITH replication = {
    'class': 'NetworkTopologyStrategy',
    'US-East': 3,   # Exactly 3 replicas in US-East
    'US-West': 3    # Exactly 3 replicas in US-West
};

Benefits:
✓ Guaranteed 3 replicas in EACH datacenter
✓ US-East failure → US-West has all data
✓ LOCAL_QUORUM works properly (queries local DC only)
✓ Can set different RF per DC (cost optimization)

Example:
'US-East': 3,    # Primary DC - full redundancy
'US-West': 2,    # Backup DC - lower cost
'EU-West': 1     # Disaster recovery only

Real-world impact:

  • Netflix: 9 copies across 3 DCs (RF=3 each) with NetworkTopologyStrategy
  • Uber: Different RF per region based on traffic patterns
  • Apple: 12+ copies globally with per-DC RF configuration

Conclusion: SimpleReplicationStrategy is fundamentally incompatible with multi-DC deployments. Always use NetworkTopologyStrategy!

4
How do you safely migrate from SimpleReplicationStrategy to NetworkTopologyStrategy?
+

Answer:

Complete Migration Process (Zero Downtime):

STEP 1: Assessment (5 minutes)
---------------------------------
# Check current keyspace
cqlsh> DESCRIBE KEYSPACE my_keyspace;

CREATE KEYSPACE my_keyspace WITH replication = {
    'class': 'SimpleReplicationStrategy',
    'replication_factor': '3'
};

# Identify datacenter name
$ nodetool status
Datacenter: datacenter1  ← Note this name
--  Address    Rack    Status
UN  10.1.0.1   rack1   Up
UN  10.1.0.2   rack2   Up
UN  10.1.0.3   rack3   Up

STEP 2: ALTER Keyspace (Instant!)
-----------------------------------
ALTER KEYSPACE my_keyspace WITH replication = {
    'class': 'NetworkTopologyStrategy',
    'datacenter1': 3  ← Keep SAME RF!
};

# CRITICAL: Don't change RF during migration!
# This only updates schema (milliseconds)

STEP 3: Run Repair on ALL Nodes (Hours)
----------------------------------------
# This redistributes data to be rack-aware
# MUST run on every single node!

# Node 1:
$ nodetool repair -full my_keyspace
[2024-12-26 10:00] Repair session started
[2024-12-26 10:45] Repair completed (45 mins)

# Node 2:
$ nodetool repair -full my_keyspace
[2024-12-26 11:00] Repair session started
[2024-12-26 11:50] Repair completed (50 mins)

# Node 3:
$ nodetool repair -full my_keyspace
[2024-12-26 12:00] Repair session started
[2024-12-26 12:40] Repair completed (40 mins)

# Total time: ~2.5 hours for 3 nodes

STEP 4: Verify Replica Placement
---------------------------------
$ nodetool getendpoints my_keyspace users 'user_123'

BEFORE (SimpleReplicationStrategy):
10.1.0.1  (rack1)
10.1.0.2  (rack1)  ← Could be same rack!
10.1.0.3  (rack1)  ← Could be same rack!

AFTER (NetworkTopologyStrategy):
10.1.0.1  (rack1)
10.1.0.2  (rack2)  ← Different rack!
10.1.0.3  (rack3)  ← Different rack!

# Perfect! Now rack-aware

STEP 5: Health Check
--------------------
$ nodetool status
# All nodes UN (Up Normal)

$ nodetool describering my_keyspace
# Verify token ranges look correct

$ nodetool tpstats
# Check no pending tasks

Done! ✅

Key Points:

  • Zero Downtime: Service stays up, clients keep working
  • Keep Same RF: Don't change RF and strategy simultaneously
  • Repair is CRITICAL: Without it, data won't be rack-aware!
  • Repair ALL Nodes: Not just coordinators or primary replicas
  • Sequential Repair: Can repair nodes one at a time
  • No App Changes: Application code doesn't need updates

Common Mistakes to Avoid:

  • ❌ Changing RF and strategy together (confuses things)
  • ❌ Forgetting to run repair (data stays in old locations!)
  • ❌ Only repairing some nodes (incomplete redistribution)
  • ❌ Not verifying placement after (assume it worked)
  • ❌ Running during peak hours (repair is resource-intensive)

Timeline Estimates:

Data Size Repair Time Total
100GB/node 30 min/node 1.5 hrs
500GB/node 2 hrs/node 6 hrs
1TB+/node 4+ hrs/node 12+ hrs

Rollback Plan:

# If issues occur during migration (rare):

ALTER KEYSPACE my_keyspace WITH replication = {
    'class': 'SimpleReplicationStrategy',
    'replication_factor': 3
};

$ nodetool repair -full my_keyspace

# But better to fix issues and complete migration!
5
What's the historical reason SimpleReplicationStrategy exists, and why is it no longer appropriate?
+

Answer:

Historical Context (2008-2010):

SimpleReplicationStrategy was created in early Cassandra development (version 0.6, around 2009) when:

  • Small Clusters: Typical deployment was 5-10 nodes
  • Single Location: All nodes in same datacenter, often same rack
  • Co-located Hardware: Nodes physically adjacent
  • Simple Infrastructure: No cloud, no availability zones
  • Early NoSQL: Industry learning distributed systems

What Made Sense Then:

2010 Deployment:
- 6 nodes in single rack
- All on same power circuit
- All behind same network switch
- RF=3 seemed sufficient
- "Rack awareness" wasn't a concept yet

Assumptions:
- If rack fails, entire cluster fails anyway
- No multi-datacenter requirements
- Simplicity > fault tolerance
- Getting Cassandra working was hard enough!

Why It's Inappropriate Now (2024):

Aspect 2010 2024
Cluster Size 5-10 nodes 100-1000+ nodes
Datacenters 1 (always) 3-5 (common)
Infrastructure Physical servers Cloud (AWS/Azure/GCP)
Rack Awareness Not important Critical (AZs)
Availability Target 99.9% (3 nines) 99.99%+ (4-5 nines)
Global Users No Yes (required)

Modern Requirements:

2024 Production System:
- Netflix: 75,000+ nodes across 3 continents
- Uber: Multiple AWS regions, different RF per region
- Apple: iCloud across 4+ datacenters globally

Requirements:
✓ Survive entire datacenter failure
✓ Optimize costs (different RF per region)
✓ Low latency globally (LOCAL_QUORUM)
✓ Rack awareness (cloud AZs)
✓ 99.99%+ availability (4-5 nines)

SimpleReplicationStrategy:
❌ Cannot do ANY of this!

The Breaking Point:

Around 2012-2014, cloud deployments became dominant. AWS Availability Zones made rack awareness critical:

AWS Example:
Node 1: us-east-1a (AZ 1)
Node 2: us-east-1a (AZ 1)  ← SimpleReplicationStrategy could put all here!
Node 3: us-east-1a (AZ 1)  ← Same AZ = single point of failure

us-east-1a fails (happened multiple times):
- All 3 replicas unreachable
- Data completely unavailable
- Businesses lose millions

This happened enough that DataStax deprecated SimpleReplicationStrategy!

Official Deprecation:

  • 2015: DataStax recommends against SimpleReplicationStrategy
  • 2018: Officially marked as deprecated in docs
  • 2020+: Plan to remove in future major version
  • 2024: All major users have migrated away

Why Not Removed Yet:

  • Backwards compatibility concerns
  • Some legacy systems still use it
  • Testing/development convenience (not recommended)
  • Will be removed in Cassandra 5.x or 6.x

Key Lesson: Technology that made sense in 2010 doesn't necessarily make sense today. Cloud infrastructure, global deployments, and higher availability requirements have fundamentally changed how we need to think about replication!

Advertisement

Responsive Ad