Section 3: Data Distribution

Replication Strategies

Master data placement! Learn SimpleReplicationStrategy vs NetworkTopologyStrategy with animated diagrams, real scenarios, and production best practices!

📖 Uber's Multi-Region Strategy: 99.99% Availability

Uber operates in 10,000+ cities globally with trip data stored across multiple AWS regions. Initial problem: SimpleReplicationStrategy couldn't handle multi-datacenter deployment. Solution: Migrated to NetworkTopologyStrategy with rack-aware placement across US-East (RF=3), US-West (RF=3), EU-West (RF=2). Result: When AWS US-East-1 had a 4-hour outage in 2020, Uber's service continued seamlessly using replicas in other regions. 15 million daily rides experienced zero disruption!

💻 Interactive Strategy Console

Try replication strategy commands in our live console! Experiment with different configurations.

cqlsh:demo@cassandra

Quick Examples:

cqlsh>
▶ Console ready! Try the quick examples above or write your own commands.
▶ Learn about SimpleReplicationStrategy vs NetworkTopologyStrategy.
▶ Tip: Press Ctrl+Enter to execute commands quickly!

Learning Tips

  • Experiment: Try both SimpleReplicationStrategy and NetworkTopologyStrategy
  • Multi-DC: See how to configure different RF per datacenter
  • Describe: Check existing keyspace configurations
  • Migration: Learn how to ALTER keyspace strategy safely

🎯 What are Replication Strategies?

Replication Strategies determine HOW and WHERE Cassandra places replica copies across the cluster. Think of it as the "placement algorithm" for your data.

Key Concept

While Replication Factor (RF) tells Cassandra "how many copies", the Replication Strategy tells it "where to place those copies". Together they ensure data redundancy and availability!

The Two Main Strategies

🔷

SimpleReplicationStrategy

Use Case: Single datacenter only
Placement: Simple clockwise on ring
Rack-Aware: No
DC-Aware: No
Production: ❌ NOT recommended

🌐

NetworkTopologyStrategy ⭐

Use Case: Single or multi-datacenter
Placement: Smart DC/rack-aware
Rack-Aware: Yes
DC-Aware: Yes
Production: ✅ Always use this!

🔷 SimpleReplicationStrategy (Deprecated)

SimpleReplicationStrategy places replicas by walking clockwise around the token ring. It's simple but has serious limitations!

SimpleReplicationStrategy: Token Ring (RF=3) N1 Primary Token: 0 N2 Replica 1 N3 Replica 2 N4 N5 N6 Clockwise ⚠️ Problem: No rack or datacenter awareness! All 3 replicas could be in same rack/DC → Single point of failure Algorithm: 1. Place on primary node | 2. Walk clockwise | 3. Next N-1 nodes

Configuration Example

-- SimpleReplicationStrategy (DO NOT USE IN PRODUCTION!)
CREATE KEYSPACE test_keyspace WITH replication = {
    'class': 'SimpleReplicationStrategy',
    'replication_factor': 3
};

-- How it works:
-- 1. Data with token T maps to Node N1 (primary)
-- 2. Walk clockwise on ring to find next 2 nodes (N2, N3)
-- 3. Place replicas on N1, N2, N3 (consecutive nodes)

-- Example placement:
-- Token: -8234719023847192847 maps to Node 1
-- Replica 1: Node 2 (next clockwise)
-- Replica 2: Node 3 (next clockwise)

-- This data is now on 3 consecutive nodes on the ring

Why SimpleReplicationStrategy is Dangerous

  • No Rack Awareness: All replicas could be in same physical rack
    Rack fails → All 3 replicas lost → DATA LOSS!
  • No Datacenter Awareness: Cannot specify RF per datacenter
  • No Multi-DC Support: Breaks with multiple datacenters
  • Deprecated: Will be removed in future Cassandra versions

⚠️ NEVER use SimpleReplicationStrategy in production! Always use NetworkTopologyStrategy, even for single DC!

🌐 NetworkTopologyStrategy (Recommended)

NetworkTopologyStrategy (NTS) is datacenter-aware and rack-aware, ensuring replicas are placed intelligently for maximum fault tolerance!

NetworkTopologyStrategy: Multi-DC with Rack Awareness 🏢 Datacenter: US-East RF = 3 (3 copies in this DC) Rack 1 N1 ✓ Copy 1 N4 Rack 2 N2 ✓ Copy 2 N5 Rack 3 N3 ✓ Copy 3 N6 🏢 Datacenter: US-West RF = 2 (2 copies in this DC) Rack 1 N7 ✓ Copy 1 N9 Rack 2 N8 ✓ Copy 2 N10 Rack 3 N11 Configuration CREATE KEYSPACE production WITH replication = { 'class': 'NetworkTopologyStrategy', 'US-East': 3, # 3 copies, different racks 'US-West': 2 # 2 copies, different racks }; ✅ Smart Placement Benefits ✓ Rack Awareness: Replicas in different racks (survive rack failure) ✓ DC Awareness: Different RF per datacenter (optimize costs) ✓ Fault Tolerance: Survive entire rack or datacenter failure ✓ Performance: LOCAL_QUORUM works efficiently per DC

Detailed Configuration Examples

-- Single Datacenter with NetworkTopologyStrategy (RECOMMENDED!)
-- Even for single DC, use NTS for rack awareness
CREATE KEYSPACE production WITH replication = {
    'class': 'NetworkTopologyStrategy',
    'datacenter1': 3
};

-- Multi-Datacenter with Different RF
-- Perfect for global deployments
CREATE KEYSPACE global_data WITH replication = {
    'class': 'NetworkTopologyStrategy',
    'US-East': 3,      -- Primary region, high availability
    'US-West': 3,      -- Secondary region, high availability
    'EU-West': 2,      -- Read-only region, lower cost
    'Asia-Pacific': 2  -- Read-only region, lower cost
};
-- Total: 10 copies across 4 DCs!

-- Cost-Optimized Multi-DC
-- Higher RF in primary, lower in secondary
CREATE KEYSPACE cost_optimized WITH replication = {
    'class': 'NetworkTopologyStrategy',
    'US-East': 3,      -- Primary: full RF=3
    'US-West': 2,      -- Backup: RF=2 saves storage
    'EU-West': 1       -- Disaster recovery only: RF=1
};

-- How NTS places replicas:
-- 1. Select primary node based on token
-- 2. For each datacenter:
--    a. Walk clockwise in that DC
--    b. Skip nodes in same rack as previous replicas
--    c. Place RF copies in different racks
-- 3. Result: Maximum fault tolerance!

Why NetworkTopologyStrategy is Superior

  • Rack Awareness: Automatically spreads replicas across racks
  • Datacenter Awareness: Different RF per DC for cost optimization
  • Multi-DC Support: Native support for geo-distribution
  • LOCAL_QUORUM: Fast local reads/writes without cross-DC latency
  • Future-Proof: Easy to add new datacenters later
  • Production Standard: Used by all major companies (Netflix, Uber, Apple)

🏗️ Rack Awareness Explained

Rack Awareness ensures NetworkTopologyStrategy places replicas in different physical racks, protecting against rack-level failures (power, network switch, etc).

Rack Awareness: Smart Replica Placement ❌ Without Rack Awareness Datacenter: All nodes in Rack 1 🏢 Rack 1 (SINGLE POINT OF FAILURE!) N1 Copy 1 N2 Copy 2 N3 Copy 3 💥 ✅ With Rack Awareness Datacenter: Replicas across 3 racks Rack 1 N1 Copy 1 Rack 2 N2 Copy 2 Rack 3 N3 Copy 3 🛡️ Failure Scenario: Power Failure in Rack 1 ❌ No Rack Awareness Rack 1 loses power → N1, N2, N3 all down → All 3 replicas lost → 💀 DATA UNAVAILABLE! ✅ With Rack Awareness Rack 1 loses power → Only N1 down → N2 & N3 still available → ✅ QUORUM maintained!

Configuring Rack Awareness

# cassandra-rackdc.properties file on each node
# Located at: /etc/cassandra/cassandra-rackdc.properties

# Node 1 configuration:
dc=US-East
rack=rack1

# Node 2 configuration:
dc=US-East
rack=rack2

# Node 3 configuration:
dc=US-East
rack=rack3

# Node 4 configuration:
dc=US-East
rack=rack1  # Can reuse rack names

# AWS Example: Map to Availability Zones
# Node in us-east-1a:
dc=US-East
rack=us-east-1a

# Node in us-east-1b:
dc=US-East
rack=us-east-1b

# Node in us-east-1c:
dc=US-East
rack=us-east-1c

# Azure Example: Map to Availability Zones
dc=East-US
rack=zone-1

# Verify configuration:
$ nodetool status

Datacenter: US-East
Status=Up/Down
|/ State=Normal/Leaving/Joining/Moving
--  Address      Rack    Status  State   Load
UN  10.1.0.1     rack1   Up      Normal  245 GB
UN  10.1.0.2     rack2   Up      Normal  238 GB
UN  10.1.0.3     rack3   Up      Normal  251 GB
                 ↑ Different racks = protected!

# Check replica placement:
$ nodetool getendpoints production users 12345
10.1.0.1   # rack1
10.1.0.2   # rack2
10.1.0.3   # rack3
# Perfect! All different racks

Cloud Provider Best Practices

AWS: Use Availability Zones as racks (us-east-1a, us-east-1b, us-east-1c)

Azure: Use Availability Zones as racks

GCP: Use Zones as racks (us-central1-a, us-central1-b, us-central1-c)

This ensures replicas are in different physical locations within the datacenter!

⚖️ Strategy Comparison

Complete Strategy Comparison Feature SimpleReplicationStrategy NetworkTopologyStrategy Datacenter Awareness ❌ No ✅ Yes Rack Awareness ❌ No ✅ Yes Multi-DC Support ❌ No ✅ Yes RF Per Datacenter ❌ Global only ✅ Per DC Placement Algorithm Simple clockwise Smart DC/rack-aware Fault Tolerance Poor (rack failure = loss) Excellent (survives DC/rack) Configuration Complexity Simple (1 RF value) Medium (RF per DC) LOCAL_QUORUM Support ❌ No ✅ Yes Future Expansion Difficult (must migrate) Easy (add DCs anytime) Production Ready ❌ NO ✅ YES Use Case Development/testing only ALL production Status Deprecated Recommended

Critical Decision

Always use NetworkTopologyStrategy! Even if you have a single datacenter today, NTS provides rack awareness and makes future expansion trivial. SimpleReplicationStrategy is deprecated and will be removed in future versions.

🔄 Migrating from Simple to NetworkTopologyStrategy

If you're currently using SimpleReplicationStrategy, here's how to safely migrate to NetworkTopologyStrategy!

Step-by-Step Migration Process

# STEP 1: Check current strategy
cqlsh> DESCRIBE KEYSPACE old_keyspace;

CREATE KEYSPACE old_keyspace WITH replication = {
    'class': 'SimpleReplicationStrategy',
    'replication_factor': '3'
};

# STEP 2: Find datacenter names
$ nodetool status

Datacenter: datacenter1   ← This is your DC name!
Status=Up/Down
|/ State=Normal/Leaving/Joining/Moving
--  Address    Status State   Load
UN  10.1.0.1   Up     Normal  245 GB
UN  10.1.0.2   Up     Normal  238 GB
UN  10.1.0.3   Up     Normal  251 GB

# STEP 3: Alter keyspace to NetworkTopologyStrategy
# IMPORTANT: Use the SAME RF value initially!
ALTER KEYSPACE old_keyspace WITH replication = {
    'class': 'NetworkTopologyStrategy',
    'datacenter1': 3
};
-- Changed from SimpleReplicationStrategy RF=3
-- to NetworkTopologyStrategy datacenter1=3
-- No data movement yet!

# STEP 4: Run repair on ALL nodes (CRITICAL!)
# This redistributes data based on new strategy

# On Node 1:
$ nodetool repair -full old_keyspace

# On Node 2:
$ nodetool repair -full old_keyspace

# On Node 3:
$ nodetool repair -full old_keyspace

# Monitor repair progress:
$ nodetool compactionstats
pending tasks: 0
Active compaction remaining time :        n/a

# STEP 5: Verify new replica placement
$ nodetool getendpoints old_keyspace users 12345

10.1.0.1
10.1.0.2
10.1.0.3

# Should now show rack-aware placement!

# STEP 6: Verify no issues
$ nodetool status
# All nodes should be UN (Up Normal)

# Done! No downtime required!

Migration Best Practices

  • Maintain Same RF: Don't change RF during strategy migration
  • Run Full Repair: Essential to redistribute data correctly
  • Repair All Nodes: Not just coordinator nodes
  • No Downtime: Clients can keep running during migration
  • Monitor Carefully: Watch for any UnavailableExceptions
  • Off-Peak Hours: Repair can be resource-intensive

Common Migration Scenarios

# Scenario 1: Single DC with rack awareness
# Before:
'class': 'SimpleReplicationStrategy',
'replication_factor': '3'

# After:
'class': 'NetworkTopologyStrategy',
'datacenter1': 3

# Scenario 2: Adding second datacenter
# Before (single DC with NTS):
'class': 'NetworkTopologyStrategy',
'US-East': 3

# After (adding US-West):
'class': 'NetworkTopologyStrategy',
'US-East': 3,
'US-West': 3  # New datacenter added!

# No repair needed for new DC (data streams automatically)

# Scenario 3: Converting to multi-DC from Simple
# This requires TWO steps:

# Step 1: Convert to NTS
ALTER KEYSPACE myks WITH replication = {
    'class': 'NetworkTopologyStrategy',
    'US-East': 3
};
# Run repair on all US-East nodes

# Step 2: Add second DC (later)
ALTER KEYSPACE myks WITH replication = {
    'class': 'NetworkTopologyStrategy',
    'US-East': 3,
    'US-West': 3
};
# Data automatically streams to US-West nodes

✅ Production Best Practices

1️⃣

Always Use NetworkTopologyStrategy

Even for: Single datacenter
Benefit: Rack awareness
Future-proof: Easy DC expansion
No downside: Same performance

2️⃣

Configure Rack Awareness

AWS: Map to Availability Zones
Azure/GCP: Map to Zones
Physical DC: Map to actual racks
Result: Survive rack failures

3️⃣

Use Meaningful Datacenter Names

Good: US-East, US-West, EU-West
Bad: datacenter1, dc2, test
Why: Clear, maintainable
Naming: Match cloud regions

4️⃣

Different RF Per DC

Primary DC: RF=3 (high availability)
Secondary DC: RF=2-3 (backup)
Read-only DC: RF=1-2 (cost savings)
Optimize: Balance cost vs availability

5️⃣

Plan for Geo-Distribution

Latency: Users query local DC
Compliance: Data residency rules
Disaster Recovery: Survive region failure
Example: 3 DCs, RF=3 each

6️⃣

Document Your Strategy

Track: DC names, RF values
Document: Rack mappings
Monitor: Replication health
Alert: Strategy changes

Real Configuration: Apple's Global Setup

# Apple's reported Cassandra configuration
# 75,000+ nodes across multiple regions worldwide

CREATE KEYSPACE icloud_data WITH replication = {
    'class': 'NetworkTopologyStrategy',
    
    # Primary regions (full redundancy)
    'US-East': 3,
    'US-West': 3,
    'EU-West': 3,
    
    # Secondary regions (cost-optimized)
    'Asia-Pacific': 2,
    'South-America': 2
};

# Benefits:
# ✓ 13 total copies globally (3+3+3+2+2)
# ✓ Survive entire continent outage
# ✓ <10ms latency for users (query local DC)
# ✓ Compliance with EU data residency
# ✓ 99.999% availability (5 nines!)
# ✓ Serves 1+ billion users worldwide

# Consistency Level:
# Writes: LOCAL_QUORUM (fast, available)
# Reads: LOCAL_QUORUM (fast, consistent within DC)
# Critical Writes: EACH_QUORUM (slow but globally consistent)

# Cost:
# Higher due to 13 copies, but availability is priceless for iCloud!

💼 Top 5 Interview Questions

1
What is the difference between SimpleReplicationStrategy and NetworkTopologyStrategy?
+

Answer:

SimpleReplicationStrategy:

  • Placement: Simple clockwise walk around token ring
  • Awareness: No rack or datacenter awareness
  • Configuration: Single RF value for entire cluster
  • Use Case: Development/testing only (deprecated)
  • Example: 'class': 'SimpleReplicationStrategy', 'replication_factor': 3

NetworkTopologyStrategy:

  • Placement: Smart DC/rack-aware algorithm
  • Awareness: Full rack and datacenter awareness
  • Configuration: Different RF per datacenter
  • Use Case: ALL production deployments (recommended)
  • Example: 'class': 'NetworkTopologyStrategy', 'US-East': 3, 'US-West': 2

Key Differences:

SimpleReplicationStrategy:
- Places replicas on next N-1 consecutive nodes clockwise
- Could put all replicas in same rack/DC (dangerous!)
- Cannot handle multiple datacenters
- No LOCAL_QUORUM support

NetworkTopologyStrategy:
- Places replicas in different racks within each DC
- Ensures fault tolerance at rack and DC level
- Supports multiple datacenters with different RF
- Full LOCAL_QUORUM support

Real Scenario: With SimpleReplicationStrategy, if Rack 1 fails and all 3 replicas were there → complete data loss! With NetworkTopologyStrategy, replicas are in different racks → survive rack failure.

Recommendation: Always use NetworkTopologyStrategy, even for single DC!

2
Why should you use NetworkTopologyStrategy even for a single datacenter?
+

Answer: Even with a single datacenter, NetworkTopologyStrategy provides critical advantages:

1. Rack Awareness (Primary Benefit)

Without Rack Awareness (SimpleReplicationStrategy):
- All 3 replicas could be in same physical rack
- Rack power failure → all replicas lost → DATA LOSS
- Network switch failure → same result

With Rack Awareness (NetworkTopologyStrategy):
- Replicas automatically placed in different racks
- Rack 1 fails → Replicas on Rack 2 & 3 still available
- Survives rack-level failures (power, network, etc.)

2. Future-Proof Architecture

  • Adding a second datacenter later is trivial (just ALTER keyspace)
  • No migration pain from SimpleReplicationStrategy → NetworkTopologyStrategy
  • Planning for growth from day one

3. Cloud Provider Integration

# AWS Example: Map racks to Availability Zones
# cassandra-rackdc.properties on each node:

# Node in us-east-1a:
dc=US-East
rack=us-east-1a

# Node in us-east-1b:
dc=US-East
rack=us-east-1b

# Node in us-east-1c:
dc=US-East
rack=us-east-1c

# Result: Replicas spread across 3 AZs
# Survives entire AZ failure!

4. Zero Performance Penalty

  • Same performance as SimpleReplicationStrategy
  • No additional latency
  • No extra resource usage
  • Only benefit, no downside!

5. Industry Standard

  • Used by Netflix, Uber, Apple, Instagram
  • SimpleReplicationStrategy is deprecated
  • Future Cassandra versions may remove SimpleReplicationStrategy

Configuration:

-- Single DC with NetworkTopologyStrategy
CREATE KEYSPACE production WITH replication = {
    'class': 'NetworkTopologyStrategy',
    'datacenter1': 3  -- or your actual DC name like 'US-East'
};

Bottom Line: NetworkTopologyStrategy gives you rack awareness and future flexibility with zero downside. Always use it!

3
How does NetworkTopologyStrategy place replicas across multiple datacenters?
+

Answer: NetworkTopologyStrategy uses an intelligent algorithm to ensure replicas are well-distributed:

Algorithm Steps:

1. Calculate token: hash(partition_key) = token

2. Find primary node: Node that owns this token

3. For EACH datacenter:
   a. Set RF from keyspace definition (e.g., US-East: 3)
   b. Start with primary node in this DC
   c. Walk clockwise around ring within this DC
   d. For each replica:
      - Find next node in different rack
      - Skip nodes in same rack as existing replicas
      - Add to replica set
   e. Repeat until RF replicas placed in this DC

4. Result: RF replicas per DC, each in different rack

Example Configuration:

CREATE KEYSPACE global WITH replication = {
    'class': 'NetworkTopologyStrategy',
    'US-East': 3,    # 3 replicas in US-East
    'US-West': 3,    # 3 replicas in US-West
    'EU-West': 2     # 2 replicas in EU-West
};

# For partition key user_id=12345:
# Token: -8234719023847192847

# Replicas:
US-East:
  - Node 1 (Rack 1, primary in US-East)
  - Node 2 (Rack 2, different rack)
  - Node 3 (Rack 3, different rack)

US-West:
  - Node 7 (Rack 1, primary in US-West)
  - Node 8 (Rack 2, different rack)
  - Node 9 (Rack 3, different rack)

EU-West:
  - Node 12 (Rack 1, primary in EU-West)
  - Node 13 (Rack 2, different rack)

# Total: 8 replicas across 3 datacenters!

Rack Selection Logic:

# In US-East datacenter with 6 nodes across 3 racks:
Rack 1: Node 1, Node 4
Rack 2: Node 2, Node 5
Rack 3: Node 3, Node 6

# Placing 3 replicas (RF=3):
Replica 1: Node 1 (Rack 1) - primary
Replica 2: Node 2 (Rack 2) - walk clockwise, different rack ✓
Replica 3: Node 3 (Rack 3) - walk clockwise, different rack ✓

# Node 4 skipped (same rack as Node 1)
# Node 5 skipped (same rack as Node 2)
# All 3 replicas in different racks!

Benefits:

  • Rack Isolation: Each DC's replicas in different racks
  • DC Isolation: Each DC has complete copy (if RF ≥ 1)
  • Fault Tolerance: Survive entire DC failure
  • Performance: LOCAL_QUORUM queries stay within DC (low latency)

Verification:

$ nodetool getendpoints global users 12345

10.1.0.1   # US-East, Rack 1
10.1.0.2   # US-East, Rack 2
10.1.0.3   # US-East, Rack 3
10.2.0.1   # US-West, Rack 1
10.2.0.2   # US-West, Rack 2
10.2.0.3   # US-West, Rack 3
10.3.0.1   # EU-West, Rack 1
10.3.0.2   # EU-West, Rack 2

# Perfect! Well-distributed across DCs and racks
4
How do you migrate from SimpleReplicationStrategy to NetworkTopologyStrategy safely?
+

Answer: Migration requires careful planning but zero downtime:

Step-by-Step Process:

STEP 1: Identify current configuration
--------------------------------
$ cqlsh -e "DESCRIBE KEYSPACE my_keyspace"

CREATE KEYSPACE my_keyspace WITH replication = {
    'class': 'SimpleReplicationStrategy',
    'replication_factor': '3'
};

STEP 2: Find datacenter name
----------------------------
$ nodetool status

Datacenter: datacenter1  ← Your DC name
Status=Up/Down
|/ State=Normal/Leaving/Joining/Moving
--  Address    Rack    Status  State   Load
UN  10.1.0.1   rack1   Up      Normal  245 GB
UN  10.1.0.2   rack2   Up      Normal  238 GB
UN  10.1.0.3   rack3   Up      Normal  251 GB

STEP 3: ALTER keyspace (keep same RF!)
--------------------------------------
ALTER KEYSPACE my_keyspace WITH replication = {
    'class': 'NetworkTopologyStrategy',
    'datacenter1': 3  ← Same RF as before!
};

# Important: Don't change RF during migration!
# Schema change is instant, no data movement yet

STEP 4: Run full repair on ALL nodes
------------------------------------
# This is CRITICAL - redistributes data based on new strategy

# Node 1:
$ nodetool repair -full my_keyspace
[2024-12-26 10:30:00] Starting repair...
[2024-12-26 10:45:00] Repair completed

# Node 2:
$ nodetool repair -full my_keyspace

# Node 3:
$ nodetool repair -full my_keyspace

# Monitor repair:
$ watch nodetool compactionstats

STEP 5: Verify replica placement
--------------------------------
$ nodetool getendpoints my_keyspace users 12345

10.1.0.1  # rack1
10.1.0.2  # rack2
10.1.0.3  # rack3

# Success! Now in different racks

STEP 6: Verify cluster health
-----------------------------
$ nodetool status
# All nodes should be UN (Up Normal)

$ nodetool describering my_keyspace
# Check token ranges look correct

Key Points:

  • No Downtime: Clients can keep running throughout
  • Keep Same RF: Don't change RF during strategy migration
  • Repair is Mandatory: Without repair, data may not be rack-aware
  • Repair All Nodes: Not just coordinator or primary nodes
  • Off-Peak Recommended: Repair can be resource-intensive

Common Pitfalls to Avoid:

  • ❌ Changing RF and strategy at same time
  • ❌ Forgetting to run repair (critical step!)
  • ❌ Only repairing some nodes
  • ❌ Not verifying replica placement after

Timeline:

ALTER keyspace:    Instant (milliseconds)
Full repair:       Hours (depends on data size)
  - 100GB/node:    ~30 minutes
  - 500GB/node:    ~2 hours
  - 1TB+/node:     ~4+ hours

Total migration:   3-6 hours for typical cluster
5
What happens if you configure different RF values per datacenter, and how does it affect consistency?
+

Answer: Different RF per datacenter allows cost optimization but requires careful consistency level management:

Configuration Example:

CREATE KEYSPACE mixed_rf WITH replication = {
    'class': 'NetworkTopologyStrategy',
    'US-East': 3,    # Primary: full redundancy
    'US-West': 2,    # Secondary: backup
    'EU-West': 1     # Read-only: disaster recovery only
};

# Total: 6 copies of every row across 3 datacenters

How It Works:

  • US-East: 3 copies → QUORUM=2, can tolerate 1 failure
  • US-West: 2 copies → QUORUM=2, cannot tolerate any failure
  • EU-West: 1 copy → QUORUM=1, no redundancy

Consistency Level Implications:

CL US-East (RF=3) US-West (RF=2) EU-West (RF=1)
LOCAL_ONE ✅ Works ✅ Works ✅ Works
LOCAL_QUORUM ✅ Works (need 2) ✅ Works (need 2) ✅ Works (need 1)
QUORUM ✅ Works (need 4 of 6 total)
EACH_QUORUM ✅ Need 2 ✅ Need 2 ✅ Need 1

Best Practices:

# Recommended CL choices with mixed RF:

# For writes:
INSERT INTO users VALUES (...) 
USING CONSISTENCY LOCAL_QUORUM;
# Fast, available within each DC

# For reads (local users):
SELECT * FROM users WHERE id=?
USING CONSISTENCY LOCAL_QUORUM;
# Fast, consistent within DC

# For critical global consistency:
INSERT INTO financial_transactions VALUES (...)
USING CONSISTENCY EACH_QUORUM;
# Slower, but guaranteed consistency across all DCs

Trade-offs:

  • Cost Savings: Lower RF in secondary DCs reduces storage (EU-West: 1 copy vs 3)
  • Fault Tolerance: US-East (RF=3) can lose 1 node, EU-West (RF=1) cannot lose any
  • Performance: LOCAL_QUORUM always works well regardless of RF
  • Availability: Lower RF = lower availability in that DC

Real Scenario:

# Uber's Configuration (approximate):
CREATE KEYSPACE trips WITH replication = {
    'class': 'NetworkTopologyStrategy',
    'US-East': 3,      # Primary: high traffic
    'US-West': 3,      # Primary: high traffic
    'EU-West': 2,      # Secondary: moderate traffic
    'Asia-Pacific': 2  # Secondary: growing market
};

# Write CL: LOCAL_QUORUM
# - In US-East: wait for 2 of 3
# - In EU-West: wait for 2 of 2 (both required!)

# If 1 node fails in US-East: Still works (2 remain for QUORUM)
# If 1 node fails in EU-West: Fails (only 1 remains, need 2!)

# Solution for EU-West: Use CL=ONE for higher availability
# Trade-off: Slightly less consistency, but service stays up
Advertisement

Responsive Ad