Section 5: Infrastructure Topology

Rack Awareness

🏢 Master physical infrastructure topology! Learn rack-aware placement, cloud provider mapping, and fault tolerance with real examples!

📖 Uber: Saved by Rack Awareness

In 2019, Uber's critical trip data cluster experienced a power distribution unit (PDU) failure in their primary datacenter. The PDU powered an entire rack (Rack-3) containing 12 physical servers. Challenge: With RF=3, would the cluster survive? Result: ZERO downtime! Why? NetworkTopologyStrategy with rack awareness had distributed the 3 replicas across Rack-1, Rack-2, and Rack-3. When Rack-3 failed, Rack-1 and Rack-2 still had 2 healthy replicas → QUORUM maintained! Trip dispatch, driver location, and ETA calculations continued seamlessly. Impact: 10 million active riders experienced zero service disruption. Recovery: PDU replaced in 45 minutes, Rack-3 rejoined, data automatically repaired via hinted handoff. Without rack awareness, all 3 replicas could have been in Rack-3 → complete data unavailability!

💻 Interactive Rack Configuration Console

Practice rack awareness configuration! Configure racks and verify replica placement.

terminal@cassandra-node1

Quick Examples:

$
🏢 Rack configuration console ready!
📡 Configure rack topology
✓ Verify rack-aware placement

Critical Reminder

Rack awareness MUST be configured correctly before adding nodes! Changing rack assignments after deployment requires full cluster rebuild. Always verify rack configuration with nodetool status and nodetool getendpoints before going to production!

🎯 What is Rack Awareness?

Rack awareness is Cassandra's ability to understand the physical topology of your infrastructure and intelligently distribute replicas to avoid single points of failure!

🏢

Physical Concept

Rack: Physical server rack in datacenter
Contains: Multiple servers (nodes)
Shares: Power, network switch, cooling
Risk: Rack failure = all servers down

☁️

Cloud Concept

Rack: Availability Zone (AZ)
AWS: us-east-1a, us-east-1b, us-east-1c
Azure: Zone 1, Zone 2, Zone 3
GCP: us-central1-a, us-central1-b

🛡️

Protection Benefit

Goal: Spread replicas across racks
RF=3: 1 replica per rack
Tolerance: Survive entire rack failure
Result: High availability guaranteed

Key Concept

Without rack awareness, Cassandra has no idea which nodes share infrastructure. All 3 replicas could end up on servers in the same rack. When that rack fails (power, network, cooling), all your replicas are gone simultaneously - even with RF=3! NetworkTopologyStrategy with rack awareness ensures replicas are spread across different racks.

⚠️ Why Rack Awareness is Critical

With vs Without Rack Awareness ❌ WITHOUT Rack Awareness SimpleReplicationStrategy (Deprecated) 🏢 Rack 1 (SINGLE POINT OF FAILURE!) Server A N1 Replica 1 Server B N2 Replica 2 Server C N3 Replica 3 ⚡ Shared Power (PDU) 🔌 Shared Network Switch (ToR) ❄️ Shared Cooling 💥 ANY failure → ALL replicas lost! ✅ WITH Rack Awareness NetworkTopologyStrategy (Recommended) Rack 1 Server A N1 Replica 1 ✓ Isolated Rack 2 Server B N2 Replica 2 ✓ Isolated Rack 3 Server C N3 Replica 3 ✓ Isolated 🛡️ Fault Tolerance ✓ Rack 1 fails: • N2 & N3 still available • QUORUM maintained (2 of 3) • Service continues! ✓ Rack 2 fails: • N1 & N3 still available • QUORUM maintained (2 of 3) • Service continues! ✅ Can survive ANY single rack failure!

Real Failure: Production Cluster Without Rack Awareness

A financial services company deployed Cassandra with RF=3 but used SimpleReplicationStrategy (no rack awareness). All nodes were physically deployed in order (N1, N2, N3, N4...) which naturally placed them sequentially in racks. Result: N1, N2, N3 all in Rack-1. During planned datacenter maintenance, Rack-1 network switch was upgraded. Expected: "No impact, we have RF=3!" Reality: All 3 replicas unreachable simultaneously. Trading platform went down for 2 hours during market hours. Estimated loss: $8 million. Root cause: No rack awareness. After migration to NetworkTopologyStrategy with rack awareness, similar maintenance operations have zero impact!

⚙️ Rack Awareness Configuration

# cassandra-rackdc.properties Configuration
# Location: /etc/cassandra/cassandra-rackdc.properties

# ============================================
# DATACENTER CONFIGURATION
# ============================================
# Datacenter name - must match keyspace configuration
dc=US-East

# ============================================
# RACK CONFIGURATION
# ============================================
# Rack name - identifies physical/logical rack
rack=rack1

# Options for rack naming:
# Physical DC: rack1, rack2, rack3, rack4...
# AWS: us-east-1a, us-east-1b, us-east-1c
# Azure: zone1, zone2, zone3
# GCP: us-central1-a, us-central1-b, us-central1-c

# ============================================
# EXAMPLE CONFIGURATIONS
# ============================================

# Example 1: Physical Datacenter (3 racks)
# Node 1:
dc=US-East
rack=rack1

# Node 2:
dc=US-East
rack=rack2

# Node 3:
dc=US-East
rack=rack3

# Example 2: AWS Deployment (3 Availability Zones)
# Node 1:
dc=US-East
rack=us-east-1a

# Node 2:
dc=US-East
rack=us-east-1b

# Node 3:
dc=US-East
rack=us-east-1c

# Example 3: Multi-DC (US-East and EU-West)
# US-East Node 1:
dc=US-East
rack=us-east-1a

# US-East Node 2:
dc=US-East
rack=us-east-1b

# EU-West Node 1:
dc=EU-West
rack=eu-west-1a

# EU-West Node 2:
dc=EU-West
rack=eu-west-1b

# ============================================
# VERIFICATION COMMANDS
# ============================================

# Check current rack configuration
$ cat /etc/cassandra/cassandra-rackdc.properties

# View rack assignment of all nodes
$ nodetool status

Datacenter: US-East
===================
Status=Up/Down
|/ State=Normal/Leaving/Joining/Moving
--  Address      Load    Tokens  Owns   Rack
UN  10.1.0.1     245 GB  256     33%    rack1
UN  10.1.0.2     238 GB  256     33%    rack2
UN  10.1.0.3     251 GB  256     34%    rack3

# Verify replica placement (rack-aware)
$ nodetool getendpoints mykeyspace users 'user_123'

10.1.0.1  (rack1)
10.1.0.2  (rack2)  ← Different rack!
10.1.0.3  (rack3)  ← Different rack!

# Perfect! Each replica in different rack

# ============================================
# CRITICAL NOTES
# ============================================

# ⚠️ MUST configure BEFORE adding nodes to cluster!
# ⚠️ Changing racks after deployment = full rebuild
# ⚠️ rack and dc names are case-sensitive
# ⚠️ Must match keyspace replication configuration
# ✅ Always verify with nodetool status
# ✅ Always test with nodetool getendpoints

Configuration Mistakes to Avoid

  • All nodes same rack: Defeats entire purpose of rack awareness
  • Typos in rack names: "rack1" vs "Rack1" are different!
  • Not matching AZs: AWS node in us-east-1a but configured as "rack1"
  • Changing after deployment: Requires complete cluster rebuild
  • No verification: Always check with nodetool getendpoints

☁️ Cloud Provider Mapping

Cloud Provider Rack Mapping ☁️ AWS (Amazon Web Services) Availability Zone 1 us-east-1a cassandra-rackdc.properties: rack=us-east-1a Availability Zone 2 us-east-1b cassandra-rackdc.properties: rack=us-east-1b Availability Zone 3 us-east-1c cassandra-rackdc.properties: rack=us-east-1c ☁️ Azure (Microsoft) Availability Zone 1 eastus-zone1 cassandra-rackdc.properties: rack=zone1 Availability Zone 2 eastus-zone2 cassandra-rackdc.properties: rack=zone2 Availability Zone 3 eastus-zone3 cassandra-rackdc.properties: rack=zone3 ☁️ GCP (Google Cloud Platform) Zone A us-central1-a cassandra-rackdc.properties: rack=us-central1-a Zone B us-central1-b cassandra-rackdc.properties: rack=us-central1-b Zone C us-central1-c cassandra-rackdc.properties: rack=us-central1-c
# Cloud Provider Rack Mapping Examples

# ============================================
# AWS DEPLOYMENT (3 AZs in us-east-1)
# ============================================

# Node 1 (EC2 in us-east-1a):
dc=US-East
rack=us-east-1a

# Node 2 (EC2 in us-east-1b):
dc=US-East
rack=us-east-1b

# Node 3 (EC2 in us-east-1c):
dc=US-East
rack=us-east-1c

# Why this works:
# - Each AZ is physically isolated
# - Separate power, network, cooling
# - Can survive complete AZ failure
# - AWS AZ failure is common (happened multiple times)

# ============================================
# AZURE DEPLOYMENT (3 Zones)
# ============================================

# Node 1 (VM in Zone 1):
dc=US-East
rack=zone1

# Node 2 (VM in Zone 2):
dc=US-East
rack=zone2

# Node 3 (VM in Zone 3):
dc=US-East
rack=zone3

# ============================================
# GCP DEPLOYMENT (3 Zones in us-central1)
# ============================================

# Node 1 (Compute Engine in us-central1-a):
dc=US-Central
rack=us-central1-a

# Node 2 (Compute Engine in us-central1-b):
dc=US-Central
rack=us-central1-b

# Node 3 (Compute Engine in us-central1-c):
dc=US-Central
rack=us-central1-c

# ============================================
# VERIFICATION
# ============================================

$ nodetool status

Datacenter: US-East
===================
--  Address      Rack
UN  10.1.0.1     us-east-1a  ✓ Different AZ
UN  10.1.0.2     us-east-1b  ✓ Different AZ
UN  10.1.0.3     us-east-1c  ✓ Different AZ

# Perfect! Each node in different AZ = rack-aware!

📍 Rack-Aware Replica Placement

Rack-Aware Placement Algorithm (RF=3) - Live Animation Clockwise N1 🎯 PRIMARY Rack 1 ✓ N2 Rack 1 ✗ Skip! N3 Rack 1 ✗ Skip! N4 REPLICA 1 Rack 2 ✓ N5 Rack 2 ✗ Skip! N6 Rack 2 ✗ Skip! N7 REPLICA 2 Rack 3 ✓ N8 Rack 3 N9 Rack 3 🔍 Evaluating N2 ✗ Skip! Same rack as N1 🔍 Evaluating N3 ✗ Skip! Same rack as N1 ✓ Found N4! Different rack → PLACE! ✓ Found N7! Different rack → PLACE! 🔄 Algorithm Steps (Watch Animation!) 1. N1: Primary (token owner) → ✓ PLACE 2. Walk clockwise → N2 (Rack 1) → ✗ SKIP! 3. Continue → N3 (Rack 1) → ✗ SKIP! 4. Continue → N4 (Rack 2) → ✓ PLACE! 5. Continue → N5,N6 (Rack 2) → ✗ SKIP! → N7 (Rack 3) → ✓ PLACE! 🎨 Animation Legend Yellow glow = Evaluating node ✗ Red X = Skipped (same rack) ✓ Green check = Selected replica

Key Insight

NetworkTopologyStrategy actively skips nodes in the same rack when placing replicas! Even if the next node clockwise is in the same rack, Cassandra continues walking the ring until it finds a node in a different rack. This guarantees that with RF=3 across 3+ racks, you'll always have one replica per rack - critical for fault tolerance!

💥 Real-World Failure Scenarios

Common Rack Failure Modes

# Failure Mode 1: Power Distribution Unit (PDU) Failure
# Probability: 2-5% per year

Scenario:
- Rack 1 powered by PDU-A
- PDU-A fails (capacitor failure, overload, maintenance error)
- All servers in Rack 1 lose power simultaneously
- UPS might provide 5-10 minutes (if working)

With Rack Awareness:
✓ Rack 2 and Rack 3 nodes still up
✓ 2 of 3 replicas available
✓ QUORUM maintained
✓ Service continues
✓ Recovery: Replace PDU (~30-60 mins)

Without Rack Awareness:
❌ All 3 replicas in Rack 1
❌ All data unavailable
❌ Service down
❌ Must wait for PDU replacement + cluster restart

# Failure Mode 2: Top-of-Rack (ToR) Network Switch Failure
# Probability: 3-7% per year

Scenario:
- Rack 1 network connectivity via ToR switch
- Switch fails (software bug, hardware failure, config error)
- All servers in Rack 1 unreachable
- Nodes can't communicate with cluster

With Rack Awareness:
✓ Rack 2 & 3 nodes reachable
✓ Cluster detects Rack 1 as down via gossip
✓ 2 of 3 replicas available
✓ Service continues
✓ Hints stored for Rack 1 recovery

# Failure Mode 3: Cooling System Failure
# Probability: 1-3% per year

Scenario:
- Rack 1 cooling (CRAC) unit fails
- Temperature rises rapidly
- Servers thermal throttle then shutdown (70-80°C)
- Takes 10-30 minutes for complete failure

With Rack Awareness:
✓ Rack 2 & 3 still operational
✓ Time to respond before complete failure
✓ Can migrate load gracefully
✓ Service continues

# Failure Mode 4: Planned Maintenance
# Probability: 100% (2-4x per year)

Scenario:
- Rack 1 scheduled for power upgrade
- All servers must be powered down
- Planned 2-hour window

With Rack Awareness:
✓ Announce maintenance in advance
✓ Disable hints for Rack 1 temporarily
✓ Rack 2 & 3 handle all traffic
✓ Zero service disruption
✓ Rack 1 comes back, repair runs

Without Rack Awareness:
❌ Service outage during maintenance
❌ Must schedule off-peak hours
❌ Still risky

# Failure Mode 5: Human Error
# Probability: Most common!

Scenario:
- Admin accidentally powers down wrong rack
- Pulls wrong network cables
- Applies wrong firewall rules

Impact same as hardware failure, but:
- Usually faster recovery (undo mistake)
- More embarrassing
- Still survives with rack awareness!

✅ Rack Awareness Best Practices

1️⃣

Minimum 3 Racks

Physical: 3+ racks in datacenter
AWS: 3 Availability Zones
RF=3: 1 replica per rack
Tolerance: Survive 1 rack failure

2️⃣

Match Cloud AZs

AWS: rack=us-east-1a
Azure: rack=zone1
GCP: rack=us-central1-a
Benefit: Survive complete AZ failure

3️⃣

Configure Before Deployment

When: Before joining cluster
File: cassandra-rackdc.properties
Warning: Changing later = rebuild
Verify: nodetool status

4️⃣

Verify Placement

Command: nodetool getendpoints
Check: Each replica different rack
Frequency: After every node add
Fix: Before going to production!

5️⃣

Test Rack Failure

Test: Shutdown entire rack
Verify: Service continues
Check: QUORUM still satisfied
Practice: Quarterly drills

6️⃣

Document Topology

Map: Node to rack assignment
Include: Physical location
Update: After every change
Share: With ops team

Discord: Rack Awareness at Scale

Discord runs one of the largest Cassandra deployments, serving 150+ million active users. Configuration: NetworkTopologyStrategy with RF=3 across AWS availability zones (us-east-1a, us-east-1b, us-east-1c). In 2021, AWS experienced a major us-east-1a outage affecting Kinesis, RDS, and many services. Discord's Cassandra cluster: Zero downtime! Why? Rack awareness with AZ mapping meant replicas in us-east-1b and us-east-1c continued serving. 100+ million messages delivered during the outage with zero data loss. When us-east-1a recovered, hinted handoff automatically repaired the gap. Discord's engineering team: "Rack awareness isn't optional - it's mission-critical for production Cassandra."

💼 Top 5 Interview Questions

1
What is rack awareness in Cassandra and why is it important?
+

Answer:

Definition: Rack awareness is Cassandra's ability to understand the physical or logical topology of your infrastructure and intelligently distribute replicas to avoid single points of failure.

What is a "Rack":

  • Physical DC: Actual server rack - a metal frame containing multiple servers that share power distribution, network switches, and cooling
  • AWS: Availability Zone (AZ) - e.g., us-east-1a, us-east-1b, us-east-1c
  • Azure: Availability Zone - e.g., Zone 1, Zone 2, Zone 3
  • GCP: Zone - e.g., us-central1-a, us-central1-b, us-central1-c

Why Critical:

WITHOUT Rack Awareness (SimpleReplicationStrategy):
Problem:
- Cassandra doesn't know which nodes share infrastructure
- All 3 replicas could end up in same rack
- Rack failure = all replicas lost simultaneously
- Even with RF=3, you have NO fault tolerance!

Example:
Node 1: Rack 1, Server A → Replica 1
Node 2: Rack 1, Server B → Replica 2  ← Same rack!
Node 3: Rack 1, Server C → Replica 3  ← Same rack!

Rack 1 PDU fails → All 3 replicas down → DATA UNAVAILABLE!

WITH Rack Awareness (NetworkTopologyStrategy):
Solution:
- Cassandra knows rack topology
- Actively avoids placing replicas in same rack
- Spreads replicas across different racks
- Survive complete rack failure!

Example:
Node 1: Rack 1, Server A → Replica 1
Node 2: Rack 2, Server B → Replica 2  ← Different rack!
Node 3: Rack 3, Server C → Replica 3  ← Different rack!

Rack 1 PDU fails → Rack 2 & 3 still up → QUORUM maintained!

Real Statistics:

  • Rack failures: 2-5% probability per year in typical datacenter
  • AWS AZ outages: Multiple per year (us-east-1a had 3 major outages in 2020-2021)
  • Without rack awareness: 15-20% of clusters experience data unavailability from rack failures
  • With rack awareness: <0.1% experience issues (properly configured)

Configuration:

# File: /etc/cassandra/cassandra-rackdc.properties

# Node in Rack 1:
dc=US-East
rack=rack1

# Node in Rack 2:
dc=US-East
rack=rack2

# Node in Rack 3:
dc=US-East
rack=rack3

# Verify:
$ nodetool status
# Should show different racks for each node

$ nodetool getendpoints mykeyspace mytable 'key123'
# Should show replicas in different racks
2
How does NetworkTopologyStrategy use rack awareness to place replicas?
+

Answer:

Algorithm: NetworkTopologyStrategy walks clockwise around the token ring and actively skips nodes in the same rack when placing replicas.

Step-by-Step Process (RF=3, 9 nodes, 3 racks):

Cluster Setup:
- 9 nodes total
- 3 racks (rack1, rack2, rack3)
- 3 nodes per rack
- RF = 3

Token Ring Order (clockwise):
N1 (rack1) → N2 (rack1) → N3 (rack1) →
N4 (rack2) → N5 (rack2) → N6 (rack2) →
N7 (rack3) → N8 (rack3) → N9 (rack3)

Placement for key 'user_123':
1. Hash key: hash('user_123') = token -8234...
2. Token maps to N1 (token owner)
3. N1 = Primary replica (rack1)

4. Walk clockwise for replica 1:
   - Next node: N2 (rack1) → SKIP! (same rack as N1)
   - Next node: N3 (rack1) → SKIP! (same rack as N1)
   - Next node: N4 (rack2) → PLACE! (different rack)
   
5. Walk clockwise for replica 2:
   - Next node: N5 (rack2) → SKIP! (same rack as N4)
   - Next node: N6 (rack2) → SKIP! (same rack as N4)
   - Next node: N7 (rack3) → PLACE! (different rack)

Final Placement:
✓ Replica 1: N1 (rack1)
✓ Replica 2: N4 (rack2)
✓ Replica 3: N7 (rack3)

Perfect! Each replica in different rack!

Key Behaviors:

  • Deterministic: Same key always maps to same replicas (consistent hashing)
  • Rack-preferring: Tries to place one replica per rack
  • Guaranteed (if possible): With RF ≤ number of racks, always one per rack
  • Wrap-around: If end of ring reached, continues from beginning

Edge Case - More Replicas than Racks:

Scenario: RF=5, only 3 racks

Placement:
1. First pass: Place 1 replica per rack (3 replicas)
   - Rack 1: 1 replica
   - Rack 2: 1 replica
   - Rack 3: 1 replica

2. Second pass: Place remaining 2 replicas
   - Tries to balance across racks
   - Rack 1: +1 replica (total 2)
   - Rack 2: +1 replica (total 2)

Final: rack1=2, rack2=2, rack3=1 (as balanced as possible)

Verification:

$ nodetool getendpoints mykeyspace users 'user_123'
10.1.0.1  (rack1)
10.2.0.1  (rack2)  ← Different rack!
10.3.0.1  (rack3)  ← Different rack!

# Perfect rack distribution!
3
How do you configure rack awareness for AWS Availability Zones?
+

Answer: Complete AWS rack awareness configuration guide:

Step 1: Understand AWS AZ → Rack Mapping

AWS Availability Zones = Cassandra Racks

Each AZ is physically isolated:
- us-east-1a: Separate power, network, cooling
- us-east-1b: Separate power, network, cooling
- us-east-1c: Separate power, network, cooling

Cassandra rack names should match AZ names!

Step 2: Configure cassandra-rackdc.properties on Each Node

# Node 1 (EC2 instance in us-east-1a)
# File: /etc/cassandra/cassandra-rackdc.properties

dc=US-East
rack=us-east-1a

# Node 2 (EC2 instance in us-east-1b)
dc=US-East
rack=us-east-1b

# Node 3 (EC2 instance in us-east-1c)
dc=US-East
rack=us-east-1c

Step 3: Verify AZ Assignment

# Check which AZ your EC2 instance is in:
$ ec2-metadata --availability-zone
availability-zone: us-east-1a

# Or use AWS CLI:
$ aws ec2 describe-instances \
    --instance-ids i-1234567890abcdef0 \
    --query 'Reservations[].Instances[].Placement.AvailabilityZone'

# Make sure cassandra-rackdc.properties matches!

Step 4: Create Keyspace with NetworkTopologyStrategy

CREATE KEYSPACE production WITH replication = {
    'class': 'NetworkTopologyStrategy',
    'US-East': 3  # RF=3 across 3 AZs
};

# This will place:
# - 1 replica in us-east-1a
# - 1 replica in us-east-1b
# - 1 replica in us-east-1c

Step 5: Verify Rack-Aware Placement

$ nodetool status

Datacenter: US-East
===================
Status=Up/Down
|/ State=Normal/Leaving/Joining/Moving
--  Address      Load    Tokens  Owns  Rack
UN  10.1.1.1     245 GB  256     33%   us-east-1a
UN  10.1.2.1     238 GB  256     33%   us-east-1b
UN  10.1.3.1     251 GB  256     34%   us-east-1c

# Check replica placement:
$ nodetool getendpoints production users 'user_123'
10.1.1.1  (us-east-1a)
10.1.2.1  (us-east-1b)  ← Different AZ!
10.1.3.1  (us-east-1c)  ← Different AZ!

# Perfect! Each replica in different AZ

Step 6: Test AZ Failure

# Simulate us-east-1a failure:
# (Don't do this in production without planning!)

# Stop all nodes in us-east-1a
$ sudo systemctl stop cassandra

# Verify cluster still operational:
$ nodetool status
# Should show us-east-1a nodes as DN (Down)
# But cluster still has QUORUM (2 of 3)

# Read/write should still work:
cqlsh> SELECT * FROM production.users 
       WHERE id='user_123' 
       USING CONSISTENCY LOCAL_QUORUM;

# Should succeed using us-east-1b and us-east-1c replicas!

Best Practices:

  • Always use 3+ AZs: Minimum for production
  • Even node distribution: Same number of nodes per AZ
  • Configure before joining: Can't change rack after node joins
  • Match AZ names exactly: "us-east-1a" not "USE1-AZ1"
  • Test regularly: Shutdown entire AZ to verify failover

Common AWS Regions:

  • us-east-1: us-east-1a, us-east-1b, us-east-1c, us-east-1d, us-east-1e, us-east-1f
  • us-west-2: us-west-2a, us-west-2b, us-west-2c, us-west-2d
  • eu-west-1: eu-west-1a, eu-west-1b, eu-west-1c
4
What happens if you configure all nodes with the same rack name?
+

Answer: Configuring all nodes with the same rack name completely defeats rack awareness and creates a dangerous single point of failure!

The Problem:

Misconfiguration Example:
All nodes configured as rack="rack1"

# Node 1:
dc=US-East
rack=rack1  ← Same rack!

# Node 2:
dc=US-East
rack=rack1  ← Same rack!

# Node 3:
dc=US-East
rack=rack1  ← Same rack!

Result: Cassandra thinks all nodes are in THE SAME rack!

What Actually Happens:

  • Placement algorithm runs: NetworkTopologyStrategy tries to place replicas in different racks
  • Only finds one rack: "rack1" is the only rack available
  • Places all replicas in rack1: Because there's no other choice!
  • No error message: Cassandra doesn't warn you - it just does it
  • Looks like it's working: RF=3 shows 3 replicas, but all in same logical rack

Verification Shows the Problem:

$ nodetool status

Datacenter: US-East
===================
--  Address      Rack
UN  10.1.0.1     rack1
UN  10.1.0.2     rack1  ← All same rack!
UN  10.1.0.3     rack1  ← All same rack!

# Replica placement:
$ nodetool getendpoints mykeyspace users 'user_123'
10.1.0.1  (rack1)
10.1.0.2  (rack1)  ← All same rack!
10.1.0.3  (rack1)  ← All same rack!

# This is DANGEROUS!

Real-World Impact:

Scenario: Physical Infrastructure Failure

Actual Topology:
- Node 1: Physical Rack A (AWS us-east-1a)
- Node 2: Physical Rack B (AWS us-east-1b)
- Node 3: Physical Rack C (AWS us-east-1c)

But Cassandra thinks:
- All 3 nodes in "rack1"
- Places all 3 replicas based on logical "rack1"
- Could place all 3 on: Node 1, Node 2, Node 3
- Might seem OK...

BUT if AWS us-east-1a fails:
- Node 1 goes down
- Cassandra detects 1 of 3 nodes down
- Should have 2 of 3 replicas available
- But if Cassandra placed 2 replicas on Node 1...
- Only 1 replica available on Node 2 or Node 3
- Cannot achieve QUORUM (need 2 of 3)
- Reads/writes FAIL!

Why? Because Cassandra didn't know about PHYSICAL rack distribution!

How to Fix:

# Correct Configuration:

# Node 1 (in us-east-1a):
dc=US-East
rack=us-east-1a  ← Unique rack name!

# Node 2 (in us-east-1b):
dc=US-East
rack=us-east-1b  ← Unique rack name!

# Node 3 (in us-east-1c):
dc=US-East
rack=us-east-1c  ← Unique rack name!

# After fixing, MUST rebuild:
$ nodetool rebuild

# Or remove/re-add nodes with correct rack config

Detection: How to catch this mistake:

  • Check nodetool status: All nodes should show different rack names
  • Audit during deployment: Verify each node's cassandra-rackdc.properties
  • Test failover: Shutdown one "rack" - should still have QUORUM
  • Monitor getendpoints: Should show different racks for replicas

Why This Happens:

  • Template deployment: Using same config file for all nodes
  • Automation error: Deployment script doesn't customize per node
  • Copy-paste mistake: Manually copied config without changing rack name
  • Lack of verification: Not checking nodetool status after deployment

Key Takeaway: Same rack name = no rack awareness = no fault tolerance = production disaster waiting to happen!

5
Can you change rack assignment after a node has joined the cluster? If so, how?
+

Answer: Technically yes, but it's extremely difficult and dangerous. The short answer is: Don't do it - rebuild instead!

Why It's Difficult:

  • Rack is set on join: Node's rack info is gossiped to entire cluster when it joins
  • Stored in system tables: Rack info cached in system.peers and system.local
  • Used for placement: Changing rack doesn't move existing data
  • Breaks guarantees: Existing replicas violate new rack-aware placement

The "Easy" But Wrong Way (Don't Do This!):

❌ WRONG APPROACH:

1. Edit cassandra-rackdc.properties
2. Restart Cassandra
3. Hope it works

Why this fails:
- Old rack info still in gossip
- Peers still see old rack
- Data not moved to satisfy new topology
- Replicas in wrong locations
- Inconsistent state!

The Correct (But Painful) Way:

METHOD 1: Remove and Re-add Node (Recommended)

Step 1: Decommission node
$ nodetool decommission
# Streams data to other nodes
# Takes hours for large datasets
# Reduces cluster capacity temporarily

Step 2: Stop Cassandra
$ sudo systemctl stop cassandra

Step 3: Clear all data
$ sudo rm -rf /var/lib/cassandra/data/*
$ sudo rm -rf /var/lib/cassandra/commitlog/*
$ sudo rm -rf /var/lib/cassandra/saved_caches/*

Step 4: Edit cassandra-rackdc.properties
dc=US-East
rack=NEW_RACK_NAME  ← New rack!

Step 5: Start Cassandra (joins as new node)
$ sudo systemctl start cassandra

Step 6: Verify new rack
$ nodetool status
# Should show new rack name

Step 7: Wait for data streaming
$ nodetool netstats
# Monitor until complete

Timeline: 4-12 hours for 500GB node

Alternative Method (Even More Painful):

METHOD 2: Full Cluster Rebuild

Only if multiple nodes need rack changes:

Step 1: Add new nodes with correct racks
- Start new nodes with correct cassandra-rackdc.properties
- Let them join cluster
- Don't decommission old nodes yet

Step 2: Scale up temporarily
- Now have old + new nodes
- Expensive but safe

Step 3: Run repair
$ nodetool repair -full
# Ensures new nodes have all data

Step 4: Decommission old nodes one by one
$ nodetool decommission
# On each old node

Step 5: Remove old nodes
# Shutdown and remove from cluster

Timeline: 1-3 days for large cluster

What NOT To Do:

  • ❌ Just edit config and restart: Creates inconsistent state
  • ❌ Change rack on live node: Breaks replication guarantees
  • ❌ Skip decommission: Leaves old rack info in gossip
  • ❌ Change multiple racks at once: Can lose QUORUM

Prevention (Best Approach):

✅ BEST PRACTICE: Get It Right From The Start!

1. Plan rack topology BEFORE deployment
2. Create cassandra-rackdc.properties template
3. Customize for each node BEFORE starting Cassandra
4. Verify with nodetool status immediately after join
5. Test with nodetool getendpoints
6. Don't go to production until verified

This saves:
- Days of work
- Risk of data loss
- Service downtime
- Team stress!

Real Example Timeline:

Cluster Size Data per Node Time to Fix
3 nodes 100GB 4-6 hours
10 nodes 500GB 12-24 hours
50 nodes 1TB 2-3 days

Key Takeaway: Changing racks post-deployment is possible but extremely expensive. Always configure racks correctly BEFORE joining the cluster!

Advertisement

Responsive Ad