Cassandra Cluster Topology
Master the art of designing resilient, globally-distributed Cassandra clusters! Learn datacenter layouts, rack awareness, and multi-region architectures that power the world's largest applications.
📖 The Story: The Global Delivery Company
Meet GlobalShip - a worldwide package delivery company that operates on 5 continents. They need a system that never goes down, even during hurricanes, earthquakes, or power outages.
🌍 GlobalShip's Infrastructure Challenge
The Problem: One massive warehouse in New York won't work!
Why Not?
- Hurricane hits NYC: Entire business stops! ❌
- Power outage: Can't track any packages! ❌
- Latency: Tokyo customers wait 200ms for data from NYC! ❌
- Network failure: Single cable cut = total disaster! ❌
GlobalShip's Smart Solution: Geographic Distribution!
Level 1: DATACENTERS (Continents)
- DC1: New York, USA 🇺🇸
- DC2: London, UK 🇬🇧
- DC3: Tokyo, Japan 🇯🇵
Each datacenter = Complete warehouse with full operations!
Level 2: RACKS (Buildings within Datacenter)
- NYC Datacenter:
- - Rack A: Building in Manhattan
- - Rack B: Building in Brooklyn
- - Rack C: Building in Queens
Each rack = Separate building with own power/network!
Level 3: NODES (Servers in Each Rack)
- Rack A: Node 1, Node 2, Node 3
- Rack B: Node 4, Node 5, Node 6
- Rack C: Node 7, Node 8, Node 9
Each node = Individual server storing data!
Result: Hurricane destroys NYC? London and Tokyo keep running! 🎉
🎯 This is EXACTLY Cassandra's Cluster Topology!
- Datacenters = Geographic Regions (US-East, EU-West, Asia-Pacific)
- Racks = Fault Domains (separate power/network)
- Nodes = Cassandra Servers (actual machines)
- Replication = Copies across all levels (survive any failure!)
GlobalShip's physical warehouses = Cassandra's virtual datacenters!
Same principles, different scale!
🗺️ What is Cluster Topology?
The hierarchical layout of your Cassandra cluster across datacenters, racks, and nodes.
Simple Definition
Cluster Topology: The physical and logical organization of Cassandra nodes into a hierarchical structure of datacenters and racks to ensure high availability and fault tolerance.
The 3-Level Hierarchy:
- Level 1 - Cluster: The entire Cassandra system (all nodes worldwide)
- Level 2 - Datacenters: Geographic or logical groupings (e.g., US-East, EU-West)
- Level 3 - Racks: Fault-isolated groups within datacenters
- Level 4 - Nodes: Individual Cassandra servers
Why Topology Matters:
- Disaster Recovery: Survive datacenter outages
- Low Latency: Serve users from nearest datacenter
- Fault Isolation: Rack failures don't take down cluster
- Compliance: Keep data in specific regions (GDPR, etc.)
Key Insight
Topology is about FAILURE DOMAINS!
Each level protects against different types of failures:
- Datacenter level: Protects against regional disasters (earthquakes, hurricanes)
- Rack level: Protects against power/network failures in one area
- Node level: Protects against individual server failures
Cassandra automatically distributes replicas across these levels!
🏢 Datacenter Design Strategies
How to organize your cluster into logical or physical datacenters.
🌐 Netflix's 3-Datacenter Strategy
Scenario: Netflix streams to 200+ million users globally. They can't afford downtime!
Their Setup:
- DC1: US-EAST-1 (Virginia) - 1,000 nodes
- DC2: US-WEST-2 (Oregon) - 1,000 nodes
- DC3: EU-WEST-1 (Ireland) - 500 nodes
Replication Strategy:
'US-EAST-1': 3, // 3 copies in Virginia
'US-WEST-2': 3, // 3 copies in Oregon
'EU-WEST-1': 3 // 3 copies in Ireland
}
Total copies per key: 9 replicas!
Can lose ENTIRE datacenter and still serve!
Why It Works:
- Hurricane destroys Virginia → Oregon + Ireland continue!
- US users get <5ms latency (nearest DC)
- EU users served from Ireland (low latency)
- Can deploy/upgrade one DC at a time (zero downtime!)
Datacenter Naming Strategies
Geographic Names
Best for: Multi-region deployments
DC2: us-west-2
DC3: eu-west-1
DC4: ap-southeast-1
Clear, intuitive, maps to cloud regions
Logical Names
Best for: Single location with logical separation
DC2: online-app
DC3: search
Separates workloads in same physical location
Simple Names
Best for: Testing/development
DC2: datacenter2
Default names, simple but not descriptive
Critical: Datacenter Names Cannot Change!
Once you set datacenter names in your cluster, you cannot rename them without rebuilding the entire cluster!
Choose wisely from day one:
- Use clear, descriptive names (us-east-1, not dc1)
- Follow your cloud provider's region naming
- Consider future expansion (leave room for more DCs)
- Document what each datacenter represents
🗄️ Rack Awareness: Power & Network Isolation
How Cassandra distributes replicas across racks to survive power/network failures.
⚡ The Power Outage Story
True Story from Amazon: In 2017, a routine maintenance error in US-EAST-1 triggered cascading power failures.
❌ Without Rack Awareness:
Node 1, Node 2, Node 3, Node 4, Node 5,
Node 6, Node 7, Node 8, Node 9
Power to Rack A fails → ALL 9 nodes die!
RF=3 doesn't matter - all replicas lost!
Result: COMPLETE OUTAGE! 💥
✅ With Rack Awareness:
Rack A: Node 1, Node 2, Node 3
Rack B: Node 4, Node 5, Node 6
Rack C: Node 7, Node 8, Node 9
Power to Rack A fails → Only nodes 1-3 die
But replicas exist on Rack B + Rack C!
Result: Cluster keeps running! ✅
Rack awareness saved the day!
Same hardware failure, different outcome!
How Rack Awareness Works
Configuring Rack Awareness
In cassandra-rackdc.properties:
dc=us-east-1
rack=rack1
# Node 4 configuration
dc=us-east-1
rack=rack2
# Node 7 configuration
dc=us-east-1
rack=rack3
Cassandra automatically:
- Spreads replicas across different racks
- Avoids placing multiple replicas in same rack
- Uses rack info during rebalancing operations
🌍 Multi-Region Deployments
Global architecture patterns for worldwide applications.
Active-Active (Recommended)
All regions serve traffic
Users connect to nearest datacenter for low latency. All DCs are equal.
Example:
- NYC users → US-EAST
- London users → EU-WEST
- Tokyo users → ASIA-PAC
- Latency: <20ms! ⚡
Pros:
- Lowest latency globally
- Maximum availability
- Lose entire region → others continue
Active-Passive
Primary + backup regions
One primary serves all traffic. Others are standby for disaster recovery.
Example:
- Primary: US-EAST (active)
- Backup: US-WEST (standby)
- All traffic → US-EAST
- Failover only if primary dies
Cons:
- Wasted standby capacity
- Higher latency for distant users
- Manual failover often required
Hybrid
Mix of both strategies
Different workloads use different patterns based on requirements.
Example:
- User data: Active-Active
- Analytics: Active-Passive
- Logs: Single datacenter
Use when:
- Different data types
- Mixed compliance needs
- Cost optimization required
Real-World Multi-Region Example: Apple iCloud
🍎 Apple iCloud Contact Sync
Challenge: 2 billion devices need instant contact sync globally.
Their Setup:
'class': 'NetworkTopologyStrategy',
'us-east': 3, // North America
'eu-west': 3, // Europe
'asia-pacific': 3 // Asia
};
Total: 75,000+ nodes across 3 continents
Each write replicated 9 times (3 per region)
Performance:
- Latency: 15ms average (LOCAL_QUORUM)
- Availability: 99.999% (5 nines!)
- Throughput: Millions of writes/sec
- Disaster Recovery: Can lose entire continent!
Result: You never notice when you edit a contact!
That's the power of multi-region topology!
🖥️ Interactive Topology Builder
Design your own cluster topology and see replica distribution!
Select datacenter count and replication factor, then click "Generate" to see the topology.
The simulator will show:
• Datacenter layout
• Rack distribution
• Replica placement
• Failure scenarios
Production Topology Best Practices
- Minimum 3 datacenters for true fault tolerance
- RF=3 per datacenter - industry standard
- Odd number of nodes in each datacenter (avoid split-brain)
- At least 3 racks per datacenter if possible
- Use LOCAL_QUORUM for reads/writes in multi-DC
💼 Interview Questions & Answers
Master these 30 essential cluster topology questions!
Answer:
Cluster Topology is the hierarchical organization of Cassandra nodes into a structure of datacenters, racks, and nodes to provide fault tolerance and high availability.
The 4-Level Hierarchy:
- Cluster: The complete Cassandra ring (all nodes globally)
- Datacenters: Geographic or logical groups (e.g., US-EAST, EU-WEST)
- Racks: Failure domains within datacenters (separate power/network)
- Nodes: Individual Cassandra servers
Why It Matters:
- Survive datacenter outages
- Survive rack-level power failures
- Low-latency access from anywhere
- Automatic replica distribution
Answer:
A Datacenter (DC) in Cassandra is a logical grouping of nodes that typically represents either a physical datacenter (geographic location) or a logical separation (different workload types).
Two Types:
1. Physical Datacenters (Geographic):
DC2: eu-west-1 (Ireland)
DC3: ap-southeast-1 (Singapore)
Use case: Global application, low latency
2. Logical Datacenters (Workload):
DC2: analytics (batch processing)
DC3: search (Solr indexing)
Use case: Same location, different workloads
Key Features:
- Each DC has independent replication factor
- Can configure different RF per DC
- LOCAL_QUORUM operates within single DC
- Datacenters cannot be renamed after creation!
Answer:
Rack Awareness is Cassandra's ability to distribute replicas across different racks to survive rack-level failures (power outages, network switch failures, etc.).
What is a "Rack"?
A rack represents a failure domain - typically:
- Physical server racks sharing same power circuit
- Servers connected to same network switch
- Availability zones in cloud (AWS AZ-A, AZ-B, AZ-C)
Example Scenario:
Rack A: Node 1, 2, 3
Rack B: Node 4, 5, 6
Rack C: Node 7, 8, 9
For any key:
- Replica 1 → Node in Rack A
- Replica 2 → Node in Rack B
- Replica 3 → Node in Rack C
Power fails to Rack A → Still have 2 replicas! ✅
Configuration:
dc=us-east-1
rack=rack1 # or rack2, rack3, etc.
Best Practice: Use at least 3 racks per datacenter for optimal fault tolerance with RF=3.
Answer:
NetworkTopologyStrategy is the replication strategy that enables datacenter-aware and rack-aware replica placement.
How It Works:
- Allows different RF per datacenter
- Automatically distributes replicas across racks
- Ensures replicas don't all go to same rack
- Supports multi-datacenter deployments
Example Configuration:
'class': 'NetworkTopologyStrategy',
'us-east-1': 3, // 3 replicas in US
'eu-west-1': 3, // 3 replicas in EU
'asia-pacific': 2 // 2 replicas in Asia
};
Total replicas per key: 8 copies!
When to Use:
- ✅ ALWAYS use in production!
- ✅ Even with single datacenter (future-proof)
- ✅ Any multi-datacenter setup
- ✅ When rack awareness needed
Don't Use:
- ❌ SimpleStrategy (only for testing!)
- ❌ Never in production
Answer:
If you lose an entire datacenter, the cluster continues operating using replicas in other datacenters - if you designed topology correctly!
Scenario: 3-Datacenter Deployment
DC1: US-EAST (3 nodes, RF=3)
DC2: US-WEST (3 nodes, RF=3)
DC3: EU-WEST (3 nodes, RF=3)
Total: 9 replicas per key (3 in each DC)
Disaster: Hurricane destroys US-EAST!
Immediate Impact:
- DC1 (US-EAST) completely offline ❌
- 3 replicas lost
- But 6 replicas still exist! ✅
Cluster Response:
- ✅ DC2 (US-WEST) continues serving
- ✅ DC3 (EU-WEST) continues serving
- ✅ QUORUM still achievable (2 out of 3 DCs)
- ✅ No data loss!
- ✅ Applications keep running!
What About Consistency Levels?
- LOCAL_QUORUM: Still works (uses DC2 or DC3)
- QUORUM: Still works (6 replicas, need 5)
- EACH_QUORUM: FAILS! (needs quorum in ALL DCs)
- ALL: FAILS! (needs all 9 replicas)
Recovery Steps:
- Spin up new nodes in different datacenter
- Run repair to sync data
- Update applications to avoid dead DC
- Eventually decommission dead DC
🎓 Chapter Summary: Master Cluster Topology
Congratulations! You now deeply understand Cassandra Cluster Topology!
Key Concepts Mastered:
- Topology Hierarchy: Cluster → Datacenters → Racks → Nodes
- Datacenters: Geographic or logical groupings for fault isolation
- Rack Awareness: Distribute replicas across power/network failure domains
- Multi-Region: Active-Active for global low latency
- NetworkTopologyStrategy: ALWAYS use in production!
The GlobalShip Analogy Recap:
Remember GlobalShip's warehouse strategy? Don't put everything in one NYC warehouse - spread across continents (datacenters), separate buildings (racks), with redundant inventory (replicas). Hurricane hits NYC? London and Tokyo keep delivering!
Production Best Practices:
- ✅ Use NetworkTopologyStrategy (ALWAYS!)
- ✅ Minimum 3 datacenters for true fault tolerance
- ✅ RF=3 per datacenter (industry standard)
- ✅ At least 3 racks per datacenter
- ✅ Use LOCAL_QUORUM for multi-DC reads/writes
- ✅ Odd number of nodes per datacenter
- ✅ Document your topology clearly
🚀 You can now design globally distributed, fault-tolerant Cassandra clusters!
Responsive Ad