Cassandra Network Topology
Master datacenter and rack organization! Deep dive into NetworkTopologyStrategy, snitch configuration, multi-region deployments, and the geographic distribution that powers global-scale applications.
📖 The Story: The Global Retail Chain
Imagine MegaMart - a retail giant with 10,000 stores across 5 continents. How do you organize inventory data so customers get instant updates, regardless of location?
❌ THE BAD ARCHITECTURE: Single Warehouse
The Problem: All inventory data stored in ONE central warehouse in Kansas City
├─ Tokyo Store #1234 (7,500 miles away!) → 250ms latency
├─ London Store #5678 (4,000 miles away!) → 150ms latency
├─ Sydney Store #9012 (9,000 miles away!) → 300ms latency
└─ New York Store #3456 (1,200 miles away!) → 40ms latency
Disasters:
• Tornado hits Kansas → ALL stores go offline!
• Network cable cut → GLOBAL OUTAGE!
• Peak holiday traffic → Single warehouse overwhelmed!
• Tokyo customer waits 300ms for product info! 😡
✅ THE BRILLIANT ARCHITECTURE: Regional Distribution Centers
The Solution: 5 Regional Distribution Centers with hierarchical organization
├─ 🇺🇸 North America DC (US-EAST)
├─ 🇪🇺 Europe DC (EU-WEST)
├─ 🇯🇵 Asia DC (ASIA-PAC)
├─ 🇦🇺 Australia DC (AU-SOUTH)
└─ 🇧🇷 South America DC (SA-EAST)
🏢 LEVEL 2: BUILDINGS (Racks)
Each DC has 3 buildings:
Building A (Rack-A) - Main power circuit
Building B (Rack-B) - Backup power circuit
Building C (Rack-C) - Generator power
💻 LEVEL 3: SERVERS (Nodes)
Each Building has 100 servers
Total: 5 DCs × 3 Racks × 100 Nodes = 1,500 nodes!
✨ MAGIC OF REPLICATION:
Product #ABC123 stored in:
→ US-EAST: Rack-A/Server-45, Rack-B/Server-78, Rack-C/Server-23
→ EU-WEST: Rack-A/Server-12, Rack-B/Server-56, Rack-C/Server-89
→ ASIA-PAC: Rack-A/Server-34, Rack-B/Server-67, Rack-C/Server-90
Total: 9 copies across 3 continents!
Why It's GENIUS:
- Low Latency: Tokyo customer → ASIA-PAC DC (5ms vs 250ms!)
- Fault Tolerance: Tornado in US? EU and ASIA serve customers!
- Power Failure: Rack-A loses power? Rack-B & C still running!
- Scalability: Add more buildings/servers per DC independently!
- Data Sovereignty: EU data stays in EU (GDPR compliant!)
🎯 This is EXACTLY Cassandra's Network Topology!
- Regional DCs = Cassandra Datacenters
- Buildings = Cassandra Racks
- Servers = Cassandra Nodes
- Replication = NetworkTopologyStrategy
- Power Circuits = Failure Domains
CREATE KEYSPACE inventory WITH replication = {
'class': 'NetworkTopologyStrategy',
'US-EAST': 3, // 3 copies in North America
'EU-WEST': 3, // 3 copies in Europe
'ASIA-PAC': 3 // 3 copies in Asia
};
Snitch: GossipingPropertyFileSnitch
Rack-aware: Yes (Rack-A, Rack-B, Rack-C)
Result: Global distribution with local speed! ✅
MegaMart's regional DCs = Cassandra's Network Topology!
Same brilliance, global scale!
🗺️ What is Network Topology in Cassandra?
The hierarchical organization of nodes across datacenters and racks for geographic distribution and fault isolation.
Complete Definition
Network Topology: The logical and physical organization of Cassandra nodes into a hierarchy of datacenters and racks, enabling geographic distribution, fault isolation, and optimized replication strategies.
Three-Level Hierarchy:
- Cluster: Top level - entire Cassandra deployment
- Datacenter: Geographic or logical grouping of nodes
- Rack: Failure domain within a datacenter
- Node: Individual server running Cassandra
Why Topology Matters:
- Enables multi-region deployments for global applications
- Provides fault isolation at rack and datacenter level
- Optimizes data locality for low-latency reads
- Supports disaster recovery and business continuity
- Enables compliance with data sovereignty requirements
🏢 Datacenters: Geographic Distribution
How Cassandra organizes nodes across geographic or logical boundaries.
Understanding Datacenters
Physical Datacenters: Actual geographic locations (AWS us-east-1, Google us-central1, on-prem Virginia DC)
Logical Datacenters: Workload separation (analytics DC, online-transactional DC, search DC)
Datacenter Benefits:
- Low Latency: Serve users from nearest DC (Tokyo users → ASIA-PAC DC)
- Disaster Recovery: Entire DC can fail, others continue serving
- Data Sovereignty: EU data stays in EU (GDPR, legal compliance)
- Workload Isolation: Heavy analytics don't impact online traffic
- Maintenance Windows: Upgrade one DC at a time (zero downtime)
Geographic Datacenters
Physical locations worldwide
US-WEST (Oregon)
EU-WEST (Ireland)
EU-CENTRAL (Frankfurt)
ASIA-PAC (Tokyo)
ASIA-SOUTH (Singapore)
AU-SOUTHEAST (Sydney)
Use Case: Global applications, serve users locally, 5-20ms latency worldwide
Logical Datacenters
Workload separation
ANALYTICS (batch jobs)
SEARCH (ElasticSearch mirror)
BACKUP (disaster recovery)
DEV (development/testing)
REPORTING (BI queries)
Use Case: Heavy analytics don't slow down online traffic, dedicated resources per workload
Cloud Provider Datacenters
Cloud region mapping
us-east-1, us-west-2
eu-west-1, ap-southeast-1
GCP:
us-central1, europe-west1
asia-northeast1
Azure:
eastus, westeurope
southeastasia
Use Case: Multi-cloud deployments, cloud-native applications
Critical: Datacenter Naming Rules
- Cannot Change: Once set, datacenter names are PERMANENT
- Alphanumeric Only: Letters, numbers, underscores, hyphens (no spaces)
- Case Sensitive: "US-EAST" ≠ "us-east"
- Descriptive: Use meaningful names (us-east-1, not dc1)
- Document: Keep datacenter topology documentation updated
📚 Racks: Failure Domain Isolation
The second level of hierarchy providing fault isolation within datacenters.
What is a Rack?
Physical Rack: Literal server rack with shared power supply and network switch
Logical Rack: Availability zone, separate network segment, or failure domain
Why Rack Awareness Matters:
Without Rack Awareness:
- All 3 replicas might be on same rack
- Power failure = lose all 3 copies = DATA LOST!
- Network switch failure = all replicas unreachable
With Rack Awareness:
- Replicas distributed across different racks
- Power failure on Rack-A = Rack-B & C still serve data
- Can lose ANY rack and still have all data available
Real-World Example: Amazon Power Outage 2017
💥 The Disaster: US-EAST-1 Power Failure
Date: February 28, 2017
Incident: Power distribution unit failure in one availability zone
3-node Cassandra cluster in US-EAST-1
All 3 nodes in same availability zone (same rack)
Power fails → ALL 3 nodes go offline
RF=3 but ALL replicas lost!
Result: 4-hour complete outage! 💥
Customers couldn't log in, checkout, or access data
Estimated loss: $500K/hour
COMPANY B (With Rack Awareness):
3-node Cassandra cluster in US-EAST-1
Nodes distributed across 3 availability zones:
Node 1 → us-east-1a (Rack-A)
Node 2 → us-east-1b (Rack-B)
Node 3 → us-east-1c (Rack-C)
Power fails in us-east-1a → Only Node 1 offline
Nodes 2 & 3 still serving (QUORUM intact!)
Result: ZERO downtime! ✅
Customers experienced no disruption
Traffic automatically routed to healthy nodes
Rack Configuration Best Practices
Best Practice
- Minimum 3 racks per datacenter
- Equal nodes per rack
- One replica per rack with RF=3
- Different power circuits per rack
- Different network switches
- Cloud: Map to Availability Zones
Anti-Pattern
- Single rack: No fault isolation
- Two racks: Can't survive 1 rack failure
- Uneven distribution: Some racks empty
- Default "rack1": All nodes same rack
- Ignore cloud AZs: Wasted isolation
👀 Snitch: The Topology Detective
How Cassandra learns the datacenter and rack location of every node in the cluster.
What is a Snitch?
Snitch: A component that determines the datacenter and rack for each node, enabling NetworkTopologyStrategy to make intelligent replica placement decisions.
What Snitch Tells Cassandra:
- Datacenter Name: Which DC does this node belong to?
- Rack Name: Which rack within the DC?
- Relative Distance: How "close" are two nodes?
- Routing Decisions: Which node should handle a request?
Why Snitch is Critical:
Without Snitch:
- Cassandra thinks all nodes are in same location
- Replicas might all go to same rack (no fault tolerance!)
- Queries routed randomly (high latency!)
- Can't enforce data sovereignty (legal violations!)
With Snitch:
- Replicas spread across racks and DCs intelligently
- Queries routed to nearest node (low latency)
- Data stays in correct geographic region
- Survives rack and DC failures gracefully
Types of Snitches
GossipingPropertyFileSnitch
RECOMMENDED for production
dc=US-EAST
rack=RACK-A
prefer_local=true
How it works:
• Each node reads local file
• Gossips topology to cluster
• Auto-discovers other nodes
• No centralized config!
- Pros: Simple, flexible, self-healing
- Cons: Manual file editing
- Use: 95% of deployments
Ec2Snitch / Ec2MultiRegionSnitch
AWS-specific auto-detection
• Region → Datacenter
• Availability Zone → Rack
• Instance metadata API
Example:
us-east-1a → DC: us-east-1, Rack: 1a
eu-west-1b → DC: eu-west-1, Rack: 1b
ap-south-1c → DC: ap-south-1, Rack: 1c
- Pros: Zero config, AWS-native
- Cons: AWS only, uses public IPs
- Use: AWS deployments
GoogleCloudSnitch
GCP-specific auto-detection
• Region → Datacenter
• Zone → Rack
Example:
us-central1-a → DC: us-central1, Rack: a
europe-west1-b → DC: europe-west1, Rack: b
asia-east1-c → DC: asia-east1, Rack: c
- Pros: Zero config, GCP-native
- Cons: GCP only
- Use: Google Cloud deployments
SimpleSnitch
DO NOT USE in production
• Same datacenter
• Same rack
• No topology awareness!
Result:
• All replicas on same rack
• No fault isolation
• Random query routing
- Pros: None (seriously!)
- Cons: Dangerous, no isolation
- Use: Testing ONLY
Critical: Changing Snitch in Production
WARNING: Changing snitch in running cluster is EXTREMELY dangerous!
What happens if you change snitch incorrectly:
- Cassandra thinks nodes moved to different DCs/racks
- Replication breaks (wrong replicas deleted/created)
- Data loss possible (old replicas dropped)
- Queries fail (can't find correct replicas)
- Cluster becomes inconsistent
Safe Snitch Migration Steps:
- Run full cluster repair first
- Change snitch on ONE node
- Restart that node
- Verify topology correct:
nodetool status - If correct, proceed to next node
- If wrong, revert immediately!
- Run repair after all nodes migrated
🎯 Replication Strategies: SimpleStrategy vs NetworkTopologyStrategy
Understanding how topology enables intelligent replication.
Expert: How NetworkTopologyStrategy Places Replicas
Intelligent Algorithm:
- Per Datacenter: Process each DC independently
- Walk the Ring: Start from primary replica, walk clockwise
- Skip Same Rack: If next node in same rack as existing replica, skip it
- Choose Different Rack: Continue until finding node in different rack
- Repeat: Continue until RF replicas placed in this DC
- Next DC: Repeat process for next datacenter
Example: RF=3 in DC1, RF=2 in DC2
Primary: Node 1 (Rack-A)
Walk ring → Node 2 (Rack-A) ❌ Skip! (same rack)
Walk ring → Node 3 (Rack-B) ✅ Place replica 2
Walk ring → Node 4 (Rack-B) ❌ Skip! (same rack)
Walk ring → Node 5 (Rack-C) ✅ Place replica 3
Result: Replicas on Rack-A, Rack-B, Rack-C ✅
DC2 (EU-WEST) - Need 2 replicas:
Primary: Node 10 (Rack-A)
Walk ring → Node 11 (Rack-B) ✅ Place replica 2
Result: Replicas on Rack-A, Rack-B ✅
Total: 5 copies (3 in US, 2 in EU) across different racks!
🖥️ Interactive Topology Builder
Build and visualize your own multi-datacenter topology!
Configure your cluster topology and click "Build Topology" to see the complete structure.
Try different configurations:
• 3 Datacenters × 3 Racks × 3 Nodes = 27 total nodes
• 2 Datacenters × 4 Racks × 5 Nodes = 40 total nodes
• Click "Show Replication" to see replica placement with RF=3!
🌍 Multi-Region Deployments: Global Scale
Real-world patterns for deploying Cassandra across continents.
Production Example: Global Social Network
Company: GlobalConnect - 500M users worldwide, 100M DAU
Scale: 300-node cluster across 5 datacenters
Topology Design
🇺🇸 US-EAST-1 (Virginia): 90 nodes (3 racks × 30 nodes)
🇺🇸 US-WEST-2 (Oregon): 90 nodes (3 racks × 30 nodes)
🇪🇺 EU-WEST-1 (Ireland): 60 nodes (3 racks × 20 nodes)
🇯🇵 ASIA-PAC-1 (Tokyo): 30 nodes (3 racks × 10 nodes)
🇧🇷 SA-EAST-1 (São Paulo): 30 nodes (3 racks × 10 nodes)
Total: 300 nodes across 5 continents
Replication Strategy
CREATE KEYSPACE user_profiles WITH replication = {
'class': 'NetworkTopologyStrategy',
'US-EAST-1': 3, // Primary region
'US-WEST-2': 3, // DR + West coast users
'EU-WEST-1': 3, // European users
'ASIA-PAC-1': 2, // Asian users (lighter load)
'SA-EAST-1': 2 // South America (lighter load)
};
Total: 13 copies of every user profile worldwide!
User Sessions Keyspace (time-sensitive):
CREATE KEYSPACE user_sessions WITH replication = {
'class': 'NetworkTopologyStrategy',
'US-EAST-1': 3,
'US-WEST-2': 3,
'EU-WEST-1': 3,
'ASIA-PAC-1': 3,
'SA-EAST-1': 3
};
Total: 15 copies for ultra-high availability!
Consistency Levels by Region
- Writes: LOCAL_QUORUM (low latency, DC-local consistency)
- Reads: LOCAL_ONE (ultra-fast, serve from nearest replica)
- Critical Ops: EACH_QUORUM (write to all DCs synchronously)
- Background Repair: Keeps DCs in sync asynchronously
Latency Results
| User Location | Nearest DC | Read Latency | Write Latency |
|---|---|---|---|
| New York | US-EAST-1 | 1-2ms | 3-5ms |
| San Francisco | US-WEST-2 | 1-2ms | 3-5ms |
| London | EU-WEST-1 | 2-3ms | 4-6ms |
| Tokyo | ASIA-PAC-1 | 2-3ms | 5-7ms |
| São Paulo | SA-EAST-1 | 2-4ms | 5-8ms |
Result: Sub-10ms latency for 99% of users worldwide!
💼 Interview Questions & Answers
Master these 25 essential questions on Cassandra Network Topology - from beginner to expert level!
Answer:
Network Topology is the hierarchical organization of Cassandra nodes into datacenters and racks, enabling geographic distribution and fault isolation.
Three-Level Hierarchy:
- Cluster: Top level - entire deployment
- Datacenter: Geographic or logical grouping (e.g., US-EAST, EU-WEST)
- Rack: Failure domain within a datacenter (e.g., Rack-A, Rack-B)
Why It's Critical:
- Low Latency: Serve users from nearest datacenter (Tokyo users → ASIA-PAC DC = 5ms vs US-EAST = 250ms)
- Fault Isolation: Power failure on Rack-A doesn't affect Rack-B and Rack-C
- Disaster Recovery: Entire datacenter can fail, others continue serving
- Data Sovereignty: EU data stays in EU (GDPR compliance)
- Smart Replication: NetworkTopologyStrategy spreads replicas across racks/DCs
Answer:
A Snitch is a component that determines the datacenter and rack location of each node, enabling NetworkTopologyStrategy to make intelligent replica placement decisions.
What Snitch Provides:
- Datacenter Name: Which DC does this node belong to? (e.g., US-EAST)
- Rack Name: Which rack within the DC? (e.g., RACK-A)
- Relative Distance: How "close" are two nodes for routing decisions
- Query Routing: Which node should handle a request (prefer local)
Without Snitch:
- All nodes treated as same location
- Replicas might all go to same rack (no fault tolerance!)
- Queries routed randomly (high latency)
- Can't enforce data sovereignty
With Snitch:
- Replicas intelligently spread across racks and DCs
- Queries routed to nearest node (low latency)
- Data stays in correct geographic region
- Survives rack and DC failures gracefully
Answer:
GossipingPropertyFileSnitch:
- Configuration: Read from cassandra-rackdc.properties file on each node
- Setup: Manual - you specify dc=US-EAST, rack=RACK-A
- Discovery: Gossip protocol shares topology with cluster
- Flexibility: Works anywhere - on-prem, any cloud, hybrid
- Use When: 95% of deployments, any environment, need full control
- Pros: Simple, flexible, self-healing, cloud-agnostic
- Cons: Manual file editing required
Ec2Snitch / Ec2MultiRegionSnitch:
- Configuration: Auto-detects from AWS EC2 metadata API
- Setup: Zero configuration required
- Mapping: AWS Region → Datacenter, Availability Zone → Rack
- Example: us-east-1a becomes DC: us-east-1, Rack: 1a
- Use When: 100% AWS deployment, want zero-config
- Pros: No manual config, AWS-native, auto-updates
- Cons: AWS only, uses public IPs (Ec2Snitch), not portable
Recommendation:
Use GossipingPropertyFileSnitch even on AWS for better control and portability. Only use Ec2Snitch if you absolutely need zero-config and will never migrate off AWS.
Answer:
SimpleStrategy:
- Topology Awareness: None - treats all nodes as equal
- Replication: Places replicas clockwise on ring, ignoring DC/rack
- Configuration: Single RF for entire cluster
- Example: {'class': 'SimpleStrategy', 'replication_factor': 3}
- Multi-DC Support: Not supported - can't specify per-DC RF
- Fault Tolerance: Low - all replicas might be on same rack
- Production Use: ❌ NEVER - only for testing/dev
NetworkTopologyStrategy:
- Topology Awareness: Fully aware of datacenters and racks
- Replication: Intelligently spreads replicas across different racks
- Configuration: Per-datacenter RF specification
- Example: {'class': 'NetworkTopologyStrategy', 'US-EAST': 3, 'EU-WEST': 2}
- Multi-DC Support: Fully supported - different RF per DC
- Fault Tolerance: High - guarantees replicas on different racks/DCs
- Production Use: ✅ ALWAYS - even with single DC
Key Difference:
NetworkTopologyStrategy ensures replicas are on different racks within each DC. With RF=3 and 3 racks, SimpleStrategy might put all 3 replicas on Rack-A (disaster!), while NetworkTopologyStrategy guarantees one replica per rack (safe!).
Answer:
NetworkTopologyStrategy uses an intelligent algorithm to ensure replicas are on different racks:
Algorithm Steps:
- Start: Primary replica placed based on partition key hash (token)
- Walk Ring: Move clockwise around token ring
- Check Rack: Is next node on same rack as existing replica?
- Skip or Place: If same rack → skip, if different rack → place replica
- Repeat: Continue until RF replicas placed in this DC
- Next DC: Repeat entire process for next datacenter
Example (RF=3, 3 racks):
- Primary: Token hash → Node 5 (Rack-B)
- Walk ring: Node 6 (Rack-B) ❌ Skip! Same rack as Node 5
- Continue: Node 7 (Rack-C) ✅ Place replica 2 (different rack)
- Continue: Node 8 (Rack-C) ❌ Skip! Same rack as Node 7
- Continue: Node 1 (Rack-A) ✅ Place replica 3 (different rack)
- Result: Replicas on Rack-A, Rack-B, Rack-C ✅
Guarantee:
With 3+ racks and RF=3, you're guaranteed one replica per rack. If rack fails, other 2 racks still have complete data!
Answer:
WARNING: This is EXTREMELY dangerous and can cause data loss!
What Goes Wrong:
- Topology Confusion: Cassandra thinks nodes moved to different DCs/racks
- Replication Breaks: Wrong replicas might be deleted, new ones created
- Data Loss Risk: Old replicas dropped before new ones fully replicated
- Query Failures: Can't find correct replicas, routing errors
- Consistency Issues: Cluster becomes inconsistent, repair can't fix
- Split Brain: Different nodes think data is in different places
Safe Migration Procedure:
- Plan: Document current topology, plan new topology
- Backup: Take full snapshots of entire cluster
- Repair: Run full cluster repair first (nodetool repair)
- Maintenance Window: Schedule downtime if possible
- One Node: Change snitch on ONE node only
- Update Config: Modify cassandra.yaml and cassandra-rackdc.properties
- Restart: Restart that one node
- Verify: Check topology: nodetool status, nodetool describecluster
- Monitor: Watch logs for errors, check metrics
- Iterate: If correct, proceed to next node. If wrong, REVERT!
- Final Repair: After all nodes migrated, run full repair
Best Practice:
Get snitch configuration right from the start! Changing it later is risky and time-consuming.
Answer:
Rack awareness is Cassandra's ability to understand physical failure domains (racks) and distribute replicas accordingly to prevent correlated failures.
Physical Rack Concept:
- Physical Rack: Literal server rack with shared power supply and network switch
- Failure Domain: If power fails or switch breaks, entire rack goes offline
- Cloud Equivalent: Availability Zone in AWS/GCP/Azure
Without Rack Awareness:
- All 3 replicas (RF=3) might be on same rack
- Power failure on that rack = lose ALL 3 copies = DATA LOST!
- Network switch failure = all replicas unreachable = OUTAGE!
- Single point of failure despite having replicas
With Rack Awareness (NetworkTopologyStrategy):
- Replicas automatically spread across different racks
- RF=3 + 3 racks = one replica per rack guaranteed
- Power failure on Rack-A = Rack-B and Rack-C still serve data
- Can lose ANY single rack without data loss or downtime
Real-World Example:
Amazon 2017 Power Outage: PDU failure in us-east-1a
- No Rack Awareness: All 3 nodes in us-east-1a → 4-hour outage, $500K/hour loss
- With Rack Awareness: Nodes in us-east-1a, 1b, 1c → Zero downtime! Other AZs served traffic
Best Practice:
Minimum 3 racks per datacenter. In cloud: map each rack to different Availability Zone.
Answer:
Use NetworkTopologyStrategy with per-datacenter replication factor specification:
Configuration Example:
'class': 'NetworkTopologyStrategy',
'US-EAST': 3, // 3 replicas in US-EAST
'EU-WEST': 2, // 2 replicas in EU-WEST
'ASIA-PAC': 2 // 2 replicas in ASIA-PAC
};
-- Total: 7 copies of every row!
Why Different RFs?
- US-EAST (RF=3): Primary datacenter, most users, highest traffic
- EU-WEST (RF=2): Secondary datacenter, moderate traffic
- ASIA-PAC (RF=2): Read-only analytics DC, lower requirements
Data Distribution:
- User record for user_id='alice' stored in 7 nodes total:
- US-EAST: Node 5 (Rack-A), Node 12 (Rack-B), Node 19 (Rack-C)
- EU-WEST: Node 34 (Rack-A), Node 41 (Rack-B)
- ASIA-PAC: Node 56 (Rack-A), Node 63 (Rack-B)
Consistency Levels:
- Writes: LOCAL_QUORUM (2 nodes in local DC)
- Reads: LOCAL_ONE (1 node in local DC, ultra-fast)
- Critical: EACH_QUORUM (QUORUM in every DC)
Verification:
nodetool describecluster;
nodetool status my_keyspace;
Answer:
prefer_local controls whether Cassandra prefers reading from replicas in the local datacenter.
Configuration:
dc=US-EAST
rack=RACK-A
prefer_local=true # Default and recommended
When prefer_local=true (RECOMMENDED):
- Queries prefer replicas in same datacenter
- Lower latency (1-5ms vs 100-200ms cross-DC)
- Reduced WAN bandwidth costs
- Better user experience (faster responses)
- Only contacts remote DC if local replicas unavailable
When prefer_local=false:
- Queries may go to any datacenter
- Higher latency (unpredictable)
- Higher WAN costs
- Useful for testing or special analytics workloads
Example Impact:
- User in New York, US-EAST DC:
- prefer_local=true: Read from US-EAST replica (2ms latency)
- prefer_local=false: Might read from EU-WEST replica (150ms latency!)
Interaction with Consistency Levels:
- LOCAL_QUORUM: Always reads from local DC (prefer_local irrelevant)
- QUORUM: Uses prefer_local to choose which DC's replicas
- ONE: Uses prefer_local to pick replica location
Best Practice:
Always set prefer_local=true in production. Only set false for specific analytics/reporting workloads.
Answer:
Cassandra automatically routes queries to healthy datacenters when one DC fails, with behavior depending on consistency level:
Scenario: US-EAST datacenter goes offline
With LOCAL_QUORUM (RECOMMENDED):
- US-EAST clients: Queries FAIL (local DC down, can't reach LOCAL_QUORUM)
- Solution: Load balancer redirects US-EAST traffic to US-WEST
- US-WEST clients: Unaffected, continue serving from US-WEST replicas
- EU-WEST clients: Unaffected, continue serving from EU-WEST replicas
- Benefit: Controlled failover, no cross-DC latency impact
With QUORUM (Global):
- All queries continue working!
- Needs (Total RF / 2) + 1 replicas globally
- Example: RF=9 (3 per DC × 3 DCs) → QUORUM=5
- US-EAST down (3 replicas lost) → 6 replicas remain → Still achieve QUORUM
- Downside: Higher latency (might query remote DCs)
With EACH_QUORUM:
- All writes FAIL!
- Requires QUORUM in EVERY datacenter
- US-EAST down → Can't reach QUORUM there → Write rejected
- Use case: Strong consistency, but poor availability
Automatic Failover Example:
Before DC Failure:
- US-EAST clients → US-EAST replicas (2ms)
- EU-WEST clients → EU-WEST replicas (2ms)
After US-EAST Fails:
- US-EAST clients → Redirected to US-WEST (20ms, slight increase)
- EU-WEST clients → Still EU-WEST replicas (2ms, no change)
Recovery Process:
- Detection: Gossip protocol detects US-EAST nodes down (~30 seconds)
- Routing: Coordinator stops sending queries to US-EAST
- Hinted Handoff: Writes queued for US-EAST nodes
- DC Restored: Hints replayed, nodes catch up
- Repair: Run nodetool repair to ensure consistency
Best Practice:
Use LOCAL_QUORUM for writes and LOCAL_ONE for reads. Set up load balancer to redirect traffic from failed DC to nearest healthy DC.
Answer:
Always use NetworkTopologyStrategy, even with single DC:
- Rack Awareness: Still need fault isolation within DC (different racks/AZs)
- Future Expansion: Easy to add DCs later without changing keyspace
- Best Practice: Consistent approach across all environments
- No Downside: Same performance as SimpleStrategy in single DC
- Production Ready: SimpleStrategy is for testing only
Even with single DC, use: {'class': 'NetworkTopologyStrategy', 'DC1': 3}
Answer:
Cloud Availability Zones (AZs) are perfect Cassandra racks:
- AWS: us-east-1a → DC: us-east-1, Rack: 1a
- GCP: us-central1-a → DC: us-central1, Rack: a
- Azure: eastus-1 → DC: eastus, Rack: 1
- Isolation: AZs have separate power, cooling, networking
- Best Practice: Deploy nodes across 3+ AZs in each region
- Configuration: Set rack=1a, rack=1b, rack=1c in cassandra-rackdc.properties
With 3 AZs and RF=3, you get one replica per AZ - survives any AZ failure!
Answer:
Minimum: 3 racks per datacenter
- With RF=3: One replica per rack guaranteed
- Fault Tolerance: Can lose any 1 rack, still have QUORUM (2/3)
- With 2 Racks: Can't survive single rack failure with QUORUM
- With 1 Rack: No fault isolation at all!
Math:
- 3 racks, RF=3: Rack fails → 2 replicas remain → QUORUM(2) achieved ✅
- 2 racks, RF=3: Rack fails → 1-2 replicas remain → QUORUM(2) might fail ❌
Answer:
Data Sovereignty: Legal requirement that data must stay within specific geographic boundaries.
- GDPR (EU): EU citizen data must stay in EU
- Russia: Russian citizen data must stay in Russia
- China: Chinese data must stay in China
How Cassandra Helps:
- Datacenter Isolation: Create separate EU-WEST datacenter
- Keyspace Design: EU users → eu_users keyspace with RF only in EU-WEST
- No Cross-DC: EU data never replicates to US or Asia DCs
- Local Queries: EU queries only hit EU-WEST nodes
Configuration Example:
'class': 'NetworkTopologyStrategy',
'EU-WEST': 3 -- ONLY in EU, never leaves!
};
Answer:
Scenario: Global e-commerce with 100M users, 1M products, 10K orders/second
Datacenter Design:
- US-EAST: 100 nodes (3 racks × 33 nodes) - Primary region
- US-WEST: 100 nodes (3 racks × 33 nodes) - DR + West coast
- EU-WEST: 60 nodes (3 racks × 20 nodes) - European customers
- ASIA-PAC: 40 nodes (3 racks × 13 nodes) - Asian customers
- Total: 300 nodes across 4 continents
Keyspace Strategy:
CREATE KEYSPACE products WITH replication = {
'class': 'NetworkTopologyStrategy',
'US-EAST': 3, 'US-WEST': 3,
'EU-WEST': 3, 'ASIA-PAC': 3
};
-- User Sessions (time-sensitive, high availability)
CREATE KEYSPACE sessions WITH replication = {
'class': 'NetworkTopologyStrategy',
'US-EAST': 3, 'US-WEST': 3,
'EU-WEST': 2, 'ASIA-PAC': 2
};
-- Orders (critical, strong consistency)
CREATE KEYSPACE orders WITH replication = {
'class': 'NetworkTopologyStrategy',
'US-EAST': 3, 'US-WEST': 3
};
-- Orders only in US (payment processing, legal)
Consistency Levels:
- Products: Write EACH_QUORUM, Read LOCAL_ONE (eventual consistency OK)
- Sessions: Write LOCAL_QUORUM, Read LOCAL_ONE (fast, local)
- Orders: Write QUORUM, Read QUORUM (strong consistency)
Answer:
Two Files Required:
1. cassandra.yaml:
2. cassandra-rackdc.properties:
rack=RACK-A
prefer_local=true
Important Notes:
- Both files in /etc/cassandra/ directory (Debian/Ubuntu) or /etc/dse/cassandra (DSE)
- Requires node restart after changes
- DC/rack names case-sensitive and permanent
- Each node can have different dc/rack values
Answer:
Gossip protocol is how nodes discover and share topology information:
- Every Second: Each node picks 1-3 random nodes to gossip with
- Exchange Info: Share state: DC/rack, health, schema, load
- Propagation: Information spreads exponentially through cluster
- Convergence: Entire cluster knows topology within ~30 seconds
- Self-Healing: Automatically detects and shares node changes
Example:
- Node 1 starts with dc=US-EAST, rack=RACK-A
- Node 1 gossips to Node 5: "I'm in US-EAST/RACK-A"
- Node 5 gossips to Node 12: "Node 1 is US-EAST/RACK-A"
- After few rounds: All nodes know Node 1's location
Answer:
Key Commands:
1. nodetool status:
Datacenter: US-EAST
Status=Up/Down, State=Normal/Leaving/Joining
-- Address Load Tokens Owns Rack
UN 10.0.1.5 256GB 256 33% RACK-A
UN 10.0.1.6 250GB 256 32% RACK-B
2. nodetool describecluster:
Name: ProductionCluster
Snitch: GossipingPropertyFileSnitch
Datacenter: US-EAST (3 racks)
Datacenter: EU-WEST (3 racks)
3. nodetool ring:
# Shows token ring with DC/rack for each token
4. CQL DESCRIBE:
DESCRIBE KEYSPACE my_keyspace;
# Shows replication configuration
Answer:
YES! Different DCs can have different node counts - this is common:
Common Scenarios:
- Primary/Secondary DCs:
- US-EAST (primary): 100 nodes
- US-WEST (DR): 50 nodes
- Different workloads justify different sizes
- Analytics DC:
- ONLINE (user-facing): 200 nodes
- ANALYTICS (batch jobs): 50 nodes
- Analytics runs overnight, needs fewer nodes
- Regional Traffic:
- US-EAST: 150 nodes (high traffic)
- EU-WEST: 100 nodes (moderate traffic)
- ASIA-PAC: 50 nodes (growing market)
Important Considerations:
- Token Distribution: Vnodes auto-adjust for different sizes
- Minimum Nodes: Must have ≥ RF nodes in each DC for that keyspace
- Load Balance: Smaller DC = more data per node
Example:
EU-WEST: 30 nodes with RF=3 → ~100GB per node
Both valid! EU nodes handle more data but less traffic.
Answer:
DISASTER! This breaks replication and can cause data loss.
Example Misconfiguration:
Node 2: dc=US-EAST
Node 3: dc=us-east ❌ Wrong case!
Node 4: dc=USEAST ❌ Missing hyphen!
What Goes Wrong:
- Multiple DCs Created: Cassandra sees 3 DCs: "US-EAST", "us-east", "USEAST"
- Replication Breaks: Keyspace expects RF=3 in "US-EAST", but only 2 nodes match!
- Writes Fail: Can't achieve QUORUM (need 2/3, only have 2 nodes in "US-EAST")
- Data Loss Risk: Some nodes not receiving replicas
- Query Routing Broken: Coordinator can't find enough replicas
Prevention:
- Automation: Use config management (Ansible, Chef, Puppet)
- Validation: Verify with nodetool status before adding to production
- Standards: Document exact DC/rack naming convention
- Case Sensitive: Use consistent capitalization (recommend ALL-CAPS)
Recovery:
- Fix cassandra-rackdc.properties on misconfigured nodes
- Restart nodes one at a time
- Run nodetool repair on entire cluster
- Verify topology with nodetool describecluster
Answer:
Rack-aware routing optimizes query performance by preferring replicas on nodes within the same rack:
- Lower Latency: Same rack = same network switch = sub-millisecond latency
- Less Network Hops: Avoid top-of-rack switch → core switch → another rack
- Higher Bandwidth: Within-rack links typically faster than cross-rack
- Cost Savings: Reduced cross-rack network traffic
Example: Coordinator on RACK-A prefers reading from RACK-A replica (0.5ms) vs RACK-B replica (2ms).
Answer:
YES! Common pattern for workload separation:
- Physical DCs: US-EAST, EU-WEST (geographic locations)
- Logical DCs: ANALYTICS, SEARCH (same physical location as US-EAST)
- Use Case: Separate heavy analytics from online traffic
Example Configuration:
├─ Logical DC: ONLINE (100 nodes) - user queries
└─ Logical DC: ANALYTICS (50 nodes) - batch jobs
Keyspace replication:
'ONLINE': 3, # replicate to online nodes
'ANALYTICS': 2 # replicate to analytics nodes
Benefit: Heavy analytics queries don't impact online user experience!
Answer:
Topology determines which nodes store hints when replicas are unavailable:
- Same DC Hints: Hints stored on nodes in same datacenter (by default)
- Rack Awareness: Hints prefer different rack than down node
- Cross-DC Hints: Disabled by default (configure with hinted_handoff_enabled)
Example:
- Write for Node 5 (US-EAST/RACK-A) but it's down
- Hint stored on Node 12 (US-EAST/RACK-B) - same DC, different rack
- When Node 5 returns, Node 12 replays hints
Answer:
Safe Migration Steps:
- Check Current: DESCRIBE KEYSPACE my_keyspace;
- Run Repair: nodetool repair -full my_keyspace
- Alter Keyspace:
ALTER KEYSPACE my_keyspace WITH replication = {
'class': 'NetworkTopologyStrategy',
'DC1': 3 -- Same RF as before
}; - Repair Again: nodetool repair -full my_keyspace
- Verify: nodetool status my_keyspace
Important: No downtime required! Replication happens asynchronously.
Answer:
Comprehensive DR Strategy:
1. Multi-DC Architecture:
- Primary: US-EAST (200 nodes, 3 racks)
- DR: US-WEST (200 nodes, 3 racks) - mirror of primary
- Remote DR: EU-WEST (100 nodes, 3 racks) - geographic diversity
2. Replication Configuration:
'class': 'NetworkTopologyStrategy',
'US-EAST': 3, # Primary
'US-WEST': 3, # Same-region DR
'EU-WEST': 3 # Remote DR
};
3. Consistency Levels:
- Normal: Write LOCAL_QUORUM, Read LOCAL_ONE
- DR Event: Switch to QUORUM (cross-DC) or redirect traffic
4. Failover Plan:
- Detection: Monitor US-EAST health (30 seconds to detect)
- DNS Update: Point users to US-WEST (TTL=60s)
- Traffic Shift: Load balancer redirects within 2 minutes
- RTO: 5 minutes (Recovery Time Objective)
- RPO: 0 seconds (no data loss - synchronous replication)
5. Testing:
- Quarterly DR drills - simulate US-EAST failure
- Verify US-WEST can handle full load
- Test failback procedures
🎓 Chapter Summary
You now understand Cassandra's Network Topology!
Key Concepts:
- Three-level hierarchy: Cluster → Datacenter → Rack → Node
- Datacenters enable geographic distribution
- Racks provide failure domain isolation
- NetworkTopologyStrategy is rack and DC aware
- Proper topology prevents catastrophic failures
Responsive Ad