Section 2: Architecture & Design

Cassandra Network Topology

Master datacenter and rack organization! Deep dive into NetworkTopologyStrategy, snitch configuration, multi-region deployments, and the geographic distribution that powers global-scale applications.

📖 The Story: The Global Retail Chain

Imagine MegaMart - a retail giant with 10,000 stores across 5 continents. How do you organize inventory data so customers get instant updates, regardless of location?

❌ THE BAD ARCHITECTURE: Single Warehouse

The Problem: All inventory data stored in ONE central warehouse in Kansas City

🏢 Central Warehouse (Kansas City):
├─ Tokyo Store #1234 (7,500 miles away!) → 250ms latency
├─ London Store #5678 (4,000 miles away!) → 150ms latency
├─ Sydney Store #9012 (9,000 miles away!) → 300ms latency
└─ New York Store #3456 (1,200 miles away!) → 40ms latency

Disasters:
• Tornado hits Kansas → ALL stores go offline!
• Network cable cut → GLOBAL OUTAGE!
• Peak holiday traffic → Single warehouse overwhelmed!
• Tokyo customer waits 300ms for product info! 😡

✅ THE BRILLIANT ARCHITECTURE: Regional Distribution Centers

The Solution: 5 Regional Distribution Centers with hierarchical organization

🌎 LEVEL 1: CONTINENTS (Datacenters)
├─ 🇺🇸 North America DC (US-EAST)
├─ 🇪🇺 Europe DC (EU-WEST)
├─ 🇯🇵 Asia DC (ASIA-PAC)
├─ 🇦🇺 Australia DC (AU-SOUTH)
└─ 🇧🇷 South America DC (SA-EAST)

🏢 LEVEL 2: BUILDINGS (Racks)
Each DC has 3 buildings:
  Building A (Rack-A) - Main power circuit
  Building B (Rack-B) - Backup power circuit
  Building C (Rack-C) - Generator power

💻 LEVEL 3: SERVERS (Nodes)
Each Building has 100 servers
  Total: 5 DCs × 3 Racks × 100 Nodes = 1,500 nodes!

✨ MAGIC OF REPLICATION:
Product #ABC123 stored in:
→ US-EAST: Rack-A/Server-45, Rack-B/Server-78, Rack-C/Server-23
→ EU-WEST: Rack-A/Server-12, Rack-B/Server-56, Rack-C/Server-89
→ ASIA-PAC: Rack-A/Server-34, Rack-B/Server-67, Rack-C/Server-90

Total: 9 copies across 3 continents!

Why It's GENIUS:

  • Low Latency: Tokyo customer → ASIA-PAC DC (5ms vs 250ms!)
  • Fault Tolerance: Tornado in US? EU and ASIA serve customers!
  • Power Failure: Rack-A loses power? Rack-B & C still running!
  • Scalability: Add more buildings/servers per DC independently!
  • Data Sovereignty: EU data stays in EU (GDPR compliant!)

🎯 This is EXACTLY Cassandra's Network Topology!

  • Regional DCs = Cassandra Datacenters
  • Buildings = Cassandra Racks
  • Servers = Cassandra Nodes
  • Replication = NetworkTopologyStrategy
  • Power Circuits = Failure Domains
CASSANDRA NETWORK TOPOLOGY:

CREATE KEYSPACE inventory WITH replication = {
  'class': 'NetworkTopologyStrategy',
  'US-EAST': 3, // 3 copies in North America
  'EU-WEST': 3, // 3 copies in Europe
  'ASIA-PAC': 3 // 3 copies in Asia
};

Snitch: GossipingPropertyFileSnitch
Rack-aware: Yes (Rack-A, Rack-B, Rack-C)

Result: Global distribution with local speed! ✅

MegaMart's regional DCs = Cassandra's Network Topology!
Same brilliance, global scale!

🗺️ What is Network Topology in Cassandra?

The hierarchical organization of nodes across datacenters and racks for geographic distribution and fault isolation.

Complete Definition

Network Topology: The logical and physical organization of Cassandra nodes into a hierarchy of datacenters and racks, enabling geographic distribution, fault isolation, and optimized replication strategies.

Three-Level Hierarchy:

  • Cluster: Top level - entire Cassandra deployment
  • Datacenter: Geographic or logical grouping of nodes
  • Rack: Failure domain within a datacenter
  • Node: Individual server running Cassandra

Why Topology Matters:

  • Enables multi-region deployments for global applications
  • Provides fault isolation at rack and datacenter level
  • Optimizes data locality for low-latency reads
  • Supports disaster recovery and business continuity
  • Enables compliance with data sovereignty requirements
Cassandra Network Topology Hierarchy CLUSTER: Production-Cluster 🇺🇸 DATACENTER: US-EAST Rack-A N1 N2 N3 ... Rack-B N4 N5 N6 ... Rack-C N7 N8 N9 ... 🇪🇺 DATACENTER: EU-WEST Rack-A N10 N11 Rack-B N12 N13 Rack-C N14 N15 🇯🇵 DATACENTER: ASIA-PAC Rack-A N16 N17 Rack-B N18 N19 Rack-C N20 N21 Each datacenter has 3 racks (failure domains), each rack has multiple nodes

🏢 Datacenters: Geographic Distribution

How Cassandra organizes nodes across geographic or logical boundaries.

Understanding Datacenters

Physical Datacenters: Actual geographic locations (AWS us-east-1, Google us-central1, on-prem Virginia DC)

Logical Datacenters: Workload separation (analytics DC, online-transactional DC, search DC)

Datacenter Benefits:

  • Low Latency: Serve users from nearest DC (Tokyo users → ASIA-PAC DC)
  • Disaster Recovery: Entire DC can fail, others continue serving
  • Data Sovereignty: EU data stays in EU (GDPR, legal compliance)
  • Workload Isolation: Heavy analytics don't impact online traffic
  • Maintenance Windows: Upgrade one DC at a time (zero downtime)
🌍

Geographic Datacenters

Physical locations worldwide

US-EAST (Virginia)
US-WEST (Oregon)
EU-WEST (Ireland)
EU-CENTRAL (Frankfurt)
ASIA-PAC (Tokyo)
ASIA-SOUTH (Singapore)
AU-SOUTHEAST (Sydney)

Use Case: Global applications, serve users locally, 5-20ms latency worldwide

⚙️

Logical Datacenters

Workload separation

ONLINE-APP (user facing)
ANALYTICS (batch jobs)
SEARCH (ElasticSearch mirror)
BACKUP (disaster recovery)
DEV (development/testing)
REPORTING (BI queries)

Use Case: Heavy analytics don't slow down online traffic, dedicated resources per workload

☁️

Cloud Provider Datacenters

Cloud region mapping

AWS:
us-east-1, us-west-2
eu-west-1, ap-southeast-1

GCP:
us-central1, europe-west1
asia-northeast1

Azure:
eastus, westeurope
southeastasia

Use Case: Multi-cloud deployments, cloud-native applications

Critical: Datacenter Naming Rules

  • Cannot Change: Once set, datacenter names are PERMANENT
  • Alphanumeric Only: Letters, numbers, underscores, hyphens (no spaces)
  • Case Sensitive: "US-EAST" ≠ "us-east"
  • Descriptive: Use meaningful names (us-east-1, not dc1)
  • Document: Keep datacenter topology documentation updated

📚 Racks: Failure Domain Isolation

The second level of hierarchy providing fault isolation within datacenters.

What is a Rack?

Physical Rack: Literal server rack with shared power supply and network switch

Logical Rack: Availability zone, separate network segment, or failure domain

Why Rack Awareness Matters:

Without Rack Awareness:

  • All 3 replicas might be on same rack
  • Power failure = lose all 3 copies = DATA LOST!
  • Network switch failure = all replicas unreachable

With Rack Awareness:

  • Replicas distributed across different racks
  • Power failure on Rack-A = Rack-B & C still serve data
  • Can lose ANY rack and still have all data available

Real-World Example: Amazon Power Outage 2017

💥 The Disaster: US-EAST-1 Power Failure

Date: February 28, 2017

Incident: Power distribution unit failure in one availability zone

COMPANY A (No Rack Awareness):
3-node Cassandra cluster in US-EAST-1
All 3 nodes in same availability zone (same rack)

Power fails → ALL 3 nodes go offline
RF=3 but ALL replicas lost!
Result: 4-hour complete outage! 💥
Customers couldn't log in, checkout, or access data
Estimated loss: $500K/hour

COMPANY B (With Rack Awareness):
3-node Cassandra cluster in US-EAST-1
Nodes distributed across 3 availability zones:
  Node 1 → us-east-1a (Rack-A)
  Node 2 → us-east-1b (Rack-B)
  Node 3 → us-east-1c (Rack-C)

Power fails in us-east-1a → Only Node 1 offline
Nodes 2 & 3 still serving (QUORUM intact!)
Result: ZERO downtime! ✅
Customers experienced no disruption
Traffic automatically routed to healthy nodes

Rack Configuration Best Practices

✅

Best Practice

  • Minimum 3 racks per datacenter
  • Equal nodes per rack
  • One replica per rack with RF=3
  • Different power circuits per rack
  • Different network switches
  • Cloud: Map to Availability Zones
❌

Anti-Pattern

  • Single rack: No fault isolation
  • Two racks: Can't survive 1 rack failure
  • Uneven distribution: Some racks empty
  • Default "rack1": All nodes same rack
  • Ignore cloud AZs: Wasted isolation

👀 Snitch: The Topology Detective

How Cassandra learns the datacenter and rack location of every node in the cluster.

What is a Snitch?

Snitch: A component that determines the datacenter and rack for each node, enabling NetworkTopologyStrategy to make intelligent replica placement decisions.

What Snitch Tells Cassandra:

  • Datacenter Name: Which DC does this node belong to?
  • Rack Name: Which rack within the DC?
  • Relative Distance: How "close" are two nodes?
  • Routing Decisions: Which node should handle a request?

Why Snitch is Critical:

Without Snitch:

  • Cassandra thinks all nodes are in same location
  • Replicas might all go to same rack (no fault tolerance!)
  • Queries routed randomly (high latency!)
  • Can't enforce data sovereignty (legal violations!)

With Snitch:

  • Replicas spread across racks and DCs intelligently
  • Queries routed to nearest node (low latency)
  • Data stays in correct geographic region
  • Survives rack and DC failures gracefully

Types of Snitches

📝

GossipingPropertyFileSnitch

RECOMMENDED for production

# cassandra-rackdc.properties
dc=US-EAST
rack=RACK-A
prefer_local=true

How it works:
• Each node reads local file
• Gossips topology to cluster
• Auto-discovers other nodes
• No centralized config!
  • Pros: Simple, flexible, self-healing
  • Cons: Manual file editing
  • Use: 95% of deployments
☁️

Ec2Snitch / Ec2MultiRegionSnitch

AWS-specific auto-detection

Auto-detects from AWS metadata:
• Region → Datacenter
• Availability Zone → Rack
• Instance metadata API

Example:
us-east-1a → DC: us-east-1, Rack: 1a
eu-west-1b → DC: eu-west-1, Rack: 1b
ap-south-1c → DC: ap-south-1, Rack: 1c
  • Pros: Zero config, AWS-native
  • Cons: AWS only, uses public IPs
  • Use: AWS deployments
🔧

GoogleCloudSnitch

GCP-specific auto-detection

Auto-detects from GCP metadata:
• Region → Datacenter
• Zone → Rack

Example:
us-central1-a → DC: us-central1, Rack: a
europe-west1-b → DC: europe-west1, Rack: b
asia-east1-c → DC: asia-east1, Rack: c
  • Pros: Zero config, GCP-native
  • Cons: GCP only
  • Use: Google Cloud deployments
⚠️

SimpleSnitch

DO NOT USE in production

Treats all nodes as:
• Same datacenter
• Same rack
• No topology awareness!

Result:
• All replicas on same rack
• No fault isolation
• Random query routing
  • Pros: None (seriously!)
  • Cons: Dangerous, no isolation
  • Use: Testing ONLY
How Snitch Works: Topology Discovery Process STEP 1 Node Starts Up cassandra.yaml loaded STEP 2 Snitch Activated Read local topology STEP 3 Determine Location DC: US-EAST Rack: RACK-A STEP 4 Gossip to Cluster Share topology info Configuration File cassandra-rackdc.properties # Node Configuration dc=US-EAST rack=RACK-A prefer_local=true # This node is in: # Datacenter: US-EAST # Rack: RACK-A # Prefers local reads # Restart required after changes Gossip Protocol Shares Topology Node 1 US-EAST N2 N3 N4 N5 All nodes learn each other's datacenter and rack location

Critical: Changing Snitch in Production

WARNING: Changing snitch in running cluster is EXTREMELY dangerous!

What happens if you change snitch incorrectly:

  • Cassandra thinks nodes moved to different DCs/racks
  • Replication breaks (wrong replicas deleted/created)
  • Data loss possible (old replicas dropped)
  • Queries fail (can't find correct replicas)
  • Cluster becomes inconsistent

Safe Snitch Migration Steps:

  1. Run full cluster repair first
  2. Change snitch on ONE node
  3. Restart that node
  4. Verify topology correct: nodetool status
  5. If correct, proceed to next node
  6. If wrong, revert immediately!
  7. Run repair after all nodes migrated

🎯 Replication Strategies: SimpleStrategy vs NetworkTopologyStrategy

Understanding how topology enables intelligent replication.

Feature SimpleStrategy NetworkTopologyStrategy
Topology Aware ❌ No ✅ Yes
Multi-Datacenter ❌ Not supported ✅ Fully supported
Rack Aware ❌ Ignores racks ✅ Respects racks
Replica Placement Clockwise on ring Smart: different racks/DCs
Configuration {'class': 'SimpleStrategy', 'replication_factor': 3} {'class': 'NetworkTopologyStrategy', 'DC1': 3, 'DC2': 2}
Fault Tolerance Low (no isolation) High (rack + DC isolation)
Production Use ❌ NEVER ✅ ALWAYS

Expert: How NetworkTopologyStrategy Places Replicas

Intelligent Algorithm:

  1. Per Datacenter: Process each DC independently
  2. Walk the Ring: Start from primary replica, walk clockwise
  3. Skip Same Rack: If next node in same rack as existing replica, skip it
  4. Choose Different Rack: Continue until finding node in different rack
  5. Repeat: Continue until RF replicas placed in this DC
  6. Next DC: Repeat process for next datacenter

Example: RF=3 in DC1, RF=2 in DC2

DC1 (US-EAST) - Need 3 replicas:
Primary: Node 1 (Rack-A)
Walk ring → Node 2 (Rack-A) ❌ Skip! (same rack)
Walk ring → Node 3 (Rack-B) ✅ Place replica 2
Walk ring → Node 4 (Rack-B) ❌ Skip! (same rack)
Walk ring → Node 5 (Rack-C) ✅ Place replica 3

Result: Replicas on Rack-A, Rack-B, Rack-C ✅

DC2 (EU-WEST) - Need 2 replicas:
Primary: Node 10 (Rack-A)
Walk ring → Node 11 (Rack-B) ✅ Place replica 2

Result: Replicas on Rack-A, Rack-B ✅

Total: 5 copies (3 in US, 2 in EU) across different racks!

🖥️ Interactive Topology Builder

Build and visualize your own multi-datacenter topology!

Network Topology Simulator
🌐 Topology Builder Ready!
Configure your cluster topology and click "Build Topology" to see the complete structure.

Try different configurations:
• 3 Datacenters × 3 Racks × 3 Nodes = 27 total nodes
• 2 Datacenters × 4 Racks × 5 Nodes = 40 total nodes
• Click "Show Replication" to see replica placement with RF=3!

🌍 Multi-Region Deployments: Global Scale

Real-world patterns for deploying Cassandra across continents.

Production Example: Global Social Network

Company: GlobalConnect - 500M users worldwide, 100M DAU

Scale: 300-node cluster across 5 datacenters

Topology Design

Datacenters & Nodes:
🇺🇸 US-EAST-1 (Virginia): 90 nodes (3 racks × 30 nodes)
🇺🇸 US-WEST-2 (Oregon): 90 nodes (3 racks × 30 nodes)
🇪🇺 EU-WEST-1 (Ireland): 60 nodes (3 racks × 20 nodes)
🇯🇵 ASIA-PAC-1 (Tokyo): 30 nodes (3 racks × 10 nodes)
🇧🇷 SA-EAST-1 (São Paulo): 30 nodes (3 racks × 10 nodes)

Total: 300 nodes across 5 continents

Replication Strategy

User Profiles Keyspace:
CREATE KEYSPACE user_profiles WITH replication = {
  'class': 'NetworkTopologyStrategy',
  'US-EAST-1': 3, // Primary region
  'US-WEST-2': 3, // DR + West coast users
  'EU-WEST-1': 3, // European users
  'ASIA-PAC-1': 2, // Asian users (lighter load)
  'SA-EAST-1': 2 // South America (lighter load)
};

Total: 13 copies of every user profile worldwide!

User Sessions Keyspace (time-sensitive):
CREATE KEYSPACE user_sessions WITH replication = {
  'class': 'NetworkTopologyStrategy',
  'US-EAST-1': 3,
  'US-WEST-2': 3,
  'EU-WEST-1': 3,
  'ASIA-PAC-1': 3,
  'SA-EAST-1': 3
};

Total: 15 copies for ultra-high availability!

Consistency Levels by Region

  • Writes: LOCAL_QUORUM (low latency, DC-local consistency)
  • Reads: LOCAL_ONE (ultra-fast, serve from nearest replica)
  • Critical Ops: EACH_QUORUM (write to all DCs synchronously)
  • Background Repair: Keeps DCs in sync asynchronously

Latency Results

User Location Nearest DC Read Latency Write Latency
New York US-EAST-1 1-2ms 3-5ms
San Francisco US-WEST-2 1-2ms 3-5ms
London EU-WEST-1 2-3ms 4-6ms
Tokyo ASIA-PAC-1 2-3ms 5-7ms
São Paulo SA-EAST-1 2-4ms 5-8ms

Result: Sub-10ms latency for 99% of users worldwide!

💼 Interview Questions & Answers

Master these 25 essential questions on Cassandra Network Topology - from beginner to expert level!

1 What is Network Topology in Cassandra and why is it important? ▼

Answer:

Network Topology is the hierarchical organization of Cassandra nodes into datacenters and racks, enabling geographic distribution and fault isolation.

Three-Level Hierarchy:

  • Cluster: Top level - entire deployment
  • Datacenter: Geographic or logical grouping (e.g., US-EAST, EU-WEST)
  • Rack: Failure domain within a datacenter (e.g., Rack-A, Rack-B)

Why It's Critical:

  • Low Latency: Serve users from nearest datacenter (Tokyo users → ASIA-PAC DC = 5ms vs US-EAST = 250ms)
  • Fault Isolation: Power failure on Rack-A doesn't affect Rack-B and Rack-C
  • Disaster Recovery: Entire datacenter can fail, others continue serving
  • Data Sovereignty: EU data stays in EU (GDPR compliance)
  • Smart Replication: NetworkTopologyStrategy spreads replicas across racks/DCs
2 What is a Snitch in Cassandra? Explain its role. ▼

Answer:

A Snitch is a component that determines the datacenter and rack location of each node, enabling NetworkTopologyStrategy to make intelligent replica placement decisions.

What Snitch Provides:

  • Datacenter Name: Which DC does this node belong to? (e.g., US-EAST)
  • Rack Name: Which rack within the DC? (e.g., RACK-A)
  • Relative Distance: How "close" are two nodes for routing decisions
  • Query Routing: Which node should handle a request (prefer local)

Without Snitch:

  • All nodes treated as same location
  • Replicas might all go to same rack (no fault tolerance!)
  • Queries routed randomly (high latency)
  • Can't enforce data sovereignty

With Snitch:

  • Replicas intelligently spread across racks and DCs
  • Queries routed to nearest node (low latency)
  • Data stays in correct geographic region
  • Survives rack and DC failures gracefully
3 Compare GossipingPropertyFileSnitch vs Ec2Snitch. When would you use each? ▼

Answer:

GossipingPropertyFileSnitch:

  • Configuration: Read from cassandra-rackdc.properties file on each node
  • Setup: Manual - you specify dc=US-EAST, rack=RACK-A
  • Discovery: Gossip protocol shares topology with cluster
  • Flexibility: Works anywhere - on-prem, any cloud, hybrid
  • Use When: 95% of deployments, any environment, need full control
  • Pros: Simple, flexible, self-healing, cloud-agnostic
  • Cons: Manual file editing required

Ec2Snitch / Ec2MultiRegionSnitch:

  • Configuration: Auto-detects from AWS EC2 metadata API
  • Setup: Zero configuration required
  • Mapping: AWS Region → Datacenter, Availability Zone → Rack
  • Example: us-east-1a becomes DC: us-east-1, Rack: 1a
  • Use When: 100% AWS deployment, want zero-config
  • Pros: No manual config, AWS-native, auto-updates
  • Cons: AWS only, uses public IPs (Ec2Snitch), not portable

Recommendation:

Use GossipingPropertyFileSnitch even on AWS for better control and portability. Only use Ec2Snitch if you absolutely need zero-config and will never migrate off AWS.

4 What is the difference between SimpleStrategy and NetworkTopologyStrategy? ▼

Answer:

SimpleStrategy:

  • Topology Awareness: None - treats all nodes as equal
  • Replication: Places replicas clockwise on ring, ignoring DC/rack
  • Configuration: Single RF for entire cluster
  • Example: {'class': 'SimpleStrategy', 'replication_factor': 3}
  • Multi-DC Support: Not supported - can't specify per-DC RF
  • Fault Tolerance: Low - all replicas might be on same rack
  • Production Use: ❌ NEVER - only for testing/dev

NetworkTopologyStrategy:

  • Topology Awareness: Fully aware of datacenters and racks
  • Replication: Intelligently spreads replicas across different racks
  • Configuration: Per-datacenter RF specification
  • Example: {'class': 'NetworkTopologyStrategy', 'US-EAST': 3, 'EU-WEST': 2}
  • Multi-DC Support: Fully supported - different RF per DC
  • Fault Tolerance: High - guarantees replicas on different racks/DCs
  • Production Use: ✅ ALWAYS - even with single DC

Key Difference:

NetworkTopologyStrategy ensures replicas are on different racks within each DC. With RF=3 and 3 racks, SimpleStrategy might put all 3 replicas on Rack-A (disaster!), while NetworkTopologyStrategy guarantees one replica per rack (safe!).

5 How does Cassandra determine which rack to place replicas on? ▼

Answer:

NetworkTopologyStrategy uses an intelligent algorithm to ensure replicas are on different racks:

Algorithm Steps:

  1. Start: Primary replica placed based on partition key hash (token)
  2. Walk Ring: Move clockwise around token ring
  3. Check Rack: Is next node on same rack as existing replica?
  4. Skip or Place: If same rack → skip, if different rack → place replica
  5. Repeat: Continue until RF replicas placed in this DC
  6. Next DC: Repeat entire process for next datacenter

Example (RF=3, 3 racks):

  • Primary: Token hash → Node 5 (Rack-B)
  • Walk ring: Node 6 (Rack-B) ❌ Skip! Same rack as Node 5
  • Continue: Node 7 (Rack-C) ✅ Place replica 2 (different rack)
  • Continue: Node 8 (Rack-C) ❌ Skip! Same rack as Node 7
  • Continue: Node 1 (Rack-A) ✅ Place replica 3 (different rack)
  • Result: Replicas on Rack-A, Rack-B, Rack-C ✅

Guarantee:

With 3+ racks and RF=3, you're guaranteed one replica per rack. If rack fails, other 2 racks still have complete data!

6 What happens if you change the Snitch on a running production cluster? ▼

Answer:

WARNING: This is EXTREMELY dangerous and can cause data loss!

What Goes Wrong:

  • Topology Confusion: Cassandra thinks nodes moved to different DCs/racks
  • Replication Breaks: Wrong replicas might be deleted, new ones created
  • Data Loss Risk: Old replicas dropped before new ones fully replicated
  • Query Failures: Can't find correct replicas, routing errors
  • Consistency Issues: Cluster becomes inconsistent, repair can't fix
  • Split Brain: Different nodes think data is in different places

Safe Migration Procedure:

  1. Plan: Document current topology, plan new topology
  2. Backup: Take full snapshots of entire cluster
  3. Repair: Run full cluster repair first (nodetool repair)
  4. Maintenance Window: Schedule downtime if possible
  5. One Node: Change snitch on ONE node only
  6. Update Config: Modify cassandra.yaml and cassandra-rackdc.properties
  7. Restart: Restart that one node
  8. Verify: Check topology: nodetool status, nodetool describecluster
  9. Monitor: Watch logs for errors, check metrics
  10. Iterate: If correct, proceed to next node. If wrong, REVERT!
  11. Final Repair: After all nodes migrated, run full repair

Best Practice:

Get snitch configuration right from the start! Changing it later is risky and time-consuming.

7 Explain the concept of "rack awareness" in Cassandra. ▼

Answer:

Rack awareness is Cassandra's ability to understand physical failure domains (racks) and distribute replicas accordingly to prevent correlated failures.

Physical Rack Concept:

  • Physical Rack: Literal server rack with shared power supply and network switch
  • Failure Domain: If power fails or switch breaks, entire rack goes offline
  • Cloud Equivalent: Availability Zone in AWS/GCP/Azure

Without Rack Awareness:

  • All 3 replicas (RF=3) might be on same rack
  • Power failure on that rack = lose ALL 3 copies = DATA LOST!
  • Network switch failure = all replicas unreachable = OUTAGE!
  • Single point of failure despite having replicas

With Rack Awareness (NetworkTopologyStrategy):

  • Replicas automatically spread across different racks
  • RF=3 + 3 racks = one replica per rack guaranteed
  • Power failure on Rack-A = Rack-B and Rack-C still serve data
  • Can lose ANY single rack without data loss or downtime

Real-World Example:

Amazon 2017 Power Outage: PDU failure in us-east-1a

  • No Rack Awareness: All 3 nodes in us-east-1a → 4-hour outage, $500K/hour loss
  • With Rack Awareness: Nodes in us-east-1a, 1b, 1c → Zero downtime! Other AZs served traffic

Best Practice:

Minimum 3 racks per datacenter. In cloud: map each rack to different Availability Zone.

8 How do you configure a 3-datacenter cluster with different replication factors? ▼

Answer:

Use NetworkTopologyStrategy with per-datacenter replication factor specification:

Configuration Example:

CREATE KEYSPACE my_keyspace WITH replication = {
  'class': 'NetworkTopologyStrategy',
  'US-EAST': 3, // 3 replicas in US-EAST
  'EU-WEST': 2, // 2 replicas in EU-WEST
  'ASIA-PAC': 2 // 2 replicas in ASIA-PAC
};

-- Total: 7 copies of every row!

Why Different RFs?

  • US-EAST (RF=3): Primary datacenter, most users, highest traffic
  • EU-WEST (RF=2): Secondary datacenter, moderate traffic
  • ASIA-PAC (RF=2): Read-only analytics DC, lower requirements

Data Distribution:

  • User record for user_id='alice' stored in 7 nodes total:
  • US-EAST: Node 5 (Rack-A), Node 12 (Rack-B), Node 19 (Rack-C)
  • EU-WEST: Node 34 (Rack-A), Node 41 (Rack-B)
  • ASIA-PAC: Node 56 (Rack-A), Node 63 (Rack-B)

Consistency Levels:

  • Writes: LOCAL_QUORUM (2 nodes in local DC)
  • Reads: LOCAL_ONE (1 node in local DC, ultra-fast)
  • Critical: EACH_QUORUM (QUORUM in every DC)

Verification:

DESCRIBE KEYSPACE my_keyspace;
nodetool describecluster;
nodetool status my_keyspace;
9 What is the purpose of prefer_local in cassandra-rackdc.properties? ▼

Answer:

prefer_local controls whether Cassandra prefers reading from replicas in the local datacenter.

Configuration:

# cassandra-rackdc.properties
dc=US-EAST
rack=RACK-A
prefer_local=true # Default and recommended

When prefer_local=true (RECOMMENDED):

  • Queries prefer replicas in same datacenter
  • Lower latency (1-5ms vs 100-200ms cross-DC)
  • Reduced WAN bandwidth costs
  • Better user experience (faster responses)
  • Only contacts remote DC if local replicas unavailable

When prefer_local=false:

  • Queries may go to any datacenter
  • Higher latency (unpredictable)
  • Higher WAN costs
  • Useful for testing or special analytics workloads

Example Impact:

  • User in New York, US-EAST DC:
  • prefer_local=true: Read from US-EAST replica (2ms latency)
  • prefer_local=false: Might read from EU-WEST replica (150ms latency!)

Interaction with Consistency Levels:

  • LOCAL_QUORUM: Always reads from local DC (prefer_local irrelevant)
  • QUORUM: Uses prefer_local to choose which DC's replicas
  • ONE: Uses prefer_local to pick replica location

Best Practice:

Always set prefer_local=true in production. Only set false for specific analytics/reporting workloads.

10 How does Cassandra handle queries when an entire datacenter goes offline? ▼

Answer:

Cassandra automatically routes queries to healthy datacenters when one DC fails, with behavior depending on consistency level:

Scenario: US-EAST datacenter goes offline

With LOCAL_QUORUM (RECOMMENDED):

  • US-EAST clients: Queries FAIL (local DC down, can't reach LOCAL_QUORUM)
  • Solution: Load balancer redirects US-EAST traffic to US-WEST
  • US-WEST clients: Unaffected, continue serving from US-WEST replicas
  • EU-WEST clients: Unaffected, continue serving from EU-WEST replicas
  • Benefit: Controlled failover, no cross-DC latency impact

With QUORUM (Global):

  • All queries continue working!
  • Needs (Total RF / 2) + 1 replicas globally
  • Example: RF=9 (3 per DC × 3 DCs) → QUORUM=5
  • US-EAST down (3 replicas lost) → 6 replicas remain → Still achieve QUORUM
  • Downside: Higher latency (might query remote DCs)

With EACH_QUORUM:

  • All writes FAIL!
  • Requires QUORUM in EVERY datacenter
  • US-EAST down → Can't reach QUORUM there → Write rejected
  • Use case: Strong consistency, but poor availability

Automatic Failover Example:

Before DC Failure:

  • US-EAST clients → US-EAST replicas (2ms)
  • EU-WEST clients → EU-WEST replicas (2ms)

After US-EAST Fails:

  • US-EAST clients → Redirected to US-WEST (20ms, slight increase)
  • EU-WEST clients → Still EU-WEST replicas (2ms, no change)

Recovery Process:

  1. Detection: Gossip protocol detects US-EAST nodes down (~30 seconds)
  2. Routing: Coordinator stops sending queries to US-EAST
  3. Hinted Handoff: Writes queued for US-EAST nodes
  4. DC Restored: Hints replayed, nodes catch up
  5. Repair: Run nodetool repair to ensure consistency

Best Practice:

Use LOCAL_QUORUM for writes and LOCAL_ONE for reads. Set up load balancer to redirect traffic from failed DC to nearest healthy DC.

11 Why should you use NetworkTopologyStrategy even with a single datacenter? ▼

Answer:

Always use NetworkTopologyStrategy, even with single DC:

  • Rack Awareness: Still need fault isolation within DC (different racks/AZs)
  • Future Expansion: Easy to add DCs later without changing keyspace
  • Best Practice: Consistent approach across all environments
  • No Downside: Same performance as SimpleStrategy in single DC
  • Production Ready: SimpleStrategy is for testing only

Even with single DC, use: {'class': 'NetworkTopologyStrategy', 'DC1': 3}

12 How do cloud Availability Zones map to Cassandra racks? ▼

Answer:

Cloud Availability Zones (AZs) are perfect Cassandra racks:

  • AWS: us-east-1a → DC: us-east-1, Rack: 1a
  • GCP: us-central1-a → DC: us-central1, Rack: a
  • Azure: eastus-1 → DC: eastus, Rack: 1
  • Isolation: AZs have separate power, cooling, networking
  • Best Practice: Deploy nodes across 3+ AZs in each region
  • Configuration: Set rack=1a, rack=1b, rack=1c in cassandra-rackdc.properties

With 3 AZs and RF=3, you get one replica per AZ - survives any AZ failure!

13 What is the minimum number of racks recommended per datacenter and why? ▼

Answer:

Minimum: 3 racks per datacenter

  • With RF=3: One replica per rack guaranteed
  • Fault Tolerance: Can lose any 1 rack, still have QUORUM (2/3)
  • With 2 Racks: Can't survive single rack failure with QUORUM
  • With 1 Rack: No fault isolation at all!

Math:

  • 3 racks, RF=3: Rack fails → 2 replicas remain → QUORUM(2) achieved ✅
  • 2 racks, RF=3: Rack fails → 1-2 replicas remain → QUORUM(2) might fail ❌
14 Explain data sovereignty and how Cassandra topology helps achieve it. ▼

Answer:

Data Sovereignty: Legal requirement that data must stay within specific geographic boundaries.

  • GDPR (EU): EU citizen data must stay in EU
  • Russia: Russian citizen data must stay in Russia
  • China: Chinese data must stay in China

How Cassandra Helps:

  • Datacenter Isolation: Create separate EU-WEST datacenter
  • Keyspace Design: EU users → eu_users keyspace with RF only in EU-WEST
  • No Cross-DC: EU data never replicates to US or Asia DCs
  • Local Queries: EU queries only hit EU-WEST nodes

Configuration Example:

CREATE KEYSPACE eu_users WITH replication = {
  'class': 'NetworkTopologyStrategy',
  'EU-WEST': 3 -- ONLY in EU, never leaves!
};
15 How would you design topology for a global e-commerce platform? ▼

Answer:

Scenario: Global e-commerce with 100M users, 1M products, 10K orders/second

Datacenter Design:

  • US-EAST: 100 nodes (3 racks × 33 nodes) - Primary region
  • US-WEST: 100 nodes (3 racks × 33 nodes) - DR + West coast
  • EU-WEST: 60 nodes (3 racks × 20 nodes) - European customers
  • ASIA-PAC: 40 nodes (3 racks × 13 nodes) - Asian customers
  • Total: 300 nodes across 4 continents

Keyspace Strategy:

-- Product Catalog (read-heavy, changes rarely)
CREATE KEYSPACE products WITH replication = {
  'class': 'NetworkTopologyStrategy',
  'US-EAST': 3, 'US-WEST': 3,
  'EU-WEST': 3, 'ASIA-PAC': 3
};

-- User Sessions (time-sensitive, high availability)
CREATE KEYSPACE sessions WITH replication = {
  'class': 'NetworkTopologyStrategy',
  'US-EAST': 3, 'US-WEST': 3,
  'EU-WEST': 2, 'ASIA-PAC': 2
};

-- Orders (critical, strong consistency)
CREATE KEYSPACE orders WITH replication = {
  'class': 'NetworkTopologyStrategy',
  'US-EAST': 3, 'US-WEST': 3
};
-- Orders only in US (payment processing, legal)

Consistency Levels:

  • Products: Write EACH_QUORUM, Read LOCAL_ONE (eventual consistency OK)
  • Sessions: Write LOCAL_QUORUM, Read LOCAL_ONE (fast, local)
  • Orders: Write QUORUM, Read QUORUM (strong consistency)
16 What files need to be configured for network topology on each node? ▼

Answer:

Two Files Required:

1. cassandra.yaml:

endpoint_snitch: GossipingPropertyFileSnitch

2. cassandra-rackdc.properties:

dc=US-EAST
rack=RACK-A
prefer_local=true

Important Notes:

  • Both files in /etc/cassandra/ directory (Debian/Ubuntu) or /etc/dse/cassandra (DSE)
  • Requires node restart after changes
  • DC/rack names case-sensitive and permanent
  • Each node can have different dc/rack values
17 How does gossip protocol share topology information? ▼

Answer:

Gossip protocol is how nodes discover and share topology information:

  • Every Second: Each node picks 1-3 random nodes to gossip with
  • Exchange Info: Share state: DC/rack, health, schema, load
  • Propagation: Information spreads exponentially through cluster
  • Convergence: Entire cluster knows topology within ~30 seconds
  • Self-Healing: Automatically detects and shares node changes

Example:

  • Node 1 starts with dc=US-EAST, rack=RACK-A
  • Node 1 gossips to Node 5: "I'm in US-EAST/RACK-A"
  • Node 5 gossips to Node 12: "Node 1 is US-EAST/RACK-A"
  • After few rounds: All nodes know Node 1's location
18 What commands verify cluster topology? ▼

Answer:

Key Commands:

1. nodetool status:

nodetool status

Datacenter: US-EAST
Status=Up/Down, State=Normal/Leaving/Joining
-- Address Load Tokens Owns Rack
UN 10.0.1.5 256GB 256 33% RACK-A
UN 10.0.1.6 250GB 256 32% RACK-B

2. nodetool describecluster:

nodetool describecluster

Name: ProductionCluster
Snitch: GossipingPropertyFileSnitch
Datacenter: US-EAST (3 racks)
Datacenter: EU-WEST (3 racks)

3. nodetool ring:

nodetool ring
# Shows token ring with DC/rack for each token

4. CQL DESCRIBE:

DESCRIBE CLUSTER;
DESCRIBE KEYSPACE my_keyspace;
# Shows replication configuration
19 Can you have different number of nodes in each datacenter? ▼

Answer:

YES! Different DCs can have different node counts - this is common:

Common Scenarios:

  • Primary/Secondary DCs:
    • US-EAST (primary): 100 nodes
    • US-WEST (DR): 50 nodes
    • Different workloads justify different sizes
  • Analytics DC:
    • ONLINE (user-facing): 200 nodes
    • ANALYTICS (batch jobs): 50 nodes
    • Analytics runs overnight, needs fewer nodes
  • Regional Traffic:
    • US-EAST: 150 nodes (high traffic)
    • EU-WEST: 100 nodes (moderate traffic)
    • ASIA-PAC: 50 nodes (growing market)

Important Considerations:

  • Token Distribution: Vnodes auto-adjust for different sizes
  • Minimum Nodes: Must have ≥ RF nodes in each DC for that keyspace
  • Load Balance: Smaller DC = more data per node

Example:

US-EAST: 90 nodes with RF=3 → ~33GB per node
EU-WEST: 30 nodes with RF=3 → ~100GB per node
Both valid! EU nodes handle more data but less traffic.
20 What happens if you misconfigure datacenter names across nodes? ▼

Answer:

DISASTER! This breaks replication and can cause data loss.

Example Misconfiguration:

Node 1: dc=US-EAST
Node 2: dc=US-EAST
Node 3: dc=us-east ❌ Wrong case!
Node 4: dc=USEAST ❌ Missing hyphen!

What Goes Wrong:

  • Multiple DCs Created: Cassandra sees 3 DCs: "US-EAST", "us-east", "USEAST"
  • Replication Breaks: Keyspace expects RF=3 in "US-EAST", but only 2 nodes match!
  • Writes Fail: Can't achieve QUORUM (need 2/3, only have 2 nodes in "US-EAST")
  • Data Loss Risk: Some nodes not receiving replicas
  • Query Routing Broken: Coordinator can't find enough replicas

Prevention:

  • Automation: Use config management (Ansible, Chef, Puppet)
  • Validation: Verify with nodetool status before adding to production
  • Standards: Document exact DC/rack naming convention
  • Case Sensitive: Use consistent capitalization (recommend ALL-CAPS)

Recovery:

  • Fix cassandra-rackdc.properties on misconfigured nodes
  • Restart nodes one at a time
  • Run nodetool repair on entire cluster
  • Verify topology with nodetool describecluster
21 How does rack-aware routing improve query performance? ▼

Answer:

Rack-aware routing optimizes query performance by preferring replicas on nodes within the same rack:

  • Lower Latency: Same rack = same network switch = sub-millisecond latency
  • Less Network Hops: Avoid top-of-rack switch → core switch → another rack
  • Higher Bandwidth: Within-rack links typically faster than cross-rack
  • Cost Savings: Reduced cross-rack network traffic

Example: Coordinator on RACK-A prefers reading from RACK-A replica (0.5ms) vs RACK-B replica (2ms).

22 Can you mix logical and physical datacenters in the same cluster? ▼

Answer:

YES! Common pattern for workload separation:

  • Physical DCs: US-EAST, EU-WEST (geographic locations)
  • Logical DCs: ANALYTICS, SEARCH (same physical location as US-EAST)
  • Use Case: Separate heavy analytics from online traffic

Example Configuration:

Physical DC US-EAST:
├─ Logical DC: ONLINE (100 nodes) - user queries
└─ Logical DC: ANALYTICS (50 nodes) - batch jobs

Keyspace replication:
'ONLINE': 3, # replicate to online nodes
'ANALYTICS': 2 # replicate to analytics nodes

Benefit: Heavy analytics queries don't impact online user experience!

23 What's the impact of topology on hinted handoff? ▼

Answer:

Topology determines which nodes store hints when replicas are unavailable:

  • Same DC Hints: Hints stored on nodes in same datacenter (by default)
  • Rack Awareness: Hints prefer different rack than down node
  • Cross-DC Hints: Disabled by default (configure with hinted_handoff_enabled)

Example:

  • Write for Node 5 (US-EAST/RACK-A) but it's down
  • Hint stored on Node 12 (US-EAST/RACK-B) - same DC, different rack
  • When Node 5 returns, Node 12 replays hints
24 How do you migrate from SimpleStrategy to NetworkTopologyStrategy? ▼

Answer:

Safe Migration Steps:

  1. Check Current: DESCRIBE KEYSPACE my_keyspace;
  2. Run Repair: nodetool repair -full my_keyspace
  3. Alter Keyspace:
    ALTER KEYSPACE my_keyspace WITH replication = {
      'class': 'NetworkTopologyStrategy',
      'DC1': 3 -- Same RF as before
    };
  4. Repair Again: nodetool repair -full my_keyspace
  5. Verify: nodetool status my_keyspace

Important: No downtime required! Replication happens asynchronously.

25 Design a disaster recovery strategy using multiple datacenters. ▼

Answer:

Comprehensive DR Strategy:

1. Multi-DC Architecture:

  • Primary: US-EAST (200 nodes, 3 racks)
  • DR: US-WEST (200 nodes, 3 racks) - mirror of primary
  • Remote DR: EU-WEST (100 nodes, 3 racks) - geographic diversity

2. Replication Configuration:

CREATE KEYSPACE critical_data WITH replication = {
  'class': 'NetworkTopologyStrategy',
  'US-EAST': 3, # Primary
  'US-WEST': 3, # Same-region DR
  'EU-WEST': 3 # Remote DR
};

3. Consistency Levels:

  • Normal: Write LOCAL_QUORUM, Read LOCAL_ONE
  • DR Event: Switch to QUORUM (cross-DC) or redirect traffic

4. Failover Plan:

  • Detection: Monitor US-EAST health (30 seconds to detect)
  • DNS Update: Point users to US-WEST (TTL=60s)
  • Traffic Shift: Load balancer redirects within 2 minutes
  • RTO: 5 minutes (Recovery Time Objective)
  • RPO: 0 seconds (no data loss - synchronous replication)

5. Testing:

  • Quarterly DR drills - simulate US-EAST failure
  • Verify US-WEST can handle full load
  • Test failback procedures

🎓 Chapter Summary

You now understand Cassandra's Network Topology!

Key Concepts:

  • Three-level hierarchy: Cluster → Datacenter → Rack → Node
  • Datacenters enable geographic distribution
  • Racks provide failure domain isolation
  • NetworkTopologyStrategy is rack and DC aware
  • Proper topology prevents catastrophic failures
Advertisement

Responsive Ad