Advanced Topics

Multi-Cloud Setup

Deploy Cassandra across AWS, GCP, Azure - active-active, disaster recovery, and hybrid cloud architectures!

☁️ Why Multi-Cloud Cassandra?

The Multi-Cloud Challenge 🌍

Your SaaS company serves customers globally:

  • 🌎 Users in US-East, EU-West, Asia-Pacific
  • ⚡ Need < 50ms latency for all regions
  • 🔥 Must survive entire cloud provider outage
  • 💰 Want to optimize costs across providers
  • 📜 Compliance requires data sovereignty (GDPR)
  • 🎯 Avoid vendor lock-in

Solution: Deploy Cassandra across multiple cloud providers for global availability, low latency, and resilience!

Multi-Cloud Benefits

Multi-cloud = Running infrastructure across multiple cloud providers (AWS, GCP, Azure) as a unified system.

Key Advantages:
  • 🌍 Global reach: Deploy close to users worldwide
  • 🛡️ High availability: Survive provider-level outages
  • ⚡ Low latency: Serve from nearest datacenter
  • 💰 Cost optimization: Use cheapest resources per region
  • 🔒 Data sovereignty: Keep data in required countries
  • 🚫 Avoid lock-in: Not dependent on single vendor
  • 🎯 Best-of-breed: Use each cloud's strengths

Use Cases

🌍

Global Applications

  • Netflix (worldwide streaming)
  • Uber (multi-region dispatch)
  • Gaming platforms
  • Social networks
  • E-commerce
🔥

Disaster Recovery

  • Financial services
  • Healthcare systems
  • Government
  • Critical infrastructure
  • Enterprise apps
📜

Compliance

  • GDPR (EU data residency)
  • HIPAA (US healthcare)
  • PCI-DSS (payment data)
  • Data localization laws
  • Industry regulations
💡

Hybrid Cloud

  • On-prem + cloud
  • Edge computing
  • Migration strategy
  • Legacy integration
  • Cost optimization

🏗️ Multi-Cloud Architectures

Architecture 1: Active-Active (Multi-Region)

ACTIVE-ACTIVE ARCHITECTURE ┌─────────────────┐ ┌─────────────────┐ ┌─────────────────┐ │ AWS US-EAST │◄───►│ GCP EU-WEST │◄───►│ AZURE ASIA │ │ Cassandra DC1 │ │ Cassandra DC2 │ │ Cassandra DC3 │ │ RF=3, 6 nodes │ │ RF=3, 6 nodes │ │ RF=3, 6 nodes │ └─────────────────┘ └─────────────────┘ └─────────────────┘ ▲ ▲ ▲ │ │ │ US Users EU Users Asia Users ✅ All DCs accept reads AND writes ✅ Data replicated globally ✅ Users connect to nearest DC (low latency) ✅ Survives entire DC/provider failure

Active-Active Configuration

# cassandra.yaml (each datacenter) # AWS US-EAST nodes: cluster_name: 'GlobalCassandra' endpoint_snitch: GossipingPropertyFileSnitch seed_provider: - class_name: org.apache.cassandra.locator.SimpleSeedProvider parameters: - seeds: "aws-seed1,gcp-seed1,azure-seed1" # cassandra-rackdc.properties (AWS nodes) dc=aws_us_east rack=rack1 # GCP EU-WEST nodes: dc=gcp_eu_west rack=rack1 # Azure Asia nodes: dc=azure_asia_pacific rack=rack1 # Create keyspace with replication across all DCs CREATE KEYSPACE global_app WITH REPLICATION = { 'class': 'NetworkTopologyStrategy', 'aws_us_east': 3, 'gcp_eu_west': 3, 'azure_asia_pacific': 3 }; -- Total: 9 copies of data (3 per DC)

Architecture 2: Active-Passive (Disaster Recovery)

ACTIVE-PASSIVE ARCHITECTURE ┌─────────────────┐ ┌─────────────────┐ │ AWS PRIMARY │ │ AZURE STANDBY │ │ Cassandra DC1 │───── async ──────►│ Cassandra DC2 │ │ RF=3, 9 nodes │ replication │ RF=3, 6 nodes │ └─────────────────┘ └─────────────────┘ ▲ ▲ │ │ ALL Traffic Backup only (no client traffic) Primary DC: Serves all production traffic Standby DC: Receives async replication, ready for failover Failover: Manual switch if primary fails
# Active-Passive with async replication CREATE KEYSPACE production WITH REPLICATION = { 'class': 'NetworkTopologyStrategy', 'aws_primary': 3, ← Active 'azure_standby': 3 ← Passive backup }; -- Client connects ONLY to aws_primary normally -- In disaster: Switch clients to azure_standby

Architecture 3: Hybrid Cloud (On-Prem + Cloud)

HYBRID ARCHITECTURE ┌──────────────────┐ ┌─────────────────┐ │ ON-PREMISES │ │ AWS CLOUD │ │ Cassandra DC1 │◄────────►│ Cassandra DC2 │ │ Legacy systems │ VPN │ New workloads │ │ Compliance data │ │ Elastic scale │ └──────────────────┘ └─────────────────┘ Use case: Gradual cloud migration Benefit: Keep sensitive data on-prem, scale in cloud

⚙️ Multi-Cloud Setup Guide

Step 1: Network Configuration

Critical: Cross-Cloud Networking

Challenge: Cassandra nodes across different clouds must communicate reliably.

Solutions:

  • 🔗 VPN: Site-to-site VPN between clouds
  • 🌐 Public IPs: Nodes communicate over internet (cheaper but less secure)
  • ⚡ Direct Connect: AWS Direct Connect + Azure ExpressRoute + GCP Interconnect
  • 🔒 VPC Peering: Cloud-native peering (limited cross-provider support)

AWS Setup (DC1)

# 1. Create VPC aws ec2 create-vpc --cidr-block 10.0.0.0/16 --region us-east-1 # 2. Launch EC2 instances (Cassandra nodes) aws ec2 run-instances \ --image-id ami-12345 \ --instance-type r5.2xlarge \ --count 6 \ --subnet-id subnet-123 \ --security-group-ids sg-456 # 3. Security group: Allow Cassandra ports from all DCs # - 7000 (inter-node) # - 7001 (TLS inter-node) # - 9042 (CQL) # - 7199 (JMX) # 4. Assign Elastic IPs (for cross-cloud communication) aws ec2 allocate-address --domain vpc

GCP Setup (DC2)

# 1. Create VPC gcloud compute networks create cassandra-net \ --subnet-mode=custom gcloud compute networks subnets create cassandra-subnet \ --network=cassandra-net \ --range=10.1.0.0/16 \ --region=europe-west1 # 2. Create VM instances gcloud compute instances create cassandra-node-{1..6} \ --zone=europe-west1-b \ --machine-type=n2-highmem-8 \ --subnet=cassandra-subnet # 3. Firewall rules gcloud compute firewall-rules create cassandra-internal \ --network=cassandra-net \ --allow=tcp:7000,tcp:7001,tcp:9042

Azure Setup (DC3)

# 1. Create Resource Group and VNet az group create --name cassandra-rg --location southeastasia az network vnet create \ --resource-group cassandra-rg \ --name cassandra-vnet \ --address-prefix 10.2.0.0/16 \ --subnet-name cassandra-subnet \ --subnet-prefix 10.2.0.0/24 # 2. Create VMs az vm create \ --resource-group cassandra-rg \ --name cassandra-node-1 \ --image UbuntuLTS \ --size Standard_E8s_v3 \ --vnet-name cassandra-vnet \ --subnet cassandra-subnet # 3. Network Security Group az network nsg rule create \ --resource-group cassandra-rg \ --nsg-name cassandra-nsg \ --name AllowCassandra \ --priority 100 \ --destination-port-ranges 7000 7001 9042

Step 2: Install Cassandra

# Run on ALL nodes (AWS, GCP, Azure) # 1. Install Java 11 sudo apt update sudo apt install openjdk-11-jdk -y # 2. Add Cassandra repository echo "deb https://debian.cassandra.apache.org 41x main" | \ sudo tee /etc/apt/sources.list.d/cassandra.list # 3. Install Cassandra sudo apt update sudo apt install cassandra -y # 4. Stop service (configure first) sudo systemctl stop cassandra

Step 3: Configure Each Datacenter

# Edit /etc/cassandra/cassandra.yaml on EACH node # Common settings (all nodes): cluster_name: 'GlobalCassandra' num_tokens: 16 endpoint_snitch: GossipingPropertyFileSnitch # Seeds: Include one from each DC seed_provider: - class_name: org.apache.cassandra.locator.SimpleSeedProvider parameters: - seeds: "52.1.2.3,35.4.5.6,13.7.8.9" # AWS GCP Azure # Listen addresses (use PUBLIC IP for cross-cloud) listen_address: # Node's private IP broadcast_address: # Node's public IP (for cross-cloud) # Edit /etc/cassandra/cassandra-rackdc.properties # AWS nodes: dc=aws_us_east rack=rack1 # GCP nodes: dc=gcp_eu_west rack=rack1 # Azure nodes: dc=azure_asia rack=rack1

Step 4: Start Cluster

# Start nodes ONE AT A TIME, waiting for each to join # 1. Start seed nodes first (one per DC) sudo systemctl start cassandra # 2. Check if node joined nodetool status # Should show: UN (Up/Normal) # 3. Wait 2 minutes, then start next node # 4. Repeat until all nodes running # 5. Verify 3 datacenters nodetool status # Output: # Datacenter: aws_us_east (6 nodes) # Datacenter: gcp_eu_west (6 nodes) # Datacenter: azure_asia (6 nodes)

Step 5: Create Multi-DC Keyspace

CREATE KEYSPACE global_app WITH REPLICATION = { 'class': 'NetworkTopologyStrategy', 'aws_us_east': 3, 'gcp_eu_west': 3, 'azure_asia': 3 }; CREATE TABLE global_app.users ( user_id uuid PRIMARY KEY, email text, username text, created_at timestamp ); -- Data now replicated to ALL three clouds!

🌐 Cloud Provider Specifics

AWS Recommendations

☁️ AWS Best Practices

  • Instance: r5.2xlarge or r6i.2xlarge (memory-optimized)
  • Storage: EBS gp3 (3000 IOPS, 125 MB/s)
  • Network: Enhanced networking (SR-IOV)
  • AZs: Spread nodes across 3 availability zones
  • Security: VPC, security groups, IAM roles
  • Monitoring: CloudWatch metrics + custom JMX

GCP Recommendations

☁️ GCP Best Practices

  • Instance: n2-highmem-8 (8 vCPUs, 64GB RAM)
  • Storage: SSD persistent disks (pd-ssd)
  • Network: Premium tier networking
  • Zones: Spread across 3 zones
  • Security: VPC, firewall rules, service accounts
  • Monitoring: Cloud Monitoring + Stackdriver

Azure Recommendations

☁️ Azure Best Practices

  • Instance: Standard_E8s_v3 (8 vCPUs, 64GB)
  • Storage: Premium SSD (P30 or P40)
  • Network: Accelerated networking
  • AZs: Availability zones (3 zones)
  • Security: VNet, NSG, managed identity
  • Monitoring: Azure Monitor + Application Insights

🔥 Disaster Recovery Strategies

Strategy 1: Multi-DC Active-Active

Built-in DR with Active-Active

How it works: All datacenters accept traffic. If one DC fails, others automatically handle load.

# Client configuration with fallback from cassandra.cluster import Cluster from cassandra.policies import DCAwareRoundRobinPolicy # Connect to local DC first, fallback to remote cluster = Cluster( contact_points=['aws-node1', 'gcp-node1', 'azure-node1'], load_balancing_policy=DCAwareRoundRobinPolicy( local_dc='aws_us_east', used_hosts_per_remote_dc=3 ← Fallback to remote DCs ) ) # If AWS fails: # - Driver automatically tries GCP # - Then Azure if GCP also fails # - Transparent to application!

RTO: Seconds (automatic)

RPO: Zero (synchronous replication)

Strategy 2: Regular Backups

# Automated backup script # 1. Take snapshot on all nodes nodetool snapshot -t backup-$(date +%Y%m%d) # 2. Upload to S3 (from AWS), GCS (from GCP), Blob (from Azure) aws s3 sync /var/lib/cassandra/data/ s3://my-backups/ # 3. Schedule daily with cron 0 2 * * * /scripts/backup-cassandra.sh # 4. Retention: Keep 30 days, monthly archives for 1 year

Strategy 3: Cross-Region Replication

CROSS-REGION DR ARCHITECTURE Production Region: ┌────────────────────────────────────────┐ │ AWS US-EAST (3 DCs) │ │ DC1: us-east-1a, DC2: us-east-1b │ │ DC3: us-east-1c │ └────────────────────────────────────────┘ ↓ async replication DR Region (Standby): ┌────────────────────────────────────────┐ │ AWS EU-WEST (1 DC, cold standby) │ │ Receives async replication │ └────────────────────────────────────────┘ Failover: Promote EU-WEST to active if US-EAST entire region fails

Disaster Scenarios & Response

Scenario Impact Response RTO
Single node fails None ✅ Automatic - other replicas serve < 1 sec
AZ/Zone fails Minimal ✅ Other AZs handle traffic < 30 sec
Entire DC fails Degraded Failover to other DCs 1-5 min
Cloud provider outage Serious Route to other cloud providers 5-15 min

✅ Multi-Cloud Best Practices

✅ DO These

  • Use NetworkTopologyStrategy
  • Configure LOCAL_QUORUM reads
  • Enable authentication & encryption
  • Monitor cross-DC latency
  • Use 3+ nodes per DC
  • Spread across availability zones
  • Test failover regularly
  • Automate backups
  • Document runbooks
  • Use consistent instance types

❌ DON'T Do These

  • Use SimpleStrategy (single DC only)
  • Run QUORUM across WAN
  • Mix different Cassandra versions
  • Ignore cross-cloud network costs
  • Deploy < 3 nodes per DC
  • Skip encryption over internet
  • Forget about latency
  • Use single seed per DC
  • Ignore time sync (NTP)
  • Deploy without monitoring

Performance Tips

  • 🎯 LOCAL_QUORUM: Read/write within local DC only (fast)
  • ⚡ Async replication: Cross-DC replication is async by default
  • 💰 Data transfer costs: Cross-cloud egress expensive (~$0.09/GB)
  • 🔒 Encryption: Use TLS for cross-cloud (adds latency)
  • ⏰ NTP sync: Critical for distributed operations
  • 📊 Monitoring: Track cross-DC replication lag

Cost Considerations

Multi-cloud costs more than single cloud:

  • 💰 Compute: 3x instances (3 clouds × nodes per cloud)
  • 💾 Storage: 3x data (replicated everywhere)
  • 🌐 Network: Cross-cloud egress fees
  • 👨‍💼 Operations: More complexity = more ops cost

Typical savings vs benefits:

  • ✅ Cost of downtime > multi-cloud cost
  • ✅ Vendor negotiation leverage
  • ✅ Use spot/preemptible instances where possible

🎯 Multi-Cloud Summary

You now know how to deploy Cassandra globally!

📚 Key Takeaways:

  • ☁️ Multi-cloud = AWS + GCP + Azure unified
  • 🌍 Active-active for global low latency
  • 🔥 Survives entire cloud provider outages
  • 🔗 VPN or public IPs for cross-cloud
  • 📊 NetworkTopologyStrategy with RF per DC
  • ⚡ LOCAL_QUORUM for performance
  • 🛡️ Built-in disaster recovery
  • 💰 Consider egress costs

🎉 CONGRATULATIONS! 🎉

You've completed ALL advanced Cassandra topics!

🚀 You're now a Cassandra expert! 🚀

Advertisement

Responsive Ad