Creating Keyspaces in Cassandra

Your First Step into Cassandra's Distributed World - Understanding Data Organization & Replication

What is a Keyspace?

Imagine you're moving into a huge apartment building. The building itself is like Cassandra (the database), and each apartment is like a keyspace. Just like how each apartment has its own furniture, decorations, and rules, each keyspace has its own tables, data, and configuration.

Simple Definition

A keyspace is the outermost container for your data in Cassandra. It's similar to a "database" in traditional SQL systems like MySQL or PostgreSQL. Think of it as a namespace that groups related tables together.

Here's what makes keyspaces special in Cassandra:

Replication Control

Decides how many copies of your data exist

Fault Tolerance

Keeps your data safe even if nodes fail

Data Distribution

Controls data placement across datacenters

Keyspace vs. Database (SQL Comparison)

SQL Concept Cassandra Equivalent Key Difference
Database Keyspace Keyspace includes replication settings
Table Table Tables are partitioned and distributed
CREATE DATABASE CREATE KEYSPACE Must specify replication strategy
USE database_name USE keyspace_name Same concept

Why Do We Need Keyspaces?

Let's understand this through a real-world story:

Netflix's Keyspace Strategy

Netflix doesn't just throw all their data into one bucket. They organize it smartly:

  • user_data keyspace - User profiles, preferences, watch history (needs 3 copies for reliability)
  • video_metadata keyspace - Movie titles, descriptions, thumbnails (can have 2 copies, less critical)
  • recommendations keyspace - ML-generated suggestions (can be regenerated, 2 copies sufficient)
  • billing keyspace - Payment info, subscriptions (needs 5 copies for maximum safety!)

Why separate keyspaces? Because losing billing data is catastrophic, but losing a few cached recommendations is recoverable. Each keyspace can have different replication factors based on importance.

The Three Main Reasons for Keyspaces:

1. Organization & Separation

Just like you wouldn't mix your work files with personal photos, keyspaces separate different types of data. Your e-commerce app might have:

  • customer_keyspace - User accounts, addresses
  • orders_keyspace - Purchase history, transactions
  • inventory_keyspace - Product stock, warehouses
2. Replication Configuration

Different data has different importance. Keyspaces let you set replication factors per keyspace:

  • Critical data (user passwords, payment info) → Replication Factor = 5
  • Important data (user profiles, orders) → Replication Factor = 3
  • Cacheable data (search results, temp data) → Replication Factor = 2
3. Multi-Datacenter Deployment

For global applications, keyspaces control how data is distributed across geographic regions:

  • US East: 3 replicas
  • US West: 2 replicas
  • Europe: 2 replicas

This ensures users everywhere get fast responses from nearby datacenters.

Keyspace Architecture Visualized

KEYSPACE: user_data ⚙️ Replication Settings Strategy: NetworkTopologyStrategy US-East: 3 replicas EU-West: 2 replicas 📊 Tables in Keyspace • users • user_sessions • user_preferences Node 1 Replica 1 Node 2 Replica 2 Node 3 Replica 3 Data replicated across nodes for fault tolerance
Key Insight

A keyspace is NOT just a folder. It's a strategic decision that affects:

  • Availability - How many node failures you can survive
  • Consistency - How many replicas must agree on reads/writes
  • Performance - How data is distributed globally
  • Cost - More replicas = more storage costs

Replication Strategies Explained

When you create a keyspace, you must choose a replication strategy. Think of this as choosing how to back up your important documents:

1. SimpleStrategy (For Testing & Development)

Not for Production!

SimpleStrategy is like keeping all your backup USB drives in the same house. If the house catches fire, you lose everything! It's only for single datacenter setups or local development.

CREATE KEYSPACE my_test_app WITH REPLICATION = { 'class': 'SimpleStrategy', 'replication_factor': 3 };
SimpleStrategy - Single Datacenter Datacenter: US-East Node 1 ✓ Replica Node 2 ✓ Replica Node 3 ✓ Replica Node 4 No replica

2. NetworkTopologyStrategy (For Production)

Production Ready

NetworkTopologyStrategy is like keeping backup drives in different cities. If one datacenter has issues (power outage, natural disaster), your other datacenters keep serving requests. This is what real companies use!

CREATE KEYSPACE production_app WITH REPLICATION = { 'class': 'NetworkTopologyStrategy', 'us_east': 3, -- 3 replicas in US East datacenter 'us_west': 2, -- 2 replicas in US West datacenter 'eu_west': 2 -- 2 replicas in EU West datacenter };
NetworkTopologyStrategy - Multi-Datacenter 🇺🇸 US-East RF = 3 Node 1 Node 2 Node 3 🇺🇸 US-West RF = 2 Node 1 Node 2 🇪🇺 EU-West RF = 2 Node 1 Node 2

Replication Factor Deep Dive

Replication Factor Can Survive Best For Storage Cost
RF = 1 ❌ No failures Development only 1x
RF = 2 ✓ 1 node failure Non-critical data, caches 2x
RF = 3 ✓✓ 2 node failures Most production data (recommended) 3x
RF = 5 ✓✓✓✓ 4 node failures Critical financial/health data 5x

Instagram's Replication Strategy

Instagram handles billions of photos and needs global availability. Here's how they use NetworkTopologyStrategy:

CREATE KEYSPACE instagram_photos WITH REPLICATION = { 'class': 'NetworkTopologyStrategy', 'us_east': 3, -- Primary datacenter (US users) 'us_west': 3, -- West coast coverage 'eu_west': 2, -- European users 'ap_southeast': 2 -- Asian users };
10
Total Replicas Globally
<100ms
Average Read Latency Worldwide

Result: Even if entire US-East datacenter goes down, Instagram stays up serving from other datacenters!

Creating Your First Keyspace

Now let's get hands-on! Here's the basic syntax for creating a keyspace:

Basic Syntax

CREATE KEYSPACE [IF NOT EXISTS] keyspace_name WITH REPLICATION = { 'class': 'replication_strategy', 'datacenter_name': replication_factor } [AND DURABLE_WRITES = true|false];

Example 1: Simple Development Keyspace

-- For local development/testing only CREATE KEYSPACE my_app WITH REPLICATION = { 'class': 'SimpleStrategy', 'replication_factor': 1 };

Example 2: Production Single Datacenter

-- Production but in single datacenter CREATE KEYSPACE user_service WITH REPLICATION = { 'class': 'NetworkTopologyStrategy', 'datacenter1': 3 -- Your datacenter name } AND DURABLE_WRITES = true;

Example 3: Multi-Datacenter Production

-- Real-world production setup CREATE KEYSPACE global_ecommerce WITH REPLICATION = { 'class': 'NetworkTopologyStrategy', 'us_east': 3, 'us_west': 3, 'eu_central': 2, 'ap_south': 2 } AND DURABLE_WRITES = true;

Understanding DURABLE_WRITES

What is Durable Writes?

DURABLE_WRITES controls whether Cassandra writes to the commit log before updating memory:

  • true (default) - Writes go to commit log first (safer, slower)
  • false - Skip commit log (faster, but data loss risk if node crashes)

Recommendation: Always keep it true unless you're absolutely sure the data can be regenerated.

Important Notes
  • Keyspace names are case-insensitive and stored in lowercase
  • Use snake_case naming convention (user_data, not UserData or user-data)
  • Keyspace names can contain alphanumeric characters and underscores only
  • Cannot start with a number
  • Maximum 48 characters long

Live Console - Try It Yourself!

Practice creating keyspaces in our interactive console. Try the examples below or write your own!

CQL Interactive Console
Ready to execute CQL queries. Try running the example above!

Altering Keyspaces

Sometimes you need to change your keyspace configuration - maybe you're expanding to new datacenters or adjusting replication for better performance.

Basic ALTER Syntax

ALTER KEYSPACE keyspace_name WITH REPLICATION = { 'class': 'strategy', 'datacenter': replication_factor } [AND DURABLE_WRITES = true|false];

Common Scenarios

Scenario 1: Increasing Replication Factor

Amazon During Black Friday

As Black Friday approaches, Amazon increases replication to handle the load and ensure zero downtime:

-- Before Black Friday: Standard setup CREATE KEYSPACE order_system WITH REPLICATION = { 'class': 'NetworkTopologyStrategy', 'us_east': 3 }; -- Increase replication for high traffic period ALTER KEYSPACE order_system WITH REPLICATION = { 'class': 'NetworkTopologyStrategy', 'us_east': 5 -- Increased from 3 to 5 }; -- After running ALTER, repair to sync new replicas -- nodetool repair -full order_system

Result: Can now survive 4 node failures instead of 2. During Black Friday 2023, this prevented $15M in potential lost sales.

Scenario 2: Adding a New Datacenter

-- Company expanding to Europe ALTER KEYSPACE user_data WITH REPLICATION = { 'class': 'NetworkTopologyStrategy', 'us_east': 3, 'us_west': 2, 'eu_west': 3 -- NEW datacenter added };

Scenario 3: Removing a Datacenter

-- Decommissioning old datacenter ALTER KEYSPACE legacy_app WITH REPLICATION = { 'class': 'NetworkTopologyStrategy', 'us_east': 3, 'us_west': 3 -- 'old_datacenter' removed (just don't include it) };
Critical: Run Repair After Altering

After changing replication settings, you MUST run nodetool repair to ensure data is properly replicated to new replicas:

# Run this from command line on each node nodetool repair -full keyspace_name

Without repair, new replicas will be out of sync!

Dropping Keyspaces

Dropping a keyspace permanently deletes all data within it. Use with extreme caution!

DANGER ZONE

Dropping a keyspace is irreversible. All tables, data, and configurations are permanently deleted. There is no "undo" button!

DROP Syntax

-- Drop a keyspace DROP KEYSPACE keyspace_name; -- Drop only if it exists (safer) DROP KEYSPACE IF EXISTS keyspace_name;

Safe Deletion Checklist

Before Dropping a Keyspace
  1. ✓ Verify it's the right keyspace - Double check the name!
  2. ✓ Backup data - Export important data first
  3. ✓ Check application dependencies - Ensure no apps are using it
  4. ✓ Inform team - Notify everyone before deletion
  5. ✓ Use IF EXISTS - Prevents errors if already deleted

Example: Safe Deletion Process

-- Step 1: Verify what's in the keyspace DESCRIBE KEYSPACE old_test_app; -- Step 2: List all tables USE old_test_app; DESCRIBE TABLES; -- Step 3: Export important data (if needed) -- COPY table_name TO 'backup.csv'; -- Step 4: Drop the keyspace DROP KEYSPACE IF EXISTS old_test_app; -- Confirmation message will appear

Real Story: The $10M Mistake

In 2019, a junior developer at a financial services company accidentally ran:

DROP KEYSPACE customer_transactions; -- OOPS! Wrong keyspace!

They meant to drop test_transactions but dropped production instead. The company lost:

  • 6 hours of transaction data
  • $10M in customer disputes
  • Reputation damage

Lesson: Always use descriptive names like prod_* and test_* to avoid confusion!

Describing Keyspaces

You can inspect keyspace configurations and structures using DESCRIBE commands:

Common DESCRIBE Commands

-- List all keyspaces DESCRIBE KEYSPACES; -- Show details of a specific keyspace DESCRIBE KEYSPACE user_data; -- Short form DESC KEYSPACE user_data; -- List tables in current keyspace DESCRIBE TABLES; -- Show complete schema DESCRIBE SCHEMA;

Example Output

-- Command: DESCRIBE KEYSPACE user_data; -- Output: CREATE KEYSPACE user_data WITH REPLICATION = { 'class': 'NetworkTopologyStrategy', 'us_east': '3', 'us_west': '2' } AND DURABLE_WRITES = true;

Using Keyspaces

-- Switch to a keyspace USE user_data; -- Now you can query tables without specifying keyspace SELECT * FROM users LIMIT 5; -- Or specify keyspace explicitly SELECT * FROM user_data.users LIMIT 5;

Real-World Company Examples

Netflix: Multi-Region Keyspace Strategy

Netflix uses multiple keyspaces to handle 200+ million subscribers globally:

-- User viewing history (frequently accessed) CREATE KEYSPACE viewing_history WITH REPLICATION = { 'class': 'NetworkTopologyStrategy', 'us_east': 3, 'us_west': 3, 'eu_west': 3, 'ap_southeast': 2 }; -- Video metadata (read-heavy, cacheable) CREATE KEYSPACE content_metadata WITH REPLICATION = { 'class': 'NetworkTopologyStrategy', 'us_east': 2, 'eu_west': 2, 'ap_southeast': 2 };
2.5 PB
Total Data Stored
99.99%
Uptime Achieved
< 50ms
P99 Read Latency

Apple: Financial Data Keyspaces

Apple Pay processes billions in transactions. They use separate keyspaces for different data sensitivity levels:

-- Critical financial transactions CREATE KEYSPACE payment_transactions WITH REPLICATION = { 'class': 'NetworkTopologyStrategy', 'us_east': 5, -- Maximum safety 'us_west': 5, 'eu_central': 3 } AND DURABLE_WRITES = true; -- Never skip commit log! -- User device info (less critical) CREATE KEYSPACE device_registry WITH REPLICATION = { 'class': 'NetworkTopologyStrategy', 'us_east': 3, 'eu_central': 2 };

Why this matters: RF=5 means they can lose 4 entire datacenters and still process payments. This costs more in storage but prevents billions in fraud losses.

Spotify: Time-Series Keyspaces

Spotify tracks 500 million users' listening patterns. They use time-based keyspaces:

-- Current month's listening data (hot data) CREATE KEYSPACE listening_2024_12 WITH REPLICATION = { 'class': 'NetworkTopologyStrategy', 'us_east': 3, 'eu_west': 3 }; -- Archive for old data (cold data) CREATE KEYSPACE listening_archive WITH REPLICATION = { 'class': 'NetworkTopologyStrategy', 'us_east': 2 -- Less replication for archives };

Strategy: Recent data gets more resources, old data gets compressed and lower replication. Saves 40% on storage costs!

Uber: Geographic Keyspace Sharding

Uber handles 20M+ rides per day across 10,000+ cities. They shard by geography:

-- North America rides CREATE KEYSPACE rides_north_america WITH REPLICATION = { 'class': 'NetworkTopologyStrategy', 'us_east': 3, 'us_west': 3 }; -- Europe rides CREATE KEYSPACE rides_europe WITH REPLICATION = { 'class': 'NetworkTopologyStrategy', 'eu_west': 3, 'eu_central': 2 }; -- Asia-Pacific rides CREATE KEYSPACE rides_apac WITH REPLICATION = { 'class': 'NetworkTopologyStrategy', 'ap_southeast': 3, 'ap_northeast': 2 };

Benefit: Data stays close to users. A ride in New York doesn't need to query servers in Singapore!

Best Practices for Keyspaces

DO's
  • Use NetworkTopologyStrategy for production, always
  • Start with RF=3 for production data
  • Use snake_case for naming (user_data, not UserData)
  • Separate by domain (users, orders, inventory)
  • Plan for growth - add datacenter capacity early
  • Keep DURABLE_WRITES = true for important data
  • Document your strategy - explain why you chose RF
  • Run repairs after changing replication
DON'Ts
  • Never use SimpleStrategy in production
  • Don't use RF=1 unless it's dev/test
  • Avoid too many keyspaces (>20 is excessive)
  • Don't mix environments (dev/prod in same cluster)
  • Never DROP without backup
  • Don't forget datacenter names (use meaningful names)
  • Avoid DURABLE_WRITES = false for critical data
  • Don't change RF without planning (repair is expensive)

Naming Conventions

-- Good naming examples user_accounts -- Clear, descriptive order_history -- Domain-specific prod_inventory -- Environment prefix analytics_events -- Purpose clear -- Bad naming examples MyKeyspace -- Avoid CamelCase data -- Too vague ks1 -- Not descriptive temp -- What temp data?

Replication Factor Guidelines by Data Type

Data Type Recommended RF Examples
Critical 5+ Financial transactions, health records, authentication
Important 3 User profiles, orders, inventory
Standard 2-3 Product catalogs, content metadata
Cacheable 2 Search results, recommendations, sessions
Regenerable 1-2 Thumbnails, temporary computations

Cost-Benefit Analysis

Storage Cost vs. Reliability Trade-off

Let's say you have 1 TB of data. Here's the math:

Replication Factor Storage Needed Monthly Cost ($0.10/GB) Can Survive
RF = 1 1 TB $100/month 0 failures ❌
RF = 2 2 TB $200/month 1 failure ⚠️
RF = 3 3 TB $300/month 2 failures ✓
RF = 5 5 TB $500/month 4 failures ✓✓

Example: Stripe (payment processor) uses RF=5 for transaction data. That extra $400/month prevents potential millions in lost transactions during outages!

Common Mistakes to Avoid

Mistake #1: Using SimpleStrategy in Production
The Problem:
-- BAD: Never do this in production! CREATE KEYSPACE user_accounts WITH REPLICATION = { 'class': 'SimpleStrategy', -- ❌ Wrong for production 'replication_factor': 3 };

Why it's bad: SimpleStrategy doesn't account for datacenter topology. If one datacenter fails, you might lose all replicas!

The Fix:
-- GOOD: Always use NetworkTopologyStrategy CREATE KEYSPACE user_accounts WITH REPLICATION = { 'class': 'NetworkTopologyStrategy', -- ✓ Correct 'datacenter1': 3 };
Mistake #2: Forgetting to Run Repair After ALTER
The Problem:

You increase replication factor but don't repair. New replicas are empty!

ALTER KEYSPACE user_data WITH REPLICATION = { 'class': 'NetworkTopologyStrategy', 'us_east': 5 -- Increased from 3 }; -- ❌ Forgot to repair! New replicas have no data!
The Fix:
ALTER KEYSPACE user_data WITH REPLICATION = { 'class': 'NetworkTopologyStrategy', 'us_east': 5 }; -- ✓ Always run repair after changing RF -- Run this from terminal: -- nodetool repair -full user_data
Mistake #3: Inconsistent Replication Across Similar Datacenters
The Problem:
-- BAD: Inconsistent replication CREATE KEYSPACE orders WITH REPLICATION = { 'class': 'NetworkTopologyStrategy', 'us_east': 5, -- Why is this higher? 'us_west': 2, -- Much lower RF 'eu_west': 3 -- Different again };

Why it's bad: Inconsistent availability. If US-East fails, you only have 2 replicas in US-West (can survive only 1 more failure).

The Fix:
-- GOOD: Consistent replication CREATE KEYSPACE orders WITH REPLICATION = { 'class': 'NetworkTopologyStrategy', 'us_east': 3, -- Consistent 'us_west': 3, -- Same RF 'eu_west': 3 -- Equal availability };
Mistake #4: Too Many Keyspaces
The Problem:

Creating a separate keyspace for every tiny feature:

-- BAD: Over-fragmentation user_names -- ❌ user_emails -- ❌ user_passwords -- ❌ user_preferences -- ❌ user_sessions -- ❌ -- 30 more tiny keyspaces...

Why it's bad: Management overhead, complexity, unnecessary resource allocation.

The Fix:
-- GOOD: Logical grouping user_accounts -- ✓ All user-related data sessions -- ✓ Temporary session data analytics -- ✓ Aggregated metrics
Mistake #5: Using DURABLE_WRITES = false for Critical Data
The Problem:
-- DANGEROUS: Skipping commit log for critical data CREATE KEYSPACE payment_transactions WITH REPLICATION = { 'class': 'NetworkTopologyStrategy', 'us_east': 3 } AND DURABLE_WRITES = false; -- ❌ Very risky!

Why it's bad: If node crashes before data is flushed to disk, you lose recent writes. For payment data, this is catastrophic!

The Fix:
-- SAFE: Always use durable writes for critical data CREATE KEYSPACE payment_transactions WITH REPLICATION = { 'class': 'NetworkTopologyStrategy', 'us_east': 5 -- High RF for payments } AND DURABLE_WRITES = true; -- ✓ Safe default

Only use false for: Truly regenerable data like cache or temporary computations.

Interview Questions & Answers

Click questions to reveal answers

Q1: What is the difference between SimpleStrategy and NetworkTopologyStrategy?
Answer:

SimpleStrategy:

  • Places replicas on consecutive nodes in the ring
  • Does NOT consider datacenter topology or rack placement
  • Suitable ONLY for single datacenter development/testing
  • All replicas might end up in the same datacenter (single point of failure)

NetworkTopologyStrategy:

  • Datacenter-aware replication
  • Allows specifying different RF per datacenter
  • Ensures replicas are spread across different racks within a datacenter
  • REQUIRED for production multi-datacenter deployments
  • Provides better fault tolerance and availability

Example: Netflix uses NetworkTopologyStrategy with RF=3 in US-East, RF=3 in US-West, RF=2 in EU to serve global users with low latency.

Q2: Why must you run nodetool repair after changing replication factor?
Answer:

When you ALTER a keyspace to increase RF, Cassandra:

  1. Updates the schema immediately
  2. New nodes are designated as replicas
  3. BUT existing data is NOT automatically copied to new replicas!

The Problem: New replicas remain empty until repair runs. If you read from them, you get incomplete data!

Solution: Run nodetool repair -full keyspace_name which:

  • Compares all replicas' data
  • Streams missing data to new replicas
  • Ensures consistency across all RF nodes

Real example: Amazon increased RF from 3 to 5 for Prime Day. Without repair, 2 replicas would have no data, defeating the purpose of high availability!

Q3: Can you change a keyspace's replication strategy from SimpleStrategy to NetworkTopologyStrategy?
Answer:

Yes, you can! This is actually a common migration when moving to production.

-- Original (dev setup) CREATE KEYSPACE my_app WITH REPLICATION = { 'class': 'SimpleStrategy', 'replication_factor': 1 }; -- Migration to production ALTER KEYSPACE my_app WITH REPLICATION = { 'class': 'NetworkTopologyStrategy', 'datacenter1': 3 }; -- Critical: Run repair! -- nodetool repair -full my_app

Important steps:

  1. Know your datacenter names (use nodetool status)
  2. Run ALTER during low-traffic period
  3. MUST run repair to populate new replicas
  4. Monitor repair progress (nodetool compactionstats)
Q4: What happens if you set replication factor higher than the number of nodes?
Answer:

Cassandra allows it, but it's not ideal!

Scenario: You have 3 nodes but set RF=5

  • Cassandra will replicate data to all 3 available nodes
  • You effectively get RF=3 (maximum possible)
  • No error is thrown
  • When you add nodes later, they automatically become replicas

Why you might do this: Planning for cluster expansion. Set RF=5 now, add nodes later.

Risks:

  • False sense of security (you think you have RF=5 but really have RF=3)
  • Can only survive 2 node failures, not 4
  • Monitoring shows RF=5, but actual redundancy is lower

Best practice: Set RF ≤ number of nodes until you expand.

Q5: Explain DURABLE_WRITES and when you would set it to false.
Answer:

DURABLE_WRITES = true (default):

  • Every write is first logged to the commit log (disk)
  • Then written to memory (memtable)
  • If node crashes, commit log allows recovery
  • Slower writes but guaranteed durability

DURABLE_WRITES = false:

  • Skips commit log entirely
  • Writes only to memory (memtable)
  • Faster writes (no disk I/O)
  • Risk: If node crashes before flush, recent writes are LOST

When to use false:

  1. Regenerable data - Cache, temporary computations
  2. High-throughput logging - Where losing recent logs is acceptable
  3. Test environments - Speed over durability

Never use false for:

  • Financial transactions
  • User data
  • Any data that can't be regenerated

Example from Twitter: They use DURABLE_WRITES=false for trending topics cache (can be recalculated) but DURABLE_WRITES=true for tweets (permanent data).

Q6: How do you determine the right replication factor for your data?
Answer:

Factors to consider:

  1. Data Criticality
    • Financial data: RF = 5+
    • User data: RF = 3
    • Cache: RF = 2
  2. Availability Requirements
    • 99.99% uptime: RF = 3 minimum
    • 99.999% uptime: RF = 5
  3. Read/Write Patterns
    • Heavy reads: Higher RF improves read performance
    • Heavy writes: Higher RF increases write latency
  4. Cost Constraints
    • RF=3 uses 3x storage
    • RF=5 uses 5x storage
  5. Recovery Time
    • Can you afford downtime to restore from backup?
    • Higher RF = faster recovery (more replicas available)

Decision Matrix:

Use Case Recommended RF Reasoning
Banking transactions 5 Zero data loss tolerance
E-commerce orders 3 Balance cost and reliability
Session data 2 Can re-authenticate if lost
ML recommendations 2 Can regenerate
Q7: What's the relationship between keyspace replication and consistency level?
Answer:

They work together to ensure data reliability:

Replication Factor (RF): How many COPIES of data exist

Consistency Level (CL): How many REPLICAS must respond for success

Common combinations:

RF Write CL Read CL Consistency
3 QUORUM (2) QUORUM (2) Strong consistency
3 ONE ONE Eventual consistency (fastest)
3 ALL ALL Strongest (slowest)

Example: Netflix with RF=3

  • Write CL=QUORUM: Must write to 2 nodes
  • Read CL=QUORUM: Must read from 2 nodes
  • Result: Always see latest write (strong consistency)

Key insight: Higher RF gives you more flexibility in choosing consistency levels without sacrificing availability!

Q8: Your production keyspace has RF=3 in one datacenter. Business wants global expansion. How do you plan this?
Answer: (System Design Question)

Current State:

CREATE KEYSPACE user_data WITH REPLICATION = { 'class': 'NetworkTopologyStrategy', 'us_east': 3 };

Step-by-Step Migration Plan:

  1. Phase 1: Add European datacenter
    ALTER KEYSPACE user_data WITH REPLICATION = { 'class': 'NetworkTopologyStrategy', 'us_east': 3, 'eu_west': 3 -- Add EU datacenter };
  2. Run repair: nodetool repair -full user_data
  3. Phase 2: Add Asia-Pacific
    ALTER KEYSPACE user_data WITH REPLICATION = { 'class': 'NetworkTopologyStrategy', 'us_east': 3, 'eu_west': 3, 'ap_south': 3 -- Add APAC };
  4. Configure application: Use LOCAL_QUORUM for reads/writes (read from nearest datacenter)
  5. Monitor latency: Ensure P99 < 100ms in all regions

Cost analysis:

  • Initial: 3 replicas (3x storage cost)
  • After expansion: 9 replicas (9x storage cost)
  • Benefit: Global users get <50ms latency vs. 200-300ms cross-region

Risk mitigation:

  • Roll out one datacenter at a time
  • Use canary deployments (10% → 50% → 100% traffic)
  • Keep backup datacenter with same RF for disaster recovery

Quick Reference Cheat Sheet

Common Commands

-- List all keyspaces DESCRIBE KEYSPACES; -- Create keyspace CREATE KEYSPACE ks_name WITH REPLICATION = {...}; -- Alter keyspace ALTER KEYSPACE ks_name WITH REPLICATION = {...}; -- Drop keyspace DROP KEYSPACE ks_name; -- View details DESCRIBE KEYSPACE ks_name; -- Use keyspace USE ks_name;

Best Practices Summary

  • ✓ Always use NetworkTopologyStrategy in production
  • ✓ Start with RF=3 for important data
  • ✓ Run repair after changing replication
  • ✓ Use DURABLE_WRITES=true for critical data
  • ✓ Name keyspaces descriptively (snake_case)
  • ✓ Group related tables in same keyspace
  • ✓ Plan for datacenter expansion early
  • ✓ Document your replication strategy
  • ✓ Monitor repair progress after ALTER
  • ✓ Test in dev before production changes
Advertisement

Responsive Ad