Interview Preparation

Real-World Scenarios

Master troubleshooting, performance issues, and production challenges with real Cassandra scenarios!

🎭 Real-World Cassandra Scenarios

Scenario-based questions test your production experience and problem-solving skills!

What Interviewers Test:

  • 🔍 Troubleshooting: Can you diagnose production issues?
  • ⚡ Performance: How do you optimize slow queries?
  • 💔 Failure Handling: What happens when things break?
  • 📈 Scaling: How do you handle growth?
  • ⚖️ Trade-offs: Why choose one approach over another?
  • 🛠️ Operations: Day-to-day cluster management

How to Answer Scenario Questions

✅ STAR Method: Situation → Task → Action → Result

✅ Be Specific: Use actual metrics, tools, commands

✅ Show Process: Explain your diagnostic steps

✅ Discuss Trade-offs: Why you chose this solution

✅ Mention Monitoring: How you detected the issue

⚡ Performance Issues

S1

Scenario: Your application's read latency suddenly increased from 5ms to 500ms. How do you troubleshoot?

Perfect Answer

Diagnostic Process:

  1. Check Metrics First
    # Check cluster health nodetool status nodetool tpstats # Thread pool statistics nodetool tablestats # SSTable counts
  2. Common Causes & Solutions
    • 🔥 Too Many SSTables: Run compaction
      nodetool compact keyspace_name table_name
    • 💾 Memory Pressure: Check heap usage, tune JVM
      nodetool info # Check heap usage # If heap >75%, tune GC or add memory
    • 🔍 Hot Partitions: Check partition sizes
      nodetool cfstats keyspace.table # Look for: Partition Size max/mean
    • 💿 Disk I/O: SSD degradation, check iostat
      iostat -x 1 # Check %util
  3. Application-Level Checks
    • Query patterns changed? (full table scan)
    • Consistency level increased?
    • Connection pool exhausted?

Real Example: "At Netflix, we found 100+ SSTables per partition. Switching from STCS to LCS compaction reduced p99 latency from 450ms to 8ms."

S2

Scenario: Write throughput dropped by 50%. CPU is at 100%. What's happening?

Perfect Answer

Most Likely Cause: Compaction Storm

Why This Happens:

  • 💥 Many SSTables triggered major compaction
  • 🔥 Compaction using all CPU cores
  • ✍️ Writes blocked waiting for compaction to finish

Immediate Fix:

# Check compaction status nodetool compactionstats # Temporarily throttle compaction nodetool setcompactionthroughput 16 # MB/s # Or stop compaction temporarily nodetool stop COMPACTION

Long-term Solutions:

  • ✅ Tune compaction strategy (TWCS for time-series)
  • ✅ Add more nodes to distribute load
  • ✅ Increase compaction throughput during off-peak
  • ✅ Use SSDs for better I/O

Pro Tip: Monitor compaction_throughput and adjust based on workload!

S3

Scenario: Query works fine in dev (1000 rows) but times out in production (1M rows). Why?

Classic Anti-Pattern: Unbounded Partition

Problem: Query is reading entire partition into memory!

❌ Bad Schema CREATE TABLE events ( event_type text, ← Only a few types! timestamp timestamp, data text, PRIMARY KEY (event_type, timestamp) ); -- Dev: 1000 rows per event_type -- Prod: 1,000,000 rows per event_type → MASSIVE partition!

Solutions:

✅ Solution 1: Add bucketing CREATE TABLE events ( event_type text, bucket text, ← 'YYYY-MM-DD' or 'YYYY-MM-DD-HH' timestamp timestamp, data text, PRIMARY KEY ((event_type, bucket), timestamp) ); ✅ Solution 2: Better partition key CREATE TABLE events ( user_id uuid, ← High cardinality! timestamp timestamp, event_type text, data text, PRIMARY KEY (user_id, timestamp) );

Lesson: Always consider partition size at scale!

💔 System Failures

S4

Scenario: 2 nodes in a 6-node cluster with RF=3 crashed. Writes with QUORUM are failing. Why?

Perfect Answer

Analysis:

  • Cluster: 6 nodes total, RF=3
  • Down: 2 nodes
  • Available: 4 nodes
  • QUORUM needs: 2 out of 3 replicas

Why Failing:

For some partitions, 2 out of 3 replica nodes are down! QUORUM can't be satisfied.

Example:

Partition X replicas: Node1, Node2, Node3 If Node1 and Node2 are down → only 1/3 replicas available QUORUM needs 2/3 → WRITE FAILS!

Solutions:

  1. Immediate: Restart failed nodes ASAP
  2. Temporary: Reduce to ONE (loses consistency!)
  3. Long-term: Increase RF or cluster size
    • RF=5 can lose 2 nodes and still do QUORUM
    • Or use 9 nodes instead of 6

Production Best Practice: RF=3 with QUORUM can tolerate 1 node failure. For 2+ failures, need RF=5.

S5

Scenario: After datacenter power outage, some data is missing. What happened and how to fix?

Disaster Recovery Scenario

What Happened:

  • 💥 Power outage → All nodes lost power simultaneously
  • 💾 Data in memtables (not yet flushed) was lost
  • 📝 Commit log may be corrupted or incomplete
  • ⏰ Hinted handoff data lost (stored in memory)

Immediate Actions:

# 1. Check commit log on startup # Cassandra replays commit log automatically # 2. Verify data integrity nodetool verify keyspace table # 3. Run repair on ALL nodes nodetool repair -full # 4. Check for inconsistencies SELECT * FROM table WHERE ...; (with consistency ALL)

Prevention:

  • ✅ Multi-DC setup (geographic distribution)
  • ✅ UPS for graceful shutdown
  • ✅ Regular backups (snapshots)
  • ✅ Periodic repairs (weekly)
  • ✅ Monitor with consistency level ALL occasionally

Real Story: "Instagram lost 5 minutes of data during datacenter failure. Solution: Multi-DC setup with LOCAL_QUORUM + async replication."

📈 Scaling Challenges

S6

Scenario: Your cluster is at 80% capacity. How do you add nodes without downtime?

Perfect Answer: Rolling Scale-Out

Step-by-Step Process:

  1. Prepare New Nodes
    # Configure cassandra.yaml cluster_name: 'production_cluster' seeds: 'existing_seed_ip' listen_address: 'new_node_ip' auto_bootstrap: true ← Important!
  2. Add Nodes One at a Time
    # Start new node sudo systemctl start cassandra # Monitor bootstrap progress nodetool netstats nodetool status # Should show 'UJ' → 'UN'
  3. Wait for Streaming to Complete
    • New node receives data from existing nodes
    • Wait until status shows "UN" (Up Normal)
    • Can take hours depending on data size
  4. Run Cleanup on Old Nodes
    # Remove data that now belongs to new nodes nodetool cleanup
  5. Repeat Until Desired Size

Important Notes:

  • ⚠️ Don't add multiple nodes simultaneously (streaming overhead)
  • ⏰ Add during low-traffic hours if possible
  • 📊 Monitor cluster performance during scaling
  • 🔄 Run repairs after all nodes added
S7

Scenario: User uploads spiked 10x. Schema designed for user_id partition. Now hot partitions. Quick fix?

Emergency Hot Partition Fix

Current Schema (Problem):

CREATE TABLE user_uploads ( user_id uuid, upload_id timeuuid, filename text, PRIMARY KEY (user_id, upload_id) ); -- Power users uploading 10,000+ files → HOT PARTITION!

Immediate Workaround (Application Level):

  1. Add Bucketing in Application
    # Python pseudocode bucket = hash(user_id) % 10 # Split into 10 buckets write_to_partition((user_id, bucket), upload_data) # Read requires 10 queries (fan-out) for bucket in range(10): read_from_partition((user_id, bucket))

Proper Long-term Fix:

CREATE TABLE user_uploads_v2 ( user_id uuid, bucket int, ← Add to partition key upload_id timeuuid, filename text, PRIMARY KEY ((user_id, bucket), upload_id) ); -- Migrate data gradually with dual-write pattern

Alternative: Time-based Bucketing

PRIMARY KEY ((user_id, date), upload_id) -- Each day = new partition

🔧 Data Issues

S8

Scenario: Deleted data is still appearing in queries. Why and how to fix?

Perfect Answer: Tombstones & GC Grace

Why This Happens:

  • 📝 Deletes create tombstones (not immediate removal)
  • ⏰ Tombstones persist for gc_grace_seconds (default 10 days)
  • 👻 During reads, tombstones can be "resurrected" by repair/hints

Scenario Breakdown:

Time 0: DELETE record X → Tombstone created Time 1: Node A down during delete Time 2: Query reads X from Node B and C → Returns empty (correct) Time 5: Node A comes back up, has old data for X Time 6: Repair runs → Node A's old data overwrites tombstone Time 7: Query returns X again! → ZOMBIE DATA

Solutions:

  1. Wait for GC Grace Period
    # Check setting DESCRIBE TABLE table_name; # default: gc_grace_seconds = 864000 (10 days) # After 10 days, compaction removes tombstones
  2. Run Immediate Repair
    nodetool repair keyspace table # Ensures all nodes have tombstones
  3. Force Compaction (Dangerous!)
    ⚠️ Only if you're SURE all nodes are in sync ALTER TABLE table WITH gc_grace_seconds = 0; nodetool compact ALTER TABLE table WITH gc_grace_seconds = 864000;

Prevention: Run repairs regularly (within gc_grace period)!

S9

Scenario: Counter column shows wrong count. Users complaining. What went wrong?

Counter Columns Are NOT Idempotent!

Problem: Duplicate Increments

// Application retries due to timeout UPDATE post_stats SET likes = likes + 1 WHERE post_id = 'abc'; ↓ Timeout (but write succeeded!) UPDATE post_stats SET likes = likes + 1 WHERE post_id = 'abc'; ↓ Success Result: likes incremented by 2 instead of 1!

Why It Happens:

  • ❌ Client retries on timeout
  • ❌ Load balancer duplicates request
  • ❌ Network issue causes duplicate packets
  • ❌ First write actually succeeded but timed out

Solutions:

  1. Use Lightweight Transactions (LWT)
    -- Idempotent but slower UPDATE likes_ledger SET action = 'liked' WHERE user_id = 'user1' AND post_id = 'abc' IF NOT EXISTS; -- Then count with SELECT COUNT(*)
  2. Application-Level Deduplication
    # Generate unique request ID request_id = generate_uuid() # Store in dedup table with TTL INSERT INTO processed_requests (request_id) VALUES ('uuid') USING TTL 3600;
  3. Accept Approximate Counts
    • For display purposes (like counts, views)
    • Use counters knowing they might be slightly off
    • Document the limitation

Lesson: Counters are great for approximate counts, terrible for exact billing!

🛠️ Operations

S10

Scenario: You need to replace a failed node. Walk through the exact steps.

Perfect Answer: Node Replacement Procedure

Prerequisites:

  • Failed node is completely dead (not coming back)
  • You have the failed node's IP and token
  • Replacement node has same specs

Step-by-Step:

  1. Get Failed Node Info
    nodetool status # Note the DN (Down) node's IP: 192.168.1.100
  2. Prepare Replacement Node
    # Edit cassandra.yaml cluster_name: 'production_cluster' auto_bootstrap: true # Add JVM option # -Dcassandra.replace_address=192.168.1.100
  3. Start Replacement Node
    sudo systemctl start cassandra # Monitor streaming nodetool netstats tail -f /var/log/cassandra/system.log
  4. Wait for Bootstrap
    • Node downloads data from other replicas
    • Status changes: UJ → UN
    • Can take hours for large datasets
  5. Remove JVM Flag
    # After node is UN, remove from JVM options: # -Dcassandra.replace_address=192.168.1.100
  6. Run Repair
    nodetool repair -full
  7. Remove Old Node
    # On any live node nodetool removenode 'dead_node_id'

Important: Don't restart the node after replacement without removing the JVM flag!

🎯 Design Decisions

S11

Scenario: Should you use Cassandra for a banking transaction system? Why or why not?

Thoughtful Answer: It Depends!

Arguments AGAINST (Traditional View):

  • ❌ No ACID Transactions: No multi-row atomicity
  • ❌ Eventual Consistency: Temporary inconsistencies possible
  • ❌ No Joins: Complex reporting difficult
  • ❌ Regulatory Concerns: Banks prefer traditional RDBMS

Arguments FOR (Modern View):

  • ✅ LWT (Lightweight Transactions): Paxos for critical operations
  • ✅ Tunable Consistency: Use QUORUM or ALL for important writes
  • ✅ High Availability: No downtime for regional failures
  • ✅ Massive Scale: Handle billions of transactions
  • ✅ Audit Trail: Immutable append-only = perfect audit log

Real-World Example:

Capital One uses Cassandra for:

  • ✅ Transaction history (read-heavy, immutable)
  • ✅ Fraud detection (real-time, high volume)
  • ✅ Customer activity logs
  • ❌ NOT for: Account balances (use RDBMS)

Best Answer:

// Hybrid Approach PostgreSQL: Account balances, critical state Cassandra: Transaction history, audit logs, events // Write pattern 1. Update balance in PostgreSQL (ACID) 2. Write transaction log to Cassandra (immutable audit) 3. If Cassandra fails, still succeed (eventual consistency OK)

Interview Tip: Show you understand trade-offs, not just "Cassandra is always good" or "never for banking"!

💡 Interview Tips for Scenarios

STAR Method for Scenario Questions

  • 📖 Situation: Describe the problem clearly
  • 🎯 Task: What needed to be done?
  • 🔧 Action: Your specific steps (commands, tools)
  • ✅ Result: Outcome with metrics (latency reduced by 80%)
✅

Good Answers

  • Use specific commands
  • Mention monitoring tools
  • Include metrics (before/after)
  • Discuss trade-offs
  • Show diagnostic process
  • Reference real companies
❌

Bad Answers

  • Vague responses ("I'd check logs")
  • No specific tools/commands
  • Skip diagnostic steps
  • One-dimensional thinking
  • No metrics or proof
  • Theoretical only

Key Topics to Master

Scenario Categories

  • ⚡ Performance: Slow reads/writes, compaction, hot partitions
  • 💔 Failures: Node crashes, datacenter outages, data loss
  • 📈 Scaling: Adding nodes, capacity planning, rebalancing
  • 🔧 Data Issues: Tombstones, counters, inconsistencies
  • 🛠️ Operations: Repairs, backups, upgrades, node replacement
  • 🎯 Design: When to use Cassandra, architecture decisions

🎯 You're Ready for Real-World Scenarios!

You now have practical experience handling production Cassandra challenges!

🎭 Scenarios Covered:

  • ⚡ Performance troubleshooting (latency, throughput)
  • 💔 Failure handling (node crashes, data loss)
  • 📈 Scaling operations (adding nodes, hot partitions)
  • 🔧 Data issues (tombstones, counters)
  • 🛠️ Operations (node replacement, repairs)
  • 🎯 Design decisions (when to use Cassandra)

💡 Remember:

  • 🔍 Diagnostic Process: Always show how you'd investigate
  • 🛠️ Specific Tools: Mention nodetool, cfstats, iostat
  • 📊 Use Metrics: Before/after numbers prove impact
  • ⚖️ Discuss Trade-offs: Why this solution over alternatives
  • 🏢 Real Examples: Reference Netflix, Instagram, etc.
  • 📝 STAR Method: Situation → Task → Action → Result

🎭 Practice scenarios daily - they reveal production expertise! 🚀

Advertisement

Responsive Ad