Real-World Scenarios
Master troubleshooting, performance issues, and production challenges with real Cassandra scenarios!
🎭 Real-World Cassandra Scenarios
Scenario-based questions test your production experience and problem-solving skills!
What Interviewers Test:
- 🔍 Troubleshooting: Can you diagnose production issues?
- ⚡ Performance: How do you optimize slow queries?
- 💔 Failure Handling: What happens when things break?
- 📈 Scaling: How do you handle growth?
- ⚖️ Trade-offs: Why choose one approach over another?
- 🛠️ Operations: Day-to-day cluster management
How to Answer Scenario Questions
✅ STAR Method: Situation → Task → Action → Result
✅ Be Specific: Use actual metrics, tools, commands
✅ Show Process: Explain your diagnostic steps
✅ Discuss Trade-offs: Why you chose this solution
✅ Mention Monitoring: How you detected the issue
⚡ Performance Issues
Scenario: Your application's read latency suddenly increased from 5ms to 500ms. How do you troubleshoot?
Perfect Answer
Diagnostic Process:
- Check Metrics First
# Check cluster health nodetool status nodetool tpstats # Thread pool statistics nodetool tablestats # SSTable counts
- Common Causes & Solutions
- 🔥 Too Many SSTables: Run compaction
nodetool compact keyspace_name table_name
- 💾 Memory Pressure: Check heap usage, tune JVM
nodetool info # Check heap usage # If heap >75%, tune GC or add memory
- 🔍 Hot Partitions: Check partition sizes
nodetool cfstats keyspace.table # Look for: Partition Size max/mean
- 💿 Disk I/O: SSD degradation, check iostat
iostat -x 1 # Check %util
- 🔥 Too Many SSTables: Run compaction
- Application-Level Checks
- Query patterns changed? (full table scan)
- Consistency level increased?
- Connection pool exhausted?
Real Example: "At Netflix, we found 100+ SSTables per partition. Switching from STCS to LCS compaction reduced p99 latency from 450ms to 8ms."
Scenario: Write throughput dropped by 50%. CPU is at 100%. What's happening?
Perfect Answer
Most Likely Cause: Compaction Storm
Why This Happens:
- 💥 Many SSTables triggered major compaction
- 🔥 Compaction using all CPU cores
- ✍️ Writes blocked waiting for compaction to finish
Immediate Fix:
Long-term Solutions:
- ✅ Tune compaction strategy (TWCS for time-series)
- ✅ Add more nodes to distribute load
- ✅ Increase compaction throughput during off-peak
- ✅ Use SSDs for better I/O
Pro Tip: Monitor compaction_throughput and adjust based on workload!
Scenario: Query works fine in dev (1000 rows) but times out in production (1M rows). Why?
Classic Anti-Pattern: Unbounded Partition
Problem: Query is reading entire partition into memory!
Solutions:
Lesson: Always consider partition size at scale!
💔 System Failures
Scenario: 2 nodes in a 6-node cluster with RF=3 crashed. Writes with QUORUM are failing. Why?
Perfect Answer
Analysis:
- Cluster: 6 nodes total, RF=3
- Down: 2 nodes
- Available: 4 nodes
- QUORUM needs: 2 out of 3 replicas
Why Failing:
For some partitions, 2 out of 3 replica nodes are down! QUORUM can't be satisfied.
Example:
Solutions:
- Immediate: Restart failed nodes ASAP
- Temporary: Reduce to ONE (loses consistency!)
- Long-term: Increase RF or cluster size
- RF=5 can lose 2 nodes and still do QUORUM
- Or use 9 nodes instead of 6
Production Best Practice: RF=3 with QUORUM can tolerate 1 node failure. For 2+ failures, need RF=5.
Scenario: After datacenter power outage, some data is missing. What happened and how to fix?
Disaster Recovery Scenario
What Happened:
- 💥 Power outage → All nodes lost power simultaneously
- 💾 Data in memtables (not yet flushed) was lost
- 📝 Commit log may be corrupted or incomplete
- ⏰ Hinted handoff data lost (stored in memory)
Immediate Actions:
Prevention:
- ✅ Multi-DC setup (geographic distribution)
- ✅ UPS for graceful shutdown
- ✅ Regular backups (snapshots)
- ✅ Periodic repairs (weekly)
- ✅ Monitor with consistency level ALL occasionally
Real Story: "Instagram lost 5 minutes of data during datacenter failure. Solution: Multi-DC setup with LOCAL_QUORUM + async replication."
📈 Scaling Challenges
Scenario: Your cluster is at 80% capacity. How do you add nodes without downtime?
Perfect Answer: Rolling Scale-Out
Step-by-Step Process:
- Prepare New Nodes
# Configure cassandra.yaml cluster_name: 'production_cluster' seeds: 'existing_seed_ip' listen_address: 'new_node_ip' auto_bootstrap: true ← Important!
- Add Nodes One at a Time
# Start new node sudo systemctl start cassandra # Monitor bootstrap progress nodetool netstats nodetool status # Should show 'UJ' → 'UN'
- Wait for Streaming to Complete
- New node receives data from existing nodes
- Wait until status shows "UN" (Up Normal)
- Can take hours depending on data size
- Run Cleanup on Old Nodes
# Remove data that now belongs to new nodes nodetool cleanup
- Repeat Until Desired Size
Important Notes:
- ⚠️ Don't add multiple nodes simultaneously (streaming overhead)
- ⏰ Add during low-traffic hours if possible
- 📊 Monitor cluster performance during scaling
- 🔄 Run repairs after all nodes added
Scenario: User uploads spiked 10x. Schema designed for user_id partition. Now hot partitions. Quick fix?
Emergency Hot Partition Fix
Current Schema (Problem):
Immediate Workaround (Application Level):
- Add Bucketing in Application
# Python pseudocode bucket = hash(user_id) % 10 # Split into 10 buckets write_to_partition((user_id, bucket), upload_data) # Read requires 10 queries (fan-out) for bucket in range(10): read_from_partition((user_id, bucket))
Proper Long-term Fix:
Alternative: Time-based Bucketing
🔧 Data Issues
Scenario: Deleted data is still appearing in queries. Why and how to fix?
Perfect Answer: Tombstones & GC Grace
Why This Happens:
- 📝 Deletes create tombstones (not immediate removal)
- ⏰ Tombstones persist for gc_grace_seconds (default 10 days)
- 👻 During reads, tombstones can be "resurrected" by repair/hints
Scenario Breakdown:
Solutions:
- Wait for GC Grace Period
# Check setting DESCRIBE TABLE table_name; # default: gc_grace_seconds = 864000 (10 days) # After 10 days, compaction removes tombstones
- Run Immediate Repair
nodetool repair keyspace table # Ensures all nodes have tombstones
- Force Compaction (Dangerous!)
⚠️ Only if you're SURE all nodes are in sync ALTER TABLE table WITH gc_grace_seconds = 0; nodetool compact ALTER TABLE table WITH gc_grace_seconds = 864000;
Prevention: Run repairs regularly (within gc_grace period)!
Scenario: Counter column shows wrong count. Users complaining. What went wrong?
Counter Columns Are NOT Idempotent!
Problem: Duplicate Increments
Why It Happens:
- ❌ Client retries on timeout
- ❌ Load balancer duplicates request
- ❌ Network issue causes duplicate packets
- ❌ First write actually succeeded but timed out
Solutions:
- Use Lightweight Transactions (LWT)
-- Idempotent but slower UPDATE likes_ledger SET action = 'liked' WHERE user_id = 'user1' AND post_id = 'abc' IF NOT EXISTS; -- Then count with SELECT COUNT(*)
- Application-Level Deduplication
# Generate unique request ID request_id = generate_uuid() # Store in dedup table with TTL INSERT INTO processed_requests (request_id) VALUES ('uuid') USING TTL 3600;
- Accept Approximate Counts
- For display purposes (like counts, views)
- Use counters knowing they might be slightly off
- Document the limitation
Lesson: Counters are great for approximate counts, terrible for exact billing!
🛠️ Operations
Scenario: You need to replace a failed node. Walk through the exact steps.
Perfect Answer: Node Replacement Procedure
Prerequisites:
- Failed node is completely dead (not coming back)
- You have the failed node's IP and token
- Replacement node has same specs
Step-by-Step:
- Get Failed Node Info
nodetool status # Note the DN (Down) node's IP: 192.168.1.100
- Prepare Replacement Node
# Edit cassandra.yaml cluster_name: 'production_cluster' auto_bootstrap: true # Add JVM option # -Dcassandra.replace_address=192.168.1.100
- Start Replacement Node
sudo systemctl start cassandra # Monitor streaming nodetool netstats tail -f /var/log/cassandra/system.log
- Wait for Bootstrap
- Node downloads data from other replicas
- Status changes: UJ → UN
- Can take hours for large datasets
- Remove JVM Flag
# After node is UN, remove from JVM options: # -Dcassandra.replace_address=192.168.1.100
- Run Repair
nodetool repair -full
- Remove Old Node
# On any live node nodetool removenode 'dead_node_id'
Important: Don't restart the node after replacement without removing the JVM flag!
🎯 Design Decisions
Scenario: Should you use Cassandra for a banking transaction system? Why or why not?
Thoughtful Answer: It Depends!
Arguments AGAINST (Traditional View):
- ❌ No ACID Transactions: No multi-row atomicity
- ❌ Eventual Consistency: Temporary inconsistencies possible
- ❌ No Joins: Complex reporting difficult
- ❌ Regulatory Concerns: Banks prefer traditional RDBMS
Arguments FOR (Modern View):
- ✅ LWT (Lightweight Transactions): Paxos for critical operations
- ✅ Tunable Consistency: Use QUORUM or ALL for important writes
- ✅ High Availability: No downtime for regional failures
- ✅ Massive Scale: Handle billions of transactions
- ✅ Audit Trail: Immutable append-only = perfect audit log
Real-World Example:
Capital One uses Cassandra for:
- ✅ Transaction history (read-heavy, immutable)
- ✅ Fraud detection (real-time, high volume)
- ✅ Customer activity logs
- ❌ NOT for: Account balances (use RDBMS)
Best Answer:
Interview Tip: Show you understand trade-offs, not just "Cassandra is always good" or "never for banking"!
💡 Interview Tips for Scenarios
STAR Method for Scenario Questions
- 📖 Situation: Describe the problem clearly
- 🎯 Task: What needed to be done?
- 🔧 Action: Your specific steps (commands, tools)
- ✅ Result: Outcome with metrics (latency reduced by 80%)
Good Answers
- Use specific commands
- Mention monitoring tools
- Include metrics (before/after)
- Discuss trade-offs
- Show diagnostic process
- Reference real companies
Bad Answers
- Vague responses ("I'd check logs")
- No specific tools/commands
- Skip diagnostic steps
- One-dimensional thinking
- No metrics or proof
- Theoretical only
Key Topics to Master
Scenario Categories
- ⚡ Performance: Slow reads/writes, compaction, hot partitions
- 💔 Failures: Node crashes, datacenter outages, data loss
- 📈 Scaling: Adding nodes, capacity planning, rebalancing
- 🔧 Data Issues: Tombstones, counters, inconsistencies
- 🛠️ Operations: Repairs, backups, upgrades, node replacement
- 🎯 Design: When to use Cassandra, architecture decisions
🎯 You're Ready for Real-World Scenarios!
You now have practical experience handling production Cassandra challenges!
🎭 Scenarios Covered:
- ⚡ Performance troubleshooting (latency, throughput)
- 💔 Failure handling (node crashes, data loss)
- 📈 Scaling operations (adding nodes, hot partitions)
- 🔧 Data issues (tombstones, counters)
- 🛠️ Operations (node replacement, repairs)
- 🎯 Design decisions (when to use Cassandra)
💡 Remember:
- 🔍 Diagnostic Process: Always show how you'd investigate
- 🛠️ Specific Tools: Mention nodetool, cfstats, iostat
- 📊 Use Metrics: Before/after numbers prove impact
- ⚖️ Discuss Trade-offs: Why this solution over alternatives
- 🏢 Real Examples: Reference Netflix, Instagram, etc.
- 📝 STAR Method: Situation → Task → Action → Result
🎭 Practice scenarios daily - they reveal production expertise! 🚀
Responsive Ad