Architecture Questions
Master Cassandra architecture interview questions with detailed explanations & real-world scenarios!
💼 Cassandra Architecture Interview Questions
Architecture questions test your deep understanding of how Cassandra works under the hood!
What Interviewers Look For:
- 🏗️ System Design: Understanding distributed architecture
- ⚖️ Trade-offs: CAP theorem, consistency vs availability
- 📊 Real-World: Production experience and troubleshooting
- 🔍 Deep Knowledge: Internals beyond surface level
- 💡 Problem Solving: Handling failures and edge cases
- 🎯 Best Practices: When to use what configuration
Interview Preparation Tips
✅ Draw diagrams - Visualize ring topology, data distribution
✅ Use examples - Reference Netflix, Instagram, Uber use cases
✅ Explain trade-offs - Why choose one approach over another
✅ Show production awareness - Mention monitoring, operations
✅ Ask clarifying questions - Requirements, scale, SLA
⭕ Ring Topology & Peer-to-Peer
Explain Cassandra's ring architecture. Why is it better than master-slave?
Perfect Answer
Cassandra uses a peer-to-peer ring architecture where:
- No Single Point of Failure: All nodes are equal - no master/slave
- Token Ring: 0 to 2^63-1 distributed among nodes
- Consistent Hashing: Each node owns a range of tokens
- Virtual Nodes (Vnodes): Each physical node owns multiple token ranges for better distribution
Advantages over Master-Slave:
- ✅ No Bottleneck: Any node can handle any request
- ✅ True High Availability: No master election downtime
- ✅ Linear Scalability: Just add more nodes to the ring
- ✅ Automatic Rebalancing: Vnodes distribute data evenly
- ✅ Geographical Distribution: Nodes across datacenters equally
Example: Netflix uses Cassandra's ring architecture to handle 1 trillion requests/day across global datacenters without any single point of failure.
What are Virtual Nodes (Vnodes)? Why are they important?
Perfect Answer
Vnodes divide each physical node into multiple virtual nodes (256 by default), each owning a small token range.
Benefits:
- 📊 Balanced Distribution: Data spreads evenly across cluster
- ⚡ Faster Rebuilds: Failed node's data distributed across many nodes
- 🔄 Easier Scaling: Adding/removing nodes rebalances automatically
- 💪 Heterogeneous Hardware: Powerful nodes get more vnodes
Without Vnodes: If Node A fails, only Node B (next in ring) rebuilds. Slow!
With Vnodes: All nodes participate in rebuild. Much faster!
How does the Gossip Protocol work? Why is it needed?
Perfect Answer
Gossip is a peer-to-peer communication protocol for cluster state sharing.
How it works:
- ⏰ Every second: Each node picks 1-3 random nodes
- 💬 Exchanges state: Share node health, schema, token ownership
- 🔄 Epidemic spread: Information propagates exponentially
- 📊 Convergence: Entire cluster learns state in O(log N) rounds
Information Shared:
- Node status (UP, DOWN, JOINING, LEAVING)
- Schema versions
- Load information
- Token ownership
Why needed: In a peer-to-peer system with no central coordinator, gossip ensures eventual consistency of cluster metadata across all nodes.
Example: When Node A goes down, within 3-5 seconds, all nodes know via gossip and stop routing requests to it.
🔄 Replication Strategy
Explain SimpleStrategy vs NetworkTopologyStrategy. When to use each?
SimpleStrategy
- Use: Single datacenter only
- How: Places replicas on next N nodes clockwise in ring
- Rack-aware: No
- Production: ❌ Never for production
- Good for: Dev/test environments
NetworkTopologyStrategy
- Use: Always in production
- How: RF per datacenter, rack-aware placement
- Rack-aware: Yes
- Multi-DC: ✅ Designed for it
- Benefit: Fault tolerance per DC and rack
Production Best Practice
ALWAYS use NetworkTopologyStrategy, even for single datacenter!
Reason: Makes future expansion to multi-DC easier. Changing strategy requires full data migration.
What is Replication Factor? How do you choose the right RF?
Perfect Answer
Replication Factor (RF) = Number of copies of each data row across the cluster.
Common Configurations:
- RF=1: ❌ No redundancy - data loss if node fails
- RF=2: ⚠️ Minimum for dev - survives 1 node failure
- RF=3: ✅ Production standard - survives 2 node failures
- RF=5: 🏢 High availability - survives 4 node failures
Selection Criteria:
- 💰 Cost vs Availability: RF=3 is sweet spot (3x storage)
- 📊 Read Performance: Higher RF = more replicas to read from
- 🔧 Maintenance: Higher RF = easier operations (can take nodes down)
- 🌍 Multi-DC: RF=3 per datacenter typically
Formula for fault tolerance: Can lose (RF - 1) nodes and still serve reads/writes with QUORUM
What is Hinted Handoff? How does it work?
Perfect Answer
Hinted Handoff ensures data isn't lost when a replica node is temporarily down.
How it works:
- Write arrives for partition owned by Node A, B, C (RF=3)
- Node B is temporarily down
- Coordinator writes to A and C successfully
- Coordinator stores a "hint" on another node (Node D)
- When Node B comes back up, Node D replays the hint to B
- Node B is now up-to-date
Important Details:
- ⏰ Hint timeout: 3 hours by default (max_hint_window)
- 💾 Storage: Hints stored on local disk
- 🔄 Not counted: Hints don't count toward consistency level
- ⚠️ Limitation: If node down >3 hours, run manual repair
Production Note: Hinted handoff is why temporary node failures don't cause data loss!
⚖️ Consistency Levels
Explain QUORUM consistency. Why is it the most common choice?
Perfect Answer
QUORUM = Majority of replicas must respond (RF/2 + 1 rounded down)
Why it's popular:
- ⚖️ Balance: Good consistency without sacrificing availability
- ✅ Strong Consistency: R(QUORUM) + W(QUORUM) > RF guarantees latest data
- 💪 Fault Tolerance: Works even if (RF - QUORUM) nodes are down
- 📊 Production Standard: Used by Netflix, Instagram, Uber
Example with RF=3:
- Write needs 2/3 nodes → Success even if 1 node down
- Read checks 2/3 nodes → Gets latest value (overlap guaranteed)
- Can lose 1 node and still operate normally
What's the difference between QUORUM and LOCAL_QUORUM?
QUORUM
- Scope: All datacenters
- Example: RF=3 per DC (2 DCs) → Need 4/6 replicas
- Latency: High (cross-DC)
- Use case: Single DC only
- Problem: One DC down = writes fail
LOCAL_QUORUM
- Scope: Current datacenter only
- Example: RF=3 → Need 2/3 in local DC
- Latency: Low (local DC)
- Use case: Multi-DC production
- Benefit: DC-independent operation
Multi-DC Best Practice
ALWAYS use LOCAL_QUORUM for multi-DC deployments!
Why:
- ✅ Low latency - no cross-DC wait
- ✅ DC independence - one DC down doesn't affect others
- ✅ Async replication to other DCs happens in background
How does Cassandra achieve "Tunable Consistency"? Give examples.
Perfect Answer
Cassandra lets you choose consistency level per-query, trading off consistency for availability/latency.
Consistency Spectrum:
Real-World Use Cases:
- Social Media Feed: Write(LOCAL_QUORUM) + Read(ONE) - Speed > Consistency
- Financial Transactions: Write(QUORUM) + Read(QUORUM) - Strong consistency needed
- Logging/Metrics: Write(ONE) - Speed critical, some loss acceptable
- User Profiles: Write(QUORUM) + Read(ONE) - Recent writes critical on write
Strong Consistency Formula: W + R > RF (e.g., QUORUM + QUORUM > 3)
📊 Partitioning & Data Distribution
How does Cassandra distribute data across nodes? Explain the partitioning process.
Perfect Answer
Cassandra uses Consistent Hashing with Virtual Nodes for data distribution.
Step-by-step process:
- Partition Key: Extract from PRIMARY KEY
- Hash Function: Murmur3 hash (default partitioner)
- Token: Hash produces 64-bit token (range: -2^63 to 2^63-1)
- Ring Lookup: Token mapped to node owning that range
- Replica Placement: RF additional nodes clockwise in ring
Key Benefits:
- 📊 Even Distribution: Hash ensures uniform spread
- 🎯 Fast Lookups: O(1) to find owning node
- ⚡ No Hotspots: Unless poor partition key choice
- 🔄 Dynamic: Easy to add/remove nodes
What causes "hot partitions"? How do you detect and fix them?
Perfect Answer
Hot partition = One partition receiving disproportionate traffic, overloading specific nodes.
Common Causes:
- ❌ Poor Partition Key: Low cardinality (e.g., status='active')
- ❌ Time-based: Current date/hour as partition key
- ❌ Celebrity: One user_id getting massive traffic
- ❌ Batch Jobs: Processing only recent data
Detection:
Solutions:
- ✅ Bucketing: Add bucket to partition key (e.g., (user_id, bucket))
- ✅ Random Salt: Add random suffix to spread data
- ✅ Composite Keys: Combine multiple high-cardinality columns
- ✅ Application-level: Cache hot data in Redis/Memcached
Example Fix:
⚡ Read & Write Path
Explain Cassandra's write path. Why are writes so fast?
Perfect Answer
Cassandra writes are fast because they're append-only with no reads required.
Write Path (4 steps):
- Commit Log: Sequential append to disk (durability)
- Memtable: Write to in-memory structure (speed)
- Response: Acknowledge to client immediately
- Flush: Memtable → SSTable periodically
Why it's fast:
- ⚡ No Read-Before-Write: Unlike RDBMS updates
- 📝 Sequential I/O: Commit log is append-only
- 💾 Memory Writes: Memtable in RAM
- 🔄 Async Flush: SSTable flush happens in background
- 🚫 No Locks: No locking or coordination needed
Performance:
- Can handle 10,000-100,000+ writes/sec per node
- Write latency typically <1ms locally
- Linear scalability - double nodes = double throughput
Describe the read path. What makes reads slower than writes?
Perfect Answer
Reads are complex because data might be in multiple places (Memtable + multiple SSTables).
Read Path:
- Row Cache: Check if full row cached (fastest)
- Bloom Filter: Check each SSTable's bloom filter
- Key Cache: Check for partition key location
- Partition Index: Find partition in SSTable
- Memtable: Check in-memory data
- SSTables: Read from disk (potentially multiple)
- Merge: Combine data using timestamp (last write wins)
- Read Repair: Check consistency if enabled
Why slower than writes:
- 🔍 Multiple Sources: Must check memtable + SSTables
- 💿 Disk I/O: Random reads from SSTables
- 🔀 Merge: Combine data from multiple sources
- ⚖️ Consistency: May query multiple replicas
Optimizations:
- ✅ Compaction: Fewer SSTables = faster reads
- ✅ Row Cache: Cache frequently read rows
- ✅ SSD: Fast random reads
- ✅ Partition Key Queries: Always query by partition key
What is compaction? Why is it necessary?
Perfect Answer
Compaction merges multiple SSTables into fewer, larger SSTables to improve read performance.
Why Necessary:
- 📈 SSTable Proliferation: Each memtable flush creates new SSTable
- 🐌 Read Amplification: More SSTables = slower reads
- 💀 Tombstone Removal: Deleted data needs cleanup
- 💾 Disk Space: Reclaim space from deleted/updated data
Compaction Strategies:
- SizeTieredCompactionStrategy (STCS): Default, general purpose
- LeveledCompactionStrategy (LCS): Better for read-heavy, predictable space
- TimeWindowCompactionStrategy (TWCS): Perfect for time-series data
Trade-offs:
- ⚡ CPU/Disk: Compaction uses resources
- 💾 Space: Needs 2x space during compaction
- ⏰ Timing: Can impact performance if not tuned
💔 Failure Handling
What happens when a node fails? How does Cassandra handle it?
Perfect Answer
Cassandra continues operating normally due to replication - no single point of failure!
Immediate Response (within seconds):
- Gossip Detection: Other nodes detect failure via gossip
- Mark Down: Node marked as DOWN in cluster state
- Reroute Requests: Coordinator stops sending requests to failed node
- Serve from Replicas: Other RF-1 replicas handle all requests
- Hinted Handoff: Hints stored for missed writes
Different Scenarios:
- Temporary (minutes): Hinted handoff replays writes when back
- Short-term (hours): Hinted handoff up to 3 hours
- Long-term (>3 hours): Run `nodetool repair` to sync data
- Permanent: Replace node, stream data from replicas
Impact on Operations:
- 📖 Reads: No impact if RF > consistency level
- ✍️ Writes: No impact if RF > consistency level
- ⚖️ QUORUM: Works as long as majority still up
Example: RF=3, QUORUM writes. One node fails → 2 replicas still available → Writes continue normally!
What is Read Repair? When does it happen?
Perfect Answer
Read Repair fixes inconsistencies between replicas detected during reads.
How it works:
- Client reads with QUORUM (2/3 replicas respond)
- Coordinator compares timestamps from all replicas
- If mismatch detected, coordinator returns latest version to client
- Coordinator asynchronously updates stale replicas with latest data
Two Types:
- Blocking Read Repair: Checks all replicas before responding (slower but consistent)
- Background Read Repair: Returns immediately, repairs async (faster)
When it's needed:
- Node was down during writes (hinted handoff failed)
- Write didn't reach all replicas
- Replica corruption
Alternative: `nodetool repair` - manual repair of entire node/table
💡 Interview Tips & Best Practices
How to Ace Architecture Interviews
- 🎨 Draw Diagrams: Visualize ring topology, replication, write path
- 📊 Use Numbers: "RF=3 with QUORUM means 2/3 replicas..."
- 🏢 Real Examples: "Netflix uses LOCAL_QUORUM for..."
- ⚖️ Explain Trade-offs: "QUORUM vs ALL - consistency vs availability"
- 🔧 Operations: Mention monitoring, nodetool, repairs
- ❓ Ask Questions: "What's the expected scale? Multi-DC?"
Common Mistakes to Avoid
- ❌ Too theoretical: Connect to real-world production scenarios
- ❌ One-word answers: Explain the "why" behind concepts
- ❌ Ignoring CAP: Always discuss trade-offs
- ❌ No failure scenarios: Discuss what happens when things break
- ❌ Forgetting ops: Mention monitoring, repair, backup
Sample Answer Structure
Question: "Explain how Cassandra achieves high availability."
Perfect Answer Structure:
- Core Concept: "Cassandra uses replication and peer-to-peer architecture..."
- Technical Details: "With RF=3, data exists on 3 nodes. Ring topology means..."
- Example: "Netflix runs 280+ node cluster. If one node fails..."
- Trade-offs: "Higher RF = better availability but 3x storage cost..."
- Production: "Monitor with nodetool status, run repairs weekly..."
Study Checklist
Master These Topics
- ☑️ Ring topology & peer-to-peer architecture
- ☑️ Consistent hashing & virtual nodes (vnodes)
- ☑️ Replication strategies (SimpleStrategy vs NetworkTopologyStrategy)
- ☑️ Consistency levels (ONE, QUORUM, ALL, LOCAL_QUORUM)
- ☑️ Write path (commit log, memtable, SSTable)
- ☑️ Read path (bloom filters, compaction)
- ☑️ Gossip protocol & failure detection
- ☑️ Hinted handoff & read repair
- ☑️ CAP theorem trade-offs
- ☑️ Partition keys & data distribution
🎯 You're Ready for Architecture Interviews!
You now have deep knowledge of Cassandra's architecture and can confidently answer interview questions!
📚 Key Concepts Covered:
- ⭕ Ring topology & peer-to-peer architecture
- 🔄 Replication strategies & fault tolerance
- ⚖️ Tunable consistency & CAP theorem
- 📊 Partitioning & data distribution
- ⚡ Read/write paths & performance
- 💔 Failure handling & recovery
💡 Remember:
- 🎨 Always draw diagrams to explain concepts
- 📊 Use specific numbers and examples
- 🏢 Reference production deployments (Netflix, Instagram)
- ⚖️ Discuss trade-offs and alternatives
- 🔧 Show operational awareness
- ❓ Ask clarifying questions about requirements
💼 Good luck with your Cassandra interview! 🚀
📱 Responsive Ad 📱