Write Path

WRITE PATH - Commit Log, Memtable & SSTables

✍️ Master how Cassandra handles writes: durability, performance, and the journey from commit log to SSTable!

📖 Netflix: Processing Billions of Writes Per Day

Netflix stores viewing data for 200M+ subscribers with billions of writes daily. Every time you watch, pause, rewind, or finish an episode, that's a write to Cassandra! The challenge? Durability + Performance at massive scale. Can't lose viewing progress (durability). Must handle billions of writes/day (performance). Must not slow down playback (latency). Netflix's write path requirements: 100% durability - if a write succeeds, it must NEVER be lost. Sub-5ms latency - writes must not block playback. Billions of writes/day - must scale horizontally. Cassandra's write path solution: Commit Log - append-only sequential writes to disk (durability!). Memtable - in-memory structure for fast writes (speed!). SSTable - immutable on-disk files (scalability!). Result: 3-5ms write latency at billions of writes/day, zero data loss! Netflix Engineering: "Cassandra's write path gives us the best of both worlds: sequential disk writes for durability, in-memory writes for speed. The commit log ensures we never lose a write, while memtables keep latency under 5ms. This architecture is why we can handle billions of viewing events without dropping a single one!"

🔄 Write Path Overview: The Journey of a Write

Cassandra's write path is optimized for speed AND durability - a rare combination! Let's follow a write from client to disk.

Cassandra Write Path: 3-Step Journey 1️⃣ Client Write INSERT INTO users VALUES (...) Step 1 2️⃣ Commit Log (Disk) ✓ Sequential append-only ✓ Durability guaranteed ~1ms latency Step 2 (parallel) 3️⃣ Memtable (Memory) ✓ In-memory sorted tree ✓ Fast lookups ~100μs latency SUCCESS 4️⃣ Flush to SSTable (Later) When memtable full or periodic ✓ Immutable disk file Async (doesn't block writes) async Timeline: Write Request Life Cycle T=0ms Client sends write T=1ms Commit Log append ✓ (durability guaranteed) T=1.1ms Memtable update ✓ (parallel with commit log) T=3ms Client gets SUCCESS ✓ (write complete!) T=later SSTable flush (async) (when memtable full) 🔑 Key: Write succeeds in ~3ms (commit log + memtable). SSTable flush happens later asynchronously!
📝

1. Commit Log (Durability)

Purpose: Guarantee durability (never lose writes)
Type: Append-only sequential file on disk
Speed: ~1ms (sequential I/O is fast!)
Guarantee: Once written = durable forever
Recovery: Replay on crash to rebuild memtable
Trade-off: Slight latency for bulletproof durability

💾

2. Memtable (Speed)

Purpose: Fast writes and reads
Type: In-memory sorted tree structure
Speed: ~100μs (memory is instant!)
Structure: Sorted by partition key for fast lookups
Size limit: Configurable (default 64-128MB)
Flush trigger: When full → becomes SSTable

💿

3. SSTable (Scalability)

Purpose: Persistent on-disk storage
Type: Immutable sorted file
Creation: Memtable flush (async)
Immutability: Never modified after creation
Compaction: Multiple SSTables merged periodically
Benefit: Horizontal scalability (add more disks!)

⚡ Why This Architecture is Brilliant

The magic: Cassandra achieves BOTH durability AND speed through clever separation!

Durability: Commit log ensures writes survive crashes (sequential I/O = fast!)
Speed: Memtable makes writes instant (memory = 100x faster than disk!)
Scalability: Immutable SSTables = easy compaction and parallelization

Most databases choose 2 of 3: fast, durable, scalable. Cassandra gets all 3 by using the right structure for each goal!

📝 Commit Log: The Durability Guardian

The commit log is Cassandra's crash-proof insurance policy - your data survives anything!

📋

What is Commit Log?

Definition: Append-only log of ALL writes
Location: Dedicated disk (not same as data!)
Format: Sequential binary file
Writes: Every mutation logged BEFORE ack
Purpose: Replay on crash to restore memtable
Guarantee: If write ACKed = guaranteed durable

⚡

Why Sequential Writes?

Sequential I/O: Append to end of file
Speed: 100-300 MB/s (even on HDD!)
Compare random: Only 1-2 MB/s on HDD
Why fast: No disk seeking (head stays in place)
Latency: ~1ms per write (tiny overhead!)
Lesson: Sequential = fast, random = slow

🔄

Crash Recovery

Scenario: Node crashes before memtable flush
Problem: Memtable data lost (it was in RAM!)
Solution: Replay commit log on restart
Process: Read log, re-apply writes to memtable
Result: Memtable restored, zero data loss!
Speed: ~1 minute for 1GB commit log

✂️

Truncation

When: After memtable flushed to SSTable
Why: Writes already on disk (no longer needed)
Process: Delete old commit log segments
Frequency: After each memtable flush
Disk space: Commit log stays small
Typical size: 1-4GB depending on write rate

⚙️

Configuration

commitlog_directory: Where log files stored
commitlog_sync: periodic (default) or batch
commitlog_sync_period: 10s default
commitlog_segment_size: 32MB default
Best practice: Separate disk from data!
Why: Avoid I/O contention

📊

Performance Impact

Latency added: ~1ms per write
Throughput: Up to 300 MB/s sequential
Cost: Minimal (1ms for guaranteed durability!)
SSD benefit: Even faster (500+ MB/s)
Bottleneck: Rarely (unless disk saturated)
Worth it: Absolutely! Durability is critical

💾 The Commit Log Saved Our Data

Production incident at Discord:

3am - Power outage in datacenter → 20 Cassandra nodes crash simultaneously
Memtables had 2GB of recent messages (not yet flushed)
If lost = millions of Discord messages gone forever!

Nodes restart:
• Cassandra reads commit logs on each node
• Replays all writes from past 10 minutes
• Rebuilds memtables with 2GB of data
• Takes ~90 seconds per node

Result: ZERO messages lost!

Discord Engineering: "The commit log is what makes Cassandra production-grade. Without it, that power outage would have cost us millions of messages. 1ms of latency for bulletproof durability is the best trade-off in databases!"

💾 Memtable: The In-Memory Speed Demon

Memtable is where speed magic happens - sub-millisecond writes in sorted memory!

🌳

Data Structure

Type: Sorted tree (LSM-tree component)
Sorted by: Partition key, then clustering key
Why sorted: Fast range queries and merges
Access time: O(log n) for lookup
Memory: Typically 64-128MB per table
Per-table: Each table has its own memtable

✍️

Write Process

Step 1: Write arrives (after commit log)
Step 2: Insert into sorted position
Step 3: Update metadata (size, count)
Latency: ~100 microseconds
No disk I/O: Pure memory operation!
Concurrent: Thread-safe with minimal locking

📖

Read Process

First stop: Always check memtable first
Why: Most recent data lives here
Speed: Sub-millisecond (it's in RAM!)
Hit rate: High for recent writes
Miss: Continue to SSTables
Optimization: Reduces disk I/O

📏

Size Management

Default size: 64MB (configurable)
Monitoring: Track current size continuously
Flush trigger: When threshold reached
Multiple tables: Each has separate threshold
Memory pressure: Can trigger early flush
Tuning: Larger = fewer flushes, more memory

⚠️

Flush Triggers

1. Size limit: Memtable reaches threshold
2. Time limit: Periodic flush (every hour)
3. Manual: nodetool flush command
4. Commit log full: Force flush to truncate
5. Shutdown: Flush all before stopping
Priority: Size-based is most common

🎯

Performance Impact

Write latency: ~100μs (0.1ms!)
Read latency: Sub-millisecond for hits
Memory cost: Proportional to write rate
GC impact: Can cause pauses if too large
Sweet spot: 64-128MB per table
Trade-off: Size vs flush frequency

💡 Why In-Memory is 100x Faster

RAM vs Disk speed comparison:

RAM (Memtable): ~100 nanoseconds access time = 0.0001ms
SSD (SSTable): ~100 microseconds access time = 0.1ms (1,000x slower)
HDD (SSTable): ~10 milliseconds access time = 10ms (100,000x slower!)

By keeping recent writes in memtable, Cassandra avoids disk I/O completely for hot data. This is why write latency is consistently under 5ms even at massive scale. The memtable acts as a write buffer that absorbs bursts and smooths out disk I/O!

💿 SSTable: Immutable Sorted String Table

SSTables are Cassandra's permanent storage layer - immutable, sorted, and built for scale!

📦

What is SSTable?

Name: Sorted String Table
Format: Immutable file on disk
Sorted: By partition key + clustering key
Created: When memtable flushes
Never modified: Immutable forever!
Typical size: 100MB-1GB per file

🔒

Why Immutable?

1. No locks needed: Concurrent reads safe
2. Easy caching: Cache never invalidated
3. Simple backups: File won't change mid-copy
4. Easy replication: Send file once
5. Predictable: No in-place updates
Trade-off: Need compaction to merge

📂

SSTable Components

Data.db: Actual data (large file)
Index.db: Partition index (fast lookups)
Summary.db: Sample of index (in memory)
Statistics.db: Metadata (min/max keys)
Filter.db: Bloom filter (quick negatives)
Total files: 5-10 per SSTable

🔍

Read Process

Step 1: Check bloom filter (key exists?)
Step 2: Lookup in summary (which block?)
Step 3: Read index.db (exact offset)
Step 4: Read data.db (actual value)
Latency: 1-10ms depending on cache
Optimization: Index + bloom filter critical

🔄

Compaction

Problem: Many SSTables = slow reads
Solution: Merge multiple SSTables → one
Process: Read N files, merge, write 1 new
Benefits: Fewer files, remove tombstones
Strategies: Size-tiered, Leveled, Time-window
Cost: I/O intensive (background process)

📊

Storage Efficiency

Compression: LZ4 default (2-5x reduction)
Bloom filter: Saves disk reads (90%+ hits)
Deduplication: Compaction removes old versions
Typical ratio: 3x data reduction after compaction
Disk usage: 2x logical data (with RF=3)
Trade-off: CPU for compression vs disk space

🎯 Immutability: The Secret to Cassandra's Scale

Why immutability matters:

Traditional databases (mutable files):
• Update in place = need locks = slow concurrent access
• Update in place = complex caching (cache invalidation!)
• Update in place = difficult replication (consistency issues)
• Update in place = random disk I/O = slow!

Cassandra (immutable SSTables):
• Never update = no locks = blazing concurrent reads ✓
• Never update = aggressive caching (never stale!) ✓
• Never update = simple replication (send file once) ✓
• Never update = sequential I/O = fast! ✓

The trade-off:
• Need compaction to merge old + new data
• Uses more disk temporarily (old + new files)
• But: This cost is worth it for the scalability gains!

Result: Cassandra can scale to hundreds of nodes and petabytes of data because immutability eliminates coordination overhead. This is the foundation of Cassandra's "write anywhere, read anywhere" architecture!

⚡ Flush Process: Memtable → SSTable

Flushing converts in-memory data to durable disk files - the bridge between speed and persistence!

Flush Process: Step-by-Step Memtable (Full) 64MB data ⚠️ Flush triggered! (size threshold reached) Step 1 Freeze Current ✓ Current → read-only ✓ New memtable created (writes continue!) Step 2 Sort & Write Write to disk: • Data.db (sorted data) • Index.db (partition index) • Filter.db (bloom filter) Step 3 New SSTable ✓ Immutable on disk ✓ Indexed & compressed ✓ Ready for reads Size: ~100MB Step 4 Truncate Commit Log ✓ Data safe on disk ✓ Delete old segments (free up disk space) ⏱️ Typical Flush Timeline 64MB memtable → 100MB SSTable Duration: 2-5 seconds (doesn't block writes!)
🎯

Flush Triggers

1. Size threshold: Memtable reaches 64MB
2. Time-based: Periodic flush (1 hour default)
3. Commit log pressure: Log too large
4. Memory pressure: JVM heap getting full
5. Manual: nodetool flush command
Most common: Size threshold

⚡

Non-Blocking Design

Key insight: Flush doesn't block writes!
How: Freeze current, create new memtable
Writes: Go to new memtable immediately
Old memtable: Flushed in background
Result: Zero impact on write latency
Benefit: Consistent performance

📊

Performance Impact

Duration: 2-5 seconds for 64MB
I/O impact: Sequential write = fast
CPU impact: Sorting + compression
Write latency: No impact (new memtable!)
Read latency: Slight improvement (fewer memtables)
Frequency: Every few minutes typical

✅ Why Flush is Non-Blocking

The genius of the freeze-and-switch pattern:

When memtable reaches 64MB:
1. Current memtable → frozen (read-only)
2. New empty memtable created (takes 1ms)
3. All new writes → go to new memtable
4. Old memtable → flushed to SSTable in background (2-5 seconds)

Result: Writes never wait! They go to the new memtable while the old one flushes. This is why Cassandra can maintain consistent 3-5ms write latency even during flushes. The cost of flush is paid in background I/O, not user-facing latency!

🚀 Performance Characteristics: Why Writes Are Fast

Cassandra achieves sub-5ms writes at billions/day through clever architecture!

⚡

Write Latency Breakdown

Commit log: ~1ms (sequential append)
Memtable: ~0.1ms (memory insert)
Network: ~1-2ms (depending on distance)
Replication: Parallel (not additive!)
Total P50: 3-5ms typical
Total P99: 8-15ms typical

📈

Throughput Limits

Single node: 10,000-50,000 writes/sec
Limited by: Commit log disk throughput
With SSD: Up to 100,000 writes/sec
Horizontal scaling: Linear! Add nodes = more throughput
10 nodes: 500,000 writes/sec easily
Netflix scale: Billions of writes/day

💾

Sequential vs Random I/O

Random I/O (traditional DB): 1-2 MB/s
Sequential I/O (Cassandra): 100-300 MB/s
Speed difference: 100-300x faster!
Why: No disk seeking (head doesn't move)
Commit log: Pure sequential append
Result: Disk not a bottleneck

🎯

Memory Efficiency

Memtable size: 64-128MB per table
Write buffer: Absorbs bursts effectively
GC friendly: Off-heap possible
Memory overhead: ~10-20% of dataset
Cache hit rate: 80-95% for recent data
Result: Minimal memory needs

📊

Comparison: Cassandra vs Others

Cassandra: 3-5ms writes (sequential!)
MySQL: 10-50ms writes (random I/O)
PostgreSQL: 10-30ms writes (WAL + table)
MongoDB: 5-10ms writes (similar approach)
Winner: Cassandra for write-heavy!
Reason: LSM-tree + sequential I/O

⚙️

Tuning Knobs

memtable_flush_writers: Parallel flush threads
commitlog_sync_period: Batch writes (10s default)
memtable_heap_space: Total memtable memory
concurrent_writes: 32 default (increase for high load)
commit_log_sync: periodic vs batch
Tuning impact: 2-3x throughput gain possible

⚡ Netflix Benchmark: Real Production Numbers

Netflix Cassandra cluster (production):

Hardware:
• 1,000+ nodes (m5.4xlarge AWS instances)
• 16 vCPUs, 64GB RAM per node
• EBS SSD storage (gp3)
• 10 Gbps network

Workload:
• 6 billion writes/day = 70,000 writes/sec average
• Peak: 300,000 writes/sec (prime time viewing)
• Data: Viewing progress, recommendations, metadata
• Average row size: 1KB

Performance:
• P50 latency: 2.8ms ✓
• P99 latency: 12ms ✓
• P99.9 latency: 45ms ✓
• Availability: 99.99% ✓
• Zero data loss (durability = 100%) ✓

Key insights:
• Write path never saturated (commit log headroom: 70%)
• Memtable flushes: Every 2-3 minutes per table
• SSTables: ~500 per node, compaction 24/7
• Bottleneck: Network (not disk!)

Netflix Engineering: "Cassandra's write path handles our massive scale effortlessly. The sequential commit log + in-memory memtable architecture means we're limited by network, not disk. We could easily 2x our write load without adding nodes!"

🏢 How Real Companies Optimize Write Path

💬 Discord: Handling Billions of Messages

Scale: 150M+ users, billions of messages/day
Challenge: Write latency must be sub-10ms (users notice delay!)

Write Path Optimization:
• Separate commit log disk (NVMe SSD)
• Larger memtables (128MB vs 64MB default)
• Aggressive compression (LZ4 fast mode)
• Off-heap memtables (avoid GC pauses)
• Tuned flush writers (8 threads vs 2 default)

Results:
• P50 write latency: 3.2ms ✓
• P99 write latency: 8ms ✓
• Throughput: 500k messages/sec peak
• Zero message loss during incidents ✓

Key lesson: Separate commit log disk was the biggest win! Moving commit log to dedicated NVMe SSD reduced write latency by 40% (from 5ms to 3ms). This is because commit log I/O no longer competes with SSTable compaction.

Discord Engineering: "The commit log is our durability guardian. During a datacenter power failure, we lost nothing because every message was safely in the commit log. The 1ms we pay for commit log durability is the best insurance policy in distributed systems!"

📊 Apple: IoT Time-Series at Massive Scale

Scale: 1B+ devices, trillions of events/year
Challenge: Extreme write volume (10M+ writes/sec!)

Write Path Strategy:
• Time-window compaction (perfect for time-series)
• Larger SSTables (1GB vs 100MB default)
• Batch writes from devices (reduce network overhead)
• Memtable_flush_writers = 16 (high parallelism)
• commitlog_sync = batch (trade latency for throughput)

Architecture:
• 75,000+ Cassandra nodes (largest cluster!)
• Custom hardware (NVMe for commit log)
• RF=3 across regions (durability)
• Data retention: 90 days (auto-compaction)

Results:
• 10M writes/sec sustained ✓
• 20M writes/sec peak ✓
• P99 latency: 15ms (acceptable for IoT)
• Storage: 100+ petabytes

Key insight: Batch mode commit log sync was critical! By batching writes every 50ms instead of fsync per write, throughput increased 5x. For IoT where 50ms delay is acceptable, this trade-off is perfect. Shows importance of matching write path to workload!

🛒 Uber Eats: Order Processing with Strong Durability

Scale: 100M+ orders/year, millions of users
Requirement: Cannot lose order data (financial!)

Write Path Configuration:
• commitlog_sync = periodic (default, safest)
• commitlog_sync_period = 5s (frequent fsync)
• Multiple commit log segments (faster recovery)
• QUORUM writes (strong consistency)
• Memtable_cleanup_threshold = low (eager flush)

Durability Focus:
• Commit log on RAID 1 mirrored disks
• Automatic commit log replay testing
• Monitor commit log latency (alert >5ms)
• Regular disaster recovery drills
• Backup commit logs to S3

Trade-offs:
• Write latency: 8-12ms (slower for durability)
• Throughput: 50k writes/sec per node (acceptable)
• Durability: 100% (zero order data loss!) ✓

Uber Engineering: "For order data, we prioritize durability over raw speed. The extra 5ms of commit log latency is worth it for absolute certainty that orders never disappear. We've had nodes crash mid-transaction, and every order was recovered from commit log. That's the power of append-only durability!"

✅ Best Practices: Optimizing Write Path

1️⃣

1. Separate Commit Log Disk

Best practice: Dedicated disk for commit log
Why: Avoid I/O contention with SSTables
Performance gain: 30-50% lower write latency
Recommended: NVMe SSD for commit log
Config: commitlog_directory separate path
Cost: Extra disk, big performance win

2️⃣

2. Tune Memtable Size

Default: 64MB memtable
High write load: Increase to 128-256MB
Benefit: Fewer flushes = less I/O
Trade-off: More memory, longer recovery
Monitor: Flush frequency (every 2-3 min ideal)
Config: memtable_heap_space_in_mb

3️⃣

3. Monitor Commit Log Latency

Metric: CommitLog write latency
Healthy: <2ms typical
Warning: >5ms = investigate
Alert: >10ms = disk problem!
Tools: nodetool tablestats, metrics
Action: Check disk I/O, saturation

4️⃣

4. Choose Right Compaction Strategy

STCS: Good for write-heavy, no updates
LCS: Good for read-heavy, predictable
TWCS: Perfect for time-series data
Impact: Affects flush/compaction balance
Write-heavy: STCS or TWCS recommended
Monitor: Pending compactions metric

5️⃣

5. Batch Writes When Possible

Pattern: Batch multiple writes together
Benefit: Amortize commit log overhead
Throughput: 3-5x higher with batching
Trade-off: Slight latency increase
CQL: BEGIN BATCH ... APPLY BATCH
Best for: High-volume ingestion

6️⃣

6. Test Commit Log Recovery

Why: Durability only matters if recovery works!
Test: Kill node during write load
Verify: Restart → commit log replays
Check: Zero data loss after recovery
Frequency: Test quarterly
Chaos: Include in chaos engineering

7️⃣

7. Avoid Small Writes

Problem: Many tiny writes = overhead
Solution: Batch or group related data
Example: Don't write 100 1-byte columns
Better: Write 1 100-byte blob
Reason: Commit log overhead per write
Impact: 5-10x better throughput

8️⃣

8. Monitor Memory Pressure

Metric: JVM heap usage
Healthy: <75% heap used
Warning: >85% = trigger flushes
Problem: Memory pressure = forced flush
Solution: More RAM or smaller memtables
Monitor: GC pause times (<100ms good)

9️⃣

9. Understand Trade-offs

Durability vs Speed: Commit log sync frequency
Memory vs Disk: Memtable size
Latency vs Throughput: Batching strategy
Write vs Read: Compaction strategy
Key lesson: No free lunch! Tune for your workload
Test: Benchmark before production

🚨 Common Mistakes to Avoid

  • ❌ Commit log on same disk as data: I/O contention kills performance
  • ❌ Tiny memtables: Excessive flushing → poor throughput
  • ❌ Ignoring commit log latency spikes: Early warning of disk problems
  • ❌ Using batch mode without testing: Can lose data on crash!
  • ❌ Over-tuning: Defaults are good! Only tune if measured bottleneck
  • ❌ Not testing recovery: Durability guarantees useless if recovery broken

💼 Interview Questions & Answers

1
Explain the complete Cassandra write path from client to disk

Complete Answer:

Cassandra's write path has 3 main steps: commit log, memtable, and eventually SSTable. Let me walk through the complete flow:

Step 1: Client Sends Write (T=0ms)

Client executes: INSERT INTO users (id, name, email) VALUES (123, 'Alice', 'alice@example.com')

The write is sent to the coordinator node (determined by partition key hash). The coordinator is responsible for routing the write to the appropriate replicas.

Step 2: Commit Log Write (T=0ms-1ms)

FIRST, before anything else, the write is appended to the commit log:

  • Commit log is an append-only file on disk
  • Sequential write = very fast (~1ms)
  • Guarantees durability - if node crashes, replay from log
  • fsync frequency configurable (default: every 10s batch)
  • This is the durability guarantee point!

Step 3: Memtable Write (T=1ms-1.1ms, parallel)

SIMULTANEOUSLY (parallel with commit log), write goes to memtable:

  • Memtable is an in-memory sorted tree structure
  • Insert into correct sorted position (by partition + clustering key)
  • O(log n) insertion time = ~100 microseconds
  • No disk I/O = blazing fast!
  • Each table has its own memtable

Step 4: Client Response (T=3ms)

Once commit log AND memtable succeed:

  • Send SUCCESS response to client
  • Total latency: ~3ms typical (1ms commit log + 0.1ms memtable + ~2ms network)
  • Write is now durable and queryable

Step 5: Eventual SSTable Flush (minutes later, async)

When memtable reaches size threshold (64MB default):

  • Current memtable frozen (made read-only)
  • New empty memtable created for new writes
  • Old memtable flushed to disk as immutable SSTable
  • SSTable components created: Data.db, Index.db, Filter.db, Summary.db
  • Flush takes 2-5 seconds but doesn't block writes!
  • After flush, corresponding commit log segments truncated

Why This Architecture?

  • Durability: Commit log ensures writes survive crashes
  • Speed: Memtable makes writes instant (memory = fast!)
  • Scalability: Immutable SSTables = easy compaction/replication
  • Non-blocking: Flush happens async, doesn't impact write latency

Crash Recovery:

If node crashes before memtable flush:

  • Memtable data lost (it was in RAM)
  • On restart, Cassandra replays commit log
  • Re-applies all writes to rebuild memtable
  • Result: Zero data loss! (commit log is the source of truth)

Key Insight: The separation of durability (commit log) and performance (memtable) is what makes Cassandra both fast AND reliable. Traditional databases have to choose between speed and durability, but Cassandra gets both through clever architecture!

2
Why are Cassandra writes so fast? Explain sequential vs random I/O

Complete Answer:

Cassandra writes are fast (3-5ms typical) because of two key architectural decisions: sequential I/O and in-memory writes.

Sequential vs Random I/O - The Physics:

Random I/O (traditional databases):

  • Update in place = write to specific disk location
  • Requires disk head to seek to that location
  • Seek time: 5-10ms on HDD (disk head moves physically!)
  • Throughput: 1-2 MB/s (100-200 operations/sec)
  • Example: MySQL updating a row = random write

Sequential I/O (Cassandra commit log):

  • Append to end of file = always same location (current end)
  • No seeking! Disk head stays in place
  • Latency: ~1ms on HDD, <0.1ms on SSD
  • Throughput: 100-300 MB/s on HDD, 500+ MB/s on SSD
  • 100-300x faster than random I/O!

Cassandra's Write Path Architecture:

1. Commit Log (Sequential I/O):

  • Append-only file, never updated in place
  • All writes go to end of current segment
  • Latency: ~1ms (sequential = fast!)
  • Provides durability without sacrificing speed

2. Memtable (In-Memory):

  • RAM access = 100,000x faster than disk!
  • Latency: ~100 microseconds = 0.1ms
  • No disk I/O at write time = instant
  • Sorted tree structure = efficient lookups

3. SSTable (Batched Sequential I/O):

  • Flush happens asynchronously (doesn't block writes)
  • Large sequential write (64-128MB at once)
  • Immutable = no random updates needed
  • Written once, read many times

Performance Comparison (Real Numbers):

Traditional Database (MySQL example):

  • Write to WAL (write-ahead log): 5ms (often random I/O)
  • Update table in place: 10ms (random I/O)
  • Update indexes: 5ms per index (more random I/O!)
  • Total: 20-50ms for a simple write
  • Throughput: 200-500 writes/sec per disk

Cassandra:

  • Commit log append: 1ms (sequential!)
  • Memtable insert: 0.1ms (memory!)
  • Total: 3ms including network
  • Throughput: 10,000-50,000 writes/sec per node
  • 50-100x faster than traditional approach!

Why Sequential is So Much Faster:

Physical Explanation (HDD):

  • Random write: Head seeks to track → wait for rotation → write
  • Seek time: 5-10ms
  • Rotational latency: 4ms (7200 RPM disk)
  • Transfer: 0.1ms
  • Total: ~10ms per write

Sequential write: Head stays at end of file

  • No seek: 0ms
  • No rotational wait (already at position): 0ms
  • Transfer: 0.1ms
  • Total: ~1ms per write
  • 10x faster!

Additional Cassandra Optimizations:

  1. Write batching: Commit log batches multiple writes (amortize fsync cost)
  2. Parallel writes: Commit log + memtable happen simultaneously
  3. No read-before-write: Don't need to read old value first
  4. No locking: Lock-free memtable operations
  5. Horizontal scaling: Add more nodes = linearly more throughput

Real Production Example:

Netflix writes 6 billion records/day to Cassandra:

  • Peak: 300,000 writes/sec
  • P50 latency: 2.8ms
  • P99 latency: 12ms
  • Hardware: Standard SSDs (not even high-end!)

This would be impossible with random I/O. Sequential commit log + in-memory memtable is the secret sauce!

Key Takeaway: Cassandra's write path is designed around the fundamental physics of storage: sequential I/O is 100x faster than random I/O. By using append-only commit log (sequential) and in-memory memtable (no I/O), Cassandra achieves both durability AND speed!

Advertisement

Responsive Ad