Section 3: MongoDB Architecture

πŸ’Ύ WiredTiger: MongoDB's Storage Engine

Understanding how MongoDB stores, compresses, and retrieves billions of documents with blazing speed

πŸ“– The Tale of Two Libraries

Imagine you're managing a massive library with millions of books...

πŸ“š

Old Library: The Traditional System (MMAPv1)

Memory-mapped files, database-level locking

The Problems:

  • Database-Level Locking: When someone is writing to ANY book, the ENTIRE library is locked! 😰
  • No Compression: Books take up full shelf space - wastes storage
  • Memory Mapped: Relies on OS to manage memory - unpredictable performance
  • Write Amplification: Small updates rewrite entire pages
  • No Document-Level Concurrency: Poor performance for concurrent operations

Result: Slow writes, wasted space, poor concurrency. Reading is fast but everything else suffers!

πŸ’Ύ

Modern Library: The Smart System (WiredTiger)

Document-level locking, compression, intelligent caching

The Solutions:

  • Document-Level Locking: Multiple people can read/write different books simultaneously! ⚑
  • Compression: Books compressed to 70% smaller - saves massive storage (snappy, zlib, zstd)
  • WiredTiger Cache: Smart in-memory cache for frequently accessed data - blazing fast reads
  • Journaling: Crash-safe with write-ahead logging
  • Checkpoints: Regular snapshots every 60 seconds - durability guaranteed

Result: 10x faster writes, 70% less storage, perfect concurrency! πŸš€

πŸ’‘ The Key Insight

MMAPv1 = Old Library (slow, locks entire database)
WiredTiger = Smart Library (fast, locks only documents)
WiredTiger has been MongoDB's default since version 3.2 (2015)! 🎯

πŸ’Ύ What is WiredTiger?

WiredTiger is MongoDB's high-performance storage engine that manages how data is stored on disk and in memory.
It provides document-level concurrency, compression, and crash recovery.

βœ… What WiredTiger Does

  • βœ“ Stores documents on disk
  • βœ“ Manages in-memory cache
  • βœ“ Compresses data (snappy)
  • βœ“ Handles concurrent reads/writes
  • βœ“ Ensures crash recovery
  • βœ“ Creates checkpoints

🎯 Key Metrics

  • πŸ“Š Compression: 70-80% reduction
  • ⚑ Cache Size: 50% of RAM - 1GB
  • πŸ’Ύ Checkpoint: Every 60 seconds
  • πŸ”’ Locking: Document-level
  • πŸ“ Journal: Every 50ms
  • πŸš€ Concurrency: Thousands/sec

πŸ—οΈ WiredTiger Architecture

MongoDB Server Application Layer WiredTiger Cache (In-Memory) β€’ Hot Documents β€’ Recently Accessed β€’ 50% RAM - 1GB Journal (WAL) (Write-Ahead Log) β€’ Crash Recovery β€’ Every 50ms β€’ Durability Checkpoints (Snapshots) β€’ Every 60 seconds β€’ Consistent state β€’ On-disk files Data Files (.wt) (Compressed Storage) β€’ B-Tree Structure β€’ Snappy Compression β€’ Document Storage RAM DISK

πŸ“‹ How Data Flows:

  1. Write Operation: Data first goes to WiredTiger Cache (RAM)
  2. Journal Write: Operation logged to Journal every 50ms (crash safety)
  3. Checkpoint: Every 60 seconds, cache data flushed to disk (.wt files)
  4. Read Operation: Check cache first β†’ if not found, read from disk β†’ cache it

✍️ Write Operation Flow (Animated)

0ms 1ms 50ms 60s Complete Write Request db.users.insert() Data WiredTiger Cache (RAM) Journal (WAL) Every 50ms Checkpoint Every 60 seconds πŸ“Έ Data Files (.wt) Compressed Storage Write arrives Cached + Logged Journal sync Checkpoint flush Durable on disk

⚑ Write Path Timeline:

  • 0ms: Write request arrives from application
  • ~1ms: Data written to WiredTiger Cache (RAM) + Journal log
  • 50ms: Journal synced to disk (durability guaranteed)
  • 60 seconds: Checkpoint flushes cache to .wt data files
  • Result: Write acknowledged after 1-50ms, fully durable on disk after 60s

πŸ“– Read Operation Flow (Animated)

Read Request db.users.findOne() WiredTiger Cache (Check RAM first) βœ“ CACHE HIT βœ— CACHE MISS Data Files (.wt) (Read from disk) πŸ’Ύ Load to cache, then return Result Returned Document data ⚑ Performance Comparison Cache Hit: ~100 microseconds Cache Miss: ~10-50 milliseconds

πŸ“– Read Path Logic:

  1. Step 1: Read request arrives β†’ Check WiredTiger Cache first
  2. Step 2a (Cache Hit): Document found in RAM β†’ Return immediately (~100ΞΌs) ⚑
  3. Step 2b (Cache Miss): Not in cache β†’ Read from disk .wt files (~10-50ms) 🐒
  4. Step 3: If read from disk β†’ Load into cache for future reads
  5. Result: Cache hit rate >90% is critical for performance!

πŸ“¦ Compression in Action

Uncompressed Data 1 TB (1000 GB) πŸ’° Cost: $1000/month WiredTiger Compression Snappy (Default) 300 GB 70% Saved ⚑ πŸ’° $300/month Zstd (Balanced) 250 GB 75% Saved βš–οΈ πŸ’° $250/month Zlib (Maximum) 200 GB 80% Saved πŸ—œοΈ πŸ’° $200/month

πŸ’‘ Compression Impact:

None
1000 GB
$1000/mo
Snappy ⚑
300 GB
$300/mo
Zstd βš–οΈ
250 GB
$250/mo
Zlib πŸ—œοΈ
200 GB
$200/mo

🎯 Key Features of WiredTiger

πŸ”’ 1. Document-Level Concurrency Control

Multiple clients can modify different documents simultaneously without blocking each other.

Example: User A updating document 1 doesn't block User B updating document 2 - both happen in parallel!

Technology: Multi-Version Concurrency Control (MVCC) - each operation sees a consistent snapshot

πŸ“¦ 2. Compression

Compresses data and indexes to save 60-80% storage space.

Available Algorithms:
β€’ Snappy (default) - Fast, moderate compression (70% reduction)
β€’ Zlib - Slower, high compression (80% reduction)
β€’ Zstd - Best balance (75% reduction)

Real Impact: 1TB uncompressed β†’ 300MB compressed = Save $700/month on cloud storage!

πŸ’Ύ 3. WiredTiger Cache

In-memory cache that stores frequently accessed data for lightning-fast reads.

Cache Size: max(50% of RAM - 1GB, 256MB)
Example: Server with 16GB RAM β†’ Cache = 7GB
Eviction: LRU (Least Recently Used) when cache is 80% full

πŸ“ 4. Journaling (Write-Ahead Log)

Records all modifications before applying them - ensures crash recovery.

How it works:
1. Write operation arrives
2. Logged to journal file (every 50ms or 100MB)
3. On crash: replay journal to recover uncommitted writes
4. Journal files kept for 2 checkpoints, then deleted

πŸ“¦ Compression in Detail

Algorithm Speed Compression Ratio CPU Usage Best For
Snappy (Default) πŸš€ Fastest 70% reduction Low General purpose, production
Zlib 🐒 Slowest 80% reduction High Storage-critical, archives
Zstd ⚑ Fast 75% reduction Medium Best balance
None πŸš€ Instant 0% None Testing, RAM-heavy workloads

πŸ’‘ Configuration Example:

storage:
  engine: wiredTiger
  wiredTiger:
    engineConfig:
      cacheSizeGB: 8
    collectionConfig:
      blockCompressor: snappy
    indexConfig:
      prefixCompression: true

βœ… Checkpoints & Crash Recovery

πŸ“Έ What is a Checkpoint?

A checkpoint is a snapshot of data written from WiredTiger cache to disk files, creating a consistent state.

  • Frequency: Every 60 seconds (configurable)
  • What happens: Dirty pages (modified in cache) flushed to .wt files
  • Consistency: Represents all data committed up to that point
  • During checkpoint: System remains available, no downtime

πŸ”„ Crash Recovery Process

  1. Server crashes (power failure, kill -9, etc.)
  2. On restart: MongoDB loads last checkpoint (consistent state)
  3. Journal replay: Applies all writes from journal since last checkpoint
  4. Result: All committed writes recovered, data intact! βœ…
Example Timeline:
β€’ 10:00:00 - Checkpoint created
β€’ 10:00:30 - 100 writes happen
β€’ 10:00:45 - Server crashes πŸ’₯
β€’ 10:01:00 - MongoDB starts
β€’ 10:01:05 - Loads checkpoint + replays 100 writes from journal = Full recovery! πŸŽ‰

βš–οΈ WiredTiger vs MMAPv1

Feature πŸ’Ύ WiredTiger (Modern) πŸ“š MMAPv1 (Legacy)
Locking Document-level Database-level
Compression βœ… Yes (70-80%) ❌ None
Concurrency High (thousands/sec) Low (limited)
Memory Management WiredTiger Cache (controlled) OS Memory Mapped (unpredictable)
Write Performance ⚑ 10x faster Slow
Storage Efficiency 70% less space Full size
Journaling Every 50ms Every 100ms
Default Since MongoDB 3.2 (2015) Deprecated (removed in 4.2)

⚠️ Important Note

MMAPv1 was completely removed in MongoDB 4.2 (2019). All modern MongoDB deployments use WiredTiger.

❓ Interview Questions & Answers

Q1 What is WiredTiger and why is it important for MongoDB? β–Ό

Answer:

WiredTiger is MongoDB's default storage engine that manages how data is stored on disk and in memory. It's been the default since MongoDB 3.2 (2015).

Why It's Important:

  • Document-Level Concurrency: Multiple operations can modify different documents simultaneously without blocking
  • Compression: Reduces storage by 70-80% using snappy/zlib/zstd algorithms
  • High Performance: In-memory cache (50% RAM) for frequently accessed data
  • Crash Safety: Write-ahead logging (journal) ensures durability
  • Better Concurrency: 10x better than old MMAPv1 engine

Real Impact: A MongoDB cluster with WiredTiger can handle thousands of concurrent writes while using 70% less disk space than without compression.

Q2 Explain the WiredTiger cache. How does it work? β–Ό

Answer:

WiredTiger cache is an in-memory cache that stores frequently accessed data for fast reads and writes.

How It Works:

  • Cache Size: By default, max(50% of RAM - 1GB, 256MB)
  • Example: Server with 16GB RAM β†’ Cache = 7GB
  • What's Cached: Documents, indexes, internal metadata
  • Read Flow: Check cache first β†’ if not found, read from disk β†’ cache it
  • Write Flow: Write to cache β†’ mark as "dirty" β†’ flush to disk during checkpoint
  • Eviction Policy: LRU (Least Recently Used) when cache reaches 80% full

Configuration Example:

storage:
  wiredTiger:
    engineConfig:
      cacheSizeGB: 8  # Set manually to 8GB

Performance Impact: Cache hit = microseconds, Disk read = milliseconds. Good cache hit ratio (90%+) is crucial for performance!

Q3 What is a checkpoint in WiredTiger? How does it ensure durability? β–Ό

Answer:

A checkpoint is a snapshot of data written from WiredTiger cache to disk files (.wt), creating a consistent state at a specific point in time.

Checkpoint Details:

  • Frequency: Every 60 seconds by default (configurable)
  • What Happens: Dirty pages (modified data in cache) flushed to .wt files
  • Consistency: Represents all committed writes up to that point
  • Non-Blocking: System remains available during checkpoint

How It Ensures Durability:

  1. Normal Operation: Writes go to cache + journal
  2. Checkpoint: Every 60s, cache data persisted to disk
  3. On Crash: Load last checkpoint + replay journal
  4. Result: All committed writes recovered!

Example Timeline:

10:00:00 - Checkpoint #1 created
10:00:30 - 100 writes happen (in cache + journal)
10:00:45 - Server crashes πŸ’₯
10:01:00 - Server restarts
10:01:05 - Loads checkpoint #1 + replays 100 writes = Full recovery! βœ…
Q4 What compression algorithms does WiredTiger support? When to use each? β–Ό

Answer:

WiredTiger supports 3 compression algorithms plus an option for no compression.

Compression Algorithms:

  • Snappy (Default):
    • Fastest compression/decompression
    • 70% storage reduction
    • Low CPU usage
    • Use When: Production workloads, general purpose, balanced performance
  • Zlib:
    • Highest compression ratio
    • 80% storage reduction
    • High CPU usage (slower)
    • Use When: Storage is critical (archives, large datasets), reads > writes
  • Zstd:
    • Best balance of speed and compression
    • 75% storage reduction
    • Medium CPU usage
    • Use When: Need better compression than snappy without zlib's slowness
  • None:
    • No compression overhead
    • Fastest possible
    • Use When: Testing, RAM-heavy workloads, data already compressed

Real Impact: 1TB uncompressed data:

Snappy: 300GB (70% savings)
Zlib:   200GB (80% savings)
Zstd:   250GB (75% savings)
Q5 Explain document-level concurrency control in WiredTiger. β–Ό

Answer:

Document-level concurrency means multiple clients can read and write different documents simultaneously without blocking each other.

How It Works:

  • Technology: Multi-Version Concurrency Control (MVCC)
  • Mechanism: Each transaction sees a consistent snapshot of data
  • Locking: Only the specific document being modified is locked
  • Isolation: Conflicts detected at commit time, not during operation

Example Scenario:

Thread A: Update user_id=1 (locks only doc 1)
Thread B: Update user_id=2 (locks only doc 2)
Thread C: Read user_id=1   (reads snapshot, no lock)

All three operations happen simultaneously! ⚑

Comparison with MMAPv1 (old engine):

  • MMAPv1: Database-level locking - one write blocks all others
  • WiredTiger: Document-level - thousands of concurrent writes

Performance Impact: WiredTiger can handle 10,000+ concurrent writes while MMAPv1 maxed out at ~1,000 due to lock contention.

Q6 What is the journal in WiredTiger and how does it ensure crash recovery? β–Ό

Answer:

The journal is a Write-Ahead Log (WAL) that records all write operations before they're applied to data files, ensuring durability and crash recovery.

How Journal Works:

  1. Write arrives: Operation first written to WiredTiger cache
  2. Journal log: Operation recorded to journal file on disk
  3. Flush frequency: Journal committed every 50ms or 100MB of data
  4. Acknowledgment: Write acknowledged to client only after journal commit

Crash Recovery Process:

1. Server crashes at 10:30:45
2. Last checkpoint: 10:30:00 (60 seconds ago)
3. On restart:
   - Load checkpoint (consistent state at 10:30:00)
   - Replay journal entries (10:30:00 to 10:30:45)
   - Recover all committed writes! βœ…

Journal Files:

  • Location: dbPath/journal/ directory
  • Format: WiredTigerLog.0000000001
  • Retention: Kept for 2 checkpoints, then deleted
  • Size: ~100MB per file

Configuration:

storage:
  journal:
    enabled: true
    commitIntervalMs: 50  # Default

Trade-off: Journal adds ~10% write latency but ensures zero data loss on crash. Worth it for durability! πŸ›‘οΈ