Section 3: MongoDB Architecture

πŸ”„ Read/Write Path: Data Journey Through MongoDB

Understanding how data flows from your application through MongoDB's layers to disk and back

πŸ“– The Restaurant Kitchen Story

Imagine MongoDB as a restaurant kitchen. Let's see how orders (data) flow...

✍️

Write Path: Taking an Order

Customer places order β†’ Kitchen processes it

The Journey:

  1. Customer orders (Application sends write request)
  2. Waiter writes it on notepad (WiredTiger Cache - RAM)
  3. Kitchen log book updated (Journal - for safety)
  4. Every minute, notepad β†’ recipe book (Checkpoint to disk)
  5. Receipt given to customer (Acknowledgment sent)

Result: Order recorded instantly (1ms), fully saved every 60s!

πŸ“–

Read Path: Serving Food

Customer wants to know order status

The Journey:

  1. Customer asks about order (Application sends read request)
  2. Check waiter's notepad first (WiredTiger Cache - super fast!)
  3. If not on notepad β†’ check recipe book (Read from disk - slower)
  4. Copy to notepad for next time (Cache it)
  5. Tell customer (Return data)

Result: Notepad (cache) = instant (100ΞΌs), Recipe book (disk) = slower (10ms)

πŸ’‘ The Key Insight

Cache (Notepad) = Fast but temporary
Disk (Recipe Book) = Slower but permanent
MongoDB balances speed and durability! 🎯

✍️ Write Path: Complete Flow

Application db.users.insert() MongoDB Driver Connection pooling mongod Process Query parser WiredTiger Cache In-Memory (RAM) ⚑ 1ms write time Journal (WAL) Write-Ahead Log Sync every 50ms Data Files (.wt) Checkpoint every 60s 0ms ~0.5ms ~1ms 50ms 60 seconds

πŸ“ Write Path Timeline:

  • 0ms: Application issues insert/update command
  • ~0.5ms: MongoDB driver sends request to mongod
  • ~1ms: mongod writes to WiredTiger Cache (RAM) - Write acknowledged here! ⚑
  • 50ms: Journal synced to disk (crash safety guaranteed)
  • 60s: Checkpoint flushes cache to .wt data files (durable on disk)

πŸ“– Read Path: Complete Flow

Application db.users.findOne() mongod Process Query router WiredTiger Cache Check RAM first βœ“ βœ— CACHE HIT (~100ΞΌs) CACHE MISS Data Files (.wt) Read from disk 🐒 10-50ms Load to cache, then return ⚑ Read Performance Comparison βœ“ Cache Hit: ~100 microseconds (0.1ms) βœ— Cache Miss: ~10-50 milliseconds (100-500x slower)

πŸ“– Read Path Logic:

  • Step 1: Application sends find/findOne query
  • Step 2: mongod checks WiredTiger Cache first (RAM)
  • Step 3a (Cache Hit): Data found β†’ Return immediately (~100ΞΌs) ⚑
  • Step 3b (Cache Miss): Not in cache β†’ Read from .wt files on disk (~10-50ms) 🐒
  • Step 4: If from disk β†’ Load into cache for future reads
  • Goal: Keep cache hit ratio >90% for best performance!

βš™οΈ Write Concerns: Durability Levels

Write concern controls when MongoDB acknowledges a write -
Trade-off between speed vs durability

w: 0 (Unacknowledged)

Behavior: Fire-and-forget, no acknowledgment

Speed: ⚑ Fastest (~0.1ms)

Durability: ⚠️ No guarantee

Use case: Logs, analytics where loss is acceptable

w: 1 (Default - Acknowledged)

Behavior: Acknowledged when written to primary's cache

Speed: ⚑ Fast (~1ms)

Durability: βœ… Good (journal ensures recovery)

Use case: Most applications, balanced approach

w: "majority" (Replica Set)

Behavior: Acknowledged when written to majority of nodes

Speed: 🐒 Slower (~10-50ms)

Durability: βœ… Excellent (survives node failures)

Use case: Critical data (banking, healthcare)

j: true (Journaled)

Behavior: Acknowledged when written to journal on disk

Speed: 🐒 Slower (~50ms)

Durability: βœ… Excellent (crash-safe)

Use case: Single node with high durability needs

πŸ’‘ Configuration Example:

// Fast, default
db.users.insertOne(doc, { writeConcern: { w: 1 } })

// Highly durable, slower
db.transactions.insertOne(doc, { 
  writeConcern: { w: "majority", j: true } 
})

πŸ“š Read Preferences: Where to Read

Read preference determines which replica set member handles read operations -
Trade-off between consistency vs performance/availability

primary (Default)

Behavior: All reads from primary node

Consistency: βœ… Strongest (always latest data)

Performance: 🐒 Primary can be bottleneck

Use case: Default for most applications

primaryPreferred

Behavior: Primary if available, else secondary

Consistency: βœ… Good (slight lag on failover)

Performance: ⚑ Better availability

Use case: High availability with consistency preference

secondary

Behavior: All reads from secondary nodes

Consistency: ⚠️ Eventual (may be stale)

Performance: ⚑ Distributes read load

Use case: Analytics, reporting, search (stale OK)

secondaryPreferred

Behavior: Secondary if available, else primary

Consistency: ⚠️ Eventual (may be stale)

Performance: ⚑ Best read distribution

Use case: Read-heavy apps, analytics dashboards

nearest

Behavior: Read from closest node (lowest latency)

Consistency: ⚠️ Eventual (may be stale)

Performance: ⚑ Lowest latency

Use case: Geo-distributed apps, multi-region

πŸ’‘ Configuration Example:

// Read from primary (default)
db.users.find().readPref("primary")

// Read from secondaries (analytics)
db.logs.find().readPref("secondary")

// Read from nearest (geo-distributed)
db.products.find().readPref("nearest")

πŸš€ Performance Optimizations

πŸ“‡ 1. Indexes

Without index: Collection scan (slow) 🐒
With index: Direct lookup (fast) ⚑

// Create index on email field
db.users.createIndex({ email: 1 })

// Query now uses index
db.users.find({ email: "john@example.com" })
// ~0.1ms instead of 100ms!

πŸ”Œ 2. Connection Pooling

Reuse connections instead of creating new ones for each request

// Connection pool configuration
const client = new MongoClient(uri, {
  maxPoolSize: 100,  // Max connections
  minPoolSize: 10    // Min connections
});

🎯 3. Projection (Select Only Needed Fields)

Don't fetch entire document if you only need specific fields

// Bad: Fetch entire document
db.users.find({ age: { $gt: 25 } })

// Good: Fetch only name and email
db.users.find(
  { age: { $gt: 25 } },
  { name: 1, email: 1, _id: 0 }
)

πŸ“¦ 4. Batch Operations

Insert multiple documents in one call instead of many individual calls

// Bad: 1000 individual inserts
for (let i = 0; i < 1000; i++) {
  db.users.insertOne(docs[i])  // 1000ms
}

// Good: One bulk insert
db.users.insertMany(docs)  // 10ms!

❓ Interview Questions & Answers

Q1 Explain the complete write path in MongoDB from application to disk. β–Ό

Answer:

Write Path Flow:

  1. Application (0ms): Issues insert/update command via MongoDB driver
  2. MongoDB Driver (~0.5ms): Sends request to mongod process over network
  3. mongod Process (~1ms): Parses query and routes to storage engine
  4. WiredTiger Cache (~1ms): Writes to in-memory cache (RAM) - Write acknowledged here!
  5. Journal (~50ms): Operation logged to write-ahead log (WAL) on disk for crash recovery
  6. Checkpoint (~60 seconds): Dirty pages flushed from cache to .wt data files on disk

Key Points:

  • Write is acknowledged after step 4 (~1ms) for speed
  • Journal ensures durability - on crash, replay journal to recover uncommitted writes
  • Checkpoint creates consistent snapshot every 60 seconds
  • Trade-off: Speed (cache) vs Durability (journal + checkpoints)
Q2 What is the difference between a cache hit and cache miss in the read path? β–Ό

Answer:

Cache Hit (Fast ⚑):

  • Data found in WiredTiger Cache (RAM)
  • Return immediately without disk access
  • Latency: ~100 microseconds (0.1ms)
  • Example: Reading a user profile that was recently accessed

Cache Miss (Slower 🐒):

  • Data NOT in cache β†’ must read from disk (.wt files)
  • After reading from disk, load into cache for future reads
  • Latency: ~10-50 milliseconds (100-500x slower than cache hit)
  • Example: Reading a document that hasn't been accessed in hours

Why Cache Hit Ratio Matters:

  • Goal: Keep cache hit ratio >90%
  • 90% hit rate = 10x faster average reads than 50% hit rate
  • Increase cache size (more RAM) to improve hit ratio
Q3 What is write concern? Explain w:1 vs w:"majority". β–Ό

Answer:

Write concern controls when MongoDB acknowledges a write operation. It's a trade-off between speed and durability.

w: 1 (Default - Acknowledged):

  • Behavior: Acknowledged when written to primary node's cache
  • Speed: Fast (~1ms)
  • Durability: Good - journal ensures recovery on crash
  • Risk: If primary fails before replicating, write could be lost
  • Use case: Most applications, balanced approach

w: "majority" (Replica Set):

  • Behavior: Acknowledged when written to majority of nodes (e.g., 2 out of 3)
  • Speed: Slower (~10-50ms depending on network)
  • Durability: Excellent - survives single node failures
  • Risk: Very low - write is durable even if primary fails
  • Use case: Critical data (banking, healthcare, financial transactions)

Example Configuration:

// Fast, default
db.users.insertOne(doc, { writeConcern: { w: 1 } })

// Highly durable, slower
db.transactions.insertOne(doc, { 
  writeConcern: { w: "majority", j: true } 
})
Q4 Explain read preferences and when to use "secondary" vs "primary". β–Ό

Answer:

Read preference determines which replica set member handles read operations. Trade-off between consistency and performance.

primary (Default):

  • Behavior: All reads from primary node
  • Consistency: Strongest - always latest data
  • Performance: Primary can become bottleneck
  • Use case: When you need guaranteed latest data (user profile updates, inventory)

secondary:

  • Behavior: All reads from secondary nodes
  • Consistency: Eventual - may be stale (lag 0-few seconds)
  • Performance: Distributes read load across secondaries
  • Use case: Analytics, reporting, search, dashboards (where stale data is acceptable)

When to use secondary:

  • Read-heavy workloads that don't need latest data
  • Reporting and analytics queries
  • Search functionality
  • Background batch processing

When to use primary:

  • Transactional data requiring latest state
  • Read-your-own-writes scenarios
  • Critical operations (money transfers, inventory checks)

Example:

// Primary for critical reads
db.inventory.findOne({ sku: "ABC123" }).readPref("primary")

// Secondary for analytics
db.orders.aggregate([...]).readPref("secondary")
Q5 How does MongoDB ensure data durability in the write path? β–Ό

Answer:

MongoDB uses multiple mechanisms to ensure data durability:

1. Journal (Write-Ahead Log):

  • All write operations logged to journal before applying
  • Journal synced to disk every 50ms (configurable)
  • On crash: Replay journal entries to recover uncommitted writes
  • Location: dbPath/journal/ directory

2. Checkpoints:

  • Every 60 seconds, flush WiredTiger cache to disk
  • Creates consistent snapshot of all data
  • After checkpoint, journal entries before that point can be deleted

3. Replica Sets (w: "majority"):

  • Write acknowledged only after majority of nodes have it
  • Survives single node failures
  • Example: 3-node cluster needs 2 nodes to acknowledge

Crash Recovery Process:

1. MongoDB crashes at 10:30:45
2. Last checkpoint: 10:30:00
3. On restart:
   - Load checkpoint (consistent state at 10:30:00)
   - Replay journal (10:30:00 to 10:30:45)
   - All committed writes recovered! βœ…

Trade-offs:

  • More durability = Higher latency (journal sync, replication)
  • Less durability = Lower latency (cache-only writes)
  • Configure via writeConcern based on use case
Q6 What performance optimizations can you apply to MongoDB reads and writes? β–Ό

Answer:

Read Optimizations:

  • Indexes: Create indexes on frequently queried fields
    db.users.createIndex({ email: 1 })  // 100x faster lookups
  • Projections: Fetch only needed fields
    db.users.find({}, { name: 1, email: 1 })  // Less data transferred
  • Covered Queries: Query + projection covered entirely by index (no document fetch)
    db.users.createIndex({ email: 1, name: 1 })
    db.users.find({ email: "x" }, { name: 1, _id: 0 })  // Covered!
  • Increase Cache Size: More RAM = higher cache hit ratio
    wiredTiger.cacheSizeGB: 16  // Larger cache
  • Read from Secondaries: Distribute load for analytics/reporting
    db.logs.find().readPref("secondary")

Write Optimizations:

  • Batch Operations: Insert many docs at once
    db.users.insertMany(docs)  // 100x faster than individual inserts
  • Lower Write Concern: For non-critical writes
    db.logs.insertOne(doc, { w: 1 })  // Faster than w: "majority"
  • Reduce Index Count: Each index slows writes
    // Remove unused indexes
    db.users.dropIndex("oldIndex")
  • Connection Pooling: Reuse connections
    const client = new MongoClient(uri, { maxPoolSize: 100 })

General Optimizations:

  • Shard for Scale: Distribute data across servers
  • Monitor Cache Hit Ratio: Keep >90%
  • Use Profiler: Identify slow queries