Section 7: Advanced Topics

πŸ“ˆ MongoDB Observability

Master monitoring, metrics, logs, and performance debugging for production MongoDB

🚨 The 3 AM Production Crisis

It's 3 AM. Your phone rings. "The app is down! Users can't login!" You check MongoDB - it's running. CPU looks fine. But queries are taking 30 seconds instead of 30ms.

The Problem: No monitoring = blind debugging for 2 hours!

With Observability: Check dashboard β†’ See slow query spike β†’ Find missing index β†’ Fix in 5 minutes! ✨

This is why observability matters: You can't fix what you can't see!

πŸ” What is Observability?

Observability = The ability to understand your system's internal state by examining its external outputs.

The Three Pillars of Observability:

  1. Metrics: Numeric measurements over time (CPU, memory, query latency)
  2. Logs: Timestamped records of events (errors, warnings, operations)
  3. Traces: Request flow through distributed systems

πŸ“Š Key MongoDB Metrics to Monitor

1. Performance Metrics

// Check current operations
db.currentOp()

// Get server status
db.serverStatus()

// Key metrics:
{
  opcounters: { query: 1500, insert: 500, update: 300 },
  connections: { current: 52, available: 948 },
  globalLock: { currentQueue: { total: 0, readers: 0, writers: 0 } }
}

What to watch: Operations per second, connection count, lock queue depth

2. Resource Utilization

// Memory usage
db.serverStatus().mem
// Returns: { resident: 1024, virtual: 2048, mapped: 512 }

// Disk I/O
db.serverStatus().wiredTiger.cache
// Cache hit ratio should be > 90%

Critical thresholds: Memory > 80%, Disk I/O wait > 20%, Cache miss rate > 10%

3. Replication Lag

// Check replica set status
rs.status()

// Replication lag (in seconds)
rs.printReplicationInfo()

// Ideal: < 1 second
// Warning: > 10 seconds
// Critical: > 60 seconds

πŸ› οΈ Monitoring Tools

1. MongoDB Atlas (Cloud Native)

// Built-in monitoring with:
- Real-time performance charts
- Automated alerts
- Query profiler
- Index recommendations
- Custom dashboards

2. Prometheus + Grafana (Open Source)

# Install MongoDB exporter
docker run -d -p 9216:9216 \
  percona/mongodb_exporter:0.40 \
  --mongodb.uri=mongodb://localhost:27017

# Prometheus config
scrape_configs:
  - job_name: 'mongodb'
    static_configs:
      - targets: ['localhost:9216']

3. mongostat (CLI Tool)

$ mongostat --host localhost:27017 -n 100 1

insert query update delete getmore command dirty  used flushes vsize  res qrw arw
   500   150    300     50     100     850  3.2% 64.1%       0 2.5G 1.8G 0|0 1|0

πŸ“ Log Analysis

Enable Slow Query Logging

// Log queries slower than 100ms
db.setProfilingLevel(1, { slowms: 100 })

// Check profiler data
db.system.profile.find().sort({ ts: -1 }).limit(10)

// Example slow query log:
{
  "op": "query",
  "ns": "mydb.users",
  "millis": 1500,
  "planSummary": "COLLSCAN",  // ⚠️ Collection scan!
  "execStats": { "nReturned": 10, "totalDocsExamined": 1000000 }
}

Log Levels

// Set log verbosity
db.setLogLevel(1, "query")  // 0-5 (0=default, 5=debug)

// Important log components:
- query: Query execution
- write: Write operations  
- command: Command execution
- network: Network activity
- replication: Replication events
⚠ Warning:

High log verbosity (4-5) can impact performance. Use in development only!

πŸ”¬ Query Profiling

Profiling Levels

// Level 0: Profiler off
db.setProfilingLevel(0)

// Level 1: Log slow queries only
db.setProfilingLevel(1, { slowms: 100 })

// Level 2: Log ALL queries (dev/debug only!)
db.setProfilingLevel(2)

// Check current level
db.getProfilingStatus()
// { was: 1, slowms: 100, sampleRate: 1.0 }

Analyze Slow Queries

// Find slowest queries
db.system.profile.find({ millis: { $gt: 1000 } })
  .sort({ millis: -1 })
  .limit(5)

// Find COLLSCAN queries (missing indexes)
db.system.profile.find({ 
  "planSummary": /COLLSCAN/ 
}).limit(10)

// Find queries examining too many docs
db.system.profile.find({ 
  "execStats.totalDocsExamined": { $gt: 10000 } 
})

🚨 Alerting Strategy

Critical Alerts (Immediate Action)

  • πŸ”΄ Primary Node Down: Replica set has no primary
  • πŸ”΄ Disk Space < 10%: Database will stop accepting writes
  • πŸ”΄ Replication Lag > 60s: Data inconsistency risk
  • πŸ”΄ Connection Pool Exhausted: App can't connect

Warning Alerts (Investigate Soon)

  • 🟑 CPU > 80%: May need scaling
  • 🟑 Slow Query Spike: Missing indexes?
  • 🟑 Cache Hit Rate < 90%: Need more memory
  • 🟑 Disk I/O Wait > 20%: Disk bottleneck

Sample Alert Rule (Prometheus)

groups:
  - name: mongodb_alerts
    rules:
      - alert: MongoDBDown
        expr: mongodb_up == 0
        for: 1m
        labels:
          severity: critical
        annotations:
          summary: "MongoDB instance is down"
          
      - alert: HighReplicationLag
        expr: mongodb_replset_member_replication_lag > 60
        for: 5m
        labels:
          severity: warning
        annotations:
          summary: "Replication lag exceeds 60 seconds"

⭐ Observability Best Practices

  1. Monitor What Matters: Focus on user-impacting metrics (latency, errors, saturation)
  2. Set Meaningful Alerts: Avoid alert fatigue - only alert on actionable issues
  3. Use Dashboards: Create role-specific dashboards (DevOps, DBAs, Developers)
  4. Enable Profiling in Production: Level 1 with slowms=100 is safe
  5. Retain Logs: Keep 30+ days for trend analysis
  6. Monitor Replica Sets: Track lag, elections, failovers
  7. Track Index Usage: Find unused indexes to remove
  8. Baseline Your Metrics: Know what "normal" looks like
  9. Test Alerts: Simulate failures to verify alert delivery
  10. Document Runbooks: What to do when each alert fires

πŸ’Ό Interview Questions & Answers

Q1 What are the three pillars of observability? β–Ό

1. Metrics: Quantitative measurements (CPU, latency, throughput)

2. Logs: Timestamped event records (errors, warnings, operations)

3. Traces: Request paths through distributed systems

Why all three? Metrics show "what", logs show "why", traces show "where"

Q2 How do you identify slow queries in production? β–Ό

Method 1: Database Profiler

db.setProfilingLevel(1, { slowms: 100 })
db.system.profile.find({ millis: { $gt: 1000 } }).sort({ millis: -1 })

Method 2: Monitor Logs

grep "slow query" /var/log/mongodb/mongod.log

Method 3: currentOp() for Active Queries

db.currentOp({ "secs_running": { $gt: 5 } })
Q3 What metrics indicate MongoDB needs more memory? β–Ό

Key indicators:

  • Low Cache Hit Ratio: < 90% means frequent disk reads
  • High Page Faults: Memory swapping to disk
  • Disk I/O Wait: > 20% indicates memory pressure
  • WiredTiger Cache Full: Cache evicting data too frequently
// Check cache statistics
db.serverStatus().wiredTiger.cache

// Healthy: bytes currently in cache close to maximum bytes configured
Q4 How do you monitor replication lag? β–Ό

Method 1: rs.status()

rs.status().members.forEach(m => {
  if (m.state === 2) {  // SECONDARY
    print(m.name + ": " + m.optimeDate)
  }
})

Method 2: rs.printReplicationInfo()

rs.printReplicationInfo()
// Shows: configured oplog size, log length start to end, oplog first/last event times

Acceptable lag: < 1 second (normal), < 10 seconds (acceptable), > 60 seconds (critical)

Q5 What's the difference between monitoring and observability? β–Ό

Monitoring: Watching predefined metrics for known failure modes

  • Example: Alert when CPU > 80%
  • Limitation: Only catches expected problems

Observability: Ability to understand ANY system state, even unknown failures

  • Example: Correlation of metrics, logs, traces to debug new issues
  • Benefit: Can investigate unexpected problems

In practice: You need both - monitoring for known issues, observability for unknown ones

Q6 How do you prevent alert fatigue? β–Ό

Best practices:

  • Alert on symptoms, not causes: Alert on "users can't login" not "CPU is high"
  • Set proper thresholds: Use baselines from historical data
  • Use alert aggregation: Group related alerts
  • Add context: Include runbook links in alerts
  • Regular review: Disable noisy alerts that don't lead to action
  • Escalation policies: Warning β†’ Critical path with delays

Rule of thumb: If an alert doesn't require immediate action, it's not an alert - make it a dashboard metric instead