Section 7: Advanced Topics

πŸ’Ύ Backup, Restore & Disaster Recovery

Protecting Your Data: Complete Guide to MongoDB Backup Strategies, Recovery Procedures, and Disaster Recovery Planning

πŸ“– The $2 Million Data Loss Incident

Meet David, CTO of a FinTech startup processing millions of transactions daily. At 3 AM on a Tuesday, a junior developer accidentally ran a production script that dropped the main database. 2.8 million records vanished instantly.

The Problem: Their last backup was 3 weeks old and corrupted. No one had tested restores. They spent $2M on recovery, compensation, and legal fees. 30% of customers left.

The Lesson: "Your backup strategy is only as good as your last successful restore test."

Why Backup & Disaster Recovery Matter

🚨 Data Loss Statistics

93% of companies that lose data for 10+ days file for bankruptcy within one year.

  • Human Error (40%): Accidental deletion, wrong commands
  • Hardware Failure (30%): Disk crashes, server failures
  • Software Issues (15%): Database corruption, bugs
  • Malicious Activity (10%): Ransomware, sabotage
  • Natural Disasters (5%): Floods, fires, power outages

MongoDB Backup Strategies

πŸ“¦

1. mongodump

What: Logical backup in BSON format

Best for: Databases <100GB

Pros: Easy, portable, selective

Cons: Slower for large DBs

πŸ“Έ

2. Filesystem Snapshots

What: Storage-level snapshots

Best for: Large production DBs

Pros: Very fast, minimal impact

Cons: Requires LVM/EBS

☁️

3. Atlas Cloud Backup

What: Fully managed by MongoDB

Best for: Production apps

Pros: Automated, PITR, reliable

Cons: Atlas-only, paid

Using mongodump

# Backup entire database
mongodump --uri="mongodb://localhost:27017" --out=/backup/$(date +%Y%m%d)

# Backup with compression
mongodump --uri="mongodb://localhost:27017" --gzip --out=/backup/compressed

# Backup specific database
mongodump --db=myDatabase --out=/backup/mydb

# Production backup script
#!/bin/bash
BACKUP_DIR="/backups/mongodb"
DATE=$(date +%Y%m%d_%H%M%S)
mongodump --uri="mongodb://user:pass@localhost:27017" --gzip --out=$BACKUP_DIR/$DATE

# Automated with cron
0 2 * * * /scripts/mongodb_backup.sh

βœ… mongodump Best Practices

  • Always use --gzip for compression
  • Schedule during off-peak hours
  • Store on separate storage
  • Test restores monthly
  • Encrypt sensitive backups
  • Use --oplog for point-in-time consistency

Restore Operations

# Basic restore
mongorestore --uri="mongodb://localhost:27017" /backup/20241213

# Restore specific database
mongorestore --db=myDatabase /backup/20241213/myDatabase

# Restore with gzip
mongorestore --gzip /backup/20241213

# Restore to different DB name
mongorestore --nsFrom="prod.*" --nsTo="restore.*" /backup/20241213
πŸ’‘ Restore Procedure

1. Verify backup integrity
2. Stop application traffic
3. Perform restore
4. Rebuild indexes if needed
5. Verify data integrity
6. Test application connectivity
7. Resume traffic gradually

Disaster Recovery Planning

πŸ’‘ RTO vs RPO

RTO (Recovery Time Objective): Maximum acceptable downtime
Example: "We can tolerate 2 hours downtime"

RPO (Recovery Point Objective): Maximum acceptable data loss
Example: "We can lose max 1 hour of data"

Cold Standby

RTO: 24-72 hours

RPO: 6-24 hours

Cost: Low

Warm Standby

RTO: 1-4 hours

RPO: 1-6 hours

Cost: Medium

Hot Standby

RTO: 5-30 minutes

RPO: <1 hour

Cost: High

Best Practices

βœ… 3-2-1 Backup Rule

  • 3 copies of your data
  • 2 different storage types
  • 1 copy off-site

⚠️ Common Mistakes to Avoid

  • Not testing restores regularly
  • Storing backups on same server
  • No backup verification
  • Single backup location
  • Manual processes
  • No encryption

πŸ’Ό Interview Questions

Q1 What backup strategies exist for MongoDB? β–Ό

Answer:

1. mongodump: Logical backup for <100GB DBs

2. Filesystem Snapshots: Fast, for large production DBs

3. Atlas Cloud: Managed with PITR

4. Delayed Replica: Protection from logical errors

Selection depends on: DB size, RTO/RPO, budget, infrastructure

Q2 Explain RTO vs RPO β–Ό

RTO: Maximum downtime tolerated (e.g., 2 hours)

RPO: Maximum data loss tolerated (e.g., 1 hour)

Example: Trading platform needs RTO <1min, RPO <1sec (expensive). Blog can have RTO 24hr, RPO 12hr (cheap).

Q3 What is the 3-2-1 backup rule? β–Ό

3 copies: 1 production + 2 backups

2 media types: Local disk + Cloud (S3)

1 off-site: Different region/datacenter

Protects against all failure scenarios: disk, server, disaster

Q4 How ensure backup consistency in replica set? β–Ό

1. Backup from secondary (no primary impact)

2. Use db.fsyncLock() before snapshot

3. Include oplog (mongodump --oplog)

4. Use filesystem snapshots (atomic)

5. db.fsyncUnlock() after completion

Q5 How test backups are restorable? β–Ό

1. Monthly restore tests: Pick random backup, restore to test env

2. Automated verification: Check file integrity, sizes

3. Quarterly DR drills: Full team, production-sized restore

4. Continuous monitoring: Alert on anomalies

Document results: time, issues, improvements