Section 3: MongoDB Architecture

πŸ” Replica Sets

Master High Availability & Data Replication - Learn how MongoDB ensures your data is always available through interactive examples and real-world scenarios

πŸ“– Meet Sarah - The E-commerce Developer

πŸ‘©β€πŸ’» Sarah's Challenge

Sarah runs a successful e-commerce platform with 10,000 daily orders. One Friday night at 9 PM (peak shopping time), her single database server crashed due to a hardware failure. πŸ’₯

❌ The Disaster:
  • Website went completely down
  • Lost $50,000 in sales during 2-hour outage
  • 1,000+ angry customers on social media
  • Took 2 hours to restore from backup
  • Lost trust of premium customers

Sarah knew she needed a solution that would keep her platform running even if servers failed. That's when she discovered MongoDB Replica Sets.

βœ… The Solution: Replica Sets
  • Setup 3 MongoDB servers (1 Primary + 2 Secondary)
  • Automatic data replication to all servers
  • When primary fails, secondary automatically promoted
  • Failover time: Less than 10 seconds!
  • Zero data loss, zero downtime

πŸŽ‰ Result: 3 months later, her server crashed again - but customers didn't even notice! The replica set automatically switched to a healthy server in 8 seconds.

πŸ€” What is a Replica Set?

πŸ“š Simple Definition

A Replica Set is a group of MongoDB servers that maintain the same data. Think of it like having multiple copies of your notebook - if one gets lost, you still have the others!

πŸ”„
3+
Minimum Nodes
⚑
<10s
Failover Time
🎯
99.99%
Availability
πŸ“Š
50
Max Members

🎯 Key Benefits

βœ“ High Availability

Automatic failover in seconds

βœ“ Data Redundancy

Multiple copies prevent data loss

βœ“ Read Scalability

Distribute reads across secondaries

πŸ—οΈ Replica Set Components

Live Replica Set Architecture
πŸ‘‘
PRIMARY
Read + Write
πŸ“‹
SECONDARY
Read Only
πŸ“‹
SECONDARY
Read Only
βš–οΈ
ARBITER
Voting Only

πŸ‘‘ Primary Node

Receives all write operations. Only one primary at a time. Replicates data to secondaries.

πŸ“‹ Secondary Node

Maintains copy of data. Can serve read operations. Can become primary during failover.

βš–οΈ Arbiter Node

Lightweight node. Participates in elections. Does not hold data. Used for tie-breaking votes.

βš™οΈ How Replication Works

1
Client Writes to Primary

All write operations (insert, update, delete) must go through the PRIMARY node.

db.orders.insertOne({
  orderId: "ORD-12345",
  customer: "John Doe",
  total: 599.99
})

πŸ“ This write is recorded in the primary's oplog (operations log).

2
Oplog Replication

Secondary nodes continuously pull operations from the primary's oplog and apply them.

πŸ“š What is Oplog?

Oplog (operations log) is a special capped collection that stores all write operations in order. It's like a journal of everything that happened to your data.

3
Secondaries Apply Changes

Each secondary independently applies the operations from the oplog to maintain an identical copy.

// Automatic replication - no code needed!
// Secondary nodes sync continuously
Replication Lag: ~100-500ms (typical)
4
Heartbeat & Health Check

Every 2 seconds, members send heartbeats to each other to check if everyone is healthy.

πŸ’“
2s
Heartbeat Interval
⏱️
10s
Timeout

🚨 Automatic Failover Process

⚑ Lightning-Fast Recovery

When the primary fails, MongoDB automatically elects a new primary in typically 10-12 seconds. Your application reconnects automatically!

πŸ’₯

Failover Scenario Simulation

Step 1: Primary Crashes

At 2:00 PM, the primary server experiences a hardware failure and goes offline.

Step 2: Detection (2s)

Secondaries notice missed heartbeats after 2 seconds.

Step 3: Election (5-8s)

Remaining members vote for a new primary. The most up-to-date secondary wins.

Step 4: New Primary (10s total)

A secondary is elected as new primary and starts accepting writes. Your app reconnects automatically!

// Timeline Visualization
00:00 - Primary Online     [πŸ‘‘ βœ“] [πŸ“‹ βœ“] [πŸ“‹ βœ“]
00:02 - Primary Crashes    [πŸ‘‘ βœ—] [πŸ“‹ βœ“] [πŸ“‹ βœ“]
00:04 - Detection          [πŸ‘‘ βœ—] [πŸ“‹ ?] [πŸ“‹ ?]
00:10 - New Primary!       [πŸ“‹ βœ“] [πŸ‘‘ βœ“] [πŸ“‹ βœ“]

Total Downtime: ~10 seconds
Data Loss: ZERO βœ…

πŸ› οΈ Setup Your First Replica Set

πŸ’‘ Beginner-Friendly Tutorial

Follow these steps to create a 3-node replica set on your local machine. We'll simulate production environment!

Step 1: Start MongoDB Instances

Terminal - Start 3 MongoDB Processes
# Terminal 1 - Start node on port 27017
mongod --replSet myReplicaSet --port 27017 --dbpath /data/db1

# Terminal 2 - Start node on port 27018
mongod --replSet myReplicaSet --port 27018 --dbpath /data/db2

# Terminal 3 - Start node on port 27019
mongod --replSet myReplicaSet --port 27019 --dbpath /data/db3

Step 2: Initialize Replica Set

mongosh - Initialize Configuration
# Connect to any node
mongosh --port 27017

# Initialize replica set
rs.initiate({
  _id: "myReplicaSet",
  members: [
    { _id: 0, host: "localhost:27017" },
    { _id: 1, host: "localhost:27018" },
    { _id: 2, host: "localhost:27019" }
  ]
})

Step 3: Verify Status

πŸ”₯ Try it Live - Check Replica Set Status
πŸ“Š Output:
Click "Run Query" to see replica set status...

🌍 Real-World Scenarios

πŸ›’

E-Commerce Platform

Challenge: 10,000 concurrent users during flash sale. Can't afford downtime or data loss.
Solution: 5-node replica set (1 primary + 4 secondaries) across 3 availability zones. Read preference set to "nearest" for faster reads.
// Connection String
mongodb://node1,node2,node3,node4,node5/?replicaSet=ecommerceRS

// Read Preference
db.products.find().readPref("nearest")

Result: 99.99% uptime. Zero downtime during Black Friday sale. Handled 50K orders/minute.

πŸ“±

Social Media App

Challenge: Global users need low latency. Must handle millions of posts/day.
Solution: Geographically distributed replica set - nodes in US, Europe, Asia. Write to primary in US, read from nearest secondary.
// Geo-Distributed Setup
Primary: us-east-1 (Virginia)
Secondary: eu-west-1 (Ireland)  
Secondary: ap-south-1 (Mumbai)

// Users read from nearest location
Response Time: 50-100ms (global average)
🏦

Banking Application

Challenge: Must ensure strong consistency. Zero data loss acceptable.
Solution: Write concern "majority" - write only acknowledged after majority of nodes confirm.
// Strong Consistency
db.transactions.insertOne(
  { 
    from: "ACC1", 
    to: "ACC2", 
    amount: 10000 
  },
  { writeConcern: { w: "majority" } }
)

// Ensures data safe on majority before success

✨ Best Practices & Production Tips

βœ… DO's

  • Use odd number of nodes (3, 5, 7)
  • Distribute nodes across availability zones
  • Monitor replication lag
  • Set appropriate write concerns
  • Use majority read concern for critical reads
  • Regular backup of all nodes

❌ DON'Ts

  • Don't use even number of voting nodes
  • Don't put all nodes in same datacenter
  • Don't ignore replication lag warnings
  • Don't use secondaries for critical reads without read concern
  • Don't skip monitoring and alerts

⚠️ Common Pitfalls

  • Network latency between nodes > 50ms
  • Insufficient oplog size
  • Not planning for growth
  • Ignoring priority and votes configuration

πŸ“Š Recommended Configurations

Use Case Configuration Write Concern
Small App/Startup 3 nodes (1P + 2S) w: 1
Production App 5 nodes (1P + 4S) w: "majority"
Critical/Financial 7 nodes (1P + 6S) w: "majority", j: true
Global App 5 nodes across 3 regions w: "majority"

❓ Interview Questions & Answers

Q1 What is a Replica Set in MongoDB? How does it ensure high availability? β–Ό

Answer:

A Replica Set is a group of MongoDB instances that maintain the same dataset. It consists of:

  • Primary Node: Receives all write operations
  • Secondary Nodes: Maintain copies of primary's data through replication
  • Arbiter (Optional): Participates in elections but doesn't hold data

High Availability Mechanism:

  • Automatic failover: If primary fails, secondaries elect a new primary (typically 10-12 seconds)
  • Data redundancy: Multiple copies prevent data loss
  • Continuous replication: Secondaries sync data asynchronously
  • Self-healing: Failed nodes automatically rejoin when recovered
Q2 Explain the automatic failover process in MongoDB Replica Sets. β–Ό

Answer:

Failover Process (Step-by-Step):

  1. Detection (2s): Secondaries detect primary failure through missed heartbeats (sent every 2 seconds)
  2. Election Initiation (1-2s): Secondaries that can see a majority of members call for an election
  3. Voting (5-8s): Members vote for the most up-to-date secondary. A candidate needs majority votes to win
  4. New Primary (Total ~10s): Winning secondary becomes primary and starts accepting writes
  5. Application Reconnection: MongoDB drivers automatically reconnect to the new primary

Key Points:

  • Total downtime: 10-12 seconds (typical)
  • No data loss if using appropriate write concerns
  • Completely automatic - no manual intervention needed
Q3 What is the purpose of an Arbiter in a Replica Set? β–Ό

Answer:

An Arbiter is a lightweight member of a replica set with these characteristics:

  • Does not hold data: Only participates in elections
  • Voting only: Provides a vote in elections to maintain odd number of voters
  • Low resource usage: Minimal CPU/memory/disk requirements
  • Tie-breaker: Prevents split-brain scenarios in 2-node setups

When to use:

  • When you have even number of data-bearing nodes (e.g., 2 or 4)
  • To maintain odd number of voters without the cost of full secondary
  • Example: 2 data nodes + 1 arbiter = 3 voters (odd number)

Best Practice: For production, prefer using actual secondaries instead of arbiters when possible for better redundancy.

Q4 What is replication lag and how do you monitor it? β–Ό

Answer:

Replication Lag: The time delay between when data is written to primary and when it's replicated to secondaries.

Typical Values:

  • Normal: 100-500ms
  • Warning: > 1 second
  • Critical: > 10 seconds

Monitoring Methods:

  1. rs.status(): Check "optimeDate" difference between primary and secondaries
  2. rs.printReplicationInfo(): Shows oplog size and coverage
  3. rs.printSecondaryReplicationInfo(): Detailed lag info for each secondary
  4. MongoDB Atlas: Built-in monitoring dashboards
  5. Third-party tools: Datadog, New Relic, Prometheus

Causes of High Lag:

  • Network latency between nodes
  • Heavy write load on primary
  • Slow disk I/O on secondaries
  • Secondary handling too many reads
Q5 Explain Write Concerns and Read Preferences in Replica Sets. β–Ό

Answer:

Write Concern: Level of acknowledgment requested from MongoDB for write operations.

Common Write Concerns:

  • w: 1 - Acknowledged by primary only (fastest, less safe)
  • w: "majority" - Acknowledged by majority of nodes (recommended for production)
  • w: <number> - Acknowledged by specific number of nodes
  • j: true - Wait for journal commit (durability)
// Example
db.orders.insertOne(
  { orderId: 123, total: 599 },
  { writeConcern: { w: "majority", j: true, wtimeout: 5000 } }
)

Read Preference: Where to route read operations.

  • primary - Read from primary only (default, most consistent)
  • primaryPreferred - Primary if available, else secondary
  • secondary - Read from secondary only
  • secondaryPreferred - Secondary if available, else primary
  • nearest - Read from nearest node (lowest latency)
// Example
db.products.find().readPref("secondary")
Q6 What is the minimum recommended number of nodes for a production Replica Set? β–Ό

Answer:

Minimum: 3 nodes (1 Primary + 2 Secondaries)

Why 3?

  • Majority for elections: Need majority (2 out of 3) to elect new primary
  • Fault tolerance: Can survive 1 node failure and still maintain quorum
  • Data redundancy: 3 copies of data protect against loss
  • Odd number: Prevents split-brain scenarios

Production Recommendations:

  • Startup/Small: 3 nodes
  • Medium/Production: 5 nodes (can survive 2 failures)
  • Critical/Enterprise: 7 nodes across multiple regions

Important: Always use odd number of voting members (3, 5, 7) to ensure clear majority in elections.

Q7 How would you troubleshoot a Replica Set where a secondary is stuck in RECOVERING state? β–Ό

Answer:

Common Causes & Solutions:

  1. Check Oplog Coverage:
    • Use rs.printReplicationInfo() on primary
    • If secondary is behind oplog window, it needs initial sync
    • Solution: Increase oplog size or resync from backup
  2. Network Issues:
    • Check connectivity between secondary and primary
    • Verify firewall rules and network latency
    • Check for packet loss or timeouts
  3. Disk Space:
    • Check if secondary has sufficient disk space
    • Monitor I/O performance
  4. Resource Constraints:
    • Check CPU/Memory usage
    • Look for long-running operations

Troubleshooting Steps:

# 1. Check status
rs.status()

# 2. Check logs
db.adminCommand({getLog: "global"})

# 3. Check replication info
rs.printSecondaryReplicationInfo()

# 4. If needed, resync
# Stop secondary, delete data, restart
# It will automatically perform initial sync