🎨 MongoDB Schema Patterns Cheatsheet
Complete guide to schema design patterns, relationships, and best practices with real-world examples
Schema Design Philosophy
MongoDB vs SQL mindset
SQL: Design schema first, then optimize queries
MongoDB: Design schema around how you'll query the data!
Ask yourself: "How will this data be accessed?" not "How should I normalize this?"
• Multiple tables for relationships
• JOIN operations everywhere
• Schema first, queries later
• One "correct" schema
• Denormalize for read performance
• Minimize lookups/joins
• Query patterns drive schema
• Many valid schemas
| Relationship Type | Cardinality | Best Pattern | Example |
|---|---|---|---|
| One-to-Few | 1 : 1-10 | Embedded Documents | User has 3 addresses |
| One-to-Many | 1 : 100s | Reference (Child → Parent) | Blog has 500 comments |
| One-to-Squillions | 1 : Millions | Reference (Parent → Child) | Server has millions of log entries |
| Many-to-Many | N : M | Two-way References or Embedded IDs | Students ↔ Courses |
Embedded Document Pattern
One-to-Few relationships
- Data is always accessed together
- Child data has no meaning without parent
- One-to-Few relationship (< 100 embedded docs)
- Embedded data doesn't grow unbounded
- Need atomic updates
- Embedded array grows unbounded (could exceed 16MB limit)
- Child data needs to be accessed independently
- Many-to-Many relationships
- Child data updated frequently (write amplification)
Reference Pattern
One-to-Many relationships
Polymorphic Pattern
Different document types in same collection
- Single collection for related but different entities
- Easy to query across all types
- Flexible schema per type
- Common fields can be indexed
Bucket Pattern
Time-series and high-volume data
- Hourly buckets: High-frequency data (every minute/second)
- Daily buckets: Medium-frequency data (every hour)
- Monthly buckets: Low-frequency data (daily)
- Keep bucket size under 16MB document limit
- Pre-compute aggregations (avg, min, max) for better performance
Subset Pattern
Partial data for performance
- Large arrays that would make documents too big
- Only need recent/top items most of the time
- Want to avoid loading unnecessary data
- Examples: Comments, reviews, notifications, messages
Computed Pattern
Pre-calculated values for performance
- Order totals: subtotal, tax, shipping, total
- Statistics: avgRating, totalReviews, viewCount
- Counts: followerCount, likeCount, commentCount
- Derived fields: fullName, ageInYears, daysUntilExpiry
Design Decision Framework
How to choose the right pattern
- Identify access patterns - How will data be queried?
- Determine relationship cardinality - One-to-few? One-to-many?
- Consider data growth - Will embedded arrays grow unbounded?
- Evaluate read vs write ratio - Read-heavy? Optimize for reads
- Check document size - Stay well under 16MB limit
- Decide atomicity needs - Need atomic updates?
- Plan for scaling - Will you need sharding later?
- Consider duplication tradeoffs - Storage vs query performance
| Question | Yes → Use This | No → Use This |
|---|---|---|
| Data always accessed together? | Embedded Documents | References |
| One-to-Few (< 100)? | Embedded Array | References |
| Child data needs independent access? | References | Embedded |
| Unbounded array growth? | References (Child → Parent) | Embedded |
| Read performance critical? | Denormalize, Computed Pattern | Normalize, References |
| High-volume time-series? | Bucket Pattern | Regular Documents |
| Large arrays, need subset? | Subset Pattern | Full Embedded |
| Different entity types? | Polymorphic Pattern | Separate Collections |
Unlike SQL, MongoDB allows you to optimize your schema for YOUR specific use case. The same data can be modeled differently depending on how you'll query it. Start with your access patterns, not with normalization!