Schema Design: Embedding vs Referencing
Schema Design: Embedding vs. Referencing
In MongoDB, relationships are modeled via two primary approaches: Embedding (Denormalization) and Referencing (Normalization).
Comparison Matrix: Embedding vs Referencing
| Decision Factor | Embedding (Denormalized) | Referencing (Normalized) |
|---|---|---|
| Data Structure | Subdocuments and arrays inside parent document | Stores ID references (foreign keys) pointing to another collection |
| Read Performance | Ultra-fast; single disk lookup fetches all related data | Requires multiple queries or an aggregation $lookup join |
| Write / Update Cost | Updates may require atomic document rewriting | Updates to child entities touch only their isolated documents |
| Data Duplication | Data may be duplicated across multiple parent documents | Normalized; single source of truth for each entity |
| Document Size Risk | Risk of exceeding the 16 MB limit if array grows unbounded | No size risk; references scale infinitely across documents |
| Cardinality Fit | One-to-One, One-to-Few (< 100 children) | One-to-Many, One-to-Squillions, Many-to-Many |
Rules of Thumb
- Embed by default if related entities are always queried and displayed together (e.g. user address, line items on an invoice).
- Reference if related data is accessed independently, updated frequently, or grows unbounded (e.g. log entries, social followers).