Aggregation Pipeline Explained
The Aggregation Pipeline Framework
The Aggregation Pipeline is a high-performance framework that processes documents through a series of sequential transformation stages. Each stage transforms documents as they pass through, streaming results to the next stage.
Core Pipeline Stages & SQL Equivalents
| Pipeline Stage | SQL Equivalent | Description & Operation |
|---|---|---|
$match |
WHERE / HAVING |
Filters documents to allow only matches to pass to the next stage |
$group |
GROUP BY |
Groups documents by an _id key and computes aggregations ($sum, $avg, $min, $max) |
$project |
SELECT (column list) |
Reshapes documents by adding, renaming, or suppressing specific fields |
$sort |
ORDER BY |
Reorders all input documents by the specified field direction (1 or -1) |
$limit / $skip |
LIMIT / OFFSET |
Restricts document count or skips preceding results for pagination |
$unwind |
CROSS JOIN LATERAL |
Deconstructs an array field from input documents to output a document for each element |
$lookup |
LEFT OUTER JOIN |
Performs an equality join to another collection in the same database |
Optimization Best Practices
- Place
$matchand$sortas early as possible in the pipeline to utilize indexes. - Use
$projector$addFieldsto strip unnecessary fields early, minimizing memory utilization across pipeline stages.