Group users by `email`, count occurrences, and return only emails that appear more than once (duplicates).
Problem Statement
Examples
Input: users collection: +-----+------------------+ | _id | email | +-----+------------------+ | 1 | alice@test.com | | 2 | bob@test.com | | 3 | alice@test.com | | 4 | charlie@test.com | | 5 | bob@test.com | +-----+------------------+
Output: +----------------+-------+ | _id | count | +----------------+-------+ | alice@test.com | 2 | | bob@test.com | 2 | +----------------+-------+
Explanation: Documents containing the matching array elements are selected from the collection.
Complexity
Time Complexity: -
Space Complexity: -
Hints
Editorial & Approach
Problem Overview & Intuition
To solve "Find Duplicate Emails", we query the MongoDB document store. The goal is to group users by `email`, count occurrences, and return only emails that appear more than once (duplicates). Using an aggregation pipeline, the database engine filters and structures the BSON documents efficiently.
Step-by-Step Approach
- Identify Target Collection: Access the collection through the
dbinstance. - Construct Query / Pipeline: Build the aggregation stages ($match, $group, $sort, etc.).
- Resolve Cursor: Invoke
.toArray()to transform the query cursor into the required array of documents.
Optimal Implementation (MongoDB)
function solve(db) {
return db.users.aggregate([
{ $group: { _id: "$email", count: { $sum: 1 } } },
{ $match: { count: { $gt: 1 } } }
]);
}
Complexity Analysis
Key Considerations & Edge Cases
- Empty Collections: If no documents match, the query cleanly returns an empty array
[]. - Missing / NULL Fields: Missing fields in documents are handled safely without throwing runtime exceptions.
- Type Coercion: BSON types (ObjectId, Numbers, Strings) are compared strictly according to MongoDB specifications.