Explorer
MongoDB

Group users by `email`, count occurrences, and return only emails that appear more than once (duplicates).

Problem Statement

<p>Group <code>users</code> by <code>email</code>, count occurrences, and return only emails that appear more than once (duplicates).</p>

Examples

Input: users collection: +-----+------------------+ | _id | email | +-----+------------------+ | 1 | alice@test.com | | 2 | bob@test.com | | 3 | alice@test.com | | 4 | charlie@test.com | | 5 | bob@test.com | +-----+------------------+

Output: +----------------+-------+ | _id | count | +----------------+-------+ | alice@test.com | 2 | | bob@test.com | 2 | +----------------+-------+

Explanation: Documents containing the matching array elements are selected from the collection.

Complexity

Time Complexity: -

Space Complexity: -

Hints

šŸ’” Hint 1: This problem requires grouping documents by a key and aggregating values. Use db.<collection>.aggregate([...]) with a $group stage. šŸ’” Hint 2: In the $group stage, set _id to the grouping field (e.g. "$category" or "$dept") and use accumulator operators like $sum, $avg, $max, or $min. šŸ’” Hint 3: Remember to call .toArray() at the end of the aggregation pipeline to return the resolved documents array.

Editorial & Approach

Problem Overview & Intuition

To solve "Find Duplicate Emails", we query the MongoDB document store. The goal is to group users by `email`, count occurrences, and return only emails that appear more than once (duplicates). Using an aggregation pipeline, the database engine filters and structures the BSON documents efficiently.

Step-by-Step Approach

  1. Identify Target Collection: Access the collection through the db instance.
  2. Construct Query / Pipeline: Build the aggregation stages ($match, $group, $sort, etc.).
  3. Resolve Cursor: Invoke .toArray() to transform the query cursor into the required array of documents.

Optimal Implementation (MongoDB)

function solve(db) {
  return db.users.aggregate([
    { $group: { _id: "$email", count: { $sum: 1 } } },
    { $match: { count: { $gt: 1 } } }
  ]);
}

Complexity Analysis

Time Complexity O(N) pipeline traversal through aggregation stages.
Space Complexity O(M) intermediate document buffer in aggregation pipeline.

Key Considerations & Edge Cases

  • Empty Collections: If no documents match, the query cleanly returns an empty array [].
  • Missing / NULL Fields: Missing fields in documents are handled safely without throwing runtime exceptions.
  • Type Coercion: BSON types (ObjectId, Numbers, Strings) are compared strictly according to MongoDB specifications.

Find Duplicate Emails

Hard

Group users by email, count occurrences, and return only emails that appear more than once (duplicates).

Example Scenarios
1Example 1
Input:
users collection
_idemail
1alice@test.com
2bob@test.com
3alice@test.com
4charlie@test.com
5bob@test.com
Output:
_idcount
alice@test.com2
bob@test.com2
Explanation:

Documents containing the matching array elements are selected from the collection.

MongoDB Editor
Loading Editor...
Query Results (JSON)

Run your MongoDB code to see results here.