Explorer
SQL

Find all email addresses that appear more than once in the `users` table.

Problem Statement

<p>Find all email addresses that appear more than once in the <code>users</code> table.</p>

Examples

Input: users table: +----+------------------+ | id | email | +----+------------------+ | 1 | alice@test.com | | 2 | bob@test.com | | 3 | alice@test.com | | 4 | charlie@test.com | | 5 | bob@test.com | +----+------------------+

Output: +----------------+-----+ | email | cnt | +----------------+-----+ | alice@test.com | 2 | | bob@test.com | 2 | +----------------+-----+

Explanation: The records are grouped by category and the aggregate calculation is applied to produce the summary result.

Complexity

Time Complexity: -

Space Complexity: -

Hints

šŸ’” Hint 1: Identify the grouping dimension(s) and which columns require aggregate functions (such as COUNT, SUM, AVG, MIN, or MAX). šŸ’” Hint 2: Add the GROUP BY clause for all non-aggregated columns listed in the SELECT projection. šŸ’” Hint 3: If filtering groups, use HAVING; otherwise use WHERE before grouping: SELECT <group_col>, <AGG>(...) FROM <table> GROUP BY <group_col>;

Editorial & Approach

Problem Overview & Intuition

To solve "Find Duplicates (GROUP BY + HAVING)", we query the relational database engine using declarative SQL. The goal is to find all email addresses that appear more than once in the `users` table. By formulating an optimal execution plan with appropriate projection and filtering, the database engine executes the query with minimal overhead.

Step-by-Step Approach

  1. Analyze Schema: Identify the target tables, necessary foreign keys, and expected output columns.
  2. Construct Filtering & Logic: Apply grouping & aggregates to isolate the requested data.
  3. Format & Order: Ensure columns match the expected project schema in order.

Optimal Implementation (SQL)

SELECT email, COUNT(*) AS cnt FROM users GROUP BY email HAVING COUNT(*) > 1;

Complexity Analysis

Time Complexity O(N log N) for sorting or partitioning rows.
Space Complexity O(N) for intermediate group hash tables or window buffers.

Key Considerations & Edge Cases

  • Empty Tables: The query executes safely returning zero rows without syntax error.
  • NULL Values: Columns containing NULL values are properly handled by standard ANSI SQL semantics.
  • Case Sensitivity: String comparisons and keywords adhere to PostgreSQL/standard SQL rules.

Find Duplicates (GROUP BY + HAVING)

Hard

Find all email addresses that appear more than once in the users table.

Example Scenarios
1Example 1
Input:
users table
idemail
1alice@test.com
2bob@test.com
3alice@test.com
4charlie@test.com
5bob@test.com
Output:
emailcnt
alice@test.com2
bob@test.com2
Explanation:

The records are grouped by category and the aggregate calculation is applied to produce the summary result.

SQL Editor
Loading Editor...
Query Results

Run a query to see results here.