Find all email addresses that appear more than once in the `users` table.
Problem Statement
Examples
Input: users table: +----+------------------+ | id | email | +----+------------------+ | 1 | alice@test.com | | 2 | bob@test.com | | 3 | alice@test.com | | 4 | charlie@test.com | | 5 | bob@test.com | +----+------------------+
Output: +----------------+-----+ | email | cnt | +----------------+-----+ | alice@test.com | 2 | | bob@test.com | 2 | +----------------+-----+
Explanation: The records are grouped by category and the aggregate calculation is applied to produce the summary result.
Complexity
Time Complexity: -
Space Complexity: -
Hints
Editorial & Approach
Problem Overview & Intuition
To solve "Find Duplicates (GROUP BY + HAVING)", we query the relational database engine using declarative SQL. The goal is to find all email addresses that appear more than once in the `users` table. By formulating an optimal execution plan with appropriate projection and filtering, the database engine executes the query with minimal overhead.
Step-by-Step Approach
- Analyze Schema: Identify the target tables, necessary foreign keys, and expected output columns.
- Construct Filtering & Logic: Apply grouping & aggregates to isolate the requested data.
- Format & Order: Ensure columns match the expected project schema in order.
Optimal Implementation (SQL)
SELECT email, COUNT(*) AS cnt FROM users GROUP BY email HAVING COUNT(*) > 1;
Complexity Analysis
Key Considerations & Edge Cases
- Empty Tables: The query executes safely returning zero rows without syntax error.
- NULL Values: Columns containing NULL values are properly handled by standard ANSI SQL semantics.
- Case Sensitivity: String comparisons and keywords adhere to PostgreSQL/standard SQL rules.