Explorer
SQL

Delete duplicate rows from `users` (by email), keeping only the row with the smallest id. Return remaining rows.

Problem Statement

<p>Delete duplicate rows from <code>users</code> (by email), keeping only the row with the smallest id. Return remaining rows.</p>

Examples

Input: users table: +----+----------------+ | id | email | +----+----------------+ | 1 | alice@test.com | | 2 | bob@test.com | | 3 | alice@test.com | | 4 | bob@test.com | +----+----------------+

Output: +----+----------------+ | id | email | +----+----------------+ | 1 | alice@test.com | | 2 | bob@test.com | +----+----------------+

Explanation: The query retrieves the requested records satisfying all problem requirements.

Complexity

Time Complexity: -

Space Complexity: -

Hints

šŸ’” Hint 1: Identify the grouping dimension(s) and which columns require aggregate functions (such as COUNT, SUM, AVG, MIN, or MAX). šŸ’” Hint 2: Add the GROUP BY clause for all non-aggregated columns listed in the SELECT projection. šŸ’” Hint 3: If filtering groups, use HAVING; otherwise use WHERE before grouping: SELECT <group_col>, <AGG>(...) FROM <table> GROUP BY <group_col>;

Editorial & Approach

Problem Overview & Intuition

To solve "Delete Duplicates Keeping One", we query the relational database engine using declarative SQL. The goal is to delete duplicate rows from `users` (by email), keeping only the row with the smallest id. return remaining rows. By formulating an optimal execution plan with appropriate projection and filtering, the database engine executes the query with minimal overhead.

Step-by-Step Approach

  1. Analyze Schema: Identify the target tables, necessary foreign keys, and expected output columns.
  2. Construct Filtering & Logic: Apply grouping & aggregates to isolate the requested data.
  3. Format & Order: Sort the resulting records according to specified order criteria.

Optimal Implementation (SQL)

DELETE FROM users WHERE id NOT IN (SELECT MIN(id) FROM users GROUP BY email); SELECT * FROM users ORDER BY id;

Complexity Analysis

Time Complexity O(N log N) for sorting or partitioning rows.
Space Complexity O(N) for intermediate group hash tables or window buffers.

Key Considerations & Edge Cases

  • Empty Tables: The query executes safely returning zero rows without syntax error.
  • NULL Values: Columns containing NULL values are properly handled by standard ANSI SQL semantics.
  • Case Sensitivity: String comparisons and keywords adhere to PostgreSQL/standard SQL rules.

Delete Duplicates Keeping One

Hard

Delete duplicate rows from users (by email), keeping only the row with the smallest id. Return remaining rows.

Example Scenarios
1Example 1
Input:
users table
idemail
1alice@test.com
2bob@test.com
3alice@test.com
4bob@test.com
Output:
idemail
1alice@test.com
2bob@test.com
Explanation:

The query retrieves the requested records satisfying all problem requirements.

SQL Editor
Loading Editor...
Query Results

Run a query to see results here.