Delete duplicate rows from `users` (by email), keeping only the row with the smallest id. Return remaining rows.
Problem Statement
Examples
Input: users table: +----+----------------+ | id | email | +----+----------------+ | 1 | alice@test.com | | 2 | bob@test.com | | 3 | alice@test.com | | 4 | bob@test.com | +----+----------------+
Output: +----+----------------+ | id | email | +----+----------------+ | 1 | alice@test.com | | 2 | bob@test.com | +----+----------------+
Explanation: The query retrieves the requested records satisfying all problem requirements.
Complexity
Time Complexity: -
Space Complexity: -
Hints
Editorial & Approach
Problem Overview & Intuition
To solve "Delete Duplicates Keeping One", we query the relational database engine using declarative SQL. The goal is to delete duplicate rows from `users` (by email), keeping only the row with the smallest id. return remaining rows. By formulating an optimal execution plan with appropriate projection and filtering, the database engine executes the query with minimal overhead.
Step-by-Step Approach
- Analyze Schema: Identify the target tables, necessary foreign keys, and expected output columns.
- Construct Filtering & Logic: Apply grouping & aggregates to isolate the requested data.
- Format & Order: Sort the resulting records according to specified order criteria.
Optimal Implementation (SQL)
DELETE FROM users WHERE id NOT IN (SELECT MIN(id) FROM users GROUP BY email); SELECT * FROM users ORDER BY id;
Complexity Analysis
Key Considerations & Edge Cases
- Empty Tables: The query executes safely returning zero rows without syntax error.
- NULL Values: Columns containing NULL values are properly handled by standard ANSI SQL semantics.
- Case Sensitivity: String comparisons and keywords adhere to PostgreSQL/standard SQL rules.