Bird
Raised Fist0
SQLquery~10 mins

How GROUP BY changes query execution in SQL - Visual Walkthrough

Choose your learning style10 modes available

Start learning this pattern below

Jump into concepts and practice - no test required

or
Recommended
Test this pattern10 questions across easy, medium, and hard to know if this pattern is strong
Concept Flow - How GROUP BY changes query execution
Start with full table data
Scan all rows
Group rows by specified column(s)
Aggregate each group (e.g., COUNT, SUM)
Return one row per group with aggregated values
End
The query first reads all rows, groups them by the chosen column(s), then calculates aggregates for each group, and finally returns one result row per group.
Execution Sample
SQL
SELECT department, COUNT(*) AS employee_count
FROM employees
GROUP BY department;
This query counts how many employees are in each department by grouping rows by department.
Execution Table
StepActionInput RowsGroups FormedAggregated ResultOutput Rows
1Scan all rows from employees table10 rowsNone yetNone yetNone yet
2Group rows by department10 rows3 groups: Sales, HR, ITNone yetNone yet
3Count employees in each groupGrouped rows3 groupsSales: 4, HR: 3, IT: 3None yet
4Return one row per group with countsAggregated data3 groupsSame as previous3 rows
5Query endsN/AN/AN/A3 rows
💡 All rows processed and grouped; output contains one row per group with aggregated counts.
Variable Tracker
VariableStartAfter Step 2After Step 3Final
Input Rows10 rows from employees10 rowsGrouped into 3 groupsGrouped into 3 groups
GroupsNone3 groups formed3 groups with counts3 groups with counts
Aggregated ResultNoneNoneCounts calculatedCounts calculated
Output RowsNoneNoneNone3 rows returned
Key Moments - 3 Insights
Why does the number of output rows decrease after GROUP BY?
Because rows are grouped by the specified column(s), multiple input rows combine into one group, so the output has one row per group instead of one per input row (see execution_table step 4).
Does GROUP BY change the original data in the table?
No, GROUP BY only organizes rows during query execution; it does not modify the original table data (see execution_table step 1 vs step 2).
What happens if you use GROUP BY without an aggregate function?
The query will return one row per group but without meaningful aggregation, often causing errors or unexpected results depending on SQL rules (not shown in this trace).
Visual Quiz - 3 Questions
Test your understanding
Look at the execution_table, at which step are the rows grouped by department?
AStep 1
BStep 2
CStep 3
DStep 4
💡 Hint
Check the 'Groups Formed' column in execution_table rows.
According to variable_tracker, how many groups exist after Step 3?
ANone
B1 group
C3 groups
D10 groups
💡 Hint
Look at the 'Groups' row after Step 3 in variable_tracker.
If the employees table had 5 departments instead of 3, how would the output rows change?
AOutput rows would be 5
BOutput rows would stay 3
COutput rows would be 10
DOutput rows would be 1
💡 Hint
Refer to execution_table step 4 where output rows equal number of groups.
Concept Snapshot
GROUP BY groups rows by one or more columns.
Aggregates like COUNT or SUM calculate values per group.
Output returns one row per group.
Original table data is not changed.
Without aggregates, GROUP BY may cause errors or unexpected results.
Full Transcript
This visual execution trace shows how the SQL GROUP BY clause changes query execution. First, the query scans all rows from the employees table. Then, it groups these rows by the department column, forming three groups. Next, it calculates the count of employees in each group. Finally, the query returns one row per group with the aggregated counts. The number of output rows is less than the input rows because multiple rows combine into groups. GROUP BY organizes data during query execution but does not modify the original table. This trace helps beginners see step-by-step how grouping and aggregation work together in SQL.

Practice

(1/5)
1. What does the GROUP BY clause do in an SQL query?
easy
A. It filters rows based on a condition.
B. It groups rows that have the same values in specified columns.
C. It deletes duplicate rows from the result.
D. It sorts the rows in ascending order.

Solution

  1. Step 1: Understand the purpose of GROUP BY

    The GROUP BY clause collects rows with the same values in specified columns into groups.
  2. Step 2: Compare with other SQL clauses

    Sorting is done by ORDER BY, filtering by WHERE, and removing duplicates by DISTINCT, not GROUP BY.
  3. Final Answer:

    It groups rows that have the same values in specified columns. -> Option B
  4. Quick Check:

    GROUP BY groups rows by column values [OK]
Hint: GROUP BY groups rows by column values, not sorting or filtering [OK]
Common Mistakes:
  • Confusing GROUP BY with ORDER BY
  • Thinking GROUP BY filters rows
  • Assuming GROUP BY removes duplicates
2. Which of the following is the correct syntax to group rows by the column department?
easy
A. SELECT department, COUNT(*) FROM employees WHERE department;
B. SELECT department, COUNT(*) FROM employees ORDER BY department;
C. SELECT department, COUNT(*) FROM employees GROUP BY department;
D. SELECT department, COUNT(*) FROM employees HAVING department;

Solution

  1. Step 1: Identify correct GROUP BY usage

    The GROUP BY clause must follow the FROM clause and specify the column to group by, here 'department'.
  2. Step 2: Check each option's syntax

    SELECT department, COUNT(*) FROM employees GROUP BY department; uses GROUP BY correctly. SELECT department, COUNT(*) FROM employees ORDER BY department; uses ORDER BY which sorts, not groups. SELECT department, COUNT(*) FROM employees WHERE department; uses WHERE incorrectly. SELECT department, COUNT(*) FROM employees HAVING department; uses HAVING without GROUP BY, which is invalid.
  3. Final Answer:

    SELECT department, COUNT(*) FROM employees GROUP BY department; -> Option C
  4. Quick Check:

    GROUP BY syntax: SELECT ... FROM ... GROUP BY column [OK]
Hint: GROUP BY follows FROM and lists columns to group by [OK]
Common Mistakes:
  • Using ORDER BY instead of GROUP BY
  • Using WHERE to group rows
  • Using HAVING without GROUP BY
3. Given the table sales with columns region and amount, what is the output of this query?
SELECT region, SUM(amount) FROM sales GROUP BY region;
medium
A. A list of regions with the total sales amount per region.
B. A list of all sales amounts without grouping.
C. An error because SUM() cannot be used with GROUP BY.
D. A list of regions sorted by amount.

Solution

  1. Step 1: Understand GROUP BY with aggregate functions

    The query groups rows by 'region' and calculates the sum of 'amount' for each group.
  2. Step 2: Analyze the output

    The result shows each region once with the total sales amount summed from all rows in that region.
  3. Final Answer:

    A list of regions with the total sales amount per region. -> Option A
  4. Quick Check:

    GROUP BY + SUM() = totals per group [OK]
Hint: GROUP BY with SUM() gives totals per group [OK]
Common Mistakes:
  • Expecting all rows without grouping
  • Thinking SUM() causes error with GROUP BY
  • Confusing grouping with sorting
4. What is wrong with this query?
SELECT department, employee_name, COUNT(*) FROM employees GROUP BY department;
medium
A. You cannot select employee_name without grouping by it or using an aggregate function.
B. COUNT(*) cannot be used with GROUP BY.
C. GROUP BY must come before SELECT.
D. The query is correct and will run without errors.

Solution

  1. Step 1: Check columns in SELECT with GROUP BY

    When using GROUP BY on 'department', all selected columns must be grouped or aggregated.
  2. Step 2: Identify the error

    'employee_name' is neither grouped nor aggregated, causing a syntax error.
  3. Final Answer:

    You cannot select employee_name without grouping by it or using an aggregate function. -> Option A
  4. Quick Check:

    Non-grouped columns must be aggregated [OK]
Hint: All selected columns must be grouped or aggregated [OK]
Common Mistakes:
  • Selecting non-grouped columns without aggregation
  • Thinking COUNT(*) is invalid with GROUP BY
  • Misplacing GROUP BY clause
5. You want to find departments with more than 5 employees. Which query correctly uses GROUP BY and HAVING to achieve this?
hard
A. SELECT department, COUNT(*) FROM employees GROUP BY department WHERE COUNT(*) > 5;
B. SELECT department, COUNT(*) FROM employees HAVING COUNT(*) > 5 GROUP BY department;
C. SELECT department, COUNT(*) FROM employees WHERE COUNT(*) > 5 GROUP BY department;
D. SELECT department, COUNT(*) FROM employees GROUP BY department HAVING COUNT(*) > 5;

Solution

  1. Step 1: Understand HAVING clause usage

    HAVING filters groups after aggregation, so it must come after GROUP BY.
  2. Step 2: Check query order and syntax

    SELECT department, COUNT(*) FROM employees GROUP BY department HAVING COUNT(*) > 5; correctly places HAVING after GROUP BY with condition COUNT(*) > 5. Other options misuse HAVING or WHERE clauses.
  3. Final Answer:

    SELECT department, COUNT(*) FROM employees GROUP BY department HAVING COUNT(*) > 5; -> Option D
  4. Quick Check:

    HAVING filters groups after GROUP BY [OK]
Hint: Use HAVING after GROUP BY to filter groups [OK]
Common Mistakes:
  • Using WHERE to filter aggregated results
  • Placing HAVING before GROUP BY
  • Confusing WHERE and HAVING clauses