Lesson 113 — Describing and Comparing Distributions: Shape, Centre, Spread and Outliers

Strand: Statistics | Descriptor: AC9M7ST02 | Duration: 45 minutes

Learning Intentions

  • To describe a data distribution by its shape, centre, spread and outliers.
  • To compare two distributions using all four features.

Success Criteria

I can:

  1. Describe a distribution’s shape as symmetric, left-skewed or right-skewed.
  2. State the centre using an appropriate measure, and the spread using the range.
  3. Identify outliers and comment on their effect.
  4. Compare two data sets systematically using all four features.

Warmup

(6 minutes — three shapes, projected)

Show three dot plots without labels:

  • A: a single hump in the middle, tailing off evenly on both sides.
  • B: most values bunched at the left, with a long thin tail stretching right.
  • C: most values bunched at the right, with a long tail stretching left.
  1. Describe each in your own words.
  2. In which would the mean and median be about equal?
  3. In which would the mean be pulled above the median?

Answers: 1. A is even/balanced; B has a right tail; C has a left tail; 2. A — the balanced shape; 3. B — the long right tail drags the mean up (Lesson 109’s outlier effect, now as shape).

The vocabulary, named: A is symmetric; B is right-skewed (tail to the right); C is left-skewed. The skew is named for where the tail points, not where the bulk sits — the commonest error in this topic.

Activities

Activity 1 — Explicit Instruction: the Four-feature Description (14 min)

Every distribution description covers four things, in this order:

FeatureWhat to sayTools
ShapeSymmetric, left-skewed, right-skewed; one peak or severalThe display
CentreWhere the data sits typicallyMedian (or mean if symmetric)
SpreadHow widely values varyRange
OutliersAny values far from the restThe display; judgement

The memory hook: Shape, Centre, Spread, Outliers — describe in that order, every time.

I do — describe a distribution fully. Test marks out of :

 Stem | Leaf
    1 | 8
    2 | 4 7 9
    3 | 1 3 3 5 6 8 9
    4 | 0 2 4 5
    5 | 0

Key: 3 | 1 means 31

The description, modelled sentence by sentence:

“The distribution is roughly symmetric with a single peak in the s. The centre is a median of . The spread runs from to , a range of . There are no clear outliers, though the mark of sits somewhat below the rest.”

Emphasise: four sentences, one per feature. Vague description (“it’s kind of spread out”) earns nothing; each sentence carries a number or a named shape.

Skew and the mean–median gap — the diagnostic:

(This is why Lesson 110’s Station D asked what a mean far above a median implies.)

We do — describe together: daily hours of screen time for students:

 Stem | Leaf
    0 | 5 8 9
    1 | 0 2 2 3 5 5 6 8
    2 | 0 1 4
    3 |
    4 | 8

Key: 1 | 2 means 1.2 hours

(Shape: right-skewed with one peak in the -hour range. Centre: median h. Spread: to h, range h. Outlier: h sits far from the rest — note the empty stem showing the gap.)

Activity 2 — Comparing Two Distributions (14 min)

Pairs. The skill the descriptor names: comparing.

Two classes sat the same -mark test.

Class A:

Class B:

  1. Build a back-to-back stem-and-leaf plot.
  2. Calculate the median and range for each.
  3. Write a four-feature comparison — shape, centre, spread, outliers.
  4. Which class performed better? Justify with more than one feature.

Socratic scaffolding:

PromptPurpose
Build the plot with a shared stem column.Aligns the scales — the whole point of back-to-back (Lesson 111 E4).
Medians: fourteen values each, so?Average the th and th: A ; B .
Ranges?A: . B: .
Shapes?A is tightly bunched and roughly symmetric; B is more spread with a tail to the left.
Q4 — which class is “better”?B has the higher median and the top mark, but A is far more consistent. “Better” depends on whether you value the typical result or the spread.
The honest comparison sentence?”Class B has a slightly higher median ( vs ) but a much wider spread (range vs ); Class A’s results are more consistent.”

The comparison template to record:

“Both distributions … [shape]. Class A’s centre is … while Class B’s is … Class A’s spread is … compared with Class B’s … [Outliers, if any]. Overall, …”

Activity 3 — Inquiry: Same Summary, Different Shape (9 min)

Pairs.

Two data sets, each of nine values, both have a median of and a range of .

Set P: Set Q:

  1. Verify the median and range for each.
  2. Draw a dot plot of each.
  3. Describe each shape.
  4. What does this show about summarising a data set with just two numbers?

Socratic scaffolding:

PromptPurpose
Check both medians.Ninth value each: the th is in both ✓
Check both ranges. in both ✓
Now plot them. What is different?P has one central hump; Q has two clusters with a gap in the middle.
Name Q’s shape.Bimodal — two peaks.
So can median and range describe a distribution?No — identical summaries, completely different shapes.
What is the lesson?Always describe the shape too. Summary statistics are necessary but not sufficient.

The closing point: this is why the descriptor asks for shape, centre and spread. Two of the three is never enough.

Checks for Understanding

(5 minutes — exit ticket, collected)

Data: .

  1. Describe the shape.
  2. State the centre, choosing an appropriate measure and justifying.
  3. State the spread.
  4. Identify any outlier and its effect on the mean.
  5. Reasoning. Two data sets have the same median and range. Can you conclude they have the same shape?

Answers: 1. Right-skewed — most values cluster between and with one far value stretching a tail to the right; 2. Median — the mean of is inflated by the outlier; 3. Range (largely due to the outlier); 4. ; it pulls the mean above every other value; 5. No — Sets P and Q showed identical medians and ranges with completely different shapes.

Common Misconceptions

MisconceptionHow to pre-empt it
Naming skew by where the bulk sits.Skew is named for the tail. Repeat every time it appears.
Describing only the centre.The SCSO order, required in every description.
Vague description without numbers.Each sentence must carry a figure or a named shape.
Believing summary statistics fully describe data.Sets P and Q — identical summaries, different shapes.
Using the mean as centre for skewed data.The mean–median diagnostic, and Lesson 109’s justification work.
Comparing two data sets on one feature only.Q4 of Activity 2 requires more than one.

Enrichment — Competition-Style Problems

E1 (Kangaroo style). A data set has mean and median . What shape is suggested?

Answer

Right-skewed — the mean is pulled above the median by high values.

E2 (AMC Junior style). Two classes have the same median mark but Class X has range and Class Y has range . Write one sentence comparing them.

Answer

“Both classes have the same typical mark, but Class X’s results are far more consistent, with marks spanning only compared with Class Y’s .”

E3 (Challenge). Construct two data sets of seven values with the same mean and the same median but visibly different shapes.

Answer

E.g. (mean , median , tight single peak) and (mean , median , two clusters with a gap). Same summaries, different distributions.

E4 (Challenge). A distribution is described as “bimodal”. Explain what this means and give a realistic context where it would occur.

Answer

Two distinct peaks. Realistic contexts: heights of a mixed adult population (male and female peaks); test marks where students either understood a topic or did not; arrival times at a shop with a lunch rush and an after-work rush.

E5 (Challenge). Why does a right-skewed distribution usually have its mean above its median, in terms of how each is calculated?

Answer

The median depends only on position, so the long right tail moves it barely at all. The mean uses every value’s size, so the large tail values pull it upward. The gap between them is a direct measure of the skew.

Homework

  1. For each, name the shape: (a) mean , median (b) mean , median (c) mean , median .
  2. Describe this distribution using all four features:
 Stem | Leaf
    2 | 3 8
    3 | 1 4 5 5 7 9
    4 | 0 2 3 6
    5 | 1
    6 |
    7 | 8

Key: 3 | 4 means 34
  1. Two data sets: P has median , range ; Q has median , range . Write a two-sentence comparison.
  2. Data: . (a) Describe the shape. (b) State the centre with justification. (c) State the spread. (d) Identify any outlier.
  3. Construct a data set of eight values that is clearly left-skewed, and state its mean and median to demonstrate the skew.
  4. Two classes’ results: A — median , range ; B — median , range . Which class would you say did better? Justify using both features.
  5. Sketch a dot plot for a bimodal distribution and suggest a realistic context for it.
  6. Reasoning. Explain why skew is named for the tail rather than the bulk of the data.
  7. Reasoning. Explain why a description with only a centre is incomplete, using an example.
  8. Challenge. Construct two data sets of nine values with identical median, mean and range, but different shapes. Explain how you did it.

Answers: Q1 — (a) symmetric (b) right-skewed (c) left-skewed. Q2 — right-skewed with one peak in the s; median (fourteen values: average of the th, , and th, — recount from the plot); range ; the value is a clear outlier, with an empty stem showing the gap. Q3 — “Both sets share the same typical value of . However, P’s values are tightly clustered (range ) while Q’s are widely scattered (range ), so P is far more consistent.” Q4 — (a) right-skewed (b) median , since the mean of is inflated by the (c) range (d) . Q6 — B has the higher median but nearly three times the spread; A is more consistent. “Better” depends on whether typical performance or reliability matters more. Q8 — the tail is what distinguishes the shape; the bulk sits near the centre in every unimodal distribution, so it cannot name the direction. Q9 — e.g. two classes both with median , one ranging and one : the centre alone hides a vast difference. Q10 — e.g. and : both have median , range , and mean — one bimodal, one single-peaked. Build by keeping the extremes and the middle fixed while rearranging the remaining values in balanced pairs.