Lesson 117 — Analysing and Interpreting Data Distributions

Strand: Statistics | Descriptor: AC9M7ST03 | Duration: 45 minutes

Block note. Stage 3 of the investigation. Students bring the data collected in Lesson 116 and produce their displays and summary statistics here, ready to report in Lesson 118.

Learning Intentions

  • To display and summarise a collected data set.
  • To interpret the distribution in terms of the original question.

Success Criteria

I can:

  1. Choose and construct an appropriate display for my data.
  2. Calculate all four summary measures and select the most representative.
  3. Describe my distribution using shape, centre, spread and outliers.
  4. Answer my original question using evidence from the analysis.

Warmup

(6 minutes — the analysis checklist, individually)

Before touching your data, list what you will produce this lesson. Then compare with the class list.

The checklist to record:

  1. An appropriate display (with key/labels).
  2. Range, median, mean and mode — all four calculated.
  3. A choice of the most representative measure, with justification.
  4. A four-feature description (shape, centre, spread, outliers).
  5. An answer to the original question, in a sentence.

The step students most often skip is 5. Analysis that never returns to the question is arithmetic, not investigation.

Activities

Activity 1 — Choose and Construct (14 min)

Individually or in the original pairs.

Step 1 — choose the display. Use the Lesson 114 decision guide:

Your dataDisplay
Few distinct values, discreteDot plot or column graph
Many values, want individuals keptStem-and-leaf plot
Many values, continuousGrouped column graph
Two samples to compareBack-to-back stem-and-leaf

Step 2 — construct it properly. The non-negotiables:

  • A title stating what is shown.
  • Axis labels with units, or a key for a stem-and-leaf plot.
  • Empty stems or intervals shown, not skipped.
  • A vertical axis starting at zero for column graphs.

Step 3 — calculate all four measures, showing working:

Circulating prompts:

PromptPurpose
Does your display have a title, labels and a key?The three habitual omissions.
Did you order the data before finding the median?Lesson 108’s standing habit.
Your mode — does any value actually repeat?With continuous data, often none.
Does your frequency total match your sample size?The Lesson 116 check, carried forward.
Which measure will you choose, and why?Prevents calculating four and selecting none.

Activity 2 — Describe and Interpret (16 min)

The analytical heart of the project.

Part A — the four-feature description. Write four sentences, one per feature, each carrying a number or a named shape (Lesson 113’s SCSO order).

Part B — answer the question. Return to your Lesson 115 question and answer it directly, citing evidence:

Template: “Our question was … Our data shows that … [centre, with measure named]. The values ranged from … to …, and the distribution was … [shape]. This suggests that …”

Part C — the prediction check. Compare with the prediction you made in Lesson 115:

  1. Was your prediction supported, contradicted, or partly both?
  2. If contradicted, what might explain the difference?
  3. Does a contradicted prediction mean the investigation failed?

The Q3 answer matters and should be discussed with the class: no. A prediction is a hypothesis; data testing it is exactly what an investigation is for. Being wrong in an interesting way is a better outcome than being right by accident. (Same principle as Lessons 74–77: models and predictions are judged by honesty, not by being flattered.)

Socratic scaffolding for pairs whose data looks “boring”:

PromptPurpose
Is a tightly clustered distribution a failure?No — consistency is a finding. Report it.
Nothing surprising happened. What can you still say?The centre, the spread, the shape, and the absence of outliers are all results.
Would you have predicted this spread?Turns a flat result into a comparison with expectation.
What would a different sample look like?Opens the limitations discussion for Lesson 118.

Activity 3 — Inquiry: what Does the Shape Tell You? (9 min)

Pairs, applied to their own data.

Look at your display, not your numbers.

  1. Is your distribution symmetric, left-skewed or right-skewed?
  2. Does that shape make sense for what you measured? Why might this variable behave that way?
  3. Are there outliers? Are they errors, or genuine?
  4. What would have to be true for the shape to look completely different?

Discussion targets — why real variables have characteristic shapes:

  • Reaction times: right-skewed. There is a physical floor (nobody reacts in ms) but no ceiling — a lapse in attention produces an arbitrarily long time. Floors with no ceilings produce right skew.
  • Heights and arm spans: roughly symmetric. Most people cluster near a typical value with similar-sized deviations both ways.
  • Number of siblings: right-skewed. Most families have ; a few have many more.
  • Test marks out of a maximum: often left-skewed, because the maximum acts as a ceiling.

The generalisable idea: the shape is not random — it usually reflects a physical or practical constraint on the variable. Naming the constraint explains the shape.

Checks for Understanding

(5 minutes — exit ticket, collected with the analysis)

  1. Name the display you chose and give one reason.
  2. State your median, mean, range and mode.
  3. Which measure best represents your data, and why?
  4. Describe your distribution’s shape in one sentence.
  5. Reasoning. Was your prediction supported? Explain what you learned either way.

Answers: 1–5. Student’s own, marked against the analysis. Q5 is marked on the quality of the reasoning, not on whether the prediction held.

Common Misconceptions

MisconceptionHow to pre-empt it
Calculating statistics without returning to the question.Checklist item 5, and Part B’s template.
Calculating four measures and choosing none.Circulating prompt; the choice must be justified.
Treating a contradicted prediction as failure.Part C’s Q3, discussed with the class.
Believing an unremarkable result is not a result.The “boring data” scaffolding — consistency is a finding.
Displays without titles, labels or keys.The non-negotiables list, marked.
Describing shape without asking why.Activity 3 links shape to physical constraints.

Enrichment — Competition-Style Problems

E1 (Kangaroo style). A data set’s mean is and its median is . What shape is suggested, and which measure would you report as typical?

Answer

Right-skewed; report the median (), since the mean is inflated by the high tail.

E2 (AMC Junior style). Two samples of reaction times: Group A median ms, range ; Group B median ms, range . What can you conclude?

Answer

The typical reaction time is essentially the same, but Group B is far less consistent — probably containing one or more lapses. The centres alone would have hidden this.

E3 (Challenge). Why would you expect a distribution of “time taken to solve a puzzle” to be right-skewed?

Answer

There is a lower limit — nobody solves it in negative time, and the fastest possible solve is bounded by the puzzle itself — but no upper limit, since someone can take arbitrarily long or get stuck. A floor with no ceiling produces a right tail.

E4 (Challenge). A student’s data has a mode but the mode is far from both the mean and the median. What might this indicate?

Answer

Possibly a bimodal distribution, or a rounding artefact where many values were recorded at one convenient figure (e.g. everyone reporting ” hours” of sleep). Worth checking the raw data and the collection method before drawing conclusions.

Homework

  1. Finalise your display, with title, labels, key and all intervals or stems shown.
  2. Present all four summary measures with working shown.
  3. Write your four-feature description in four sentences.
  4. Answer your original question in two sentences, citing your evidence.
  5. Compare your result with your Lesson 115 prediction and explain any difference.
  6. Explain why your distribution has the shape it does, referring to the variable itself.
  7. If you collected two samples, write a two-sentence comparison covering both centre and spread.
  8. Reasoning. Explain why a contradicted prediction is not a failed investigation.
  9. Reasoning. Your data has no mode. Explain what this tells you about the variable.
  10. Challenge. Sketch what you would expect your distribution to look like if the sample were students instead of , and explain your reasoning.

Answers: Q1–7 — student’s own, marked against the project folder. Q8 — the purpose of collecting data is to test the prediction, not confirm it; a contradiction is informative and often points to a faulty assumption worth examining. Q9 — no value repeated, which is typical of precisely-measured continuous data; the mode is not a useful measure here. Q10 — the shape would become smoother and clearer (less lumpiness from chance), the centre would probably stay similar if the small sample was representative, and the range would likely widen as extreme values become more likely.