Lesson 109 — Problem Solving and Consolidation: Analysing and Reporting Distributions

Strand: Statistics | Descriptor: AC9M8ST02 | Duration: 45 minutes

Learning Intentions

  • To consolidate analysing shape, centre and spread from primary and secondary sources.
  • To apply the full statistical reporting process to a real, multi-part investigation.

Success Criteria

I can:

  1. Calculate and interpret mean, median, mode, range and IQR for a real data set.
  2. Identify the shape of a distribution and justify the appropriate measures of centre and spread.
  3. Classify a data source as primary or secondary and explain why that matters for a report.
  4. Write and justify a complete statistical report and recommendation for a real decision.

Warmup

(6 minutes — believe it or not, pairs)

  1. A data set has mean and median . Is the distribution more likely skewed left or skewed right?
  2. A data set’s range is but its IQR is only . What does this suggest is happening in the data?
  3. True or false: “A secondary source is always less reliable than data you collect yourself.”

Answers: 1. Right-skewed — the mean is pulled well above the median by a tail of high values; 2. There is likely at least one extreme outlier stretching the range, while the middle of the data is tightly clustered; 3. False — reliability depends on how carefully the data was collected and how large/representative the sample is, not simply on whether it is primary or secondary.

Today’s focus: applying everything from Lessons 107–108 to real, layered problems.

Activities

Activity 1 — Mixed Problem Circuit (16 min)

Stations. Every answer needs the calculation shown, not just the result.

Station A — Calculate. “Number of goals scored by a school soccer team in its last matches (primary data, coach’s records)”: .

  1. Find the mean, median, mode, range and IQR.
  2. Describe the shape, with a reason.

Station B — Identify shape.

Data setFrequency pattern
PPeaks in the middle, tapers evenly on both sides
QMost values high, small tail stretching to low values
RValues spread roughly evenly across the whole range, no clear peak
  1. Name the shape of each of P, Q and R.
  2. For Q, would you expect the mean to be above or below the median? Explain.

Station C — Classify and justify.

  1. “A student measures the height of every plant in the school garden themselves.”
  2. “A student uses Bureau of Meteorology archives for the past year’s rainfall.”
  3. For each, state whether the source is primary or secondary, and give one reason this choice matters for how much you’d trust a report based on it.

Station D — Critique.

  1. A report says: “The average score was . Most students did well.” Identify two things missing from this report.

(Answers: 1. Mean ; median ; mode ; range ; , , IQR . 2. Right-skewed — mean above median, pulled up by the . 3. P — symmetric; Q — skewed left; R — uniform. 4. Below — a left tail of low values pulls the mean down below the median. 5. Primary. 6. Secondary. 7. Primary data lets you check the method yourself; secondary data relies on trusting someone else’s collection method and sample — worth noting explicitly in a report. 8. No spread reported (so we don’t know how consistent scores were); no sample size or source stated; “did well” from a mean alone could hide a skewed distribution with some very low scores.)

Activity 2 — Applied Problem: the Wellbeing Committee Investigation (16 min)

Pairs. A single connected problem, worked in stages.

Your school’s Wellbeing Committee wants to know whether Year 8 students at this school sleep less than Year 8 students nationally, and whether a “sleep awareness” campaign is worth running.

Primary data — nightly sleep hours for a class survey of students (sorted):

Secondary data — a published national wellbeing survey of Year 8 students reports: mean h, median h, IQR h ( h), roughly symmetric shape.

  1. Calculate the mean, median and IQR of the school’s primary sample.
  2. Compare the school sample’s centre and spread with the national secondary figures.
  3. Describe the shape of the school sample, and justify your description.
  4. Write a recommendation for the Wellbeing Committee, using both data sets and stating any limitations.

Socratic scaffolding (Polya cycle):

PromptPurpose
Understand the problem. What exactly is the committee trying to decide?Whether this school’s students sleep less than the national picture, using the best available evidence from both sources.
What do you know from Lessons 107–108 about comparing distributions?Compare shape, then centre (justified by shape), then spread — not just one number in isolation.
Devise a plan. How will you calculate the school’s statistics?Sort the data (already done), then apply the mean, median and IQR formulas from Lesson 107.

| Carry out the plan. |

|

| Look back. Does your comparison make sense, and what should the committee be cautious about? | School mean ( h) and median ( h) are both noticeably below the national figures ( h, h), and the school’s spread (IQR h) is wider than the national figure ( h) — but the school sample is only from one class, far smaller than the national , so the comparison should be treated as suggestive, not conclusive. |

Reference recommendation (teacher, one valid version): “This school’s sample shows lower average sleep (mean h, median h) and more variable sleep (IQR h) than the national figures (mean h, median h, IQR h). This supports investigating a sleep awareness campaign, though the school sample is small (, one class) compared with the national sample (), so a larger school-wide sample should be collected before drawing a firm conclusion.”

Checks for Understanding

(7 minutes — exit ticket, collected)

  1. A data set has mean and median . Describe the likely shape.
  2. Calculate the mean, median and range for: .
  3. Explain why the median is a better “typical value” than the mean for the data set in Q2.
  4. A school survey (primary, ) finds median screen time h; a national report (secondary, ) finds median screen time h. Which median would you trust more as an estimate for “all Australian Year 8 students,” and why?
  5. Reasoning. Explain why a complete report should state both a measure of centre and a measure of spread, using an example of what could be missed by reporting only one.

Answers: 1. Skewed left — mean well below median, pulled down by a tail of low values; 2. Mean ; median ; range ; 3. The mean is inflated by the outlier , while the median is unaffected and better reflects most of the values, which cluster between and ; 4. The national report — its far larger, broader sample ( vs ) makes it more representative of the whole population, whereas the school’s small sample only really describes that school; 5. Two data sets can share the same centre but have very different spreads (e.g. Lesson 107’s household size example, or Lesson 108’s commute-time comparison) — reporting only the centre would hide important differences in consistency or variability.

Common Misconceptions

MisconceptionHow to pre-empt it
Comparing only the means of two samples, ignoring shape and spread.The Wellbeing Committee task explicitly requires comparing shape, centre and spread.
Treating a small primary sample as equally reliable as a large secondary sample for general claims.Exit Q4 and the “look back” step of the Polya scaffold — sample size affects how far a conclusion can be generalised.
Believing skewed data has no valid mean, only a median.The mean can still be calculated and is meaningful — it’s just not the best single summary of “typical” for skewed data.
Assuming any conclusion drawn from one small sample is a proven fact.Reference recommendation explicitly flags the sample-size limitation rather than stating a firm conclusion.
Ignoring the “critique” step — accepting a report’s claim without checking what evidence backs it.Station D and CFU Q5 both require identifying what is missing before trusting a claim.

Enrichment — Competition-Style Problems

E1 (Kangaroo style). Two data sets each have values. Set A has mean , median ; Set B has mean , median . Which set is more likely to contain an unusually high outlier? Explain.

Answer

Set B — its mean is far above its median, a strong sign that one or more high values are pulling the mean upward, while Set A’s mean and median being equal is more consistent with a symmetric distribution (though not a guarantee, as in Lesson 108’s E3).

E2 (AMC Junior style). A sample of values has IQR , with . A second sample of the same variable, collected from a much larger secondary source, has IQR . Explain what this difference in IQR suggests about the two samples, assuming both are genuinely random.

Answer

The smaller primary sample shows more spread in its middle than the larger secondary sample — this could reflect genuine differences in the two populations, or simply that a small sample is more prone to showing an unusually wide (or narrow) spread by chance than a much larger one.

E3 (Challenge). A report claims: “Since our sample’s median () matches the national median () exactly, our sample must be identical to the national data.” Explain the flaw in this reasoning.

Answer

Matching medians only means the two data sets share the same centre — they could have completely different shapes and spreads (e.g. one tightly clustered, one widely spread; one symmetric, one skewed). A single matching statistic never proves two distributions are “identical.”

E4 (Investigation). A school wants to compare Year 8 and Year 10 reaction times using an app (secondary data, published by the app company) against their own stopwatch trial (primary data). Propose one advantage and one disadvantage of each source for this specific investigation.

Answer

App data (secondary): advantage — likely a much larger sample, giving a more stable estimate; disadvantage — the school cannot verify exactly how or under what conditions the app’s data was collected, or whether its users are similar to their own students. Stopwatch data (primary): advantage — collected under known, controlled conditions specific to this school; disadvantage — likely a much smaller sample, more prone to chance variation, and possibly less precise than app-based timing.

Homework

  1. A data set has mean and median . What does this suggest, and what would you still want to check before concluding the distribution is symmetric?
  2. Calculate the mean, median, mode, range and IQR for: .
  3. Describe the shape of the Q2 data set, with justification.
  4. A council’s primary survey ( residents) finds support a new bike lane; a state government secondary report ( residents across many suburbs) finds support. Which figure is more useful for predicting state-wide support, and why?
  5. Reasoning. Explain why comparing two samples’ shapes (not just their centres) can change how you interpret a difference in their means.
  6. Reasoning. A student writes: “Our sample’s IQR was bigger than the national IQR, so our data collection must have been done badly.” Explain what is wrong with this conclusion.
  7. Challenge. A school of students takes a primary sample of for a wellbeing survey and finds mean sleep h. A secondary national report of finds mean sleep h with IQR h. Design one improvement to the school’s data collection that would make a future comparison more reliable, and explain why.

Answers: Q1 — this is consistent with a roughly symmetric distribution, but does not prove it; you’d want to check the actual spread and shape (e.g. a dot plot) since matching mean and median can still occur in some non-symmetric cases. Q2 — mean ; median ; mode ; range ; , , IQR . Q3 — mildly right-skewed: mean slightly above median, with a tail toward the higher value . Q4 — the state report; its far larger, more geographically spread sample better represents the whole state, whereas the council’s small local sample may reflect only that particular area’s views. Q5 — e.g. two samples could have the same mean, but if one is symmetric and one is heavily skewed, the “typical” value each mean represents is different, so a raw comparison of means alone can be misleading without also considering shape. Q6 — a larger IQR does not necessarily mean the data collection was flawed; it could reflect genuine variability in a smaller or more diverse sample, or simply that a smaller sample is more prone to showing an unusually wide spread by chance — poor method and natural variation are not the same thing. Q7 — e.g. increase the sample size beyond (ideally sampling across the whole school, not just one class) using a random sampling technique from Lesson 106, since a larger, more representative sample would give a more stable, trustworthy estimate to compare against the national figures.