Lesson 107 — Explicit Instruction: Analysing Distributions from Primary and Secondary Sources
Strand: Statistics | Descriptor: AC9M8ST02 | Duration: 45 minutes
Learning Intentions
- To distinguish primary and secondary data sources.
- To analyse the shape, centre and spread of a data distribution, choosing statistics appropriate to its shape.
Success Criteria
I can:
- Distinguish primary data (collected first-hand) from secondary data (already collected and published by someone else), with examples.
- Describe the shape of a distribution as symmetric, skewed left, skewed right or uniform, from raw data or a frequency table.
- Calculate the mean, median and mode of a data set, and identify which best represents its centre.
- Calculate the range and interquartile range (IQR), and explain why IQR is less affected by outliers than range.
- Identify a likely outlier in a data set and describe its effect on the mean.
Warmup
(6 minutes — primary or secondary? pairs, mini whiteboards)
For each, decide: is this primary data (collected first-hand by the person using it) or secondary data (already collected and published by someone else)?
- A class measures each other’s arm span with a tape measure.
- A student uses ABS Census data to find the average household size in their suburb.
- A researcher downloads five years of Bureau of Meteorology rainfall records for an assignment.
- A group of students counts the number of cars of each colour in the school car park themselves.
Answers: 1. Primary — collected first-hand by the class; 2. Secondary — already collected and published by the ABS; 3. Secondary — collected and published by the Bureau of Meteorology; 4. Primary — collected first-hand by the students.
The key idea: primary data means you controlled the collection; secondary data means someone else did, and you are relying on their methods and choices.
Activities
Activity 1 — Explicit Instruction: Shape, Centre and Spread (15 min)
I do — a primary data set. “Number of sit-ups completed in one minute by 15 Year 8 students in a PE fitness test (primary data — collected first-hand by the class).”
Sorted data:
Median (8th of 15 values)
Notice: the mean (
The key contrast: the range (
Shape reference table:
| Shape | What it looks like | Best measure of centre |
|---|---|---|
| Symmetric | Roughly mirror-image either side of the centre | Mean or median (similar values) |
| Skewed right (positive) | Cluster of low/typical values, long tail stretching to high values | Median (mean is pulled up) |
| Skewed left (negative) | Cluster of high/typical values, long tail stretching to low values | Median (mean is pulled down) |
| Uniform | Values spread roughly evenly across the range, no clear peak | Mean or median (similar values) |
We do — a secondary data set. “Average weekly screen time (hours) reported for a sample of 12 Australian teenagers, from a published national youth wellbeing survey (secondary data).”
Sorted data:
Together: describe the shape (right-skewed — mean above median, tail toward high values) and note the source is secondary.
You do. “Daily minutes of exercise for 10 students, downloaded from a fitness app’s published usage statistics (secondary data).”
Sorted data:
Calculate the mean, median, mode, range and IQR, and classify the data source.
(Answers: mean
Activity 2 — Reading Shape from Frequency Tables (13 min)
We do — quick shape identification. Two classes’ test scores out of
Class A:
| Score | 10–11 | 12–13 | 14–15 | 16–17 | 18–19 | 20 |
|---|---|---|---|---|---|---|
| Frequency | 2 | 4 | 9 | 8 | 5 | 2 |
Class B:
| Score | 4–6 | 7–9 | 10–12 | 13–15 | 16–18 | 19–20 |
|---|---|---|---|---|---|---|
| Frequency | 1 | 2 | 3 | 5 | 10 | 9 |
Discuss: Class A peaks in the middle and tapers off similarly on both sides — roughly symmetric, so mean and median would be similar. Class B has most students scoring high, with a small tail down to low scores — skewed left (negative skew), so the median better represents a “typical” score than the mean, which the low tail would pull down.
You do — a deeper secondary example. “Household size for a sample of 50 dwellings in a suburb, from published ABS Census QuickStats (secondary data).”
| Household size | 1 | 2 | 3 | 4 | 5 | 6 |
|---|---|---|---|---|---|---|
| Frequency | 12 | 17 | 8 | 8 | 3 | 2 |
Since mean (
Activity 3 — Inquiry: Comparing a Primary and a Secondary Source (7 min)
Pairs.
Your class collects primary data on time spent on homework last night, from your own
students. A national wellbeing survey (secondary data) reports the distribution of homework time for thousands of Australian Year 8 students.
- Why might the two distributions differ, even if both were collected carefully?
- Which source better describes your class? Which better describes Year 8 students in Australia generally?
- Could you use both together? How?
Socratic scaffolding:
| Prompt | Purpose |
|---|---|
| Understand: what is each data set actually describing? | Our class ( |
| What might cause the two distributions to differ, even if nothing was done “wrong”? | Different population, different sample size, different collection time and context (e.g. subject load, school culture). |
| Which source better describes “our class”? | The primary data — it was collected from exactly this population. |
| Which source better describes “Year 8 students in Australia”? | The secondary data — a much larger, broader sample intended to represent the whole population. |
| Could you use both together? How? | Compare our class’s shape, centre and spread against the national figures to see whether our class is fairly typical or unusual. |
| Look back: is one source always “better” than the other? | No — it depends on the question being asked; primary data is often more locally relevant, secondary data more relevant for broad claims. |
Checks for Understanding
(4 minutes — exit ticket, collected)
- Classify as primary or secondary: “A group of students records the colour of
cars passing the school gate.” - A data set has mean
and median . Describe the likely shape, and explain your reasoning. - Calculate the IQR for the data set:
. - Reasoning. Explain why the median is often preferred over the mean for reporting income distributions, which are typically strongly right-skewed.
Answers: 1. Primary — collected first-hand by the students; 2. Right-skewed — the mean (
Common Misconceptions
| Misconception | How to pre-empt it |
|---|---|
| Believing the mean is always the “best” or only measure of centre. | Contrast symmetric and skewed shapes explicitly in the shape reference table; require justification for every choice of centre. |
| Believing “secondary” automatically means “less trustworthy” than primary. | Emphasise that large, well-designed secondary sources (like the ABS) are often more reliable than a small primary sample — the key question is how the data was collected, not just who collected it. |
| Treating range and IQR as measuring the same thing. | Direct comparison in Activity 1: range ( |
| Assuming an outlier should always be deleted from a data set. | Outliers should be investigated and reported, not automatically removed, unless known to be a genuine error. |
| Confusing “skewed left” with “most of the data is on the left.” | Skew direction names the direction of the tail, not where most data sits — Class B in Activity 2 has most scores high, with a tail (and skew) to the left. |
| Judging shape from mean and median alone, without checking the actual data or graph. | Mean-versus-median gives a useful clue, but a full shape description should always be checked against the frequency table or dot plot. |
Enrichment — Competition-Style Problems
E1 (Kangaroo style). A data set of
Answer
Let the sorted values be
E2 (AMC Junior style). A data set has
Answer
Since
E3 (Challenge). A data set of
Answer
Yes. Mean equalling median does not guarantee symmetry — a distribution can have unusual values balancing out on both sides of the centre without the overall shape being a mirror image. This shows that comparing mean and median is a useful clue to shape, but not a complete substitute for examining the actual data or a graph.
E4 (Investigation). Household sizes in two suburbs both have median
Answer
Suburb X’s household sizes are tightly clustered around
Homework
- Classify each as primary or secondary, with a one-sentence reason: (a) A student times how long it takes
classmates to solve a puzzle. (b) A student uses AFL match statistics published on the league’s official website. (c) A council reads its own traffic sensor data, collected specifically for this report. (d) A researcher uses a United Nations data set on global life expectancy. - For the data set
, calculate the mean, median, mode, range and IQR. - Describe the shape of the data set in Q2, justifying your answer using the mean and median.
- A distribution has mean
and median . Describe its likely shape. - Reasoning. Explain, using an example, why a company reporting “average salary” might prefer to use the mean even when the median would give a more typical picture.
- Reasoning. A sample of
house prices ( ‘000\text{s} 420, 450, 460, 470, 480, 490, 510, 1200$. Calculate the mean and median, and explain which better represents a “typical” house price in this sample. - Challenge. A data set of
numbers has and (IQR ). A new value of is added, becoming the new maximum, making numbers in total. Explain whether the IQR is likely to change much, using the definition of IQR.
Answers: Q1 — (a) primary; (b) secondary; (c) primary; (d) secondary. Q2 — mean