Lesson 107 — Explicit Instruction: Analysing Distributions from Primary and Secondary Sources

Strand: Statistics | Descriptor: AC9M8ST02 | Duration: 45 minutes

Learning Intentions

  • To distinguish primary and secondary data sources.
  • To analyse the shape, centre and spread of a data distribution, choosing statistics appropriate to its shape.

Success Criteria

I can:

  1. Distinguish primary data (collected first-hand) from secondary data (already collected and published by someone else), with examples.
  2. Describe the shape of a distribution as symmetric, skewed left, skewed right or uniform, from raw data or a frequency table.
  3. Calculate the mean, median and mode of a data set, and identify which best represents its centre.
  4. Calculate the range and interquartile range (IQR), and explain why IQR is less affected by outliers than range.
  5. Identify a likely outlier in a data set and describe its effect on the mean.

Warmup

(6 minutes — primary or secondary? pairs, mini whiteboards)

For each, decide: is this primary data (collected first-hand by the person using it) or secondary data (already collected and published by someone else)?

  1. A class measures each other’s arm span with a tape measure.
  2. A student uses ABS Census data to find the average household size in their suburb.
  3. A researcher downloads five years of Bureau of Meteorology rainfall records for an assignment.
  4. A group of students counts the number of cars of each colour in the school car park themselves.

Answers: 1. Primary — collected first-hand by the class; 2. Secondary — already collected and published by the ABS; 3. Secondary — collected and published by the Bureau of Meteorology; 4. Primary — collected first-hand by the students.

The key idea: primary data means you controlled the collection; secondary data means someone else did, and you are relying on their methods and choices.

Activities

Activity 1 — Explicit Instruction: Shape, Centre and Spread (15 min)

I do — a primary data set. “Number of sit-ups completed in one minute by 15 Year 8 students in a PE fitness test (primary data — collected first-hand by the class).”

Sorted data:

Median (8th of 15 values) . Mode .

Notice: the mean () is higher than the median () — it is being pulled upward by the single high value, . This is a sign of right skew (also called positive skew): most values cluster together, with a smaller number stretching out toward high values.

The key contrast: the range () is far larger than the IQR (), because the range is entirely determined by the two most extreme values. The IQR, based on the middle of the data, barely notices the outlier at — this is why IQR is the preferred measure of spread when a distribution is skewed or has outliers.

Shape reference table:

ShapeWhat it looks likeBest measure of centre
SymmetricRoughly mirror-image either side of the centreMean or median (similar values)
Skewed right (positive)Cluster of low/typical values, long tail stretching to high valuesMedian (mean is pulled up)
Skewed left (negative)Cluster of high/typical values, long tail stretching to low valuesMedian (mean is pulled down)
UniformValues spread roughly evenly across the range, no clear peakMean or median (similar values)

We do — a secondary data set. “Average weekly screen time (hours) reported for a sample of 12 Australian teenagers, from a published national youth wellbeing survey (secondary data).”

Sorted data:

Together: describe the shape (right-skewed — mean above median, tail toward high values) and note the source is secondary.

You do. “Daily minutes of exercise for 10 students, downloaded from a fitness app’s published usage statistics (secondary data).”

Sorted data:

Calculate the mean, median, mode, range and IQR, and classify the data source.

(Answers: mean ; median ; mode ; range ; , , IQR ; secondary data; right-skewed — the stretches the mean above the median, and the range is three times the IQR.)

Activity 2 — Reading Shape from Frequency Tables (13 min)

We do — quick shape identification. Two classes’ test scores out of (primary data, collected by the teacher).

Class A:

Score10–1112–1314–1516–1718–1920
Frequency249852

Class B:

Score4–67–910–1213–1516–1819–20
Frequency1235109

Discuss: Class A peaks in the middle and tapers off similarly on both sides — roughly symmetric, so mean and median would be similar. Class B has most students scoring high, with a small tail down to low scores — skewed left (negative skew), so the median better represents a “typical” score than the mean, which the low tail would pull down.

You do — a deeper secondary example. “Household size for a sample of 50 dwellings in a suburb, from published ABS Census QuickStats (secondary data).”

Household size123456
Frequency12178832

Since mean () median (), this is right-skewed — most households are small, with a long tail toward larger households.

Activity 3 — Inquiry: Comparing a Primary and a Secondary Source (7 min)

Pairs.

Your class collects primary data on time spent on homework last night, from your own students. A national wellbeing survey (secondary data) reports the distribution of homework time for thousands of Australian Year 8 students.

  1. Why might the two distributions differ, even if both were collected carefully?
  2. Which source better describes your class? Which better describes Year 8 students in Australia generally?
  3. Could you use both together? How?

Socratic scaffolding:

PromptPurpose
Understand: what is each data set actually describing?Our class (, one specific group) versus a national sample (thousands of students, many schools).
What might cause the two distributions to differ, even if nothing was done “wrong”?Different population, different sample size, different collection time and context (e.g. subject load, school culture).
Which source better describes “our class”?The primary data — it was collected from exactly this population.
Which source better describes “Year 8 students in Australia”?The secondary data — a much larger, broader sample intended to represent the whole population.
Could you use both together? How?Compare our class’s shape, centre and spread against the national figures to see whether our class is fairly typical or unusual.
Look back: is one source always “better” than the other?No — it depends on the question being asked; primary data is often more locally relevant, secondary data more relevant for broad claims.

Checks for Understanding

(4 minutes — exit ticket, collected)

  1. Classify as primary or secondary: “A group of students records the colour of cars passing the school gate.”
  2. A data set has mean and median . Describe the likely shape, and explain your reasoning.
  3. Calculate the IQR for the data set: .
  4. Reasoning. Explain why the median is often preferred over the mean for reporting income distributions, which are typically strongly right-skewed.

Answers: 1. Primary — collected first-hand by the students; 2. Right-skewed — the mean () is well above the median (), suggesting a small number of high values are pulling the mean upward; 3. Median ; ; ; IQR ; 4. A small number of very high incomes pull the mean well above what is “typical,” while the median is unaffected by extreme values and better represents the income of a typical person.

Common Misconceptions

MisconceptionHow to pre-empt it
Believing the mean is always the “best” or only measure of centre.Contrast symmetric and skewed shapes explicitly in the shape reference table; require justification for every choice of centre.
Believing “secondary” automatically means “less trustworthy” than primary.Emphasise that large, well-designed secondary sources (like the ABS) are often more reliable than a small primary sample — the key question is how the data was collected, not just who collected it.
Treating range and IQR as measuring the same thing.Direct comparison in Activity 1: range () versus IQR () for the same skewed data set.
Assuming an outlier should always be deleted from a data set.Outliers should be investigated and reported, not automatically removed, unless known to be a genuine error.
Confusing “skewed left” with “most of the data is on the left.”Skew direction names the direction of the tail, not where most data sits — Class B in Activity 2 has most scores high, with a tail (and skew) to the left.
Judging shape from mean and median alone, without checking the actual data or graph.Mean-versus-median gives a useful clue, but a full shape description should always be checked against the frequency table or dot plot.

Enrichment — Competition-Style Problems

E1 (Kangaroo style). A data set of numbers, sorted in order, has median . If the two smallest numbers are removed, is the new median guaranteed to increase, decrease, or could it do either?

Answer

Let the sorted values be , so the median is . Removing and leaves (7 values), whose median is the 4th value, . Since the data is sorted, . So the new median is always greater than or equal to — it can never decrease.

E2 (AMC Junior style). A data set has and IQR . Using the rule that a value is an outlier if it lies more than beyond or , determine whether a value of would be classed as an outlier.

Answer

Since , the value is classed as an outlier.

E3 (Challenge). A data set of values has mean and median , but the distribution is not symmetric. Is this possible? Explain.

Answer

Yes. Mean equalling median does not guarantee symmetry — a distribution can have unusual values balancing out on both sides of the centre without the overall shape being a mirror image. This shows that comparing mean and median is a useful clue to shape, but not a complete substitute for examining the actual data or a graph.

E4 (Investigation). Household sizes in two suburbs both have median people, but Suburb X has IQR while Suburb Y has IQR . What does this tell you about the two suburbs, even though their “typical” household size is the same?

Answer

Suburb X’s household sizes are tightly clustered around — most households are quite similar in size. Suburb Y has far more variety, ranging from small households to large ones. Same centre, very different spread — a reminder that spread matters just as much as centre when describing a distribution.

Homework

  1. Classify each as primary or secondary, with a one-sentence reason: (a) A student times how long it takes classmates to solve a puzzle. (b) A student uses AFL match statistics published on the league’s official website. (c) A council reads its own traffic sensor data, collected specifically for this report. (d) A researcher uses a United Nations data set on global life expectancy.
  2. For the data set , calculate the mean, median, mode, range and IQR.
  3. Describe the shape of the data set in Q2, justifying your answer using the mean and median.
  4. A distribution has mean and median . Describe its likely shape.
  5. Reasoning. Explain, using an example, why a company reporting “average salary” might prefer to use the mean even when the median would give a more typical picture.
  6. Reasoning. A sample of house prices (‘000\text{s}420, 450, 460, 470, 480, 490, 510, 1200$. Calculate the mean and median, and explain which better represents a “typical” house price in this sample.
  7. Challenge. A data set of numbers has and (IQR ). A new value of is added, becoming the new maximum, making numbers in total. Explain whether the IQR is likely to change much, using the definition of IQR.

Answers: Q1 — (a) primary; (b) secondary; (c) primary; (d) secondary. Q2 — mean ; median ; mode ; range ; , , IQR . Q3 — mild right skew: mean () is slightly above median (), consistent with the high value pulling it up. Q4 — mean below median suggests skewed left (negative skew), with a tail toward low values. Q5 — e.g. a company with a few very highly paid executives and many lower-paid staff can report a high mean salary that looks impressive, even though the median (the salary of a typical employee) is much lower; using the mean in a skewed distribution can be misleading. Q6 — mean ; median ; the median better represents a typical price, since the mean is inflated by the 1{,}200{,}000Q_1Q_350%$ of data), not on the maximum or minimum, so adding one extreme value at the top is unlikely to change the IQR by much — this is exactly why IQR is described as resistant to outliers, unlike the range, which would jump considerably.