Lesson 116 — Making Inferences from Sample Data

Strand: Statistics | Descriptor: AC9M8ST04 | Duration: 45 minutes

Block note. Stage 4 (Infer) of the investigation. In Lesson 115, the stratified sample of responses gave a combined sample proportion of meeting the screen-time guideline. Today we use this sample proportion to make a responsible inference about all Year 8 students at our school — the population we could not fully survey.

Learning Intentions

  • To use a sample proportion to estimate a population proportion and population count.
  • To make inferences responsibly, recognising that a sample estimate will rarely match the true population value exactly.

Success Criteria

I can:

  1. Explain what it means to “make an inference” in a statistical investigation.
  2. Use a sample proportion to estimate a population proportion and a population count.
  3. Explain why a point estimate must be treated as approximate, using sample-size reasoning from Lessons 111–112.
  4. Compare the reliability of an inference from a small subgroup with an inference from the full combined sample.

Warmup

(5 minutes — rapid-fire, mini whiteboards)

  1. A sample finds from a population of . Estimate the count meeting the condition.
  2. A sample finds from a population of . Estimate the count.
  3. Recall from Lesson 115 that our combined sample gave . Make a first guess: about how many of the Year 8 students meet the guideline?
  4. Two students both estimate the count using , but round differently to and students. Does this difference matter? Explain.

Answers: 1. ; 2. ; 3. Roughly (exact calculation comes in Activity 1); 4. Not much — since a person must be a whole number, some rounding judgement is unavoidable, and the estimate should be described as “about ”, not stated as if it were exact to the individual.

Activities

Activity 1 — Explicit Instruction: from Sample Proportion to Population Inference (11 min)

What “inference” means: using what we know from a sample — which we can fully measure — to draw a reasoned conclusion about a population, which we cannot.

I do — our investigation.

Say aloud: “We estimate that about of Year 8 students — around of the — meet the screen-time guideline. This single best guess is called a point estimate. It is our best available estimate, not a guaranteed exact fact.”

Key idea to anchor: a different sample of different students would almost certainly give a slightly different — exactly what Lessons 111–112 showed about repeated samples of the same size.

We do — a café example. A café surveys diners; say they prefer takeaway. The café serves roughly diners a week. Estimate how many of the prefer takeaway.

You do:

  1. A sample gives from a population of . Estimate the count.
  2. A sample of gives successes; the population is . Estimate the count.
  3. A sample of gives ; the population is . Estimate the count.

(Answers: 1. ; 2. , estimate ; 3. .)

Activity 2 — Guided Inquiry: how Much Can We Trust Our Estimate? (16 min)

Pairs, using the two-way table from Lesson 115.

Class-by-class population estimates, if each class’s own were (wrongly) trusted alone:

ClassEstimated population count ()
A
B
C
D
E
F
G
Combined (41 responses)
  1. Calculate the range of the class-level estimates in the table above.
  2. This range runs from (Class C) to (Class D) — a difference of students, out of a population of only . Explain why this does not mean the true population count could genuinely be anywhere in that range.
  3. Using ideas from Lessons 111–112, explain why the combined estimate of (from responses) is far more trustworthy than any single class’s estimate (from only responses).
  4. A classmate claims: “Class D found , so at least Year 8 students must be heavy screen users.” Identify the flaw in this reasoning.
  5. Suggest one change that would make the combined estimate itself even more reliable.

Circulating prompts:

PromptPurpose
How many responses is each class estimate built on? How many is the combined estimate built on?Surfaces the versus contrast directly.
What did Lessons 111–112 show about the range of across many same-size small samples, even from one identical population?Connects today’s spread to already-known sampling variation, not a real difference between classes.
Would you bet on Class D’s estimate or the combined estimate to be closer to the true value? Why?Forces an explicit reliability judgement, not just a calculation.
What word should replace “must be” in the flawed claim?Builds the “estimate”, “suggests”, “likely” vocabulary needed for Lesson 117.

(Answers: 1. ; 2. The spread reflects ordinary sampling variation from very small subgroup samples ( each) — not genuinely different populations; 3. The combined estimate pools nearly seven times as much data, so, exactly as in Lessons 111–112, it is far less likely to have been thrown off by chance; 4. “Must be” overclaims certainty from a sample of only students — Class D’s high result could easily be ordinary sampling variation, not a genuine subgroup difference; 5. E.g. survey more students per class, survey on more than one day, or repeat the whole investigation and combine results.)

Activity 3 — Inquiry: Responsible Inference Language (8 min)

Pairs.

Sort each statement as appropriately hedged or overclaiming, then rewrite any overclaiming statement responsibly.

  1. “Exactly Year 8 students meet the guideline.”
  2. “Based on our sample, we estimate that about of Year 8 students meet the guideline.”
  3. “All Year 8 students spend too much time on screens.”
  4. “Our sample suggests roughly a third of Year 8 students meet the guideline, though the true figure could be somewhat higher or lower.”

Answers: 1. Overclaiming — rewrite: “We estimate that about Year 8 students meet the guideline.” 2. Appropriately hedged. 3. Overclaiming and unsupported by the data — the investigation only measured whether students meet a specific threshold, not a value judgement about “too much”; rewrite: “Our sample suggests a substantial proportion of Year 8 students exceed the recommended screen-time guideline.” 4. Appropriately hedged.

Checks for Understanding

(5 minutes — exit ticket, collected)

  1. A sample of gives . Estimate the count in a population of .
  2. Explain the difference between a sample proportion and a population proportion.
  3. Our investigation’s combined sample gave (). Estimate how many of the Year 8 students meet the guideline.
  4. Why is the combined -response estimate more trustworthy than any one class’s estimate alone?
  5. Reasoning. A friend claims: “Exactly Year 8 students meet the guideline.” Explain what is wrong with this claim and rewrite it responsibly.

Answers: 1. ; 2. The sample proportion is calculated from the students actually surveyed; the population proportion is the (unknown) true value for every one of the students — the sample proportion is only an estimate of it; 3. students; 4. It is built from nearly seven times as much data, making it far less likely to have been distorted by ordinary sampling variation; 5. It states a sample-based estimate as if it were an exact, guaranteed fact; better: “We estimate that about Year 8 students meet the guideline, based on our sample.”

Common Misconceptions

MisconceptionHow to pre-empt it
Stating a point estimate as an exact, certain fact (“exactly students”).Model “estimate”, “about”, “based on our sample” in every worked example; challenge any absolute language.
Trusting a small subgroup’s estimate (e.g. one class) as much as the combined estimate.Activity 2’s side-by-side table showing the spread against the stable combined estimate of .
Believing the wide class-to-class range means the true population value is genuinely uncertain across that whole range.Explicit link back to Lessons 111–112: small- subgroup spread is expected noise, not real variation.
Multiplying by the wrong total (sample size instead of population size, or vice versa) when scaling an estimate.Always label = population, = sample size, before calculating; check the answer is sensible relative to .
Assuming a larger sample makes an estimate exactly correct, rather than just more reliable.Reinforce “more trustworthy”, never “guaranteed correct” — mirrors the language used in Lessons 111–112.

Enrichment — Competition-Style Problems

E1 (Kangaroo style). A sample of gives . Estimate the count in a population of .

Answer

E2 (AMC Junior style). A researcher estimates that out of a population of meet a condition, based on a sample proportion . Find as a percentage.

Answer

E3 (Challenge). If instead of the full -response sample, our investigation had used only Class C’s data ( Y out of ), find the resulting population estimate, and state by how much it differs from the true combined estimate of .

Answer

This differs from the combined estimate of by students — over half of the combined estimate — showing how misleading a single small class’s data would be if trusted alone.

E4 (Challenge). A population has . Two different subgroup samples give estimates of and students respectively. Find the two subgroups’ sample proportions, and explain why averaging these two proportions would not be a valid way to find the true combined estimate (refer to Lesson 115).

Answer

Averaging proportions () ignores each subgroup’s actual sample size — Lesson 115 established that combined proportions must be found by adding raw counts (successes and totals) first, not by averaging percentages, unless the subgroups happen to be exactly equal in size.

Homework

  1. A sample of gives . Estimate the count in a population of .
  2. A sample of has successes; the population is . Estimate the count.
  3. A sample of gives . Estimate the count in a population of .
  4. Using our investigation’s combined figures ( Y out of , population ), show the full calculation for the estimated population count.
  5. Explain, in one sentence, why a point estimate is described as a “best guess” rather than a certainty.
  6. State two reasons the combined -response estimate is more trustworthy than Class G’s estimate alone ().
  7. Reasoning. A school of students takes a sample of and finds . A second school of the same size takes a sample of and also finds . Both estimate the same population count. Explain why the second school’s estimate should still be considered more trustworthy.
  8. Reasoning. Explain why rounding a population count estimate (e.g. to ) is reasonable, but rounding the sample proportion too early, before scaling up, can introduce unnecessary error. Illustrate with our investigation’s numbers.
  9. Challenge. A population of students has an unknown true proportion. Two independent samples of size each are combined into one sample of . Sample 1 gives ; Sample 2 gives . Find the combined estimate for a population of , and explain why this differs from simply averaging the two population estimates found from each sample alone.

Answers: Q1 — . Q2 — , estimate . Q3 — . Q4 — ; estimate . Q5 — because a different sample would very likely give a slightly different result, so the figure is our best available approximation, not a guaranteed exact value. Q6 — the combined estimate is built from far more data (41 vs 5 responses) and, per Lessons 111–112, larger samples show far less sampling variation, making the estimate more stable and reliable. Q7 — even though both give the same point estimate, the second school’s larger sample () is far less likely to have been thrown off by ordinary sampling variation than the first school’s small sample (), so it deserves more confidence despite an identical result. Q8 — using the full unrounded fraction () before multiplying by keeps the calculation accurate; rounding early (e.g. to ) and then multiplying can shift the final estimate slightly, whereas rounding only the final count is a reasonable, harmless simplification since a person must be a whole number. Q9 — combined counts: successes out of , ; estimate . This differs from averaging the two separate population estimates ( and , averaging to — coincidentally equal here only because the two samples were equal in size; with unequal sample sizes the two methods would give different answers, and combining raw counts first is the mathematically correct approach.