Lesson 118 — Problem Solving and Consolidation: The Statistical Investigation

Strand: Statistics | Descriptor: AC9M8ST04 | Duration: 45 minutes

Block note. The final lesson of the Lessons 113–118 investigation. Together the class has posed a precise proportion question (113), designed an ethical and fair stratified sample (114), collected and organised responses (115), made a responsible inference — , roughly of the Year 8 students (116) — and written a report acknowledging uncertainty and limitations (117). Today consolidates all five stages, first by reviewing our own investigation, then by applying the full cycle to a brand-new context.

Learning Intentions

  • To consolidate all five stages of a statistical investigation into one coherent understanding.
  • To apply the investigation cycle, end to end, to a new context.

Success Criteria

I can:

  1. Explain the purpose of each of the five stages of a statistical investigation.
  2. Solve short problems from any stage of the cycle, correctly and efficiently.
  3. Apply the full cycle — pose, plan, collect, infer, report — to a new investigation.
  4. Judge the reliability of an estimate and communicate a finding responsibly.

Warmup

(6 minutes — which stage failed? pairs)

Each scenario shows an investigation going wrong. Name the stage (pose, plan, collect, infer, report) where the failure happened.

  1. “Do you think screens are bad?” was used as the survey question.
  2. The survey was only sent to the school’s chess club members.
  3. A decline was recorded as a “no” answer.
  4. A class of students’ result was presented as the final answer for all students.
  5. The report stated ” of students are definitely lazy” based on an exercise survey.

Answers: 1. Pose — an opinion question, not a measurable, testable variable; 2. Plan — a biased, unrepresentative sampling method; 3. Collect — a decline must be excluded, not coded as “no”; 4. Infer — a small subgroup wrongly treated as if it were the trustworthy, combined estimate; 5. Report — overclaiming certainty and drawing an unsupported value judgement.

Activities

Activity 1 — Mixed Review Circuit: All Five Stages (15 min)

Stations, working through our own investigation. Every answer needs the working shown.

Station A — Pose.

  1. “Do students care about the environment?” — is this a good proportion question? If not, rewrite it.
  2. Write a precise proportion question of your own about a topic of your choice.

Station B — Plan.

  1. Which is the fairest method for our population of students across classes: (a) survey your own friends, (b) stratified random sample across all classes, (c) survey whoever answers an online link first? Justify.
  2. Identify the ethical issue: “Record each student’s name next to their answer, so we can double-check later if needed.”

Station C — Collect.

  1. Tally this raw data and find : Y, N, Y, Y, N, N, Y.
  2. A group of students is surveyed but one declines. State the correct denominator for calculating .

Station D — Infer.

  1. A sample gives from ; the population is . Estimate the population count.
  2. Explain why an estimate built from is less trustworthy than one built from , referring to our own investigation.

Station E — Report.

  1. Spot the overclaim: “The data proves students don’t meet the guideline.” Rewrite it responsibly.
  2. Name one limitation of a self-reported survey, and its likely effect.

(Answers: 1. No — “care about” is an opinion, not measurable; e.g. “What proportion of students place a recycling bin out at home every fortnight?” 2. Student’s own, checked against Lesson 113’s four properties. 3. (b) — it guarantees every class is represented, unlike (a) friendship-group bias or (c) voluntary-response bias. 4. Breaks anonymity; names are unnecessary for a proportion investigation and could discourage honest answers. 5. Y, N, . 6. (the surveyed minus the decline). 7. . 8. A sample of carries far less sampling variation than a sample of , per Lessons 111–112, so its estimate is far more likely to be close to the true population value. 9. “Proves” overclaims; rewrite: “We estimate that approximately students do not meet the guideline, based on our sample.” 10. E.g. students may over- or under-report their own habits, shifting the estimate in an unknown direction.)

Activity 2 — Applied Problem: a New Investigation, Start to Finish (17 min)

Pairs. A single connected problem spanning all five stages.

A neighbouring school ( students across six classes of ) investigates: “What proportion of students bring a reusable water bottle to school on a typical day?” A stratified random sample of students per class ( selected) is surveyed. One selected student in Class U declines, leaving responses.

  1. Pose check: is this question precisely defined? What would need clarifying about “reusable water bottle” for two surveyors to always agree?
  2. Plan check: name the sampling method used, and one ethical safeguard it should include.
  3. Collect: tally each class, complete a two-way table, and calculate the combined sample proportion for all responses.
  4. Infer: estimate how many of the students bring a reusable water bottle.
  5. Compare: find the lowest and highest class-level estimates. How do they compare with the combined estimate — and which should the school trust?
  6. Report: write one complete, appropriately hedged sentence stating the finding.

Two-way table (complete the blanks):

ClassSurveyedRespondedYN
P55???
Q55???
R55???
S55???
T55???
U54???
Total3029???

Socratic scaffolding for parts 5–6 (Polya cycle):

PromptPurpose
Understand: what exactly are you comparing?The reliability of a single small class’s estimate against the combined estimate, not just which number looks bigger.
What do you know from Lessons 111–112 and 116 about small-sample estimates?Subgroup samples of only vary widely by chance, even from one identical population.
Devise a plan: which figure should the school’s report lead with?The combined estimate — built from nearly six times as much data as any single class.
Carry it out: draft the report sentence.Must include the estimate, the word “estimate” or “approximately”, and the sample size.
Look back: does your sentence pass Lesson 117’s overclaiming check?Final quality check before submitting the answer.

(Answers: 1. Needs a precise operational definition, e.g. “carries a bottle intended for refilling, on that specific day” — excludes single-use bottles and distinguishes “owns one” from “brought it today.” 2. Stratified random sampling; should include anonymity (no names recorded) and the right to decline. 3. P: Y/N (); Q: Y/N (); R: Y/N (); S: Y/N (); T: Y/N (); U: Y/N (, ). Totals: Y ; N ; responded ✓. Combined . 4. students. 5. Lowest: Class S, ; highest: Classes P/R/T, — a spread of students around a combined estimate of ; the combined estimate is far more trustworthy, since it draws on nearly six times the data of any one class. 6. E.g. “Based on a stratified sample of responses, we estimate that approximately of the school’s students bring a reusable water bottle on a typical day.“)

Checks for Understanding

(7 minutes — exit ticket, collected)

  1. Name the five stages of a statistical investigation, in order.
  2. Our own investigation’s combined sample gave (, population ). State the estimated population count.
  3. Explain why a stratified sample is generally fairer than a convenience sample.
  4. Rewrite responsibly: ” of students definitely don’t recycle.”
  5. Reasoning. Using ideas from across Lessons 113–117, explain why a single class’s data should never be reported as if it represents the whole Year 8 population.

Answers: 1. Pose, plan, collect, infer, report; 2. students; 3. It deliberately samples from every subgroup (class), rather than relying on chance or ease of access, which removes the specific risk of a whole subgroup being missed entirely; 4. E.g. “We estimate that approximately of students surveyed do not recycle.”; 5. A single class is a very small sample (), so its proportion can swing widely from the true population value purely by ordinary sampling variation (Lessons 111–112); only the much larger combined sample gives a reliable enough estimate to report about the whole population (Lesson 116), and reporting otherwise would overclaim certainty the data does not support (Lesson 117).

Common Misconceptions

MisconceptionHow to pre-empt it
Confusing a proportion question with a mean question when posing the investigation.Warmup Q1’s direct scenario; revisit Lesson 113’s four-property test.
Choosing a convenience sample because it is faster, without weighing fairness.Station B’s justified-choice requirement in Activity 1.
Coding a decline as a “no” response when tallying data.Station C and Activity 2’s two-way table both require excluding declines from the denominator.
Treating a sample-based point estimate as an exact, certain population fact.Station D and Activity 2 Q4–5 both require “estimate” language and reliability reasoning.
Reporting findings with no acknowledgement of uncertainty or limitations.Station E and the Socratic “look back” step both require passing Lesson 117’s overclaiming check.
Believing a large spread between small subgroups reflects genuine differences rather than sampling noise.Activity 2 Q5’s explicit -to- spread contrasted with the stable combined estimate of .

Enrichment — Competition-Style Problems

E1 (Kangaroo style). A stratified sample of students (six from each of six classes) finds successes overall. Estimate the count in a population of .

Answer

E2 (AMC Junior style). A report states a sample of gave , and a second, independent sample of from the same population gave . Which estimate should be trusted more, and by how much data is it supported?

Answer

The sample should be trusted more — it is supported by four times as much data as the sample, and per Lessons 111–112, larger samples show far less sampling variation.

E3 (Challenge). A population of is sampled with a stratified method across groups of each (). Four groups give , and the fifth gives . Find the combined sample proportion, and explain why the combined estimate is much closer to than to .

Answer

The combined estimate sits closer to because four of the five groups agree at that value; the single outlying group () has the same weight (one-fifth of the data) as any other group, and cannot pull the combined estimate all the way to its own value.

E4 (Challenge). A school report claims: “Our sample of out of students is too small to trust.” A rival report claims: “Our sample of out of students is guaranteed accurate.” Evaluate both claims using this unit’s ideas.

Answer

Neither is quite right. A well-chosen random or stratified sample of is not “too small to trust” outright — it gives a usable, if less precise, estimate; size affects reliability, not validity. The second claim overclaims — even a sample of out of (a very large sample) is not “guaranteed accurate”; it only makes the estimate considerably more likely to be close to the true value, exactly as established in Lessons 111–112 and 116.

E5 (Investigation). Design, in outline, a full five-stage investigation plan (one sentence per stage) to answer: “What proportion of students at a -student school travel to school by bus?” Include a plausible stratified sample size and a report sentence with appropriate hedging.

Answer

Student’s own; should include: Pose — a precise yes/no test for “travels by bus” and a stated population of ; Plan — a stratified random sample (e.g. across year levels or classes) with anonymity and consent built in; Collect — an honest tally excluding any declines; Infer — scaling the combined sample proportion by to estimate a population count; Report — a sentence using “estimate” or “approximately” and naming at least one limitation (e.g. self-reported travel method, single-day snapshot).

Homework

  1. Name the five stages of a statistical investigation and, in one sentence each, what happens at each stage.
  2. A stratified sample of across groups of finds successes. Estimate the population count for .
  3. Identify the flaw and name the affected stage: “The survey asked, ‘Don’t you agree screens are harmful?‘”
  4. A class of students gives ; the combined sample of students gives . Which should a report cite as the school’s estimate, and why?
  5. Rewrite responsibly: “This proves that of Year 8 eats breakfast.”
  6. State one ethical principle that must apply throughout data collection, and explain why it matters even after the data has been organised and reported.
  7. Reasoning. Explain why “acknowledging uncertainty” is not the same as “being unsure of your method” — use our own investigation as your example.
  8. Reasoning. A friend says the whole five-stage process is “overkill” for a simple yes/no question. Explain, using at least two stages, why skipping steps could produce a misleading result even for a simple question.
  9. Challenge. A school of students, in classes of , plans a stratified sample of students per class (). After collection, students across two different classes decline. Given Y responses out of the total responded, find the combined , the estimated population count, and write a complete, appropriately hedged one-sentence report including the sample size and at least one limitation.

Answers: Q1 — pose (write a precise proportion question, and define population/variable), plan (choose a fair, ethical sampling method), collect (gather and organise data honestly), infer (use the sample proportion to estimate the population), report (communicate the finding, acknowledging uncertainty). Q2 — ; estimate . Q3 — a leading question, priming the respondent toward “yes”; affects the Pose/Plan stage. Q4 — the combined sample () — it is built from eight times as much data as the single class (), making it far more reliable, per Lessons 111–112 and 116. Q5 — e.g. “Our sample suggests that approximately of Year 8 students eat breakfast on a typical school morning.” Q6 — e.g. anonymity — even after reporting, a poorly anonymised dataset could allow individual students to be identified from combinations of details, so privacy must be protected throughout, not just at collection. Q7 — acknowledging uncertainty means honestly stating that a sample-based estimate could differ from the true population value, which is a mathematical fact about any well-run sample-based investigation, not a sign the method itself was flawed; our investigation’s stratified sampling and careful data handling were sound, and the uncertainty statement simply reflects that a sample is not a census. Q8 — e.g. skipping the Plan stage could produce a biased sample (e.g. convenience sampling) that answers the question for the wrong group entirely; skipping the Report stage’s uncertainty acknowledgement could see a shaky small-sample result presented as if it were certain, misleading decision-makers even though the underlying data collection was fine. Q9 — responded ; ; estimate students; e.g. “Based on a stratified sample of responses (out of selected, with declines), we estimate that approximately of the school’s students meet the condition; as with any sample-based estimate, the true figure could reasonably be somewhat higher or lower.”