Lesson 128 — Comparing Predictions About Outcomes with Observed Results

Strand: Probability | Descriptor: AC9M7P02 | Duration: 45 minutes

Learning Intentions

  • To compare predicted probabilities with observed experimental results.
  • To judge whether a difference is ordinary variation or evidence of something else.

Success Criteria

I can:

  1. Calculate a predicted frequency and compare it with an observed one.
  2. Express the difference as both a count and a proportion.
  3. Judge whether a difference is within ordinary variation.
  4. Recognise when a difference is large enough to warrant investigation.

Warmup

(6 minutes — predicted versus observed, mini whiteboards)

For each, state the predicted count and the difference from the observed.

ExperimentTrialsPredictedObservedDifference
Coin — heads
Die — sixes
Spinner ( red)
Coin — heads

Answers: , diff ; , diff ; , diff ; , diff .

The ranking question: which difference is the most striking? Most students say the last — correctly, but the reason matters. from is a relative frequency of ; from is against . Both are proportionally large, but the -trial result is far harder to explain by chance.

Activities

Activity 1 — Explicit Instruction: Judging a Difference (14 min)

Two ways to express a difference, and both are needed:

MeasureFormulaExample
Absoluteobserved predicted heads
Proportionalobserved relative frequency probability

Why both? A difference of means very different things in trials and in .

The judging principle — the rough rule to teach:

Differences shrink, proportionally, as trials increase. A relative frequency within about of the prediction is unremarkable at trials; at trials, that same gap would be surprising.

I do — three comparisons, judged aloud:

Case 1. A die rolled times gives sixes (predicted ).

Judgement: trials is a small run; a gap of is unremarkable. No cause for concern.

Case 2. A die rolled times gives sixes (predicted ).

Judgement: the same proportional gap, but across trials. Chance alone almost never produces this. Strong evidence the die is loaded.

Emphasise the pairing. Cases 1 and 2 have identical relative frequencies and opposite conclusions. The trial count is what separates them — the single most important idea in this lesson.

Case 3. A coin tossed times gives heads (predicted ).

Judgement: strikingly close. Exactly what a fair coin does.

We do — judge these together:

  1. Coin, tosses, heads.
  2. Die, rolls, sixes.
  3. Spinner (), spins, reds.
  4. Coin, tosses, heads.

(Answers: 1. , — unremarkable in trials; 2. , — slightly low but ordinary; 3. , over trials — a large gap on many trials, worth investigating; 4. , — remarkably close, exactly as expected.)

Activity 2 — Comparison Circuit (14 min)

Pairs. Every answer gives both differences and a judgement with a reason.

Set A — compute and judge.

  1. Die, rolls, sixes.
  2. Coin, tosses, heads.
  3. Spinner (), spins, hits.
  4. Bag (), draws, reds.

Set B — same proportion, different scale. All four have relative frequency for a fair coin:

  1. heads in .
  2. heads in .
  3. heads in .
  4. heads in .

Rank them from least to most surprising, and explain your ranking.

Set C — use the class’s own data. From Lesson 126’s pooled coin experiment:

  1. State the class’s total tosses and heads.
  2. Compute the absolute and proportional differences from prediction.
  3. Judge the result.

Socratic scaffolding for Set B:

PromptPurpose
All four have the same relative frequency. Are they equally surprising?No.
Which is least surprising? in — this happens constantly with fair coins.
Which is most? in — chance essentially never does this.
So what makes a result surprising?The size of the deviation and the number of trials together.
A one-line summary?The same proportion becomes more convincing evidence with more trials.

(Answers: 1. predicted ; , — ordinary; 2. predicted ; , — ordinary; 3. predicted ; , — ordinary; 4. predicted ; , over trials — noticeable, worth more data; 5–8. ranked from least to most surprising; 9–11. class’s own.)

Activity 3 — Inquiry: the Suspicious Dice (9 min)

Pairs, then class discussion.

Four students each test a die for fairness by counting sixes.

StudentRollsSixesPredicted
Ana
Ben
Cara
Dev
  1. Complete the predicted column.
  2. Find each relative frequency.
  3. Rank the four by how convincing their evidence of loading is.
  4. Whose die would you actually suspect? Whose would you clear?
  5. What should Ana do next?

Socratic scaffolding:

PromptPurpose
Ana’s relative frequency? — twice the expected !
Is that convincing?No — rolls is far too few. Getting sixes in is uncommon but entirely possible.
Ben’s? — above expectation, on a decent number of trials. Suspicious.
Cara’s? — almost exactly . Her die looks fair.
Dev’s? — essentially perfect. Clearly fair.
Ranking?Ben’s evidence is the most convincing of loading; Ana’s looks most extreme but proves least.
Q5: what should Ana do?Roll several hundred more times. Only more data can distinguish luck from loading.

The closing point: the most extreme-looking result came from the weakest evidence. Judging a result requires looking at the trial count first, not the deviation.

Checks for Understanding

(5 minutes — exit ticket, collected)

  1. A die rolled times gives sixes. Find the predicted count and both differences.
  2. Judge the Q1 result, with a reason.
  3. A coin tossed times gives heads; another tossed times gives heads. Which is stronger evidence of bias?
  4. A spinner with is spun times, landing red times. Comment.
  5. Reasoning. Explain why the same relative frequency can be unremarkable in one experiment and alarming in another.

Answers: 1. Predicted ; absolute ; proportional ; 2. Above expectation on a reasonable number of trials — worth more data before concluding, but not yet alarming; 3. The second — the same proportion () from a hundred times as many trials is far harder to explain by chance; 4. Predicted ; observed — essentially perfect agreement; 5. Because the number of trials determines how much variation chance can produce: small runs vary widely, large runs barely at all.

Common Misconceptions

MisconceptionHow to pre-empt it
Judging by the absolute difference alone.Both measures required in every answer.
Judging by the proportional difference alone.Cases 1 and 2 — identical proportions, opposite conclusions.
Treating an extreme small-sample result as strong evidence.Ana’s die in the inquiry.
Believing an exact match proves fairness.Dev’s near-perfect result is consistent with fairness but does not prove it.
Expecting results to match predictions exactly.Every case shows a gap; the question is how big.
Concluding bias from a single experiment.Every judgement ends with “collect more data” where relevant.

Enrichment — Competition-Style Problems

E1 (Kangaroo style). A coin tossed times gives heads. Find both differences.

Answer

Predicted ; absolute ; proportional — unremarkable.

E2 (AMC Junior style). Two dice are tested for sixes: one gives over rolls, the other over rolls. Which is more likely loaded?

Answer

The second. is a smaller deviation but comes from rolls, where chance produces almost no variation. from rolls is well within luck.

E3 (Challenge). A spinner is claimed to have . In spins there are wins. Assess the claim.

Answer

Predicted ; observed — a shortfall of , and a relative frequency of against . Over trials that gap is substantial and unlikely from chance. The claimed probability appears overstated.

E4 (Challenge). Why is a result exactly matching the prediction not itself suspicious in one experiment, but would be if it happened every time across fifty experiments?

Answer

One exact match is a perfectly ordinary outcome. But chance produces variation, so fifty consecutive exact matches would be extraordinarily unlikely — suggesting the data was fabricated or the process is not random at all. Real data is slightly messy; suspiciously perfect data is a known marker of fraud.

E5 (Challenge). A die is rolled times with these counts: , , , , , . Is the die fair?

Answer

Predicted each; every count lies within of it. This is exactly the pattern a fair die produces — close but not identical. No evidence of bias.

Homework

  1. Complete the table:

    ExperimentTrialsProbabilityPredictedObservedAbsolute diffProportional diff
    Coin — heads
    Die — fives
    Spinner — red
    Bag — blue
  2. For each row in Q1, write a one-sentence judgement.

  3. Rank from least to most surprising, all for a fair coin: heads in ; in ; in ; in . Explain your ranking.

  4. A die rolled times gives sixes. (a) Predicted count. (b) Both differences. (c) Your judgement.

  5. Two students test the same spinner (). One gets from spins; the other from . Whose result is more concerning?

  6. A coin tossed times gives exactly heads. Is this proof the coin is fair? Explain.

  7. Reasoning. Explain why the trial count must be considered before the size of a deviation.

  8. Reasoning. A student says “my die gave sixes in rolls, so it’s loaded.” Explain the flaw.

  9. Reasoning. Explain why data that matches predictions too perfectly can be suspicious.

  10. Challenge. Design an experiment to test whether a spinner is fair, stating how many trials you would use and what result would convince you of bias.

Answers: Q1 — coin: predicted , , ; die: predicted , , ; spinner: predicted , , ; bag: predicted , , . Q2 — all four are within ordinary variation; none warrants suspicion. Q3 — same order as listed: more trials make the same proportion more surprising. Q4 — (a) (b) ; (c) a substantial excess over trials — worth investigating, and more rolls would settle it. Q5 — the second: from spins is a smaller deviation but from far more data, so chance is a much less plausible explanation than for from . Q6 — no. An exact match is one ordinary outcome among many; it is consistent with fairness but does not prove it, and a biased coin could produce it too. Q7 — the trial count determines how much variation chance can generate; without it, a deviation cannot be interpreted at all. Q8 — rolls is far too few; sixes in arises from fair dice regularly. Q9 — genuine random data varies; results matching predictions exactly and repeatedly suggest the data was invented rather than collected. Q10 — e.g. spin at least times, compare each sector’s relative frequency with , and treat a consistent deviation of more than about across that many trials as convincing — while noting that a single sector being slightly off proves little.