Lesson 115 — Collecting and Organising Sample Data

Strand: Statistics | Descriptor: AC9M8ST04 | Duration: 45 minutes

Block note. Stage 3 (Collect) of the investigation. The stratified sample designed in Lesson 114 — six students randomly selected from each of the seven Year 8 classes — has now been surveyed. One selected student exercised their right to decline, leaving responses. Today’s job is to organise this raw data honestly and calculate the sample proportion, ready for the inference stage in Lesson 116.

Learning Intentions

  • To organise raw survey data into frequency and two-way tables.
  • To calculate a sample proportion from organised data, and recognise data quality issues.

Success Criteria

I can:

  1. Tally raw yes/no responses into a frequency table without error.
  2. Calculate a sample proportion as a fraction, decimal and percentage.
  3. Organise data into a two-way table by subgroup (e.g. class) and compare subgroup proportions.
  4. Identify and handle data quality issues honestly, including a decline.

Warmup

(5 minutes — spot the data-entry error, pairs)

A pair tallies raw Y/N responses and reports: ” said yes, said no.”

  1. What is wrong with this report?
  2. The correct tally shows yes and no. What is the sample proportion who said yes?
  3. Why does checking that counts add to the total sample size matter before calculating any proportion?

Answers: 1. — the counts do not add to the total sample size, so at least one tally was miscounted; 2. ; 3. An error here would make every later calculation — the proportion, the inference, the report — wrong, even if the arithmetic from that point on is perfect.

Activities

Activity 1 — Explicit Instruction: from Raw Data to a Sample Proportion (14 min)

I do — tallying Class A’s raw responses. Six students were surveyed: Y = meets the screen-time guideline, N = does not.

ResponseTallyCount
Y||
N||||
Total

The habit to model aloud: always check before trusting the proportion.

We do — Class B’s raw responses, together:

(Tally: Y, N, total . .)

You do — Class C’s raw responses:

Tally the responses, check the total, and calculate .

(Answer: Y, N, total . .)

Activity 2 — Guided Practice: Organising the Full Sample (14 min)

Pairs. The remaining four classes’ raw data, plus one important quality issue.

  1. Tally each class and check its total against the number surveyed (remember: Class G has only usable responses).
  2. Complete the two-way table below.
  3. Calculate the overall sample proportion for all responses.
  4. Calculate each class’s individual proportion.
  5. Which class had the highest proportion meeting the guideline? Which had the lowest?

Two-way table — complete the blank cells:

ClassSurveyedRespondedMeets guideline (Y)Does not (N)
A6624
B6633
C6615
D66???
E66???
F66???
G65???
Total4241???

Circulating prompts:

PromptPurpose
Does your row’s Y and N add to the number responded, not the number surveyed?The Class G decline must be handled without inflating or shrinking the wrong total.
What do you do with a decline — count it as “N”, ignore it, or something else?It must be excluded from both the numerator and denominator, honestly reported as a decline, never coded as a “no”.
Does your overall total () match the sum of each class’s responded count?The same cross-check as the warmup, now at full scale.

(Answers: D: Y, N, . E: Y, N, . F: Y, N, . G: Y, N, . Totals: Y ; N ; total responded ✓. Overall . Highest: Class D (). Lowest: Class C ().)

Activity 3 — Inquiry: Does the Class-to-class Spread Mean Anything? (7 min)

Pairs, using the completed two-way table.

Class proportions range from (Class C) to (Class D) — a huge spread, from a sample of only per class.

  1. Does this spread mean some classes are genuinely much more screen-aware than others?
  2. Connect this to Lessons 111–112: what do you expect from any set of same-size samples of only , even from an identical population?
  3. Which is a more trustworthy estimate of the whole Year 8 proportion: one class’s , or the combined -student ? Why?

Answers: 1. Not necessarily — with only students per class, ordinary sampling variation alone can easily produce swings this large, exactly as seen with the small- samples in Lesson 111; 2. A wide range of proportions, purely from chance, even if every class had an identical true rate; 3. The combined -student estimate — pooling all classes gives a much larger effective sample than any one class alone, and Lessons 111–112 showed that larger samples produce far less erratic proportions.

The teaching point to close on: organising data by subgroup is valuable for spotting genuine patterns, but a Year 8 statistician must ask, before concluding a subgroup is different, whether ordinary sampling variation alone could explain what is seen. That question is answered properly in Lesson 116.

Checks for Understanding

(5 minutes — exit ticket, collected)

  1. A class’s raw data is: Y, N, Y, Y, N. Tally it and find .
  2. Explain, in one sentence, how a “decline” should be handled when calculating a sample proportion.
  3. Two subgroups of a two-way table show and . Does this necessarily mean the subgroups are genuinely different? Justify briefly.
  4. If Class A ( Y out of ) and Class B ( Y out of ) are combined into one group of , find the combined .
  5. Reasoning. Explain why checking that tallied counts add to the correct total is an essential step, not just a tidiness habit.

Answers: 1. Y, N, ; 2. A decline is excluded from both the numerator and the denominator — it is not counted as either a “yes” or a “no”; 3. Not necessarily — with small subgroup sample sizes, this spread is consistent with ordinary sampling variation, as seen in Lessons 111–112; a genuine difference cannot be assumed without further reasoning; 4. Combined: Y out of , ; 5. An uncaught tallying error propagates into every later stage of the investigation — the proportion, the inference and the final report would all be built on a false number, even if every later calculation is done correctly.

Common Misconceptions

MisconceptionHow to pre-empt it
Coding a decline as a “no” response.Explicit rule in Activity 2 and the exit ticket: declines are excluded entirely.
Believing any spread between subgroups reflects a real difference.Activity 3’s direct link back to Lessons 111–112’s sampling-variation ideas.
Averaging subgroup proportions instead of combining raw counts when merging groups.Exit Q4 and the Activity 2 total both require adding counts first, not averaging percentages.
Skipping the total-count check before calculating a proportion.Warmup’s deliberate planted error.
Treating “surveyed” and “responded” as always equal.Class G’s decline forces the distinction explicitly in the two-way table.

Enrichment — Competition-Style Problems

E1 (Kangaroo style). A class’s raw tally is Y and N, but the class has students. What has gone wrong, and by how much are the counts short?

Answer

— the counts are short of the class size, meaning one response was lost or miscounted before it could be classified.

E2 (AMC Junior style). Two classes are combined: Class X has from students; Class Y has from students. Find the combined .

Answer

E3 (Challenge). A survey of students has declines. Explain why reporting "" (treating declines as “no”) would be misleading, and state the correct calculation.

Answer

Treating declines as “no” wrongly assumes we know their answer, when we do not — it silently inflates the “no” count and could bias the proportion in either direction. The correct calculation uses only the actual responses: , not .

E4 (Challenge). Seven classes each contribute exactly responses (total , no declines) to a combined proportion of exactly (i.e. Y overall). If six of the seven classes’ Y-counts are , find the seventh class’s Y-count.

Answer

. Since the total must be , the seventh class has — but a class of cannot have “yes” responses. This is a genuine contradiction, showing the stated overall proportion of exactly is impossible with these six classes’ data — a good discussion point about checking a claimed total against the parts.

Homework

  1. Tally this raw data and find : N, N, Y, N, Y, Y, N, N.
  2. A group of students has declines. Write the correct denominator for calculating , and explain why.
  3. Combine two subgroups: Group 1 has Y out of ; Group 2 has Y out of . Find the combined .
  4. A two-way table shows Boys: (); Girls: (). Find the combined for all students.
  5. Explain why a -percentage-point difference between two subgroups of each should not automatically be called “a real difference between boys and girls”, referencing Lessons 111–112.
  6. A class’s tally shows Y and N, but students were surveyed. Identify the error and what should be checked.
  7. Reasoning. Explain why organising data into a two-way table by class is useful even though — as this lesson showed — differences between classes may just be sampling noise.
  8. Reasoning. Explain the difference between the number surveyed and the number who responded, and why an honest report must state both.
  9. Challenge. A stratified sample of (eight per class, six classes) has declines spread across three different classes. If the overall count is Y out of the responses received, find the overall , and explain what additional information you would need to calculate individual class proportions.

Answers: Q1 — Y, N, . Q2 — denominator (the surveyed minus the declines), since declines cannot be classified as either yes or no. Q3 — combined: Y out of , . Q4 — combined: Y out of , . Q5 — with only per subgroup, ordinary sampling variation alone can easily produce a -point gap even if the true underlying rate is identical for both groups, exactly as seen with same-size samples in Lessons 111–112. Q6 — ; two responses are missing from the tally and must be traced and re-counted before the proportion can be trusted. Q7 — it lets the class check for patterns worth investigating further (e.g. very large, consistent differences), even though small, ordinary-looking differences should not be over-interpreted without more evidence. Q8 — “surveyed” is how many were selected to take part; “responded” is how many actually provided usable data (excluding declines); reporting only one of the two could hide how much data was lost to non-response, which is itself useful information about the investigation’s reliability. Q9 — overall (since responses); to calculate individual class proportions, you would need to know exactly how many of the eight selected per class responded (i.e. which three classes lost a respondent) and how the “yes” responses were distributed across the six classes.