Read complete course notes

All guides

AP Statistics

Fourteen connected lessons on design, variation and inference, with explainable visual tools.

Focused inference and design guide aligned to the fall-2026 revision. It is not the complete five-unit course. Removed slope-inference, geometric-distribution and goodness-of-fit topics are not core lessons.

First-year algebra, ratios, square roots and basic graph interpretation.

Course outline

  1. Population, sample and parameter

    Keep the target population separate from observed data.

  2. Sampling and assignment do different jobs

    Separate generalizability from causal evidence.

  3. Variation shrinks, but not linearly

    Explain how sample size changes a sample proportion’s standard error.

  4. A p-value is a conditional probability

    Interpret evidence under the null, not probability the null is true.

  5. Confidence is a procedure, not a personal probability

    Interpret intervals with population and repeated-sampling context.

  6. A complete inference argument

    Connect procedure, conditions, calculation and conclusion.

  7. Correlation describes linear form, not cause

    Read scatterplots without turning correlation into causation.

  8. Residuals compare observation with prediction

    Explain a regression error in the response variable’s units.

  9. Conditional probability changes the denominator

    Read two-way tables with the conditioning group fixed.

  10. Errors and power are different questions

    Explain a testing mistake relative to the truth and decision.

  11. Describe the distribution, then add the context

    Describe shape, outliers, center and spread together, using the variable and its units.

  12. A z-score compares across scales

    Standardize values to compare them and to find areas under a normal model.

  13. Name the condition, not just the check mark

    Match each condition for a proportion interval to what it protects.

  14. Chi-square: what the statistic adds up

    Compute and interpret a goodness-of-fit statistic with its degrees of freedom.

Sources and curriculum note

Fall 2026 / May 2027 course version. Earlier exam errors are used only where the skill remains in the revised course. Lessons 11 to 14 draw on the course description for the SOCS and condition-labelling points; the arithmetic is standard and was checked by calculation.

Complete course reading notes

Read every lesson below. The interactive reader above contains the same explanations, with visual tools and quizzes.

1. Population, sample and parameter

Learning goal: Keep the target population separate from observed data.

A statistical question asks about variability, not simply one fact. Identify who or what you want to describe, the variable measured and the time/context. If you ask about the proportion of students at a school using an app this month, the population is those students, not all teenagers everywhere.

A parameter summarizes a population, often unknown. A statistic summarizes the sample you actually observed. For a binary outcome, p is the population proportion and p̂ is the sample proportion. For a quantitative variable, μ is the population mean and x̄ is the sample mean. Do not define p as a count.

The inference target must stay consistent from hypotheses to conclusion. College Board’s 2025 report flags missing context and unclear definitions of parameters. A precise first sentence can prevent several later mistakes: "Let p be the proportion of all students at this school who used the app this month."

Work through one example from start to finish. Suppose a school librarian wants to know how many of the school's students used a booking kiosk this term. Population: every enrolled student this term. Variable: used the kiosk, yes or no. Parameter: p, the proportion of all students who used it. If she surveys a random sample of 100 students and 24 say yes, the statistic is p̂ = 0.24. The numbers here are invented for teaching.

Notice what changes between samples and what does not. Another random sample of 100 would probably give a different p̂, but p stays the same fixed unknown value. That is why p̂ is a statistic with variability and p is a target. When you write any later step, such as hypotheses or an interval, point back to this first sentence so a reader never has to guess what the symbol means.

Worked example

In an original model, 28 of 100 randomly sampled library patrons used a booking kiosk. Define p and p̂.

  1. p is the proportion of all patrons of this library who used its booking kiosk during the defined month.
  2. p̂=28/100=0.28 is the observed sample proportion.
  3. The sample gives evidence about p; it does not establish p=0.28.
Practice problem and solution

Hypothetical audit: 12 of 40 first-week customers and 6 of 20 second-week customers report an issue. Pool the samples and enter p̂. In your reasoning, identify the denominator and explain why this is not the population parameter.

Pool 18 issues among 60 sampled customers: 18/60=0.30. This is a sample statistic; the population proportion remains unknown.

Mental model: Define the population parameter in words before symbols.

Common trap: A sample statistic is not automatically the population truth.

2. Sampling and assignment do different jobs

Learning goal: Separate generalizability from causal evidence.

Random sampling helps a sample represent a defined population. Random assignment allocates treatment conditions and helps balance confounding variables, supporting a causal comparison. Neither phrase can substitute for the other.

An observational study does not impose treatments. It can reveal associations, but an observed difference may arise from confounding. An experiment deliberately assigns treatments; good control, replication and randomization matter. A randomized experiment on volunteers can support causation in those conditions without automatically representing everyone.

Bias can come from undercoverage, nonresponse, leading questions or self-selection. A larger sample reduces sampling variability, but it does not repair a biased selection process. State the design and then match the strength of the conclusion to it.

Compare two short designs. Design one: a teacher asks for volunteers to try a new study method and then compares their scores with everyone else's. Design two: the teacher lists all students, randomly assigns half to the new method, and compares the two groups. Design one is an observational comparison with self-selection, so motivated students may be over-represented in the method group. Design two uses random assignment, so the groups should be similar apart from the method.

Now add sampling. If the students in design two were chosen at random from the whole school, conclusions could reach the whole school. If they were volunteers, the causal claim can stand for those volunteers but generalizing needs caution. A useful habit is to write two separate sentences: one about who the results apply to, and one about whether the difference can be attributed to the treatment.

Worked example

Volunteers are randomly assigned to two learning tools. Tool A users score higher. What is justified?

  1. Random assignment helps support a causal effect of the tool under the study conditions.
  2. Volunteer selection may limit extension to all students.
  3. Do not claim a random sample unless the recruitment actually used one.
Practice problem and solution

Hypothetical school: volunteers choose to enter a study, then a random draw assigns tutoring A or B. Mean improvement is 8 versus 5 points. Enter the comparison supported by assignment. In reasoning calculate the difference, defend the assignment-based interpretation and explain the population limit from recruitment.

The difference is 3 points. Random assignment supports a causal comparison of methods in the experiment if implementation is sound; volunteer recruitment does not establish representativeness of all students.

Mental model: Sampling asks whom you studied; assignment asks what happened to them.

Common trap: Confusing random sampling with random treatment assignment.

3. Variation shrinks, but not linearly

Learning goal: Explain how sample size changes a sample proportion’s standard error.

The sampling distribution describes how a statistic varies across repeated samples of the same size. It is not a histogram of the individual observations within one sample. For independent binary trials with population proportion p, the standard deviation of p̂ is √[p(1−p)/n].

For a null model in a proportion test, use p₀ in that formula. To halve standard error, multiply n by four, not two. Greater precision still cannot compensate for biased selection or invalid independence assumptions.

A normal approximation needs enough expected successes and failures under the relevant model. In an AP significance-test setup, check np₀ and n(1−p₀) are at least 10. For sampling without replacement, the usual independence check is that the sample is no more than 10% of the population. State randomness separately. These are conditions, not magic guarantees that every dataset is well behaved.

A worked check of the formula. Use an invented case with p = 0.5 and n = 100. The standard deviation of p̂ is √(0.5 × 0.5 / 100) = √0.0025 = 0.05. With n = 400, it is √(0.25 / 400) = √0.000625 = 0.025. Quadrupling the sample size halved the spread, which is the square-root pattern in action.

This pattern explains why polls get expensive fast. Moving from a small precision gain to a large one needs many more respondents, and each extra respondent helps less than the one before. It also explains why two samples of the same size can still disagree: the standard deviation describes typical spread, not a guarantee. When you describe a sampling distribution, name the statistic (p̂ or x̄), the sample size and what is being repeated, so the reader knows you mean repeated samples and not repeated observations.

Worked example

Under p₀=0.5 and n=100, calculate SE. What n halves it?

  1. SE=√[0.5(0.5)/100]=0.05.
  2. At n=400, SE=√(0.25/400)=0.025.
  3. Four times the sample size halves SE.
  4. Expected successes and failures are 50 each at n=100, so the large-count check is met.
Practice problem and solution

Hypothetical design: p₀=0.5 and the initial n=100 gives SE₀=0.05. Funding allows only one quarter as many observations. Enter the new SE₀; show the variance calculation and explain the size change.

New n=25. SE₀=√(0.5×0.5/25)=0.10, twice 0.05 because n is quartered.

Mental model: Sampling distributions concern statistics across repeated samples.

Common trap: Treating √n as n, or using a normal model before conditions.

4. A p-value is a conditional probability

Learning goal: Interpret evidence under the null, not probability the null is true.

A p-value is the probability, assuming the null model is true, of obtaining a result at least as extreme as the observed statistic in the direction specified by the alternative. It is not the probability that H₀ is true, the probability the result happened by chance in every sense, or the size of the effect.

Choose the alternative from the research question before looking at the sample. "Greater than" means an upper-tail test; "different" means a two-sided test. A two-sided p-value counts extremeness in both directions. Alpha is the decision threshold selected for the procedure.

If p-value ≤ α, reject H₀ in favor of evidence for Hₐ. Otherwise fail to reject H₀. Failing to reject does not prove equality. Use uncertainty-aware language: "The data provide convincing statistical evidence…" rather than "We proved…". The 2025 report directly flags proof language and inequality mistakes.

A short example of the logic. An invented company claims 60% of its deliveries arrive on time. A customer group samples 100 deliveries and sees 52 on time. The question is: if the true rate really were 0.60, how often would a random sample of 100 give a result as low as 0.52 or lower? If that happens rarely, the sample is hard to explain under the claim. If it happens fairly often, the sample is consistent with the claim.

That is all a p-value measures. It starts from the null model and asks about the data, not the other way round. It cannot tell you the probability that the company's claim is true, because the claim is not a random event in this framework. It also says nothing on its own about whether a gap of eight points matters to customers. Pair the p-value with the estimated difference, and always write the conclusion in the context of deliveries.

Worked example

A test gives p-value 0.037 with α=0.05. State the conclusion.

  1. 0.037 < 0.05, so reject the null hypothesis.
  2. State the evidence in terms of the alternative and target population.
  3. Avoid "H₀ has only a 3.7% chance of being true"; that reverses the conditioning.
Practice problem and solution

Hypothetical tests of the same null give p=0.032. At α=0.01, enter the decision. In reasoning contrast the α=0.05 decision, distinguish evidence thresholds from effect magnitude and explain why the first decision is not proof of no effect.

At 0.01, 0.032>0.01, so fail to reject. At 0.05 it is below threshold, so reject. Neither threshold tells the size of the effect; failure to reject does not prove absence.

Mental model: p-value = probability of data extremeness under a null model.

Common trap: Reversing probability conditioning or claiming proof.

5. Confidence is a procedure, not a personal probability

Learning goal: Interpret intervals with population and repeated-sampling context.

A confidence interval combines a point estimate with a margin of error. Under valid conditions, a 95% confidence method captures the fixed population parameter in about 95% of repeated samples. After an interval is calculated, the parameter is either in that interval or not; the usual frequentist interpretation is not a 95% probability that this fixed parameter moved into it.

In practical language, "We are 95% confident that the population proportion lies between…" is acceptable when the population and parameter are specified. It does not say that 95% of individual observations fall in that range, or that 95% of sample means are inside this one interval.

Higher confidence makes an interval wider if data are unchanged. Larger n tends to make it narrower. A narrow but biased interval can still miss the target systematically. The confidence level concerns the procedure’s long-run coverage under assumptions, not protection against a bad sample.

An example sentence in full. Suppose a random sample of 120 students gives an interval for the proportion who commute by bike. A good interpretation reads: "We are 95% confident that the proportion of all students at this school who commute by bike is between the lower and upper endpoints." Each part carries information: the confidence level, the word proportion, the population and the interval itself. Remove the population and the sentence could refer to anyone.

Common wrong versions to avoid: "There is a 95% chance the sample proportion is in the interval" (the sample proportion is the center of the interval, so it is certainly inside); "95% of students are in the interval" (confuses individuals with a parameter); and "the interval is 95% likely to be right" with no population. Check the conditions first, because the whole interpretation depends on a random sample and a sampling distribution that is close to normal.

Worked example

p̂=0.4 with standard error 0.03. Using a multiplier 2 for this teaching approximation, calculate an interval.

  1. Margin of error = 2(0.03)=0.06.
  2. Endpoints are 0.4−0.06=0.34 and 0.4+0.06=0.46.
  3. This approximate interval concerns a population proportion, not individual binary responses.
  4. The multiplier 2 is a rounded teaching approximation, not a universal exact critical value.
Practice problem and solution

Hypothetical survey: x̄=20 minutes, SE=2, multiplier=2. Enter the interval’s lower endpoint. Show both endpoints and assess whether a population mean of 25 is consistent with this interval.

Margin=2×2=4; interval [16,24]. A population mean of 25 is outside this interval; it is not an interval for individual commutes.

Mental model: Name the parameter and explain what the method covers.

Common trap: Interpreting a mean interval as the distribution of individuals.

6. A complete inference argument

Learning goal: Connect procedure, conditions, calculation and conclusion.

An inference solution is a chain of claims. Define the parameter, choose hypotheses, name a suitable procedure, check conditions, compute the statistic and p-value, compare to alpha and conclude in context. A calculator result alone is not an argument.

For a one-proportion z-test, z=(p̂−p₀)/√[p₀(1−p₀)/n]. Use the null proportion for the test’s standard error. A confidence interval usually estimates its standard error from observed data instead. Do not swap these formulas without understanding the target.

Always distinguish statistical evidence from practical size. A tiny effect can yield a small p-value in a huge sample. A useful conclusion reports the estimated difference as well as uncertainty. Current fall-2026 Statistics curriculum differs from the 2025 course; removed inference-on-slopes and goodness-of-fit topics are not core activities here.

Here is a skeleton you can reuse for any significance test. Step one: define the parameter in context. Step two: state H₀ and Hₐ with the same parameter. Step three: name the procedure and check each condition with numbers from the problem, such as a random sample, 10% of the population and expected counts of at least 10. Step four: compute the test statistic and p-value. Step five: compare to α and write a conclusion that mentions the context and the type of evidence.

Graders and teachers look for each link, so skipping one costs more than a small arithmetic slip would. A common failure is writing correct numbers with no conditions, or a correct conclusion that does not refer to the parameter. Treat the layout as a checklist you tick off, and keep the wording plain. A clear, short paragraph beats a long one that wanders.

Worked example

In a random sample of 100 students, 30 use an app. Test H₀:p=0.2 vs Hₐ:p>0.2. Compute z, assuming independence.

  1. p̂=0.30 and SE₀=√[0.2(0.8)/100]=0.04.
  2. z=(0.30−0.20)/0.04=2.5.
  3. Large-count values are 20 and 80; randomness was given and independence was assumed for this worked example.
  4. A p-value and alpha comparison are still needed for a formal decision. z by itself is not the final answer.
Practice problem and solution

Hypothetical sample: 60 successes in 100 trials; H₀:p=0.5. Enter z for a one-proportion z-test. Show p̂ and null SE, check expected success/failure counts, and explain which additional sampling conditions are needed.

p̂=0.60; null expected counts 50 and 50; SE₀=√(0.25/100)=0.05; z=2. Random sampling and independence still need design evidence, including the 10% condition when sampling without replacement.

Mental model: An inference answer is a contextual chain, not a lone p-value.

Common trap: Skipping conditions because the calculator gives a number.

7. Correlation describes linear form, not cause

Learning goal: Read scatterplots without turning correlation into causation.

A scatterplot shows pairs of quantitative variables. Describe direction, form, strength and unusual points in context before calculating a summary. Correlation describes linear association, not every kind of relationship.

A correlation near zero can coexist with a strong curved pattern. Correlation has no units and is unchanged by positive linear unit conversions. Its sign changes if one axis is reversed.

Causation needs study-design support and plausible alternative explanations. Confounding can induce an association between two variables even when changing one would not cause the other to change.

An example of describing a scatterplot in full. Suppose an invented plot shows hours studied against quiz score for 30 students. A complete description says: the association is positive and roughly linear, moderately strong, with one student who studied many hours but scored low. Each of direction, form, strength and unusual points appears, and the context (hours and quiz scores) is named.

Then ask the cause question. If students who study more also tend to be students who sleep more or who started with stronger background knowledge, those hidden factors could account for part of the pattern. Only an experiment that assigns study time at random could separate the effect of studying from those other factors. Also remember that a plot with a tight curve can have a correlation close to zero, so look at the picture before trusting r. A single outlier can also raise or lower r a lot.

Finally, check units before comparing. Converting hours to minutes would change the slope you see on the axis, but the correlation would stay the same, because r has no units. Reporting r alongside a plot, and saying what each axis measures, makes the summary honest and easy to check.

Worked example

A graph has a U-shaped relationship and r≈0. Explain.

  1. The association is nonlinear.
  2. Correlation summarizes linear association.
  3. Near-zero r does not imply no relationship.
Practice problem and solution

Hypothetical model ŷ=12−1.5x: a case has x=4 and observed y=8. Enter its predicted response. In your reasoning calculate its residual and explain whether the model underpredicts or overpredicts this case.

Prediction=12−1.5×4=6. Residual=8−6=2, so the model underpredicts the observed response by two units.

Mental model: Describe form before using a linear summary.

Common trap: Reading correlation as a causal effect.

8. Residuals compare observation with prediction

Learning goal: Explain a regression error in the response variable’s units.

A least-squares line predicts a response from an explanatory variable. A residual is observed y minus predicted y. Positive residual means the observation lies above the line; negative means below.

Residuals use the response variable’s units. They are not the x-distance to the line and not a percentage unless the response itself uses percentage units. Extrapolation outside the observed input range can be unreliable.

A residual plot helps inspect whether a linear model is suitable. Curved structure suggests unmodelled form; a fan shape suggests changing spread. No pattern is reassuring but does not prove every model assumption.

A small invented example. A line predicts quiz score from hours studied as ŷ = 50 + 6x. For a student who studied 4 hours, the prediction is 50 + 24 = 74. If that student actually scored 80, the residual is 80 − 74 = +6 points, so the student scored 6 points above what the line predicts. Another student who studied 2 hours has prediction 62; a score of 58 gives a residual of −4.

The sign and the units carry the meaning. Residuals are in quiz points here, because the response is a score. Plotting residuals against x helps you see whether errors are scattered evenly around zero. If the points form a smile or frown, a straight line is the wrong shape. If the spread widens, predictions are less reliable for larger x. Also keep predictions within the range of the data you actually observed.

One more habit: say what the residual means in a sentence. "This student scored 6 points above the prediction for 4 hours of study" ties the number to the context and shows you understand the sign.

Worked example

A model predicts 15 minutes, but the actual commute is 18 minutes. Residual?

  1. Observed minus predicted =18−15.
  2. Residual is +3 minutes.
  3. The model underpredicted this observation.
Practice problem and solution

Hypothetical model ŷ=5+4x: a case has x=5 and observed y=21. Enter its residual, then explain the direction and size of the prediction error.

Prediction=25; residual=21−25=−4. The model overpredicts by four response units.

Mental model: A residual is a signed prediction error.

Common trap: Dropping units or using a horizontal distance.

9. Conditional probability changes the denominator

Learning goal: Read two-way tables with the conditioning group fixed.

P(A|B) is the fraction of B cases that also meet A. The denominator is B, not the entire sample and not A. Conditioning narrows the reference group.

P(A|B) and P(B|A) generally differ. A high probability of a positive result among affected people does not directly tell you the probability of being affected among positive results. Base rates matter.

Independence means conditioning on one event does not change the probability of the other, when the conditional probability is defined. Mutually exclusive nonzero-probability events are not independent because one excludes the other.

Suppose 80 of 100 cases are unaffected and 20 are affected. If 10 affected and 40 unaffected cases test positive, the positive-result group has 50 members. P(positive|affected)=10/20=0.50, while P(affected|positive)=10/50=0.20. The same numerator produces different answers because the conditioning groups differ.

Build the table slowly with invented numbers. Suppose 200 students are classified by whether they took a prep course and whether they passed. Say 80 took the course and 60 of them passed; 120 did not and 66 of them passed. Then P(pass | course) = 60/80 = 0.75, and P(pass | no course) = 66/120 = 0.55. The rates differ, so pass and course are not independent in this sample.

Switch the condition and the result changes: of the 126 who passed, 60 took the course, so P(course | pass) = 60/126 ≈ 0.48. Each question uses a different denominator, and that is what you must name before dividing. Write the condition in words first, for example "among those who passed", then pick the row or column total that matches it. This habit avoids most conditional probability slips on two-way tables.

Worked example

Of 100 students, 40 play a sport; 10 of those 40 also play music. Find P(music|sport).

  1. Condition on sport, so the eligible group has 40 students.
  2. Ten in that group play music.
  3. The conditional probability is 10/40=0.25.
Practice problem and solution

Hypothetical table: A∩B=12, A only=18, B only=18, neither=52. Enter P(A|B). In your reasoning also find P(B|A) and explain why conditioning changes the denominator.

B total=30, so P(A|B)=12/30=0.4. A total=30, so P(B|A)=0.4 here; equal values follow from equal group totals, not a general symmetry rule.

Mental model: Conditioning chooses the reference group.

Common trap: Reversing conditional probabilities.

10. Errors and power are different questions

Learning goal: Explain a testing mistake relative to the truth and decision.

A Type I error rejects a true null hypothesis. A Type II error fails to reject a false null. These are defined using the unknown real state and the testing decision, not whether the sample statistic looks large.

Alpha controls the Type I error rate for the procedure under its model assumptions. Power is the probability of rejecting a false null at a particular alternative. Power depends on effect size, sample size, variability and alpha.

Lowering alpha reduces willingness to reject but can lower power with other factors fixed. A non-significant result does not establish no effect; the study may have limited power.

Picture the two kinds of mistake with a medical-style analogy that uses invented numbers only for illustration. A test is set up to decide whether a new teaching method changes scores. If the method truly has no effect but the sample happens to look impressive and you reject H₀, that is a Type I error. If the method truly helps but the sample is too small or noisy to show it, and you fail to reject H₀, that is a Type II error.

You can influence the risks. A larger sample and a larger real effect both increase power. Less variable data also helps. Raising α makes rejection easier, which increases power but also the Type I rate. In a written answer, describe an error in context: "Concluding the new method improves scores when it actually does not" is stronger than just the label. Never claim the null is true after failing to reject.

Worked example

H₀ says a process meets a target. A test rejects H₀ although it actually meets the target. Name the error.

  1. The null is true.
  2. The decision rejects it.
  3. This is a Type I error; explain it as falsely flagging the compliant process.
Practice problem and solution

Hypothetical screen: H₀ says not defective. One defective item passes inspection; one sound item is flagged. Enter the error type for the defective item. In reasoning classify both cases from truth and decision, and explain which change in detection probability would increase power at a specified defective alternative.

Passing a defective item fails to reject a false H₀: Type II. Flagging a sound item rejects a true H₀: Type I. Higher probability of rejecting at the specified defective alternative means higher power.

Mental model: Decision uncertainty has two distinct error directions.

Common trap: Using a non-significant finding as proof of no effect.

11. Describe the distribution, then add the context

Learning goal: Describe shape, outliers, center and spread together, using the variable and its units.

When a question asks you to describe a distribution of quantitative data, a complete answer covers four things: the shape, any outliers or gaps, the center and the spread. Many students memorize the four letters of an acronym and then name only one or two of them. The College Board course description says exactly this: students often struggle to explain all the elements, and in particular they neglect unusual features such as gaps or outliers.

The same document stresses that all data has context. That means naming the variable of interest and its units. "The median is 15.5" is incomplete. "The median commute time is 15.5 minutes" can be read by someone who has not seen the graph.

Choose the summary that fits the shape. For a roughly symmetric distribution without outliers, the mean and standard deviation describe center and spread well. For a skewed distribution, or one with outliers, the median and the interquartile range (IQR) are more resistant, so they are better choices. A single extreme value can pull the mean a long way but barely moves the median.

One common rule marks a value as an outlier if it lies more than 1.5 IQR below the first quartile or above the third quartile. Check which rule your course states; the rule is a convention for flagging values, not a proof that the value is an error. Always say what the outlier is, and consider that it may be a real observation worth keeping. Do not delete it without a reason from the context.

Finally, compare distributions with comparative words, such as "higher median" and "greater spread", not two separate lists of numbers.

Worked example

Describe the commute times 12, 13, 13, 14, 15, 15, 16, 17, 18, 19, 21, 45 (minutes, invented).

  1. Order the data: it is already in order. The median is (15 + 16)/2 = 15.5 minutes.
  2. The lower half is 12 to 15, so Q1 = 13.5. The upper half is 16 to 45, so Q3 = 18.5. IQR = 5.
  3. Upper fence = 18.5 + 1.5(5) = 26. The value 45 is above it, so it is an outlier.
  4. The distribution is skewed right with one outlier. Because of this, report the median and IQR.
  5. Answer: the commute times are skewed right with an outlier at 45 minutes; the median is 15.5 minutes and the IQR is 5 minutes.
Practice problem and solution

Invented data: 3, 4, 4, 5, 6, 6, 7, 8, 9, 30 (hours). Using Q1 = 4 and Q3 = 8 and the 1.5 IQR rule, enter the upper fence, and say whether 30 is an outlier.

IQR = 8 − 4 = 4. Upper fence = 8 + 1.5(4) = 14. Since 30 is above 14, it is an outlier.

Mental model: Shape, outliers or gaps, center and spread, each with units and in context.

Common trap: Naming a shape and a mean while ignoring an outlier.

12. A z-score compares across scales

Learning goal: Standardize values to compare them and to find areas under a normal model.

A z-score tells you how many standard deviations a value lies from the mean: z = (x − μ)/σ. A z-score of 1.5 means the value is one and a half standard deviations above the mean. A negative z-score means below the mean. Because z-scores have no units, they let you compare values measured on different scales.

Suppose a student scores 620 on a test with mean 500 and standard deviation 100, and 28 on another test with mean 21 and standard deviation 5. (These numbers are invented.) The raw numbers cannot be compared directly. The first z-score is (620 − 500)/100 = 1.2. The second is (28 − 21)/5 = 1.4. The second score is relatively higher, because it lies further above its own mean in standard deviation units.

To turn a z-score into a proportion, you need a model for the shape of the distribution. If the values are approximately normal, the area to the right of z = 1.5 is about 0.067, so about 6.7% of values would be above it. If the distribution is strongly skewed, that area is not reliable. The College Board course description lists relying on vague references to the normal distribution as a common error, so state why the model is reasonable, for example from a roughly symmetric, unimodal graph of the data.

Also practice working backwards. Given the mean, standard deviation and a z-score, the value is x = μ + zσ. Write this relationship down before substituting; it prevents the frequent slip of using σ where μ belongs.

Worked example

Test A has mean 500 and SD 100; test B has mean 21 and SD 5 (invented). A student scores 620 on A and 28 on B. Which score is relatively higher?

  1. Standardize A: z = (620 − 500)/100 = 1.2.
  2. Standardize B: z = (28 − 21)/5 = 1.4.
  3. Compare the z-scores, not the raw scores.
  4. The B score is further above its mean in SD units, so it is relatively higher.
Practice problem and solution

Invented distribution: mean 70, SD 8. Enter the value with z = −1.25, and explain what the negative sign means.

x = μ + zσ = 70 + (−1.25)(8) = 60. The negative sign means below the mean.

Mental model: z = (x − μ)/σ counts standard deviations from the mean and puts different scales side by side.

Common trap: Using normal-model areas without checking that the shape is roughly normal.

13. Name the condition, not just the check mark

Learning goal: Match each condition for a proportion interval to what it protects.

Before using a one-sample z-interval for a proportion, you check conditions. The course description reports a common error here: students mislabel the conditions, for example by associating the 10% condition with the randomization condition. Each condition protects something different, so each needs its own label and its own evidence.

The random condition says the sample came from a random sample or a randomized process. It protects against bias and is what lets you generalize to the population. A calculation cannot supply it; you have to read the study design.

The 10% condition says the sample is no more than 10% of the population when sampling without replacement. It protects the independence assumption behind the standard error formula. If you sample a large share of a small population, the observations are no longer close to independent. The check is a calculation: population size at least 10 times the sample size.

The large counts condition says that both n·p̂ and n(1 − p̂) are at least 10. It protects the approximately normal shape of the sampling distribution, which the interval relies on. Use the counts of successes and failures in the sample.

The interval is p̂ ± z*·√(p̂(1 − p̂)/n). Writing the numbers is not enough. State each condition, give the evidence, and say in the context of the problem what the interval means. Use the words of the problem, such as 'students at this school', not just 'the population'. The interval estimates the population proportion; it does not describe the share in the sample, which you already know.

Worked example

In a random sample of 200 students from a school of 3,000 (invented), 62 said yes. Check the conditions and build a 95% confidence interval for the school proportion.

  1. Random: the sample was random, so results can be generalized to the school.
  2. 10%: 3,000 is at least 10 times 200, so independence is reasonable.
  3. Large counts: 200(0.31) = 62 and 200(0.69) = 138 are both at least 10.
  4. Interval: 0.31 ± 1.96·√(0.31·0.69/200) = 0.31 ± 0.064.
  5. We are 95% confident the true school proportion is between about 0.246 and 0.374.
Practice problem and solution

Hypothetical sample size n = 150. Enter the smallest population size that satisfies the 10% condition, and name what the condition protects.

The population must be at least 10 times the sample: 10(150) = 1500. It protects independence.

Mental model: Random protects against bias; 10% protects independence; large counts protects normality.

Common trap: Calling the 10% check a randomization check.

14. Chi-square: what the statistic adds up

Learning goal: Compute and interpret a goodness-of-fit statistic with its degrees of freedom.

A chi-square goodness-of-fit test asks whether the counts in one categorical variable match a stated distribution. The null hypothesis names the claimed proportions, for example that all six faces of a die are equally likely. The alternative says that at least one proportion differs from the claim.

For each category, the expected count is the sample size times the claimed proportion. The test statistic adds up how far the observed counts are from the expected counts, scaled by the expected counts: χ² = Σ (O − E)²/E. Every term is nonnegative, so χ² is zero only when every observed count equals its expected count. Large values are evidence against the claim.

The degrees of freedom are the number of categories minus 1. With six categories, df = 5. This is not the sample size or the number of categories; one count is determined once the others and the total are known. You use the df, not the sample size, to find the p-value from the chi-square distribution.

The usual condition is that the data come from a random sample or randomized experiment and that every expected count is at least 5. Check expected counts, not observed counts. Write them down.

When you conclude, link the p-value to the claim in context. A large p-value means the observed counts are consistent with the claimed distribution. It does not prove that the die is fair; it only says the data give no convincing evidence that it is unfair. Standard deviations and the 1.96 multiplier do not appear in this test, so keep it separate from z-procedures in your notes.

Worked example

A die is rolled 60 times with counts 8, 12, 9, 11, 10, 10 (invented). Test whether the die is fair.

  1. H0: each face has probability 1/6. Ha: at least one probability differs.
  2. Expected count for each face is 60(1/6) = 10, so the condition holds.
  3. χ² = (4 + 4 + 1 + 1 + 0 + 0)/10 = 1.0.
  4. df = 6 − 1 = 5, which gives a p-value of about 0.96.
  5. The p-value is large, so there is no convincing evidence that the die is unfair.
Practice problem and solution

Invented counts: 14, 6, 10 across three categories, with expected counts 10, 10, 10. Enter the chi-square statistic.

χ² = (14−10)²/10 + (6−10)²/10 + (10−10)²/10 = 1.6 + 1.6 + 0 = 3.2.

Mental model: χ² adds (O − E)²/E over categories; df is categories minus 1.

Common trap: Reading a large p-value as proof that the claim is true.