Read complete course notes

All guides

Econometrics II: Causal Inference

13 lessons on identification, omitted variables, controls, difference-in-differences, regression discontinuity, instruments and uncertainty.

A targeted problem-solving course, not a replacement for a full second course in econometrics. Examples are simplified population or toy calculations.

Econometrics I: multiple regression, hypothesis tests and basic probability.

Course outline

  1. Association is not an effect

    Separate association, identification and estimation using potential outcomes.

  2. The variable you left out

    Compute omitted-variable bias in a population linear projection.

  3. Controls that block the path

    Decide when adding a control removes bias and when it hides the effect.

  4. Differences in differences

    Compute a difference-in-differences estimate and state what it assumes.

  5. Testing the trend you assumed

    Use pre-period data to probe parallel trends and read the limits of the test.

  6. A jump at the cutoff

    Estimate a regression discontinuity effect by comparing fitted lines at the cutoff.

  7. A nudge that moves x

    Compute an instrumental-variables estimate and check its assumptions.

  8. A coefficient is not a marginal effect

    Compute a logit marginal effect and see why it varies with x.

  9. How sure is the slope

    Compute an OLS slope and its standard error by hand.

  10. Neighbors who share a shock

    See how clustering reduces the effective sample size.

  11. Staggered adoption: when old controls mislead

    See why a two-way fixed-effects regression can mix good and bad comparisons when treatment timing differs.

  12. Synthetic control: weighting donors to match the past

    Build a comparison unit from a weighted mix of donor units and read the post-period gap.

  13. Power and the minimum detectable effect

    Size a study so it can detect the effect you care about, and see why clustering raises the bar.

Sources and curriculum note

Checked and extended October 7, 2026. Calculations were checked in Python. These are population or toy-sample illustrations, not findings. Staggered-adoption and synthetic-control lessons cite the papers listed in the sources.

Complete course reading notes

Read every lesson below. The interactive reader above contains the same explanations, with visual tools and quizzes.

1. Association is not an effect

Learning goal: Separate association, identification and estimation using potential outcomes.

Each unit has two potential outcomes: Y(1) if treated and Y(0) if not. We see only one. The causal effect for a unit is Y(1) - Y(0), and the average treatment effect (ATE) is its mean. A simple difference in means estimates the ATE only if treatment is unrelated to the potential outcomes, as in a well-run randomized experiment.

In observational data, people choose treatment for reasons that also affect outcomes, so the raw difference mixes the effect with selection. Identification asks what assumption lets the data reveal the effect. Estimation asks how to compute it from a sample. Keep these three ideas apart: association, identification, estimation.

A four-person example makes selection concrete. Untreated outcomes are 2, 4, 6, 8 and every person gains 3 from treatment, so the ATE is 3. If the two people with the highest untreated outcomes choose treatment, the treated mean is (9 + 11) / 2 = 10 and the control mean is (2 + 4) / 2 = 3. The raw difference is 7, more than double the effect, and the extra 4 is selection. Randomization breaks the link between choosing treatment and having a high untreated outcome.

Worked example

In a randomized trial, treated mean outcome is 70 and control mean is 62. What do you estimate and what does it rely on?

  1. The difference in means is 70 - 62 = 8.
  2. Randomization makes assignment independent of potential outcomes, so the estimate targets the ATE.
  3. Sampling noise is separate; a standard error would measure it.
  4. The estimate is 8 outcome units, conditional on the randomization being valid.
Practice problem and solution

Hypothetical randomized trial has treated mean 54, control mean 47.5 and SE for their difference 2. Enter the estimated average effect. Construct estimate±2SE and explain why randomization, rather than the size of the gap, supports a causal interpretation.

Effect=54−47.5=6.5; illustrative estimate±2SE interval=[2.5,10.5]. Randomization supports comparable treatment groups; interpretation still depends on implementation and relevant design assumptions.

Mental model: Name the identifying assumption before reading a gap as an effect.

Common trap: Estimating without stating the assumption.

2. The variable you left out

Learning goal: Compute omitted-variable bias in a population linear projection.

Suppose the true model is y = b1 x + b2 z + u with Cov(x,u) = 0. If you regress y on x and omit z, the population slope is b1 + b2 x Cov(x,z)/Var(x). The second term is the omitted-variable bias. Its sign is the sign of b2 times the sign of Cov(x,z).

Example: b1 = 1, b2 = 2, Cov(x,z) = 3 and Var(x) = 4. The short-regression slope is 1 + 2 x 3/4 = 2.5, so the bias is +1.5. This is a population calculation, not a finite-sample estimate and not proof that any particular dataset is confounded.

Use the formula as a sign test. If b2 is positive and Cov(x,z) is positive, the short regression overstates the effect. If b2 is negative and Cov(x,z) is positive, it understates it. If Cov(x,z) is zero, omitting z does not bias the slope on x, though it may still raise the standard error. This is why a randomized treatment, which is unrelated to every other variable on average, needs no controls for unbiasedness.

Worked example

True model y = x + 2z + u, Cov(x,u) = 0, Cov(x,z) = 3, Var(x) = 4. What slope does a regression of y on x (with intercept) recover?

  1. The omitted term is b2 z, so the short slope is b1 + b2 Cov(x,z)/Var(x).
  2. Plug in: 1 + 2 x 3 / 4.
  3. That is 1 + 1.5 = 2.5.
  4. The bias is +1.5.
Practice problem and solution

True model y = 2x + 3z + u, Cov(x,u) = 0, Cov(x,z) = 2, Var(x) = 5. What slope does a regression of y on x alone recover? In your reasoning: Separate the omitted-variable bias from the true slope and predict the slope if Cov(x,z) were −2 instead.

2 + 3 x 2/5 = 2 + 1.2 = 3.2. Bias=3×2/5=1.2; with covariance −2, bias=−1.2 and slope=0.8, holding the other assumptions fixed.

Mental model: Bias = effect of the omitted variable times how it moves with x.

Common trap: Applying the formula when Cov(x,u) is not 0.

3. Controls that block the path

Learning goal: Decide when adding a control removes bias and when it hides the effect.

A confounder affects both the treatment and the outcome, so controlling for it removes bias. A mediator lies on the path from treatment to outcome, so controlling for it removes part of the effect you may want. Variables determined after treatment are called post-treatment variables and are risky controls.

Toy linear model: m = 0.5 x and y = 1 x + 2 m + u. The total effect of x on y is the direct 1 plus the indirect 0.5 x 2 = 1, so 2. The structural direct effect is 1 and total effect is 2. Here m and x are perfectly collinear: a regression on both cannot identify their separate coefficients. A controlled-regression interpretation needs independent variation and appropriate causal assumptions.

A collider is the opposite problem: a variable caused by both treatment and outcome, or by two unrelated causes. If talent and looks are unrelated in a population, but a casting agency picks people with a high sum, then within the selected group talent and looks are negatively associated. Controlling for or conditioning on a collider creates a spurious link. Draw the causal graph first, then decide which variables to add, instead of adding every available variable.

Worked example

In the toy model m = 0.5 x and y = x + 2m + u (no confounding), what are the total and direct effects of x on y?

  1. Direct effect is the coefficient on x: 1.
  2. Indirect effect is 0.5 x 2 = 1.
  3. Total = 1 + 1 = 2.
  4. These are structural effects. Because m=0.5x deterministically, x and m are collinear; regression on both cannot separately recover their coefficients.
Practice problem and solution

Toy model: m = 0.4 x and y = 0.5 x + 3 m + u (no confounding). What is the total effect of x on y? In your reasoning: Identify direct and mediated effects and explain why this structural calculation does not identify separate regression coefficients under perfect collinearity.

Direct 0.5 + indirect 0.4 x 3 = 0.5 + 1.2 = 1.7. This structural calculation does not identify separate regression coefficients when m is perfectly collinear with x. The indirect effect is 1.2 and direct effect 0.5. Perfect collinearity prevents separately identifying their regression coefficients from x and m alone.

Mental model: Choose controls by causal role, not by R-squared.

Common trap: Controlling for something x causes.

4. Differences in differences

Learning goal: Compute a difference-in-differences estimate and state what it assumes.

A difference-in-differences (DID) design compares the change over time in a treated group to the change in a comparison group. The comparison group's change stands in for what would have happened to the treated group without treatment. This is the parallel-trends assumption: absent treatment, the two groups would have changed by the same amount.

Treated mean goes from 50 to 70 and control mean from 45 to 55. The treated change is 20, the control change is 10, so DID = 10. The counterfactual treated post-mean is 50 + 10 = 60. The estimate also needs no anticipation, no spillovers to the control group and a stable composition of the groups.

In regression form, DID is the coefficient d in y = a + b treat + c post + d treat x post. Using the means above, a = 45, b = 5, c = 10 and d = 10. The same number results from subtracting two differences, so the regression is a convenient way to get standard errors and add controls. If the treated and comparison groups differ in level, that is allowed: DID only needs the changes to be parallel, which is weaker than equal levels.

Worked example

Treated means 50 (pre) and 70 (post); control means 45 and 55. Compute the DID estimate and the counterfactual treated post-mean.

  1. Treated change = 70 - 50 = 20.
  2. Control change = 55 - 45 = 10.
  3. DID = 20 - 10 = 10 outcome units.
  4. Counterfactual treated post = 50 + 10 = 60, so the effect is 70 - 60 = 10, if parallel trends hold.
Practice problem and solution

Treated means 62 (pre) and 80 (post); control means 50 (pre) and 56 (post). Compute the DID estimate. In your reasoning: Compute post-only and pre-only gaps as well as DID, and state the identifying parallel-trends assumption.

Treated change 18, control change 6, DID = 12. Post gap=80−56=24; pre gap=62−50=12; DID=12. Interpretation requires parallel untreated trends, not merely a baseline gap correction.

Mental model: DID = change in treated minus change in control, under parallel trends.

Common trap: Reading the post-only gap as the effect.

5. Testing the trend you assumed

Learning goal: Use pre-period data to probe parallel trends and read the limits of the test.

Parallel trends concerns the untreated outcome after treatment, which we never see. We can only check pre-treatment periods: if the gap between groups is constant before treatment, the assumption is more believable. A diverging pre-trend is a warning.

If the gap is 5 in periods -2, -1 and 0, and 9 after treatment, the DID estimate assuming the gap would have stayed 5 is 9 - 5 = 4. Passing a pre-trend check does not prove the assumption holds after treatment, and a low-powered check can miss real differences. Report the check as supporting evidence, not as a proof.

An event-study plot shows the estimated gap in each period relative to one reference period, usually the last pre-treatment period. Estimates before treatment should be near zero with confidence bands that include zero, and a jump should appear after treatment. If the pre-period estimates drift upward, the treated group may have been on a different trend, and the post-treatment gap partly reflects that drift. Be careful: confidence bands that are wide can include zero while hiding a meaningful trend.

Worked example

The treated-minus-control gap is 5, 5, 5 in three pre-periods and 9 in the post period. Estimate the effect and say what supports it.

  1. Pre-period gaps are constant at 5.
  2. Assuming the gap would have stayed 5, the post-period counterfactual gap is 5.
  3. Estimate = 9 - 5 = 4.
  4. The constant pre-trend supports the assumption but cannot prove it.
Practice problem and solution

Hypothetical pre-period treated-minus-control gaps are 8,8,8; the post gap is 15. Enter the DID estimate. Predict the counterfactual post gap and state a plausible time-varying confounder that would invalidate the interpretation.

Under parallel untreated trends, counterfactual gap remains 8. DID=15−8=7. A simultaneous policy affecting only the treated group could invalidate that counterfactual. Stable pre-gaps do not prove the post-period assumption.

Mental model: Test the assumption in the pre-period, and report its limits.

Common trap: Calling a pre-trend test proof.

6. A jump at the cutoff

Learning goal: Estimate a regression discontinuity effect by comparing fitted lines at the cutoff.

In a regression discontinuity design, treatment switches on when a running variable crosses a cutoff, such as a test score. People just below and just above the cutoff are assumed to be similar except for treatment. The effect is the jump in the outcome at the cutoff.

To estimate it, fit a line on each side using data near the cutoff and compare the two fitted values at the cutoff itself. Example: cutoff 5, left fit y = 30 + 2x gives 40 at 5; right fit y = 20 + 5x gives 45 at 5. The jump is 5. The result is local to the cutoff and relies on units not being able to precisely manipulate the running variable.

Bandwidth is the main choice. A narrow window around the cutoff keeps units similar but leaves few observations and noisy estimates. A wide window gives more data but depends on the line fitting well far from the cutoff, which can add bias. Report results for several bandwidths. Also look for bunching: if many units sit just above the cutoff, they may have manipulated the running variable, and the design is not credible.

Worked example

Cutoff c = 5. Left fit y = 30 + 2x; right fit y = 20 + 5x. Estimate the effect at the cutoff.

  1. Left fit at x = 5: 30 + 10 = 40.
  2. Right fit at x = 5: 20 + 25 = 45.
  3. Jump = 45 - 40 = 5.
  4. This is local to units near the cutoff.
Practice problem and solution

Cutoff c = 4. Left fit y = 12 + 3x; right fit y = 10 + 4x. Estimate the jump at the cutoff. In your reasoning: Compute both fitted limits at the cutoff and explain why this is local rather than a global effect.

Left: 12 + 12 = 24. Right: 10 + 16 = 26. Jump = 2. The limits are evaluated at c=4: left 24 and right 26. Local continuity of untreated potential outcomes supports comparison near the cutoff, not across the full score range.

Mental model: Compare fitted values at the cutoff, with a stated window.

Common trap: Evaluating each fit at a different x.

7. A nudge that moves x

Learning goal: Compute an instrumental-variables estimate and check its assumptions.

When x is correlated with the error, regression is biased. An instrument z can help if it (1) moves x (relevance), (2) affects y only through x (exclusion) and (3) is unrelated to unobserved factors (independence). With one instrument and a constant effect, the IV estimate is the ratio of the reduced-form effect (y on z) to the first-stage effect (x on z).

If a unit increase in z raises x by 1.5 and raises y by 3, the IV estimate is 3 / 1.5 = 2. Exclusion and independence cannot be fully tested with the data. A weak first stage makes the ratio unstable, so report the first stage too.

A binary-instrument example. Mean x is 0.7 when z = 1 and 0.2 when z = 0, so the first stage is 0.5. Mean y is 12 when z = 1 and 6 when z = 0, so the reduced form is 6. The IV (Wald) estimate is 6 / 0.5 = 12. With heterogeneous effects this is a local average treatment effect for compliers, the units whose x responds to z, and not necessarily the effect for everyone. A common rough screen is a first-stage F statistic above 10, but it is not a guarantee.

Worked example

The instrument raises x by 1.5 per unit and y by 3 per unit. Compute the IV estimate and state which assumptions still need design-based justification.

  1. First stage = 1.5; reduced form = 3.
  2. IV = 3 / 1.5 = 2.
  3. Relevance is checked by the first stage.
  4. Exclusion and independence require design-based justification. The stated ratio alone cannot determine which assumption is most doubtful.
Practice problem and solution

Hypothetical binary instrument groups have mean x=1.4 versus 0.6 and mean y=8.1 versus 5.7. Enter the Wald IV estimate. Show both differences and explain why a direct effect of the instrument on y would defeat its interpretation.

First stage=1.4−0.6=0.8; reduced form=8.1−5.7=2.4; ratio=3. A direct instrument effect on y violates exclusion, so the ratio cannot isolate the exposure effect under that violation.

Mental model: IV scales the instrument's effect on y by its effect on x.

Common trap: Dividing in the wrong order.

8. A coefficient is not a marginal effect

Learning goal: Compute a logit marginal effect and see why it varies with x.

In a logit model P(y=1|x) = 1/(1 + exp(-(a + b x))). The coefficient b is the change in the log-odds for a unit change in x, not the change in probability. The change in probability depends on where you are: the marginal effect is b x p x (1 - p), with p evaluated at the chosen x.

Take a = -1 and b = 0.8. At x = 1 the index is -0.2 and p = 0.450. The marginal effect is 0.8 x 0.450 x 0.550 = 0.198, about 20 percentage points per unit of x. The effect is largest where p is near 0.5 and shrinks near 0 or 1. This is the same coefficient giving different probability effects at different x.

Compare effects at two points. With a = -1 and b = 0.8, the odds ratio is exp(0.8) = 2.23, so each unit of x multiplies the odds by 2.23 at every x. But at x = 3 the index is 1.4, p = 0.802 and the marginal effect is 0.8 x 0.802 x 0.198 = 0.127, smaller than 0.198 at x = 1. Average marginal effects average this quantity over the sample, and marginal effects at the mean evaluate it at the average x. These can differ.

Worked example

In the logit with a = -1, b = 0.8, compute p and the marginal effect at x = 1.

  1. Index = -1 + 0.8 x 1 = -0.2.
  2. p = 1 / (1 + exp(0.2)) = 0.450.
  3. Marginal effect = 0.8 x 0.450 x 0.550 = 0.198.
  4. About 0.198, i.e., 19.8 percentage points per unit of x at x = 1.
Practice problem and solution

Hypothetical logit coefficient b=1.2 with p=0.5 at one covariate value. Enter the local marginal effect there; compare it with p=0.2 and explain why b is not a constant probability effect.

At p=0.5: 1.2×0.5×0.5=0.30. At p=0.2: 1.2×0.2×0.8=0.192. The probability derivative depends on p; b is the log-odds coefficient.

Mental model: Convert coefficients into effects on the scale you care about, and say where.

Common trap: Reading b as a probability change.

9. How sure is the slope

Learning goal: Compute an OLS slope and its standard error by hand.

For simple regression y on x, the slope is Sxy/Sxx and the intercept makes the line pass through the means. Residuals are the vertical gaps. The error variance estimate is s^2 = SSR / (n - 2), and the slope's standard error is sqrt(s^2 / Sxx).

Data x = 1, 2, 3, 4 and y = 2, 3, 5, 6. The means are 2.5 and 4, Sxx = 5 and Sxy = 7, so the slope is 1.4 and the intercept 0.5. Fitted values 1.9, 3.3, 4.7, 6.1 give residuals 0.1, -0.3, 0.3, -0.1, SSR = 0.2, s^2 = 0.1 and SE = sqrt(0.1/5) = 0.141. A reproducible analysis keeps the data, code and software version together, so the number can be recomputed.

The coefficient of determination is R squared = 1 - SSR / SST. For the data above SST = 4 + 1 + 1 + 4 = 10, so R squared = 1 - 0.2 / 10 = 0.98. The t statistic for the slope is 1.4 / 0.141 = 9.9. A high R squared does not mean the slope is causal; it only measures fit. With n = 4 the t distribution with 2 degrees of freedom is wide, so a critical value is far above 1.96. This is a toy data set.

Worked example

Fit y on x for x = 1, 2, 3, 4 and y = 2, 3, 5, 6. Find the slope and its standard error.

  1. Means: x-bar = 2.5, y-bar = 4. Sxx = 5, Sxy = 7, slope = 1.4, intercept = 0.5.
  2. Residuals: 0.1, -0.3, 0.3, -0.1, so SSR = 0.2.
  3. s^2 = SSR / (n - 2) = 0.1.
  4. SE = sqrt(0.1 / 5) = 0.141.
Practice problem and solution

Fit y on x for x = 1, 2, 3, 4 and y = 1, 3, 4, 7. Find the standard error of the slope (3 decimals). In your reasoning: Show the fitted line, residuals, SSR, residual degrees of freedom and Sxx; distinguish the slope estimate from its standard error.

Slope = 1.9 (Sxy = 9.5, Sxx = 5), intercept -1.0, SSR = 0.7, s^2 = 0.35, SE = sqrt(0.35/5) = 0.2646. The fitted line is −1+1.9x; residuals are 0.1,0.2,−0.7,0.4. Their squared sum is 0.70; n−2=2, Sxx=5. SE quantifies uncertainty about slope 1.9 under the stated model.

Mental model: An estimate comes with its sampling uncertainty.

Common trap: Reporting a slope without a standard error.

10. Neighbors who share a shock

Learning goal: See how clustering reduces the effective sample size.

Observations in the same group, such as students in a classroom, often share unobserved shocks, so their errors are correlated. Treating them as independent makes standard errors too small. A common approximation for equal cluster size m and within-cluster correlation rho is the design effect DEFF = 1 + (m - 1) rho.

The effective sample size is N / DEFF. With m = 20, rho = 0.05 and N = 1000: DEFF = 1 + 19 x 0.05 = 1.95 and effective n = 513. Using the naive n would overstate precision. In practice you cluster the standard errors at the level where treatment or shocks are shared. The approximation shown is a teaching tool and assumes equal clusters.

A larger cluster and stronger correlation hurt more. With m = 50 and rho = 0.1, DEFF = 1 + 49 x 0.1 = 5.9, and N = 1,000 gives an effective sample size of about 169. If few clusters exist, even clustered standard errors can be unreliable, and methods such as the wild cluster bootstrap are used. A common caution is to be careful when the number of clusters is below about 30 to 50, which is a rough guide and not a rule.

Worked example

N = 1000 students in classrooms of m = 20 with rho = 0.05. Find the design effect and effective sample size.

  1. DEFF = 1 + (m - 1) rho = 1 + 19 x 0.05 = 1.95.
  2. Effective n = 1000 / 1.95.
  3. That is 512.8, about 513.
  4. Standard errors that ignore clustering would be too small by about a factor of sqrt(1.95) = 1.4.
Practice problem and solution

N = 950 observations in clusters of m = 10 with rho = 0.1. What is the effective sample size? In your reasoning: Show the design effect and effective size; compare the effective size if rho=0 and explain the equal-cluster-size approximation.

DEFF = 1 + 9 x 0.1 = 1.9. 950 / 1.9 = 500. With rho=0, DEFF=1 and effective size is 950. The stated formula uses equal cluster size and a common intracluster correlation; it is an approximation to information loss.

Mental model: Count independent information, not rows.

Common trap: Treating clustered rows as independent.

11. Staggered adoption: when old controls mislead

Learning goal: See why a two-way fixed-effects regression can mix good and bad comparisons when treatment timing differs.

In many applications, groups adopt treatment at different times. The standard two-way fixed-effects regression with a single treatment dummy is then a weighted average of many two-group, two-period comparisons. Goodman-Bacon (2021) shows how to decompose it.

Some of those comparisons use groups that are already treated as controls for groups treated later. If the treatment effect grows over time, the control group's own effect is changing during the window, and it contaminates the comparison. The estimate can even have the wrong sign if effects are heterogeneous.

A simple identity shows the problem: if the control group is already treated and its effect grows by delta over the window, a two-group comparison estimates the true effect minus delta. With a true effect of 10 and a growing control effect of 4, the estimate is 6.

Newer estimators, such as Callaway and Sant'Anna (2021), compare each treatment-timing group with never-treated or not-yet-treated units and then aggregate the group-time effects with stated weights. In practice, plot an event study, check the decomposition and report which comparison you rely on.

Worked example

The true effect is 10 and an already treated control's effect grows by 4 over the window. What does a two-group comparison estimate?

  1. Estimate = true effect minus the control's growth.
  2. 10 - 4.
  3. = 6.
  4. Understates the effect.
Practice problem and solution

True effect 10 and the control's effect grows by 4. What is the estimate? Enter a number.

10 - 4 = 6.

Mental model: With staggered timing, use clean comparisons and check the decomposition.

Common trap: Assuming a single treatment dummy always gives an average effect.

12. Synthetic control: weighting donors to match the past

Learning goal: Build a comparison unit from a weighted mix of donor units and read the post-period gap.

When one unit is treated, for example a state that passes a law, we can build a synthetic version of it. A synthetic control is a weighted average of untreated donor units, with weights chosen so its pre-treatment outcomes match the treated unit as closely as possible. Abadie, Diamond and Hainmueller (2010) applied it to California's tobacco control program.

The effect in each post-treatment period is the treated unit's outcome minus the synthetic control's outcome. If the synthetic control tracks the treated unit well before treatment, the post gap is a credible estimate of the effect, assuming nothing else changed for the treated unit at that time.

With two donors and one pre-period, the weight w on donor 1 solves w x D1 + (1 - w) x D2 = T, so w = (T - D2) / (D1 - D2), limited to between 0 and 1. If T is outside the donors' range, no mix matches it and the method fails.

Check fit, try leaving out donors, and run placebo tests that apply the method to untreated units. If many placebos show gaps as large as the treated unit's, the result is not strong evidence.

Worked example

Treated pre-period value 70, donors 90 and 50. Weight on donor 1?

  1. w = (T - D2) / (D1 - D2).
  2. (70 - 50) / (90 - 50).
  3. 20 / 40.
  4. 0.5.
Practice problem and solution

Treated 70, donors 90 and 50. Weight on donor 1?

(70 - 50) / (90 - 50) = 0.5.

Mental model: Weights make the synthetic unit match the pre-period. The post gap is the estimate.

Common trap: Reading a large gap without checking pre-fit or placebos.

13. Power and the minimum detectable effect

Learning goal: Size a study so it can detect the effect you care about, and see why clustering raises the bar.

Power is the probability of detecting an effect that exists. With a two-sided 5 percent test and 80 percent power, the minimum detectable effect (MDE) is about (1.96 + 0.84) = 2.8 standard errors. Smaller true effects are often missed.

For a difference in means with n units per arm and outcome standard deviation sigma, the standard error is sigma x sqrt(2 / n). With sigma = 10 and n = 200 per arm, SE = 10 x sqrt(0.01) = 1, so the MDE is about 2.8.

Clustering multiplies the variance by the design effect, so the standard error rises by the square root of DEFF. With DEFF = 1.95 the standard error rises by about 40 percent, to 1.40, and the MDE to about 3.9. To keep the MDE, you would need roughly DEFF times as many units.

Before collecting data, decide the smallest effect worth acting on, compute the MDE and compare. If the MDE is larger than any plausible effect, the study is unlikely to be informative and a null result will say little.

Worked example

sigma = 10, n = 200 per arm. Standard error and MDE?

  1. SE = 10 x sqrt(2/200).
  2. = 1.
  3. MDE = 2.8 x 1.
  4. 2.8.
Practice problem and solution

sigma = 10, n = 200 per arm. MDE using 2.8 standard errors, to 1 decimal?

SE = 1, so 2.8 x 1 = 2.8.

Mental model: MDE is about 2.8 SEs. Clustering raises SE by the square root of DEFF.

Common trap: Reading a null result as proof of no effect.