Demonstration 1 of 4
A before and after comparison counts the trend as an effect
If an outcome rose after an intervention, how much of the rise did the intervention cause?
Both series share the same upward trend and seasonal wave. The before and after change of the treated series mixes that trend with the effect; the control series, untouched by the intervention, measures the trend alone, and difference in differences subtracts it.
Scroll sideways for the whole equation
y-bar is a mean of the outcome; T marks the treated series and C the control series; pre is periods 0 to 79, post the chosen number of periods from 80 on. The constructed treated series gets exactly 8 extra units from period 80.
Predict first. With all 40 post periods, will the plain before and after change of the treated series be close to the true 8?
Choose an example
Scroll sideways for the whole figure
Constructed data: the chapter notebook's seeded treated and control series (seed 20260940, first cell), true effect 8 from period 80.
Calculated values
- Treated, mean before
- 121.84
- Treated, mean after
- 145.45
- Control, mean before
- 109.93
- Control, mean after
- 125.24
- Estimate shown
- 23.61
- True effect
- 8
- Estimate minus truth
- 15.61
With 40 periods after the intervention, before and after = 145.45 - 121.84 = 23.61, far from the true effect of 8. The treated series would have risen anyway: the control rose 125.24 - 109.93 = 15.31 with no treatment at all.
Worked steps
- Treated change: 145.45 - 121.84 = 23.61.
- Control change: 125.24 - 109.93 = 15.31.
- Difference of the changes: 23.61 - 15.31 = 8.30.
Use the idea
Before crediting an intervention with a rise, find a comparison series that faced the same conditions without the intervention and subtract its change over the same periods.
Where the conclusion applies
The constructed generator makes parallel trends true by design: treated equals 12 plus the control plus noise, plus 8 after period 80. With real data parallel trends is an assumption that cannot be checked after the intervention.
Check your understanding: A treated store's weekly sales went from 100 to 120 after a promotion; a similar store without it went from 80 to 90. What is the difference in differences estimate?
Chapter 22 source: section "Section 4: The Full Methodology".
Demonstration 2 of 4
The counterfactual is a line projected from before
Where does the 'what would have happened' series come from, and how much does it depend on the fit?
The method learns, before the intervention, how the treated series relates to the control, then applies that relation to the control after the intervention. The gap between what happened and the projection is the effect estimate, which inherits every error in the fitted relation.
Scroll sideways for the whole equation
y-hat T is the counterfactual for the treated series in period t, built from the control series y C with intercept a and slope b fitted before period 80; tau-hat is the mean gap between observed and counterfactual over the 40 post periods.
Predict first. Fitted on all 80 pre periods, will the mean effect land within 0.5 of the true 8?
Choose an example
Scroll sideways for the whole figure
Constructed data: the chapter notebook's seeded series (seed 20260940) and its pre period least squares counterfactual (fourth code cell); the shorter windows refit the same line.
Calculated values
- Periods fitted
- 80
- Intercept a
- 10.63
- Slope b
- 1.01
- Fit error before, RMSE
- 1.72
- Mean effect after
- 8.12
- True effect
- 8
Fitted on periods 0 to 79, the counterfactual for period 100 is 10.63 + 1.01 x 123.97 = 135.84, against an observed 145.08. Averaged over periods 80 to 119 the estimated effect is 8.12 against the true 8. The effect is only as good as the projected line: it is a model of what did not happen, not an observation.
Worked steps
- Fit treated on control over periods 0 to 79: a = 10.63, b = 1.01.
- Project to period 100: 10.63 + 1.01 x 123.97 = 135.84.
- Effect at period 100: 145.08 - 135.84 = 9.24.
- Mean of the 40 post period effects: 8.12.
Use the idea
Report the pre period fit and the window used next to any effect estimate, and check how far the estimate moves when the window changes.
Where the conclusion applies
The relation between the series must stay the same after the intervention, apart from the effect itself. Here the generator guarantees it; in practice nothing does.
Check your understanding: A counterfactual fitted before an intervention is 50 + 0.5 x control. After the intervention the control is 120 and the treated outcome 118. What is the estimated effect?
Chapter 22 source: section "Section 4: The Full Methodology".
Demonstration 3 of 4
Two worlds with the same past and different effects
If the treated and control series moved in parallel before the intervention, is the estimate safe?
Both estimators learn only from the periods before the intervention and attribute every post period change at the treated unit to the intervention. A shock that starts with the treatment leaves the past untouched, so it passes every check on the past and lands in the estimate.
Scroll sideways for the whole equation
The shock is an unrecorded change of 0, 3 or 6 units that hits only the treated series from period 80, at the same time as the intervention whose true effect is 8.
Predict first. With a treated only shock of 6, will a perfect pretrend fit protect the difference in differences estimate?
Choose an example
Scroll sideways for the whole figure
Constructed data: the chapter notebook's seeded series (seed 20260940) and its unrecorded shock case (shock 6 from period 80).
Calculated values
- Unrecorded shock
- 6
- Estimator
- difference in differences
- Estimate without shock
- 8.30
- Estimate with shock
- 14.30
- True intervention effect
- 8
- Pre period values changed
- none
The difference in differences reads 8.30 + 6 = 14.30 with an unrecorded shock of 6, while the intervention's true effect stays 8. Every value before period 80 is identical in both cases, so no pretrend check can tell them apart. Only knowledge of what else changed at the treated unit can.
Worked steps
- Estimate without the shock: 8.30.
- The shock adds 6 to every treated value from period 80 and nothing before.
- Estimate with the shock: 8.30 + 6 = 14.30.
Use the idea
Before reporting an effect, list what else changed at the treated unit when the intervention started; a clean pretrend plot does not answer that question.
Where the conclusion applies
The shock is constructed to start exactly at period 80 and to touch only the treated series, the notebook's case with a shock of 6; the value 3 is added here.
Common wrong turn: Parallel pretrends prove parallel trends
Check your understanding: A difference in differences estimate is 5. You learn that a competitor closed beside the treated store in the same week, adding about 2 units. What estimate is left for the intervention?
Chapter 22 source: section "Section 5: What the Counterfactual Actually Is".
Demonstration 4 of 4
Three comparisons, one known truth
With several control series, does it matter how you combine them into a counterfactual?
Each comparison is a different guess at the missing series. Averaging all controls equally assumes each moves in parallel with the treated unit; synthetic control requires a weighted average to match the treated level; the regression allows an intercept and free weights. The visible signal is the pre period fit.
Scroll sideways for the whole equation
The treated series is built as 12 plus 0.6 times the first control plus 0.4 times the second, plus noise, plus 8 from the intervention; the third control has no trend. w are synthetic control weights, nonnegative and summing to one.
Predict first. Which comparison comes within 0.5 of the true effect of 8?
Choose an example
Scroll sideways for the whole figure
Constructed data: the companion's seeded workshop example for this chapter (three controls, true effect 8 from month 100), analysed by the chapter's applied tool as in the notebook's workshop cell.
Calculated values
- Comparison
- equal weight difference in differences
- Estimated effect
- 9.91
- True effect in the generator
- 8
- Estimate minus truth
- 1.91
- Fit error before, RMSE
- 1.59
The equal weight difference in differences estimates 9.91, so the error is 9.91 - 8 = 1.91: the equal weight average includes a control with no trend, so the comparison drifts away from the treated path. Its pre period fit error is 1.59, larger than the regression counterfactual's. The generator's truth is known here; in real data only the pre period fit is visible.
Worked steps
- Pre period fit error (RMSE): 1.59.
- Mean gap after the intervention: 9.91.
- Against the generator's true 8: 9.91 - 8 = 1.91.
Use the idea
Show the pre period fit of every comparison you tried, choose the design before seeing the post period, and treat a poor pre period fit as a reason to distrust the effect.
Where the conclusion applies
A good pre period fit is necessary, not sufficient: the shock in the previous demonstration fits the past perfectly and is still wrong. The true effect is known only because the example is constructed.
Check your understanding: One comparison fits the pre period with error 3.3 and estimates 9.4; another fits with error 1.0 and estimates 7.9. Which deserves more trust, and does the fit prove it right?
Chapter 22 source: section "When an Experiment Is Not Available".