Demonstration 1 of 4
Spending that follows demand inflates the apparent return
If a brand spends most when customers are already buying, what does a regression of sales on spend measure?
The terracotta bar regresses sales on exposure alone. When exposure rises with demand, it absorbs credit for sales the demand produced. Adding the demand variable removes that bias, but only because this simulation knows and measures the confounder. The randomized bar needs no such knowledge: treatment was assigned by chance.
Scroll sideways for the whole equation
S is sales, T the advertising exposure, D the underlying demand (a common cause), c how strongly demand moves sales, and epsilon noise. The link is how strongly exposure follows demand. The true advertising effect is 2 in every state.
Predict first. With exposure following demand at 0.8 and demand effect 5, what will the unadjusted regression report?
Choose an example
Scroll sideways for the whole figure
Constructed data: the chapter notebook's seeded confounding experiment (seed 20260938, 3,000 draws, cell 5), with the demand link and demand effect varied on the same draws.
Calculated values
- Spend follows demand (link)
- 0.8
- Demand effect on sales
- 5
- Unadjusted estimate
- 4.49
- Adjusted estimate
- 2.01
- Randomized estimate
- 2.27
- Expected bias, unadjusted
- 2.44
The expected bias of the unadjusted slope is 5 x 0.8 / 1.64 = 2.44, and the simulation gives 4.49 - 2 = 2.49. Spending rises when demand is high, so the unadjusted regression credits the advertising with sales the demand produced. The randomized campaign estimates 2.27 whatever the link, because a coin decided who was treated.
Worked steps
- Expected bias: 5 x 0.8 / 1.64 = 2.44.
- Unadjusted slope in 3,000 simulated weeks: 4.49.
- Adding demand as a control: 2.01. Random assignment: 2.27.
Use the idea
Before reading an attribution figure as a return on spend, ask what decided when the money was spent. If the answer is the forecast of demand, the figure includes demand, and a randomized or geographic test is the check.
Where the conclusion applies
The generator, the true effect of 2 and the demand effect are teaching values from the notebook. Real confounders are rarely measured as cleanly as the demand variable here, so the adjusted bar is a best case.
Check your understanding: Exposure follows demand at 0.5 and demand moves sales by 4. What bias does the unadjusted slope carry?
Chapter 20 source: section "Section Two: Geo-Experiments and the Randomized Approach".
Demonstration 2 of 4
Advertising carries over, and the data window does not start the business
If spend keeps working after the week it runs, what happens when a model ignores the spend before its first week?
Both lines get the same spend inside the window: 20 per week. The solid line also carries the stock built by twelve earlier weeks at 100. The dashed line pretends the business began in week 1. The gap shrinks by the decay factor every week.
Scroll sideways for the whole equation
x is the spend in week t, a the advertising stock, and lambda the share of last week's stock that carries into this week (the Koyck geometric decay).
Predict first. At decay 0.8, after twelve weeks of spend at 100, how far off is a zero start in week 1 of the window?
Choose an example
Scroll sideways for the whole figure
Constructed data: the chapter notebook's initial state example (cell 12): 12 weeks at 100 then 16 at 20, with the companion's adstock function; no random draws.
Calculated values
- Decay lambda
- 0.8
- Stock before the window
- 465.64
- True stock, week 1
- 392.51
- Model stock, week 1
- 20.00
- Error, week 1
- 372.51
- Error, week 8
- 78.12
The model starts from zero. Twelve weeks at 100 leave a stock of 465.64; in week 1 the true stock is 20 + 0.8 x 465.64 = 392.51. Starting from zero misses 0.8 x 465.64 = 372.51 in week 1 and still 78.12 in week 8, an error that fades only as fast as the advertising itself.
Worked steps
- Stock entering the window: 465.64.
- Week 1 true stock: 20 + 0.8 x 465.64 = 392.51.
- Zero start week 1 stock: 20. Gap 0.8 x 465.64 = 372.51, shrinking by the factor 0.8 each week.
Use the idea
When fitting a mix model, collect spend from before the first sales week and use it to initialise each channel's stock. Without it, vary the starting stock and report how much the early weeks' attribution moves.
Where the conclusion applies
Geometric decay with a fixed lambda is the simplest carryover; the chapter notes that modern tools also allow delayed peaks. The spend schedule is the notebook's illustration, not a real campaign.
Check your understanding: Decay 0.5, previous stock 40, spend this week 10. What is the new stock?
Chapter 20 source: section "Section Four: The Methodology".
Demonstration 3 of 4
Where a brand sits on its saturation curve
Does every extra unit of spend buy less than the one before it?
The curve rises toward a ceiling. K marks where it reaches half the maximum. With alpha 1 the first unit of spend is the most productive; with alpha above 1 the curve starts flat and steepens before saturating, so below K extra spend can buy more per unit, not less.
Scroll sideways for the whole equation
x is the spending level, K the half-saturation point (the spend at which the response reaches half its maximum) and alpha the curvature. The response is shown as a percent of its maximum.
Predict first. With K = 160 and alpha 3, does the step from 100 to 150 add less response than the step from 50 to 100?
Choose an example
Scroll sideways for the whole figure
Constructed data: response curves computed from the chapter's Hill function with assumed parameters, extending the notebook's half-saturation values (30, 80, 160) of cell 3; no random draws.
Calculated values
- Half-saturation K
- 80
- Curvature alpha
- 1
- Response at 50
- 38.46
- Response at 100
- 55.56
- Response at 150
- 65.22
- Gain 50 to 100
- 17.10
- Gain 100 to 150
- 9.66
At spend 100 the response is 100 x 100^1 / (100^1 + 80^1) = 55.56 percent of the maximum. Moving from 50 to 100 adds 55.56 - 38.46 = 17.10; from 100 to 150 adds 65.22 - 55.56 = 9.66, so the second 50 buys less than the first, so returns are already diminishing here. With alpha 1 the curve is concave everywhere (the Michaelis-Menten case).
Worked steps
- Response at 100: 100 x 100^1 / (100^1 + 80^1) = 55.56.
- Gain from 50 to 100: 55.56 - 38.46 = 17.10.
- Gain from 100 to 150: 65.22 - 55.56 = 9.66.
Use the idea
Before cutting or raising a channel, read where current spend sits relative to the estimated K, and check whether the fitted alpha is above 1; a rule of diminishing returns from the first dollar applies only to the alpha 1 curve.
Where the conclusion applies
The curves use assumed K and alpha values, as in the notebook's illustration; they are not estimated from any brand. Fitted curves carry wide uncertainty when spend has varied little.
Check your understanding: With K = 100 and alpha 1, what share of the maximum response does a spend of 300 reach?
Chapter 20 source: section "Section Four: The Methodology".
Demonstration 4 of 4
A good sales fit can still assign the wrong credit
If a mix model predicts sales well, can you trust how it splits the credit between channels?
Each line follows one approach's channel A estimate as the fitting window grows. The faint lines are the other approaches. Knowing the true background does not help, because the two channels are nearly the same series; a wrong fixed background gives a stable but wrong answer; and two priors each pull the estimate to where they started.
Scroll sideways for the whole equation
y is weekly sales, b the background (level, season and trend), A and B the two channels' activity, and beta A and beta B their effects: 12 and 8 in the simulation. The channels move almost together.
Predict first. Fitting everything jointly by least squares, will the channel A estimate stay near 12 as more weeks are added?
Choose an example
Scroll sideways for the whole figure
Constructed data: the chapter notebook's seeded attribution stability experiment (seed 20260938, cell 10): 220 weeks, true effects 12 and 8, fits ending at weeks 104, 130, 156 and 182, each scored on the next 26 weeks.
Calculated values
- Joint OLS: mean channel A
- 2.61
- Joint OLS: spread across fits (sd)
- 2.66
- Joint OLS: holdout MAE
- 1.51
Joint OLS: the mean channel A estimate misses the truth by 2.61 - 12 = -9.39, with a spread of 2.66 across the four fits, so it is unstable and far from 12. Yet the holdout errors of all five approaches are nearly equal (1.51, 1.49, 1.49, 1.50, 1.48): the sales forecast cannot tell which credit split is right, because the two channels move almost together.
Worked steps
- Mean channel A over four fits: 2.61; truth 12.
- Spread across fits: 2.66.
- Holdout MAE of the five approaches: 1.51, 1.49, 1.49, 1.50, 1.48.
Use the idea
Report attribution beside a refit table and beside the answers under other defensible backgrounds or priors. If the split moves while the forecast error does not, the data have not decided the split; an experiment can.
Where the conclusion applies
The simulation's truth is known, which is never true of a business. The prior rows are exact Gaussian updates with a fixed background and known noise, not a full Bayesian mix model; priors A and B disagree on purpose and are not recommendations.
Check your understanding: Two fits give channel A coefficients of 15.8 and 4.0 with holdout errors of 1.50 and 1.48. Which split does the holdout favour?
Chapter 20 source: section "When a Good Fit Gives an Unstable Explanation".