What does it really mean to "control for" a variable?
A multivariate regression returns a coefficient, but a multivariate scatter plot does not exist. The Frisch–Waugh–Lovell (FWL) theorem resolves the tension: the coefficient on one regressor in a multiple regression equals the slope of a simple two-variable regression, once both variables have been residualized on the other controls (partialling out).
In the post's fast-food example (— simulated restaurants, each handing out 100 coupons in one day and counting how many come back that month), ignoring neighborhood income gives a slope of sales on coupons of — (p = —): the wrong sign. After partialling income out of both variables, the slope becomes —, exactly the multiple-regression coefficient, and within one standard error (—) of the true effect —.
Partialling out on the post's 50 restaurants
The loop cycles through three views of the same restaurants. First the naive scatter of sales against coupons. Then partial out income: regress coupons on income and keep the residuals (dashed lines), the part of coupon redemption that income cannot explain; sales are residualized on income the same way. Finally, regress the residuals on each other: residualized sales on residualized coupons, whose slope is the multiple-regression coefficient. That last regression is what the post calls FWL Step 2. The post's Step 1 residualizes coupons only and regresses raw sales on them; it appears in Tab 3.
Glossary (open a card if a term is unfamiliar)
FWL theorem
Partialling out
Confounder
Omitted variable bias
Conditional vs. marginal effect
Backdoor path
Simpson's paradox
Degrees-of-freedom correction
DML bridge
Confounding Lab — make the sign flip yourself
The simulator keeps the post's income and coupon equations: income ~ N(50, 10²), coupons = 60 + π · income + N(0, 5²), and sales = 10 + β · coupons + γ · income + N(0, 3²), with the true coupon effect fixed at β = 0.20 (the post's day-of-week term is left out: it is unrelated to coupons, so it cannot bias anything). The lab draws its own random samples, so its default draw of 50 restaurants is not the post's sample (Tab 1 uses the post's exact restaurants). Move the sliders and compare the naive slope of sales on coupons with the FWL slope. Map out where the naive slope has the wrong sign — that is when ignoring income misleads the analyst — and notice that the FWL slope stays centered on 0.20 everywhere.
Naive OLS sales ~ coupons
FWL = Full OLS controlling for income
Where the bias comes from — the omitted-variable-bias formula
δ = π·σ²I / (π²·σ²I + σ²C) with σI = — (income) and σC = — (coupon noise).
—
What to look for
- Sign flip. At the defaults (γ = 0.30, π = −0.50) the population δ is —, so the predicted bias is — and naive OLS converges to —: the wrong sign. A single sample of 50 restaurants is noisy, so the FWL slope in one draw can land well away from 0.20. Reseed a few times; Tab 4 shows it is centered on 0.20 across samples.
- Two OVB numbers, two questions. The population γ · δ says where naive OLS is headed on average; in any one sample, naive β − true β differs from it by sampling noise. The in-sample identity γ · δ = naive β − FWL β holds exactly in every draw. Reseed and watch those two rows stay equal.
- Direction matters. δ is the slope of income regressed on coupons, not the DGP slope π of coupons on income. At the defaults π = −0.50 while δ = —. |δ| peaks at |π| = 0.5, where income explains half the variance of coupons, and shrinks toward zero on either side: try π = −1.00 or π = −0.10.
- Turn off the confounding. Drag γ (or π) to 0. The predicted bias is zero, and the naive and FWL slopes differ only by γ · δ, which is now pure noise.
- Sample size. Drag n up to 500. Both standard errors shrink, but the naive slope stays near β + γ · δ: more data does not fix a misspecified model.
The post's estimates — one coefficient, several standard errors
These rows are written by the post's script.py and match the
Summary of results table (the NumPy row is omitted: it has no standard error). Every
FWL variant reproduces the full-regression coefficient exactly; that is the algebraic guarantee
of the theorem. Only the standard errors differ, and each difference has a mechanical
explanation. The vertical steel line marks the true effect β = 0.20. Toggle methods to
declutter; hover or tap a row for its SE, CI, p-value, and residual df.
Methods
Table view of the selected rows
| Method | β | SE | 95% CI | p | df |
|---|
What to look for
- Naive OLS is the only negative point estimate (—, p = —). Omitting income flips the sign.
- Full OLS and FWL Step 2 share one coefficient, —, to machine precision: the FWL identity made visible. Their SEs differ slightly (— vs. —); the card below explains why.
- FWL Step 1's CI is enormous (SE —). The main culprit is the dropped intercept: raw sales average —, and a no-intercept regression on mean-zero coupon residuals leaves that level in the residuals. Adding an intercept alone cuts the SE to — (the Step 1 + intercept row), closing — of the gap to the full model. The rest of the way to — comes from also removing income-driven variation in sales (Step 2). Untick Step 1 to zoom in on the other rows.
- Adding day-of-week (last two rows) barely moves the coupon coefficient (— → —) because, holding income fixed, day-of-week is nearly unrelated to coupons (partial correlation −0.021), and it nudges the SE from — to —.
Why do FWL Step 2 and the full regression report different SEs?
The FWL theorem is an algebraic identity: the two regressions produce the same coefficient and the same residuals, so they share the sum of squared residuals. The difference is the divisor. Step 2 is a one-regressor, no-intercept regression, so the software divides by n − 1 = —. The full regression estimated an intercept, coupons, and income, so it divides by n − 3 = —. Hence SE(Step 2) = SE(full) × √(—/—): — × — = —. The FWL route recovers the coefficient exactly; for inference, use the full model's SE, or rescale the Step 2 SE by √((n − 1)/(n − 3)).
Monte Carlo — bias vs. variance over many samples
A single sample is noisy: maybe the naive estimate was off because the random draw happened to be unlucky. To find out, rerun the whole simulation 100 times (fixed seeds 1000–1099, so the same settings always give the same result) and look at the distribution of estimates. If naive OLS were merely noisy, the orange histogram would center on β = 0.20. If it is biased, it centers near β + γ · δ, the dashed orange line.
What to look for
- Two distributions, not one. The teal histogram (FWL) centers near 0.20. The orange one (naive) centers near the dashed line β + γ · δ, which at the defaults is —: on the wrong side of zero.
- Bias versus variance. The FWL histogram is a little wider than the naive one: partialling out income discards the part of coupon variation that income explains. Unbiased but noisier beats precise but wrong. Push n to 500: both clusters tighten, the orange one around the wrong value.
- Disable the bias. Set γ or π to 0. Both histograms center on 0.20 and the naive sign-flip rate falls to about zero.
- The sign-flip rate. At the default settings (n = 100, seeds 1000–1099) the naive slope is negative in — of the 100 samples, computed live by this page. A naive analyst would almost always conclude that coupons hurt sales; the post's single sample of 50 restaurants is one such case (—).
Quiz — predict, then check
Seven questions that mirror the post's predict-then-check moments. Pick an answer; the feedback explains why it is right or wrong, using the post's own numbers. Answered 0 of 7.
Q1 · Sign of the naive slope
The true effect of coupons on sales is —. You regress sales on coupons alone, with no controls. What sign will the slope have?
Q2 · Add income as a control
Now fit sales ~ coupons + income. What happens to the coupon coefficient?
Q3 · Which regression gives the OVB δ?
The omitted-variable-bias formula says naive β − full β = γ · δ, with γ = — (income's coefficient in the full model). Which auxiliary regression delivers δ?
Q4 · Why is Step 1's SE so large?
FWL Step 1 regresses raw sales on residualized coupons with no intercept. The slope is right (—) but the SE is —, against — in the full model. What is mainly responsible?
Q5 · Step 2's SE vs. the full model's
FWL Step 2 residualizes both variables and reports SE —; the full regression reports —. What explains the gap?
Q6 · Add the sample means back
To make the residual plot readable, you add the sample means back (coupons —, sales —) and regress with an intercept. Does the coupon slope change?
Q7 · What does Double Machine Learning replace?
DML keeps FWL's logic of residualizing and then regressing residual on residual. Which part of the FWL recipe does it replace?