Downloads
Each dataset is available as a labeled Stata .dta and its source file.
⇩ Download all data (ZIP)stata_codebook.do
| Dataset | Grain | Rows | Stata | Source |
|---|---|---|---|---|
fwl_store_data | restaurant (cross-section, January) | 50 × 4 | fwl_store_data.dta | fwl_store_data.csv |
fwl_restaurant_panel | restaurant x month (balanced panel) | 100 × 6 | fwl_restaurant_panel.dta | fwl_restaurant_panel.csv |
Run stata_codebook.do in Stata once to attach long-form per-variable notes to the .dta files.
Load directly in code
Every file loads straight from GitHub (raw URLs). Swap the file name to load any dataset.
Stata
* Stata 14+ : `use` reads an https URL directly
global BASE "https://raw.githubusercontent.com/cmg777/starter-academic-v501/master/content/post/python_fwl/data/"
use "${BASE}fwl_store_data.dta", clear
describe
notesPython
!pip install -q pyreadstat
import pandas as pd
BASE = "https://raw.githubusercontent.com/cmg777/starter-academic-v501/master/content/post/python_fwl/data/"
df = pd.read_stata(BASE + "fwl_store_data.dta")
# load every dataset at once
files = ["fwl_store_data", "fwl_restaurant_panel"]
data = {f: pd.read_stata(BASE + f + ".dta") for f in files}
# pyreadstat (richest metadata) reads LOCAL files -> download first
import pyreadstat, urllib.request
urllib.request.urlretrieve(BASE + "fwl_store_data.dta", "fwl_store_data.dta")
df, meta = pyreadstat.read_dta("fwl_store_data.dta")Copy and paste this snippet in Google Colab app. https://colab.research.google.com/notebooks/empty.ipynb
R
# R : haven::read_dta auto-downloads an https URL
library(haven)
BASE <- "https://raw.githubusercontent.com/cmg777/starter-academic-v501/master/content/post/python_fwl/data/"
df <- read_dta(paste0(BASE, "fwl_store_data.dta"))Overview & sources
Companion data for the tutorial The FWL Theorem: Making Multivariate Regressions Intuitive, inspired by Courthoud (2022). A fast-food chain runs 50 simulated restaurants, one per neighborhood. Each restaurant hands out 100 discount coupons on a single day and counts how many are redeemed during the month: coupons is that redemption rate in percent, and sales is the restaurant's monthly sales. The data are built so that the answer is known in advance: a higher redemption rate raises monthly sales by exactly +0.2 per percentage point, but neighborhood income lowers redemption and raises sales, so it confounds the comparison. Ignoring income, the slope of sales on coupons is −0.1059 (p = 0.365); controlling for income it becomes +0.2673 (p = 0.031), and the Frisch–Waugh–Lovell (FWL) theorem recovers the same +0.2673 from a regression of residualized sales on residualized coupons. The appendix repeats the promotion in June and applies FWL to the two-month panel (restaurant and two-way fixed effects, first differences).
fwl_store_data is the January cross-section: 50 simulated restaurants with four numeric columns (sales, coupons, income, dayofweek), no missing values and no identifier (rows are in the order the restaurants were drawn). It is the exact input for every estimate in the main body of the post, the slides and the web app. fwl_restaurant_panel is the balanced panel of the appendix: the same 50 restaurants in January (period = 1, identical to fwl_store_data, row for row) and June (period = 2), 100 rows keyed by restaurant_id × period. Both files are written by the post's script.py (numpy default_rng(42)).
Data sources
| Source | Provides | Reference / URL |
|---|---|---|
| Simulated (this study) | fwl_store_data — 50 simulated restaurants (January) with a known confounder (income) and a known true coupon effect (+0.2); fwl_restaurant_panel — the same restaurants in January and June | Mendez, C. (2026). Data-generating process in script.py (numpy default_rng(42), PCG64): https://raw.githubusercontent.com/cmg777/starter-academic-v501/master/content/post/python_fwl/script.py |
| Scenario inspiration | The coupons, income and sales confounding example | Courthoud, M. (2022). Understanding the Frisch-Waugh-Lovell Theorem. Towards Data Science (originally published as "The FWL Theorem, Or How To Make All Regressions Intuitive"). https://towardsdatascience.com/the-fwl-theorem-or-how-to-make-all-regressions-intuitive-59f801eb3299/ |
| Method references | The Frisch–Waugh–Lovell theorem | Frisch, R., & Waugh, F. V. (1933). Econometrica, 1(4), 387–401. https://www.jstor.org/stable/1907330 — Lovell, M. C. (1963). Journal of the American Statistical Association, 58(304), 993–1010. https://doi.org/10.1080/01621459.1963.10480682 |
Cite this data
Please cite this dataset as follows.
APA
Mendez, C. (2026). Simulated fast-food restaurants (FWL tutorial) [Data set]. https://carlos-mendez.org/post/python_fwl/
Courthoud, M. (2022). Understanding the Frisch-Waugh-Lovell Theorem. Towards Data Science (originally published as "The FWL Theorem, Or How To Make All Regressions Intuitive"). https://towardsdatascience.com/the-fwl-theorem-or-how-to-make-all-regressions-intuitive-59f801eb3299/ — Frisch, R., & Waugh, F. V. (1933). Partial Time Regressions as Compared with Individual Trends. Econometrica, 1(4), 387–401. — Lovell, M. C. (1963). Seasonal Adjustment of Economic Time Series and Multiple Regression Analysis. Journal of the American Statistical Association, 58(304), 993–1010.BibTeX
@misc{mendez2026pythonfwl,
author = {Mendez, Carlos},
title = {Simulated fast-food restaurants (FWL tutorial)},
year = {2026},
howpublished = {\url{https://carlos-mendez.org/post/python_fwl/}},
note = {Data set}
}
@misc{courthoud2022fwl,
author = {Courthoud, Matteo},
title = {Understanding the {Frisch-Waugh-Lovell} Theorem},
year = {2022},
howpublished = {Towards Data Science},
note = {Originally published as ``The {FWL} Theorem, Or How To Make All Regressions Intuitive''},
url = {https://towardsdatascience.com/the-fwl-theorem-or-how-to-make-all-regressions-intuitive-59f801eb3299/}
}
@article{frisch1933partial,
author = {Frisch, Ragnar and Waugh, Frederick V.},
title = {Partial Time Regressions as Compared with Individual Trends},
journal = {Econometrica}, volume = {1}, number = {4}, pages = {387--401}, year = {1933}
}
@article{lovell1963seasonal,
author = {Lovell, Michael C.},
title = {Seasonal Adjustment of Economic Time Series and Multiple Regression Analysis},
journal = {Journal of the American Statistical Association},
volume = {58}, number = {304}, pages = {993--1010}, year = {1963}
}Variable explorer search & filter all 6 variables
Type to filter by name or label, or use the chips to filter by type. Each row shows a mini distribution. Click a header to sort.
| Variable | Type | Distribution | Label | Definition | Units | In files | Source |
|---|---|---|---|---|---|---|---|
coupons# | continuous | Coupon redemption rate, treatment (percent) | Share of the restaurant's 100 coupons, handed out on one day, that were redeemed during the month, in percent — the treatment (regressor of interest). Its true effect on sales is +0.2 per percentage point. | percent | fwl_store_data, fwl_restaurant_panel | Simulation (script.py, seed 42) | |
dayofweek# | continuous | Day-of-week index (1-7) | Day-of-week index, an integer from 1 to 7. It shifts sales (+0.5 per step) but is unrelated to coupons by construction; the post adds it as a second control. No weekday names are attached to the codes. | index 1–7 | fwl_store_data, fwl_restaurant_panel | Simulation (script.py, seed 42) | |
income# | continuous | Neighborhood income, confounder (thousands of US dollars) | Income of the restaurant's neighborhood — the confounder: it lowers the redemption rate (−0.5 per unit) and raises sales (+0.3 per unit). | thousands of US dollars | fwl_store_data, fwl_restaurant_panel | Simulation (script.py, seed 42) | |
period# | continuous | Month of the promotion (1 = January, 2 = June) | Month in which the restaurant handed out its 100 coupons: 1 = January (the main-body data), 2 = June. Panel file only. | code | fwl_restaurant_panel | Simulation (script.py, seed 42) | |
restaurant_id# | continuous | Restaurant identifier (1-50) | Restaurant identifier, 1 to 50, in the order the restaurants were drawn. Panel file only; row i of fwl_store_data is restaurant_id i. | id | fwl_restaurant_panel | Simulation (script.py, seed 42) | |
sales# | continuous | Monthly sales, outcome (thousands of US dollars) | Monthly sales of the restaurant — the outcome variable. One unit is 1,000 US dollars. | thousands of US dollars | fwl_store_data, fwl_restaurant_panel | Simulation (script.py, seed 42) |
Cross-file variable index
Which file each variable appears in (● = present).
Construction & formulas
Data-generating process
Simulated by simulate_store_data(n=50, seed=42) in script.py with
numpy's default_rng(42); N(mean, SD) denotes a normal draw.
income ~ N(50, 10)dayofweek ~ Uniform{1, 2, …, 7}(rng.integers(1, 8))coupons = 60 − 0.5 × income + u,u ~ N(0, 5)sales = 10 + 0.2 × coupons + 0.3 × income + 0.5 × dayofweek + e,e ~ N(0, 3)
The true causal effect of coupons on sales is +0.2 by construction: one
percentage point of redemption (one more of the 100 coupons redeemed) adds 0.2 thousand US
dollars (200 US dollars) of monthly sales. sales is computed from the unrounded
draws; sales, coupons and income are then rounded to 2
decimals.
The June round (panel appendix)
simulate_restaurant_panel(n=50, seed=42) replays the January draws exactly and
continues the same random stream. January's sales noise α ~ N(0, 3) is a persistent
restaurant trait that raises sales in both months; in June it also raises redemption.
incomeJun = incomeJan + N(0, 2);dayofweekJun ~ Uniform{1, …, 7}couponsJun = 60 − 0.5 × incomeJun + α + N(0, 4)salesJun = 13 + 0.2 × couponsJun + 0.3 × incomeJun + 0.5 × dayofweekJun + α + N(0, 1.5)(a summer lift of +3)
Pooled OLS with income and a June dummy gives 0.4496 (biased by the trait); restaurant fixed effects give 0.2138; two-way fixed effects and first differences both give 0.1394 (clustered SE 0.0691), which is within sampling noise of the true 0.2.
The Frisch–Waugh–Lovell (FWL) theorem
In Y = X1β1 + X2β2 + ε, let
M2 = I − X2(X2′X2)−1X2′
be the residual maker for the controls (here X2 = a constant and
income; M2 is symmetric and idempotent). Then
β̂1 = (X1′M2Y) / (X1′M2X1)
— the multivariate coefficient equals the slope of a regression of residualized Y
on residualized X1.
- Full regression
sales ~ coupons + income: coupons 0.2673 (SE 0.1203, p = 0.031); incomeγ̂ = 0.3836. - Residualize coupons only:
coupons_tilde= residuals ofcoupons ~ income(uncorrelated with income, not independent of it). Regressing rawsalesoncoupons_tildewith no intercept returns the same 0.2673, but its SE of 1.2715 (p = 0.834) is inflated mainly by the dropped intercept: with an intercept the SE is 0.1437. - Residualize both:
sales_tilde= residuals ofsales ~ income;sales_tilde ~ coupons_tilde - 1returns 0.2673 with SE 0.1178 = 0.1203 × √(47/49). The residuals are identical to the full model's; the full model spends two extra degrees of freedom (intercept + income). - By hand:
Cov(sales_tilde, coupons_tilde) / Var(coupons_tilde) = 3.9380 / 14.7320 = 0.2673. - Two controls (income + dayofweek partialled out together): full OLS and FWL both give 0.2706.
Omitted-variable-bias identity
Exact in the sample:
β̂naive = β̂full + γ̂ × δ̂, where γ̂ is income's coefficient in
the full regression and δ̂ is the slope from regressing the omitted variable
(income) on coupons:
−0.1059 = 0.2673 + 0.3836 × (−0.9730) = 0.2673 − 0.3732
The auxiliary regression must run in this direction (income ~ coupons);
regressing coupons on income instead (slope −0.3935) gives a product of −0.151 that reconciles
nothing. In the population δ = Cov(income, coupons) / Var(coupons) = −50 / 50 = −1.0, so the
naive slope converges to 0.2 + 0.3 × (−1.0) = −0.10 — confounding flips the sign of a positive
effect.
The datasets
Switch datasets with the tabs. Each shows the full variable dictionary plus a sortable statistics table with mini distributions and data coverage.
expand to search (Ctrl/⌘+F) or print across all datasets
Variable dictionary
| Variable | Label | Definition | Construction | Units | Source | Coverage |
|---|---|---|---|---|---|---|
sales continuous | Monthly sales, outcome (thousands of US dollars) | Monthly sales of the restaurant — the outcome variable. One unit is 1,000 US dollars. | sales = 10 + 0.2 × coupons + 0.3 × income + 0.5 × dayofweek + e, e ~ N(0, 3); computed from the unrounded draws, then rounded to 2 decimals. | thousands of US dollars | Simulation (script.py, seed 42) | 50 restaurants (100 restaurant-months in the panel) |
coupons continuous | Coupon redemption rate, treatment (percent) | Share of the restaurant's 100 coupons, handed out on one day, that were redeemed during the month, in percent — the treatment (regressor of interest). Its true effect on sales is +0.2 per percentage point. | coupons = 60 − 0.5 × income + u, u ~ N(0, 5); rounded to 2 decimals. Wealthier neighborhoods redeem fewer coupons. In June (panel) the persistent restaurant trait is added. | percent | Simulation (script.py, seed 42) | 50 restaurants (100 restaurant-months in the panel) |
income continuous | Neighborhood income, confounder (thousands of US dollars) | Income of the restaurant's neighborhood — the confounder: it lowers the redemption rate (−0.5 per unit) and raises sales (+0.3 per unit). | income ~ N(50, 10); in June (panel) January income + N(0, 2); rounded to 2 decimals. | thousands of US dollars | Simulation (script.py, seed 42) | 50 restaurants (100 restaurant-months in the panel) |
dayofweek continuous | Day-of-week index (1-7) | Day-of-week index, an integer from 1 to 7. It shifts sales (+0.5 per step) but is unrelated to coupons by construction; the post adds it as a second control. No weekday names are attached to the codes. | rng.integers(1, 8): uniform on the integers 1 to 7 (the upper bound 8 is exclusive). | index 1–7 | Simulation (script.py, seed 42) | 50 restaurants (100 restaurant-months in the panel) |
Distribution & statistics (click a header to sort)
| Variable | Distribution | Coverage | N | Distinct | Min | Mean | Median | Max | SD |
|---|---|---|---|---|---|---|---|---|---|
sales | 100% | 50 | 50 | 25.76 | 33.61 | 33.24 | 44.38 | 3.96 | |
coupons | 100% | 50 | 50 | 23.26 | 33.84 | 33.25 | 43.79 | 4.89 | |
income | 100% | 50 | 50 | 30.49 | 50.91 | 51.74 | 71.42 | 7.68 | |
dayofweek | 100% | 50 | 7 | 1.00 | 3.92 | 4.00 | 7.00 | 1.88 |
Variable dictionary
| Variable | Label | Definition | Construction | Units | Source | Coverage |
|---|---|---|---|---|---|---|
restaurant_id continuous | Restaurant identifier (1-50) | Restaurant identifier, 1 to 50, in the order the restaurants were drawn. Panel file only; row i of fwl_store_data is restaurant_id i. | np.arange(1, n + 1). | id | Simulation (script.py, seed 42) | 50 restaurants x 2 periods |
period continuous | Month of the promotion (1 = January, 2 = June) | Month in which the restaurant handed out its 100 coupons: 1 = January (the main-body data), 2 = June. Panel file only. | 1 for the January rows, 2 for the June rows. | code | Simulation (script.py, seed 42) | 50 restaurants x 2 periods |
sales continuous | Monthly sales, outcome (thousands of US dollars) | Monthly sales of the restaurant — the outcome variable. One unit is 1,000 US dollars. | sales = 10 + 0.2 × coupons + 0.3 × income + 0.5 × dayofweek + e, e ~ N(0, 3); computed from the unrounded draws, then rounded to 2 decimals. | thousands of US dollars | Simulation (script.py, seed 42) | 50 restaurants (100 restaurant-months in the panel) |
coupons continuous | Coupon redemption rate, treatment (percent) | Share of the restaurant's 100 coupons, handed out on one day, that were redeemed during the month, in percent — the treatment (regressor of interest). Its true effect on sales is +0.2 per percentage point. | coupons = 60 − 0.5 × income + u, u ~ N(0, 5); rounded to 2 decimals. Wealthier neighborhoods redeem fewer coupons. In June (panel) the persistent restaurant trait is added. | percent | Simulation (script.py, seed 42) | 50 restaurants (100 restaurant-months in the panel) |
income continuous | Neighborhood income, confounder (thousands of US dollars) | Income of the restaurant's neighborhood — the confounder: it lowers the redemption rate (−0.5 per unit) and raises sales (+0.3 per unit). | income ~ N(50, 10); in June (panel) January income + N(0, 2); rounded to 2 decimals. | thousands of US dollars | Simulation (script.py, seed 42) | 50 restaurants (100 restaurant-months in the panel) |
dayofweek continuous | Day-of-week index (1-7) | Day-of-week index, an integer from 1 to 7. It shifts sales (+0.5 per step) but is unrelated to coupons by construction; the post adds it as a second control. No weekday names are attached to the codes. | rng.integers(1, 8): uniform on the integers 1 to 7 (the upper bound 8 is exclusive). | index 1–7 | Simulation (script.py, seed 42) | 50 restaurants (100 restaurant-months in the panel) |
Distribution & statistics (click a header to sort)
| Variable | Distribution | Coverage | N | Distinct | Min | Mean | Median | Max | SD |
|---|---|---|---|---|---|---|---|---|---|
restaurant_id | 100% | 100 | 50 | 1.00 | 25.50 | 25.50 | 50.00 | 14.50 | |
period | 100% | 100 | 2 | 1.00 | 1.50 | 1.50 | 2.00 | 0.503 | |
sales | 100% | 100 | 94 | 25.76 | 35.34 | 35.22 | 50.53 | 4.72 | |
coupons | 100% | 100 | 99 | 21.47 | 34.17 | 34.45 | 54.63 | 5.68 | |
income | 100% | 100 | 97 | 30.49 | 50.98 | 51.73 | 71.42 | 7.70 | |
dayofweek | 100% | 100 | 7 | 1.00 | 4.08 | 4.00 | 7.00 | 1.95 |
Known limitations & caveats
- Simulated, not observed. The 50 restaurants come from a known data-generating process with a true coupon effect of +0.2 by construction. The files illustrate confounding and the FWL theorem, not a real fast-food market.
couponskeeps two decimals, so read 36.93 as about 37 of the 100 coupons. - Generated in Python with numpy's PCG64 generator (
np.random.default_rng(42)inscript.py).sales,couponsandincomeare rounded to 2 decimals, and every estimate in the post is computed on these rounded values. - Load this file to reproduce the post in R or Stata. Native simulations with seed 42 (
set.seed(42)in R,set seed 42in Stata) use different random-number generators, so they produce different draws and different estimates. Readfwl_store_data.csvorfwl_store_data.dtainstead of re-simulating. - Small sample. With n = 50 the controlled estimate (+0.2673, SE 0.1203) is close to, but not equal to, the true +0.2; the naive slope (−0.1059) sits near its population limit of −0.10.
- dayofweek is a bare 1–7 index. The simulation attaches no weekday names to the codes, and the variable enters the data-generating process linearly (slope +0.5). It is drawn independently of income and coupons (sample correlation with coupons −0.076).
- No identifier in the January file.
fwl_store_datarows are in the order the restaurants were drawn, which is the order ofrestaurant_id1–50 infwl_restaurant_panel.