← Back to the post
Interactive data dictionary

Simulated fast-food restaurants (FWL tutorial)

Fifty simulated fast-food restaurants, one per neighborhood, in which neighborhood income confounds the effect of the coupon redemption rate on monthly sales — the data behind every number in the Python FWL tutorial, plus the January + June panel of its appendix.

2
datasets
6
variables
50
restaurants
+0.2
true coupon effect

Downloads

Each dataset is available as a labeled Stata .dta and its source file.

⇩ Download all data (ZIP)stata_codebook.do

DatasetGrainRowsStataSource
fwl_store_datarestaurant (cross-section, January)50 × 4fwl_store_data.dtafwl_store_data.csv
fwl_restaurant_panelrestaurant x month (balanced panel)100 × 6fwl_restaurant_panel.dtafwl_restaurant_panel.csv

Run stata_codebook.do in Stata once to attach long-form per-variable notes to the .dta files.

Load directly in code

Every file loads straight from GitHub (raw URLs). Swap the file name to load any dataset.

Stata

* Stata 14+ : `use` reads an https URL directly
global BASE "https://raw.githubusercontent.com/cmg777/starter-academic-v501/master/content/post/python_fwl/data/"
use "${BASE}fwl_store_data.dta", clear
describe
notes

Python

!pip install -q pyreadstat
import pandas as pd
BASE = "https://raw.githubusercontent.com/cmg777/starter-academic-v501/master/content/post/python_fwl/data/"
df = pd.read_stata(BASE + "fwl_store_data.dta")

# load every dataset at once
files = ["fwl_store_data", "fwl_restaurant_panel"]
data = {f: pd.read_stata(BASE + f + ".dta") for f in files}

# pyreadstat (richest metadata) reads LOCAL files -> download first
import pyreadstat, urllib.request
urllib.request.urlretrieve(BASE + "fwl_store_data.dta", "fwl_store_data.dta")
df, meta = pyreadstat.read_dta("fwl_store_data.dta")

Copy and paste this snippet in Google Colab app. https://colab.research.google.com/notebooks/empty.ipynb

R

# R : haven::read_dta auto-downloads an https URL
library(haven)
BASE <- "https://raw.githubusercontent.com/cmg777/starter-academic-v501/master/content/post/python_fwl/data/"
df <- read_dta(paste0(BASE, "fwl_store_data.dta"))

Overview & sources

Companion data for the tutorial The FWL Theorem: Making Multivariate Regressions Intuitive, inspired by Courthoud (2022). A fast-food chain runs 50 simulated restaurants, one per neighborhood. Each restaurant hands out 100 discount coupons on a single day and counts how many are redeemed during the month: coupons is that redemption rate in percent, and sales is the restaurant's monthly sales. The data are built so that the answer is known in advance: a higher redemption rate raises monthly sales by exactly +0.2 per percentage point, but neighborhood income lowers redemption and raises sales, so it confounds the comparison. Ignoring income, the slope of sales on coupons is −0.1059 (p = 0.365); controlling for income it becomes +0.2673 (p = 0.031), and the Frisch–Waugh–Lovell (FWL) theorem recovers the same +0.2673 from a regression of residualized sales on residualized coupons. The appendix repeats the promotion in June and applies FWL to the two-month panel (restaurant and two-way fixed effects, first differences).

Two files. fwl_store_data is the January cross-section: 50 simulated restaurants with four numeric columns (sales, coupons, income, dayofweek), no missing values and no identifier (rows are in the order the restaurants were drawn). It is the exact input for every estimate in the main body of the post, the slides and the web app. fwl_restaurant_panel is the balanced panel of the appendix: the same 50 restaurants in January (period = 1, identical to fwl_store_data, row for row) and June (period = 2), 100 rows keyed by restaurant_id × period. Both files are written by the post's script.py (numpy default_rng(42)).

Data sources

SourceProvidesReference / URL
Simulated (this study)fwl_store_data — 50 simulated restaurants (January) with a known confounder (income) and a known true coupon effect (+0.2); fwl_restaurant_panel — the same restaurants in January and JuneMendez, C. (2026). Data-generating process in script.py (numpy default_rng(42), PCG64): https://raw.githubusercontent.com/cmg777/starter-academic-v501/master/content/post/python_fwl/script.py
Scenario inspirationThe coupons, income and sales confounding exampleCourthoud, M. (2022). Understanding the Frisch-Waugh-Lovell Theorem. Towards Data Science (originally published as "The FWL Theorem, Or How To Make All Regressions Intuitive"). https://towardsdatascience.com/the-fwl-theorem-or-how-to-make-all-regressions-intuitive-59f801eb3299/
Method referencesThe Frisch–Waugh–Lovell theoremFrisch, R., & Waugh, F. V. (1933). Econometrica, 1(4), 387–401. https://www.jstor.org/stable/1907330 — Lovell, M. C. (1963). Journal of the American Statistical Association, 58(304), 993–1010. https://doi.org/10.1080/01621459.1963.10480682

Cite this data

Please cite this dataset as follows.

APA

Mendez, C. (2026). Simulated fast-food restaurants (FWL tutorial) [Data set]. https://carlos-mendez.org/post/python_fwl/

Courthoud, M. (2022). Understanding the Frisch-Waugh-Lovell Theorem. Towards Data Science (originally published as "The FWL Theorem, Or How To Make All Regressions Intuitive"). https://towardsdatascience.com/the-fwl-theorem-or-how-to-make-all-regressions-intuitive-59f801eb3299/ — Frisch, R., & Waugh, F. V. (1933). Partial Time Regressions as Compared with Individual Trends. Econometrica, 1(4), 387–401. — Lovell, M. C. (1963). Seasonal Adjustment of Economic Time Series and Multiple Regression Analysis. Journal of the American Statistical Association, 58(304), 993–1010.

BibTeX

@misc{mendez2026pythonfwl,
  author       = {Mendez, Carlos},
  title        = {Simulated fast-food restaurants (FWL tutorial)},
  year         = {2026},
  howpublished = {\url{https://carlos-mendez.org/post/python_fwl/}},
  note         = {Data set}
}

@misc{courthoud2022fwl,
  author       = {Courthoud, Matteo},
  title        = {Understanding the {Frisch-Waugh-Lovell} Theorem},
  year         = {2022},
  howpublished = {Towards Data Science},
  note         = {Originally published as ``The {FWL} Theorem, Or How To Make All Regressions Intuitive''},
  url          = {https://towardsdatascience.com/the-fwl-theorem-or-how-to-make-all-regressions-intuitive-59f801eb3299/}
}
@article{frisch1933partial,
  author  = {Frisch, Ragnar and Waugh, Frederick V.},
  title   = {Partial Time Regressions as Compared with Individual Trends},
  journal = {Econometrica}, volume = {1}, number = {4}, pages = {387--401}, year = {1933}
}
@article{lovell1963seasonal,
  author  = {Lovell, Michael C.},
  title   = {Seasonal Adjustment of Economic Time Series and Multiple Regression Analysis},
  journal = {Journal of the American Statistical Association},
  volume  = {58}, number = {304}, pages = {993--1010}, year = {1963}
}

Variable explorer search & filter all 6 variables

Type to filter by name or label, or use the chips to filter by type. Each row shows a mini distribution. Click a header to sort.

VariableTypeDistributionLabelDefinitionUnitsIn filesSource
coupons#continuousmin 23.3 | median 33.3 | max 43.8Coupon redemption rate, treatment (percent)Share of the restaurant's 100 coupons, handed out on one day, that were redeemed during the month, in percent — the treatment (regressor of interest). Its true effect on sales is +0.2 per percentage point.percentfwl_store_data, fwl_restaurant_panelSimulation (script.py, seed 42)
dayofweek#continuousmin 1 | median 4 | max 7Day-of-week index (1-7)Day-of-week index, an integer from 1 to 7. It shifts sales (+0.5 per step) but is unrelated to coupons by construction; the post adds it as a second control. No weekday names are attached to the codes.index 1–7fwl_store_data, fwl_restaurant_panelSimulation (script.py, seed 42)
income#continuousmin 30.5 | median 51.7 | max 71.4Neighborhood income, confounder (thousands of US dollars)Income of the restaurant's neighborhood — the confounder: it lowers the redemption rate (−0.5 per unit) and raises sales (+0.3 per unit).thousands of US dollarsfwl_store_data, fwl_restaurant_panelSimulation (script.py, seed 42)
period#continuousmin 1 | median 1.5 | max 2Month of the promotion (1 = January, 2 = June)Month in which the restaurant handed out its 100 coupons: 1 = January (the main-body data), 2 = June. Panel file only.codefwl_restaurant_panelSimulation (script.py, seed 42)
restaurant_id#continuousmin 1 | median 25.5 | max 50Restaurant identifier (1-50)Restaurant identifier, 1 to 50, in the order the restaurants were drawn. Panel file only; row i of fwl_store_data is restaurant_id i.idfwl_restaurant_panelSimulation (script.py, seed 42)
sales#continuousmin 25.8 | median 33.2 | max 44.4Monthly sales, outcome (thousands of US dollars)Monthly sales of the restaurant — the outcome variable. One unit is 1,000 US dollars.thousands of US dollarsfwl_store_data, fwl_restaurant_panelSimulation (script.py, seed 42)

Cross-file variable index

Which file each variable appears in (● = present).

Variablefwl_store_datafwl_restaurant_panel
coupons●●
dayofweek●●
income●●
period●
restaurant_id●
sales●●

Construction & formulas

Data-generating process

Simulated by simulate_store_data(n=50, seed=42) in script.py with numpy's default_rng(42); N(mean, SD) denotes a normal draw.

The true causal effect of coupons on sales is +0.2 by construction: one percentage point of redemption (one more of the 100 coupons redeemed) adds 0.2 thousand US dollars (200 US dollars) of monthly sales. sales is computed from the unrounded draws; sales, coupons and income are then rounded to 2 decimals.

The June round (panel appendix)

simulate_restaurant_panel(n=50, seed=42) replays the January draws exactly and continues the same random stream. January's sales noise α ~ N(0, 3) is a persistent restaurant trait that raises sales in both months; in June it also raises redemption.

Pooled OLS with income and a June dummy gives 0.4496 (biased by the trait); restaurant fixed effects give 0.2138; two-way fixed effects and first differences both give 0.1394 (clustered SE 0.0691), which is within sampling noise of the true 0.2.

The Frisch–Waugh–Lovell (FWL) theorem

In Y = X1β1 + X2β2 + ε, let M2 = I − X2(X2′X2)−1X2′ be the residual maker for the controls (here X2 = a constant and income; M2 is symmetric and idempotent). Then

β̂1 = (X1′M2Y) / (X1′M2X1) — the multivariate coefficient equals the slope of a regression of residualized Y on residualized X1.

Omitted-variable-bias identity

Exact in the sample: β̂naive = β̂full + γ̂ × δ̂, where γ̂ is income's coefficient in the full regression and δ̂ is the slope from regressing the omitted variable (income) on coupons:

−0.1059 = 0.2673 + 0.3836 × (−0.9730) = 0.2673 − 0.3732

The auxiliary regression must run in this direction (income ~ coupons); regressing coupons on income instead (slope −0.3935) gives a product of −0.151 that reconciles nothing. In the population δ = Cov(income, coupons) / Var(coupons) = −50 / 50 = −1.0, so the naive slope converges to 0.2 + 0.3 × (−1.0) = −0.10 — confounding flips the sign of a positive effect.

The datasets

Switch datasets with the tabs. Each shows the full variable dictionary plus a sortable statistics table with mini distributions and data coverage.

expand to search (Ctrl/⌘+F) or print across all datasets

restaurant (cross-section, January)  50 × 4 · 50 simulated restaurants

Panel key: row order (no id column; row i is restaurant_id i in the panel) · Input for every estimate in the main body of the post: naive vs. full OLS, FWL residualization, the OVB identity and the two-control extension.

Variable dictionary

VariableLabelDefinitionConstructionUnitsSourceCoverage
sales continuousMonthly sales, outcome (thousands of US dollars)Monthly sales of the restaurant — the outcome variable. One unit is 1,000 US dollars.sales = 10 + 0.2 × coupons + 0.3 × income + 0.5 × dayofweek + e, e ~ N(0, 3); computed from the unrounded draws, then rounded to 2 decimals.thousands of US dollarsSimulation (script.py, seed 42)50 restaurants (100 restaurant-months in the panel)
coupons continuousCoupon redemption rate, treatment (percent)Share of the restaurant's 100 coupons, handed out on one day, that were redeemed during the month, in percent — the treatment (regressor of interest). Its true effect on sales is +0.2 per percentage point.coupons = 60 − 0.5 × income + u, u ~ N(0, 5); rounded to 2 decimals. Wealthier neighborhoods redeem fewer coupons. In June (panel) the persistent restaurant trait is added.percentSimulation (script.py, seed 42)50 restaurants (100 restaurant-months in the panel)
income continuousNeighborhood income, confounder (thousands of US dollars)Income of the restaurant's neighborhood — the confounder: it lowers the redemption rate (−0.5 per unit) and raises sales (+0.3 per unit).income ~ N(50, 10); in June (panel) January income + N(0, 2); rounded to 2 decimals.thousands of US dollarsSimulation (script.py, seed 42)50 restaurants (100 restaurant-months in the panel)
dayofweek continuousDay-of-week index (1-7)Day-of-week index, an integer from 1 to 7. It shifts sales (+0.5 per step) but is unrelated to coupons by construction; the post adds it as a second control. No weekday names are attached to the codes.rng.integers(1, 8): uniform on the integers 1 to 7 (the upper bound 8 is exclusive).index 1–7Simulation (script.py, seed 42)50 restaurants (100 restaurant-months in the panel)

Distribution & statistics (click a header to sort)

VariableDistributionCoverageNDistinctMinMeanMedianMaxSD
salesmin 25.8 | median 33.2 | max 44.4100%505025.7633.6133.2444.383.96
couponsmin 23.3 | median 33.3 | max 43.8100%505023.2633.8433.2543.794.89
incomemin 30.5 | median 51.7 | max 71.4100%505030.4950.9151.7471.427.68
dayofweekmin 1 | median 4 | max 7100%5071.003.924.007.001.88

restaurant x month (balanced panel)  100 × 6 · 50 restaurants x 2 periods = 100 rows

Panel key: restaurant_id x period · Input for the panel-data appendix: pooled OLS, restaurant fixed effects, two-way fixed effects and first differences, all shown to be FWL.

Variable dictionary

VariableLabelDefinitionConstructionUnitsSourceCoverage
restaurant_id continuousRestaurant identifier (1-50)Restaurant identifier, 1 to 50, in the order the restaurants were drawn. Panel file only; row i of fwl_store_data is restaurant_id i.np.arange(1, n + 1).idSimulation (script.py, seed 42)50 restaurants x 2 periods
period continuousMonth of the promotion (1 = January, 2 = June)Month in which the restaurant handed out its 100 coupons: 1 = January (the main-body data), 2 = June. Panel file only.1 for the January rows, 2 for the June rows.codeSimulation (script.py, seed 42)50 restaurants x 2 periods
sales continuousMonthly sales, outcome (thousands of US dollars)Monthly sales of the restaurant — the outcome variable. One unit is 1,000 US dollars.sales = 10 + 0.2 × coupons + 0.3 × income + 0.5 × dayofweek + e, e ~ N(0, 3); computed from the unrounded draws, then rounded to 2 decimals.thousands of US dollarsSimulation (script.py, seed 42)50 restaurants (100 restaurant-months in the panel)
coupons continuousCoupon redemption rate, treatment (percent)Share of the restaurant's 100 coupons, handed out on one day, that were redeemed during the month, in percent — the treatment (regressor of interest). Its true effect on sales is +0.2 per percentage point.coupons = 60 − 0.5 × income + u, u ~ N(0, 5); rounded to 2 decimals. Wealthier neighborhoods redeem fewer coupons. In June (panel) the persistent restaurant trait is added.percentSimulation (script.py, seed 42)50 restaurants (100 restaurant-months in the panel)
income continuousNeighborhood income, confounder (thousands of US dollars)Income of the restaurant's neighborhood — the confounder: it lowers the redemption rate (−0.5 per unit) and raises sales (+0.3 per unit).income ~ N(50, 10); in June (panel) January income + N(0, 2); rounded to 2 decimals.thousands of US dollarsSimulation (script.py, seed 42)50 restaurants (100 restaurant-months in the panel)
dayofweek continuousDay-of-week index (1-7)Day-of-week index, an integer from 1 to 7. It shifts sales (+0.5 per step) but is unrelated to coupons by construction; the post adds it as a second control. No weekday names are attached to the codes.rng.integers(1, 8): uniform on the integers 1 to 7 (the upper bound 8 is exclusive).index 1–7Simulation (script.py, seed 42)50 restaurants (100 restaurant-months in the panel)

Distribution & statistics (click a header to sort)

VariableDistributionCoverageNDistinctMinMeanMedianMaxSD
restaurant_idmin 1 | median 25.5 | max 50100%100501.0025.5025.5050.0014.50
periodmin 1 | median 1.5 | max 2100%10021.001.501.502.000.503
salesmin 25.8 | median 35.2 | max 50.5100%1009425.7635.3435.2250.534.72
couponsmin 21.5 | median 34.4 | max 54.6100%1009921.4734.1734.4554.635.68
incomemin 30.5 | median 51.7 | max 71.4100%1009730.4950.9851.7371.427.70
dayofweekmin 1 | median 4 | max 7100%10071.004.084.007.001.95

Known limitations & caveats