- Scope: Assignment 1, /model-selection
- Related: DR-002, DR-003
Context
Assignment 1 asked for the "best" logistic regression for a binary survey outcome and an interpretation of its odds ratios. The course-provided survey records whether 1,316 women reported physical domestic violence or mental abuse in the previous 12 months (369 did), with age band, education band, marital status (six levels), smoking, heavy drinking, a family member's drinking while growing up, marrying more than once and region (four levels) as candidate predictors.
Two decisions were bundled together. The first was statistical: how to search a space that grows to 66 parameters once two-way interactions are allowed. The second came with the 2026 revival: the subject matter is sensitive, so what could be published at all.
Decision
I selected the model in stages, using analysis-of-deviance tests to prune main effects, an AIC search (step()) over all two-way interactions, and then likelihood-ratio tests to remove interactions that did not earn their place, ending with an 18-parameter model; and the site publishes only model-level summaries of that path, never the survey rows.
The final model is dv ~ age + factor(ms) + smok + falc + educ + factor(reg) + factor(ms):falc.
Options considered
- Keep every main effect, no selection (model0, 18 parameters, AIC 1453.7). Simple and honest about uncertainty, but includes heavy drinking, which added little (sequential p = 0.116).
- Pure AIC search from the all-interactions model (model4, 22 parameters, AIC 1450.5). Lowest AIC on the path, but it keeps a smoking by family-alcohol interaction with p = 0.21.
- AIC search, then prune by tests (chosen; model6, 18 parameters, AIC 1451.1).
- Penalised regression (lasso or elastic net with cross-validated tuning). Better suited to a large candidate space, but outside what the subject taught and what the assignment asked for.
- A small set of pre-specified models chosen from the literature before looking at the data.
For publication: release the CSV, release a synthetic copy, or release model-level summaries only (chosen).
Why
- The assignment rewarded a parsimonious, interpretable model built with the tools the subject taught: deviance tests and AIC.
- Moving from model4 to model6 drops four parameters for an AIC cost of 0.56, well inside the range where AIC does not separate models.
- Coding age and education as scores instead of factors cost little (likelihood-ratio test against the factor coding: deviance 7.10 on 3 df, p = 0.069) and makes their odds ratios one number each.
- The survey is about family violence. The individual responses are not mine to publish, and a portfolio site gains nothing from them: deviances, coefficients and the covariance matrix reproduce every number on the page.
scripts/a1-model-selection.Rre-runs the original code privately and exports only those summaries.
What happened
- Final model: deviance 1415.09 on 1,298 df, AIC 1451.09. Smokers had 1.70 times the odds of reporting violence or abuse (95% CI 1.28 to 2.27); each higher education level multiplied the odds by 0.61 (0.48 to 0.78). Marital status interacts with a family member's drinking: for de facto partners the odds ratio against married women moves from about 2.2 to about 0.37 depending on it.
- The evidence for the exact model is weak. The last three models on the path (model4 to model6) sit within 0.6 AIC units of each other, and two of the choices were borderline: the score coding (p = 0.069) and dropping education by region (p = 0.089 against model5). A different analyst could reasonably have kept either.
- Marrying more than once (
mmo, sequential p = 0.043 in model0) was left out of the interaction search without a recorded test. Nothing in the 2023 files gives a reason, so it is recorded here as a gap rather than a choice. - The intervals on the final model ignore the selection step, so they are narrower than they should be. Post-selection inference was not covered in the subject and the 2023 write-up does not mention it.
- No calibration or discrimination check was reported, and there was no held-out data.
- The 2023 text described odds ratios below 1 as "times more" and transposed a digit in the smoking odds ratio (printed 1.7004437; it is 1.7044). The site shows the corrected reading and leaves the original files unchanged.
What I'd change
- Fix a small set of candidate models from subject knowledge before fitting, and report all of them.
- If a search is needed, use penalised logistic regression with cross-validated tuning, and report how stable the selection is (bootstrap inclusion frequencies) rather than a single winner.
- Treat p-values and intervals after selection as descriptive, or use a method that accounts for the search.
- Add a calibration plot and a discrimination measure with a confidence interval, on held-out data or with optimism correction by bootstrap.
- Keep the publication rule: summaries only. If a richer public artefact were ever needed, a synthetic data set with a documented disclosure-risk check would come before any real rows.