MAST90139 · University of Melbourne · 2023 Semester 1
GLM Playground
Six case studies from my Statistical Modelling for Data Science assignments. Logistic, log-linear, multinomial, proportional-odds and GEE models are refitted live in your browser, and checked against the R code I submitted in 2023.
Take the guided tour · three short recorded walkthroughs
The coursework
What the assignments asked
MAST90139 takes regression beyond the normal linear model: binary, count and ordered responses, contingency tables and correlated data. The three individual assignments (written in R Markdown) asked me to fit, test and interpret these models on small published data sets.
Assignment 1
Choosing a logistic model
Select a parsimonious logistic regression for a binary survey outcome with forward and backward steps, analysis of deviance and AIC, then interpret its odds ratios — including an interaction.
Assignment 2
Binomial and Poisson GLMs
A dose–response model with intervals, LD50 and goodness-of-fit tests; additive versus interaction models for survey attitudes; and log-linear tests of independence in a three-way table, with their logistic equivalents.
Assignment 3
Multicategory and correlated responses
Write down and interpret nominal (baseline-category) and ordinal (proportional-odds) models for disease severity, and a GEE model with exchangeable correlation for repeated binary measurements.
Interactive
Six case studies
Each page states the model, lets you change its inputs, and recomputes estimates, intervals and tests on the spot.
- Assignment 2 · Q1
Beetle mortality and insecticide dose
Binomial GLM, logit link
Drag the dose, watch the fitted kill probability and its interval, find the LD50 and test a quadratic term.
LD50 1.771
95% CI (1.764, 1.779); quadratic term p = 0.0035
- Assignment 2 · Q2
Education, gender and attitudes to women's role
Binomial GLM with factors and interactions
Six candidate models side by side, from saturated to the sex-specific education slope I chose.
× 0.708
odds of agreeing per extra year of school (women; men × 0.768)
- Assignment 2 · Q3
Race and the death penalty: a three-way table
Poisson log-linear models
Test four independence hypotheses, see their logistic twins, and meet Simpson's paradox.
p = 0.3903
only [VD][VP] fits: the victim's race, not the defendant's, predicts the sentence
- Assignment 3 · Q1–Q2
Coal-face exposure and lung disease severity
Multinomial vs proportional-odds logit
Slide the years of exposure and compare a nominal and an ordinal model category by category.
AIC 422.9
proportional odds beats the nominal model (425.4) with one fewer parameter
- Assignment 3 · Q3
Children's wheeze, age and maternal smoking
GEE, exchangeable working correlation
Repeated binary outcomes: robust sandwich errors versus a naive GLM, and an odds-ratio calculator.
α = 0.354
within-child correlation; robust SE for smoking 0.178 vs naive 0.123
- Assignment 1
Stepwise logistic model selection
Logistic regression, analysis of deviance
The selection path and final odds ratios, published as summaries only because the survey is sensitive.
1.70×
odds for smokers in the final 18-parameter model (summaries only)
Added in 2026
Checked, documented and explained
The coursework fitted and interpreted the models. The revival adds what I would expect of the same analysis at work: checks of each model, written reasons for each choice, and an optional AI reading aid that is labelled, reviewed and logged.
Model checking
Simulated residual envelopes, influence, over-dispersion, profile-likelihood intervals, a proportional-odds test and GEE sensitivity, every simulation seeded.
Methods and decision records
Data provenance, evaluation design, assumptions and limitations, three decision records and a model card, with the weak numbers stated.
Optional AI, your own key
Explanations grounded only on each page's numeric summary, labelled as AI-generated, checked for untraceable numbers and recorded in an exportable audit log.
Faithful ports
Same algorithms, same answers
The R functions I used — glm, nnet::multinom, MASS::polr and geepack::geeglm — are reimplemented in around 1,300 lines of dependency-free TypeScript: IRLS with Householder QR, Newton–Raphson for the multinomial and cumulative logit likelihoods, and GEE with a sandwich estimator.
A script re-runs the original model calls and records R's output; the test suite holds the ports to it, typically within 10⁻⁹. Where the 2023 write-up contained a slip, the page says so and shows the recomputed value.
49/49
parity checks on /verification
10⁻⁹
typical agreement with glm()
5
published data sets, no raw survey rows
0.028 → 0.136
smoking p-value, naive GLM → GEE
About this project
From R Markdown to an interactive site
- Subject
- MAST90139 Statistical Modelling for Data Science
- University
- University of Melbourne
- Year / semester
- 2023 Semester 1
- Format
- Three individual assignments (no team)
- Author
- Sunchuangyu (Rin) Huang
- Original stack
- R 4.2.2, R Markdown, glm, nnet, MASS, geepack, dplyr, ggplot2
- Revived stack
- Next.js 16, React 19, TypeScript ports with vitest parity tests, Tailwind CSS, KaTeX
- Hosting
- Static pages on Vercel. Fits are computed at build time and refitted in your browser in the interactive panels
- Repository
- rNLKJA/Unimelb-Master-2023-MAST90139 on GitHub, kept private because coursework/ holds University of Melbourne material and the Assignment 1 survey.
Academic integrity. The original submissions are preserved unchanged in the (private) repository under coursework/ for reference. Assignment specifications, model answers and lecture material belong to the University of Melbourne and are not reproduced on this site; the tasks above are paraphrased.
Data. The beetle, attitudes, death-penalty, pneumoconiosis and Ohio wheeze data are classic published teaching data sets, cited on each page. The Assignment 1 survey is sensitive and appears only as model summaries.