Methods
Methods, decisions and AI use
Where the data come from, how each model is fitted and checked, what it assumes, where it falls short, the decisions behind it, and exactly what the optional AI feature does.
Data provenance
Beetle mortality · A2 Q1
- Source
- Bliss (1935), Annals of Applied Biology 22, 134-167
- Size
- 481 beetles, 8 dose groups
- How it reaches the site
- Typed from the published table
Attitudes to women's role · A2 Q2
- Source
- National Opinion Research Center, University of Chicago, General Social Surveys 1974-1975
- Size
- 2,678 respondents, 24 cells
- How it reaches the site
- Typed from the published table
Death penalty · A2 Q3
- Source
- Radelet (1981), American Sociological Review 46, 918-927
- Size
- 326 defendants, 2 × 2 × 2 table
- How it reaches the site
- Typed from the published table
Pneumoconiosis · A3 Q1–Q2
- Source
- Ashford (1959), British Journal of Industrial Medicine 16, 268-278
- Size
- 371 miners, 8 exposure groups
- How it reaches the site
- Typed from the published table
Ohio wheeze · A3 Q3
- Source
- geepack::ohio (Fitzmaurice and Laird)
- Size
- 537 children, 2,148 observations
- How it reaches the site
- Aggregated by script into 32 response-pattern counts; the GEE fit is unchanged
Assignment 1 survey · A1
- Source
- Course-provided; sensitive (family violence)
- Size
- 1,316 respondents, 369 events
- How it reaches the site
- Never shipped: model-level summaries only
The five public data sets are classic published teaching tables; none was collected for this site, and none contains personal information. The Assignment 1 rows stay in the private repository (see DR-001).
Methods
The 2023 models, ported
- Binomial and Poisson GLMs by iteratively reweighted least squares with Householder QR, line by line as R's glm.fit.
- Multinomial (baseline-category) and proportional-odds models by Newton–Raphson with analytic scores and observed information.
- GEE with a sandwich covariance, matching geepack::geeglm.
- Analysis of deviance, likelihood-ratio tests, AIC, Wald and delta-method intervals.
Added in 2026
- Leverage, Cook's distance and standardised residuals; half-normal plots with simulated envelopes.
- Pearson dispersion, quasi-binomial standard errors and F tests.
- Profile-likelihood intervals for coefficients and for LD50, solved by root finding.
- LD50 under two alternative mean models (quadratic logit and complementary log-log), to show how much the estimate depends on the model.
- Brant's test and a likelihood-ratio test against separate slopes for proportional odds.
- AR(1) working correlation, QIC and CIC, and a cluster bootstrap for GEE.
- A random-intercept logistic GLMM by Gauss–Hermite quadrature.
- Wilson score intervals for every observed proportion; seeded parametric and cluster bootstraps.
Evaluation design
Numerical parity with R
Two R scripts re-run the original model calls and the 2026 checks with established packages and write reference values at full precision. The vitest suite holds every port to them, typically to 10⁻⁸ or better; 49 of 49 headline comparisons are shown on /verification.
Uncertainty on every result
95% intervals throughout: Wald or delta-method for coefficients and contrasts, profile likelihood where the likelihood may be skewed, Wilson score intervals for observed proportions, and percentile intervals for bootstraps. Sample sizes are stated with each data set.
Reproducible simulation
Every simulation uses a fixed, displayed seed (for example 20230501 for the beetle bootstrap and 20230601 for the wheeze bootstrap) and is repeated over 5 seeds to show how much Monte Carlo noise is left. Failed refits are counted and reported, never imputed.
Where two methods are compared on the same data (Wald against profile intervals, three dose-response mean models, naive GLM against GEE, three working correlations, GEE against GLMM), they are fitted to identical inputs and shown side by side, with effect sizes and intervals rather than p-values alone. The comparisons are on /diagnostics.
Assumptions
- Binomial GLMs. Independent binomial counts with a logit-linear mean. Checked with deviance and Pearson statistics, residual envelopes and influence measures.
- Log-linear models. Independent Poisson cell counts (equivalently, a multinomial table with fixed total); the models are compared by deviance.
- Proportional odds. One slope shared by every cumulative split. Checked with Brant's test and a nested likelihood-ratio test.
- GEE. A correctly specified marginal mean; the working correlation can be wrong, but the robust standard errors need many independent clusters (537 here).
- GLMM. A normally distributed random intercept per child, and conditional independence of the four visits given it.
- All models. Observations (or clusters) are independent of each other, and the published tables are accurate. Every model describes associations, not causes.
Limitations
The data are small and historical, chosen to teach methods rather than to answer current questions. The beetle straight-line model, which the assignment specified, does not fit well, and its intervals are conditional on it: under a quadratic or complementary log-log model the LD50 moves by up to 0.0086, more than the straight-line Wald half-width of 0.0076. The proportional-odds tests have little power with 82 abnormal cases. The Assignment 1 intervals ignore the model-selection step. Bootstrap results depend on Monte Carlo error, which the seed tables quantify but do not remove. The ports reproduce R closely, but optimiser-based fits in R (multinom, polr, glmer) stop early, so agreement there is 10⁻⁵ to 10⁻³ rather than to machine precision.
What I'd change
- State the estimand and the checks before fitting, then report the checks whatever they show.
- Replace stepwise selection with pre-specified models or penalised regression, and report selection stability.
- Choose the dose-response mean model on the diagnostics before reporting LD50, and show the LD50 under the competing models side by side, as the model checks now do, rather than one model's interval.
- Show predicted probabilities with intervals, which readers understand better than odds ratios.
- Add calibration and discrimination checks with intervals wherever a model is used to predict.
Decision records
Each record sets out the context, states the decision in one sentence before any argument for it, then gives the options considered, the reasons, what happened (including the weak numbers) and what I would change. Records are not edited after the fact; a changed mind gets a new record that supersedes the old one.
DR-001
How the Assignment 1 model was selected, and why only summaries are published
I selected the model in stages, using analysis-of-deviance tests to prune main effects, an AIC search (step()) over all two-way interactions, and then likelihood-ratio tests to remove interactions that did not earn their place, ending with an 18-parameter model; and the site publishes only model-level summaries of that path, never the survey rows.
DR-002
Ordinal (proportional odds) rather than nominal for pneumoconiosis severity
I used the proportional-odds (cumulative logit) model as the main summary and kept the nominal model as a check.
DR-003
GEE (population-averaged) rather than a GLMM for repeated wheeze measurements
I used a logistic GEE with an exchangeable working correlation and robust (sandwich) standard errors as the main analysis, and in 2026 added a random-intercept GLMM as a sensitivity analysis rather than a replacement.
Model card
This card covers the statistical models behind the site: three generalised linear models from Assignments 2 and 3 that are refitted live, and the Assignment 1 logistic model, which is published as summaries only. It follows the spirit of Mitchell et al. (2019), adapted to small inferential models rather than predictive systems. Every number below is computed by the site's TypeScript ports, which are tested against R (/verification), and the model checks are on /diagnostics.
Model details
| Model | Coursework | Response and covariates | Fitting | Reference implementation |
|---|---|---|---|---|
| Binomial GLM, logit link | A2 Q1 | Beetles killed out of exposed, by log-dosage (straight line; quadratic as a check) | IRLS with Householder QR | stats::glm |
| Proportional-odds cumulative logit | A3 Q1 and Q2 | Pneumoconiosis severity (normal, mild, severe) by years at the coal face | Newton–Raphson, observed information | MASS::polr |
| Logistic GEE, exchangeable working correlation | A3 Q3 | Wheeze at ages 7 to 10 by age and maternal smoking | GEE with sandwich covariance | geepack::geeglm |
| Logistic regression after stepwise selection | A1 | Reported domestic violence or abuse, by demographic and family factors | glm in R; summaries only | stats::glm, step() |
Author: Sunchuangyu (Rin) Huang. Original analysis: University of Melbourne, 2023 Semester 1. Revived: 2026.
Intended use
- Teaching and portfolio demonstration: showing how these models are specified, fitted, checked and interpreted, with the uncertainty stated.
- Reproducing and checking the 2023 coursework results.
Out of scope. None of these models should inform real decisions: not pesticide dosing (the beetle data are from 1935), not occupational health or compensation (the miner data are from 1959 and grouped), not clinical or public-health advice on smoking and childhood wheeze (an observational sample from one town in the 1980s), and not anything about individuals or groups in the Assignment 1 survey.
Training data provenance
| Data | Source | Size | How it is shipped |
|---|---|---|---|
| Beetle mortality | Bliss (1935), Annals of Applied Biology 22 | 481 beetles in 8 dose groups | Typed from the published table |
| Pneumoconiosis | Ashford (1959), British Journal of Industrial Medicine 16 | 371 miners in 8 exposure groups | Typed from the published table |
| Ohio wheeze | geepack::ohio (Fitzmaurice and Laird; Steubenville air-pollution study) | 537 children, 2,148 observations | Aggregated into 32 response-pattern counts by a script; the fit is identical |
| Assignment 1 survey | Course-provided; sensitive | 1,316 respondents, 369 events | Never shipped. Only model-level summaries are exported |
All four data sets are historical or teaching data; none was collected for this site.
Evaluation
Numerical correctness. The ports match R: coefficients and deviances of the binomial GLMs to about 10⁻¹⁰, the proportional-odds and multinomial fits to 10⁻⁶ (R's optimisers stop early), and GEE to 10⁻⁹. The 2026 diagnostics (influence measures, quasi-binomial summaries, profile intervals, Brant's test, AR(1) GEE, QIC and the GLMM) are tested against R, MASS, VGAM, brant, geepack and lme4.
Beetle dose-response (straight line).
- LD50 1.7712. 95% intervals: Wald (delta method) 1.7636 to 1.7788; profile likelihood 1.7634 to 1.7787; parametric bootstrap (B = 2,000, seed 20230501) 1.7638 to 1.7790, with endpoints varying by under 0.001 across five seeds; quasi-binomial 1.7577 to 1.7847.
- These intervals are conditional on the straight-line logit model, which the checks below reject. Under a quadratic logit model the LD50 is 1.7798 (profile likelihood 1.7703 to 1.7884), a shift of 0.0086 that is larger than the Wald half-width of 0.0076 and outside all three binomial intervals above. A complementary log-log straight line gives 1.7782 (profile likelihood 1.7699 to 1.7859). The model, not the interval method, is the main source of uncertainty.
- Fit: residual deviance 13.63 on 6 df (p = 0.034); Pearson X² 12.11 (p = 0.059). Three of the eight standardised residuals fall outside a simulated 95% half-normal envelope, for every one of five seeds; none do for the quadratic model.
- The apparent over-dispersion (Pearson dispersion 2.02) disappears once the curvature is modelled (0.997 for the quadratic), so it signals lack of fit rather than extra-binomial variation.
Pneumoconiosis (proportional odds).
- Ten more years at the coal face multiply the odds of a worse category by 2.61 (95% CI 2.07 to 3.30).
- Parallel-slopes checks: Brant X² = 0.051 on 1 df (p = 0.82); likelihood-ratio test against separate slopes 0.0002 on 1 df (p = 0.99). Both have low power with 38 mild and 44 severe cases.
Wheeze (GEE).
- Maternal smoking: odds ratio 1.30 (95% CI 0.92 to 1.85; robust SE of the log odds ratio 0.178). Age: 0.89 per year (0.82 to 0.97).
- Working-correlation sensitivity: smoking log odds ratio 0.272, 0.265 and 0.234 under independence, exchangeable and AR(1); QIC 1829.49, 1829.48 and 1830.26.
- Cluster bootstrap (1,000 children-level resamples for each of five seeds): 95% percentile interval for the smoking odds ratio from about 0.90 to 1.83 for the primary seed.
Known failure modes
- Separation. On /dose-response you can edit the counts. Counts with no deaths below some dose and all deaths above it have no finite maximum-likelihood estimate; the page detects this and withholds intervals rather than showing nonsense.
- Extrapolation. The dose-response curve is only supported between log-dosages 1.69 and 1.88. The quadratic model in particular bends sharply outside that range.
- Lack of fit. The straight-line beetle model is the one the assignment asked for, but the diagnostics show systematic curvature, and its LD50 intervals leave out the uncertainty about the mean model (see the alternative-model estimates above).
- Sparse categories. The pneumoconiosis tests of the proportional-odds assumption cannot detect modest departures.
- Small numbers of clusters. Sandwich standard errors need many clusters; 537 is plenty, but the GEE code does not apply a small-sample correction and would be optimistic with, say, 20.
- Selection. The Assignment 1 intervals ignore the model-selection step (see DR-001).
Ethical considerations
- The Assignment 1 survey concerns family violence. No individual responses leave the private repository; only aggregate model output is published, and the AI explanation feature is not offered on that page.
- The odds ratios describe associations in historical samples. They are not causal effects and should not be read as statements about any person or group today.
- The occupational data come from a period of very different working conditions and diagnostic standards.
- AI explanations are optional, use the visitor's own key, are labelled as AI-generated and are logged; they never produce the numbers. See the AI use statement.
Caveats and recommendations
Read the estimates with their intervals, check the diagnostics before trusting a model, and see the decision records (DR-001 to DR-003) for why each modelling choice was made and what I would do differently now.
AI use statement
This statement describes how artificial intelligence is and is not used on the GLM Playground site. It is informed by the Australian Government's policy for the responsible use of AI in government (Digital Transformation Agency), the transparency principles of the EU AI Act and the NIST AI Risk Management Framework. It is a voluntary statement for a personal portfolio site; it does not claim compliance or certification with any of them.
What AI does here
One optional feature, Explain this output, on five case-study pages. When you press it, a large language model writes a short plain-English explanation of the model output on that page. It is a reading aid, nothing more.
What AI never does here
- It never computes, changes or chooses any number on the site. Every estimate, interval and test comes from the TypeScript ports of the original R code, which are tested against R.
- It never sees raw data rows, never sees the Assignment 1 survey in any form (the feature is not offered on that page), and never sees anything about you.
- It does not run unless you add your own API key and press the button. There is no AI on the site by default and no AI budget behind it.
- It does not make decisions. Its output is a draft for you to accept, edit or reject.
Data sent to the provider
Only a numeric summary of the model on that page: the model description, the data source, the sample size, a list of estimates with their intervals and tests, and a few plain statements about the design. Before you press the button you can expand What would be sent to see the numbers, and Exact request text inside it to read the system prompt and message verbatim. Apart from that text, the request carries only the model id, a response schema and your key (in a request header). The request goes from your browser straight to the provider you chose (Anthropic by default, or OpenAI) with your key. This site has no server, so it never receives the key, the request or the response. The provider's own terms and data-retention policy apply to the call.
Your key
The key is kept in your browser's sessionStorage and is gone when the tab closes, unless you tick Remember on this device, which moves it to localStorage. Forget keys removes it from both. It is never written to the audit log, never sent anywhere except the provider, and anything that looks like a key is redacted before a log entry is stored.
Human in the loop and transparency
- Every AI output is labelled AI-generated, with the model, provider, response time and token usage shown next to it.
- A deterministic grounding check compares the numbers in the explanation with the numbers that were sent and flags any it cannot trace, including a number quoted with the wrong sign. It skips small whole numbers (usually counts in prose), the 95% level and conventional significance thresholds such as 0.05 and 0.001.
- You review each explanation: accept, edit or reject. Your decision is recorded.
- Every call, including failed ones, is appended to an audit log in your browser (IndexedDB) with its timestamp, feature, provider, model, the exact input (without the key), the output, latency, token usage and your decision. You can read it, export it as JSON or CSV, or clear it at /ai-log.
- Nothing AI-generated is shown unless it has been logged. If your browser blocks or fills its storage, the answer is withheld and the page says why.
Known limitations
- Language models can misread or overstate statistical results, for example by treating a non-significant result as proof of no effect. The grounding check catches invented, mis-copied or sign-flipped numbers, not wrong reasoning.
- Explanations vary from run to run and between models.
- Costs are charged to your provider account.
How this site was built
The 2026 revival of the website (code, tests and documentation) was developed with an AI coding assistant. Every statistical result on the site is tested against R, including the original 2023 analysis code; the comparisons are listed on /verification.