Skip to content
GLM Playground home

Methods

Methods, decisions and AI use

Where the data come from, how each model is fitted and checked, what it assumes, where it falls short, the decisions behind it, and exactly what the optional AI feature does.

Data provenance

  • Beetle mortality · A2 Q1

    Source
    Bliss (1935), Annals of Applied Biology 22, 134-167
    Size
    481 beetles, 8 dose groups
    How it reaches the site
    Typed from the published table
  • Attitudes to women's role · A2 Q2

    Source
    National Opinion Research Center, University of Chicago, General Social Surveys 1974-1975
    Size
    2,678 respondents, 24 cells
    How it reaches the site
    Typed from the published table
  • Death penalty · A2 Q3

    Source
    Radelet (1981), American Sociological Review 46, 918-927
    Size
    326 defendants, 2 × 2 × 2 table
    How it reaches the site
    Typed from the published table
  • Pneumoconiosis · A3 Q1–Q2

    Source
    Ashford (1959), British Journal of Industrial Medicine 16, 268-278
    Size
    371 miners, 8 exposure groups
    How it reaches the site
    Typed from the published table
  • Ohio wheeze · A3 Q3

    Source
    geepack::ohio (Fitzmaurice and Laird)
    Size
    537 children, 2,148 observations
    How it reaches the site
    Aggregated by script into 32 response-pattern counts; the GEE fit is unchanged
  • Assignment 1 survey · A1

    Source
    Course-provided; sensitive (family violence)
    Size
    1,316 respondents, 369 events
    How it reaches the site
    Never shipped: model-level summaries only

The five public data sets are classic published teaching tables; none was collected for this site, and none contains personal information. The Assignment 1 rows stay in the private repository (see DR-001).

Methods

The 2023 models, ported

  • Binomial and Poisson GLMs by iteratively reweighted least squares with Householder QR, line by line as R's glm.fit.
  • Multinomial (baseline-category) and proportional-odds models by Newton–Raphson with analytic scores and observed information.
  • GEE with a sandwich covariance, matching geepack::geeglm.
  • Analysis of deviance, likelihood-ratio tests, AIC, Wald and delta-method intervals.

Added in 2026

  • Leverage, Cook's distance and standardised residuals; half-normal plots with simulated envelopes.
  • Pearson dispersion, quasi-binomial standard errors and F tests.
  • Profile-likelihood intervals for coefficients and for LD50, solved by root finding.
  • LD50 under two alternative mean models (quadratic logit and complementary log-log), to show how much the estimate depends on the model.
  • Brant's test and a likelihood-ratio test against separate slopes for proportional odds.
  • AR(1) working correlation, QIC and CIC, and a cluster bootstrap for GEE.
  • A random-intercept logistic GLMM by Gauss–Hermite quadrature.
  • Wilson score intervals for every observed proportion; seeded parametric and cluster bootstraps.

Evaluation design

Numerical parity with R

Two R scripts re-run the original model calls and the 2026 checks with established packages and write reference values at full precision. The vitest suite holds every port to them, typically to 10⁻⁸ or better; 49 of 49 headline comparisons are shown on /verification.

Uncertainty on every result

95% intervals throughout: Wald or delta-method for coefficients and contrasts, profile likelihood where the likelihood may be skewed, Wilson score intervals for observed proportions, and percentile intervals for bootstraps. Sample sizes are stated with each data set.

Reproducible simulation

Every simulation uses a fixed, displayed seed (for example 20230501 for the beetle bootstrap and 20230601 for the wheeze bootstrap) and is repeated over 5 seeds to show how much Monte Carlo noise is left. Failed refits are counted and reported, never imputed.

Where two methods are compared on the same data (Wald against profile intervals, three dose-response mean models, naive GLM against GEE, three working correlations, GEE against GLMM), they are fitted to identical inputs and shown side by side, with effect sizes and intervals rather than p-values alone. The comparisons are on /diagnostics.

Assumptions

  • Binomial GLMs. Independent binomial counts with a logit-linear mean. Checked with deviance and Pearson statistics, residual envelopes and influence measures.
  • Log-linear models. Independent Poisson cell counts (equivalently, a multinomial table with fixed total); the models are compared by deviance.
  • Proportional odds. One slope shared by every cumulative split. Checked with Brant's test and a nested likelihood-ratio test.
  • GEE. A correctly specified marginal mean; the working correlation can be wrong, but the robust standard errors need many independent clusters (537 here).
  • GLMM. A normally distributed random intercept per child, and conditional independence of the four visits given it.
  • All models. Observations (or clusters) are independent of each other, and the published tables are accurate. Every model describes associations, not causes.

Limitations

The data are small and historical, chosen to teach methods rather than to answer current questions. The beetle straight-line model, which the assignment specified, does not fit well, and its intervals are conditional on it: under a quadratic or complementary log-log model the LD50 moves by up to 0.0086, more than the straight-line Wald half-width of 0.0076. The proportional-odds tests have little power with 82 abnormal cases. The Assignment 1 intervals ignore the model-selection step. Bootstrap results depend on Monte Carlo error, which the seed tables quantify but do not remove. The ports reproduce R closely, but optimiser-based fits in R (multinom, polr, glmer) stop early, so agreement there is 10⁻⁵ to 10⁻³ rather than to machine precision.

What I'd change

  1. State the estimand and the checks before fitting, then report the checks whatever they show.
  2. Replace stepwise selection with pre-specified models or penalised regression, and report selection stability.
  3. Choose the dose-response mean model on the diagnostics before reporting LD50, and show the LD50 under the competing models side by side, as the model checks now do, rather than one model's interval.
  4. Show predicted probabilities with intervals, which readers understand better than odds ratios.
  5. Add calibration and discrimination checks with intervals wherever a model is used to predict.

Decision records

Each record sets out the context, states the decision in one sentence before any argument for it, then gives the options considered, the reasons, what happened (including the weak numbers) and what I would change. Records are not edited after the fact; a changed mind gets a new record that supersedes the old one.

Model card

This card covers the statistical models behind the site: three generalised linear models from Assignments 2 and 3 that are refitted live, and the Assignment 1 logistic model, which is published as summaries only. It follows the spirit of Mitchell et al. (2019), adapted to small inferential models rather than predictive systems. Every number below is computed by the site's TypeScript ports, which are tested against R (/verification), and the model checks are on /diagnostics.

Model details

ModelCourseworkResponse and covariatesFittingReference implementation
Binomial GLM, logit linkA2 Q1Beetles killed out of exposed, by log-dosage (straight line; quadratic as a check)IRLS with Householder QRstats::glm
Proportional-odds cumulative logitA3 Q1 and Q2Pneumoconiosis severity (normal, mild, severe) by years at the coal faceNewton–Raphson, observed informationMASS::polr
Logistic GEE, exchangeable working correlationA3 Q3Wheeze at ages 7 to 10 by age and maternal smokingGEE with sandwich covariancegeepack::geeglm
Logistic regression after stepwise selectionA1Reported domestic violence or abuse, by demographic and family factorsglm in R; summaries onlystats::glm, step()

Author: Sunchuangyu (Rin) Huang. Original analysis: University of Melbourne, 2023 Semester 1. Revived: 2026.

Intended use

  • Teaching and portfolio demonstration: showing how these models are specified, fitted, checked and interpreted, with the uncertainty stated.
  • Reproducing and checking the 2023 coursework results.

Out of scope. None of these models should inform real decisions: not pesticide dosing (the beetle data are from 1935), not occupational health or compensation (the miner data are from 1959 and grouped), not clinical or public-health advice on smoking and childhood wheeze (an observational sample from one town in the 1980s), and not anything about individuals or groups in the Assignment 1 survey.

Training data provenance

DataSourceSizeHow it is shipped
Beetle mortalityBliss (1935), Annals of Applied Biology 22481 beetles in 8 dose groupsTyped from the published table
PneumoconiosisAshford (1959), British Journal of Industrial Medicine 16371 miners in 8 exposure groupsTyped from the published table
Ohio wheezegeepack::ohio (Fitzmaurice and Laird; Steubenville air-pollution study)537 children, 2,148 observationsAggregated into 32 response-pattern counts by a script; the fit is identical
Assignment 1 surveyCourse-provided; sensitive1,316 respondents, 369 eventsNever shipped. Only model-level summaries are exported

All four data sets are historical or teaching data; none was collected for this site.

Evaluation

Numerical correctness. The ports match R: coefficients and deviances of the binomial GLMs to about 10⁻¹⁰, the proportional-odds and multinomial fits to 10⁻⁶ (R's optimisers stop early), and GEE to 10⁻⁹. The 2026 diagnostics (influence measures, quasi-binomial summaries, profile intervals, Brant's test, AR(1) GEE, QIC and the GLMM) are tested against R, MASS, VGAM, brant, geepack and lme4.

Beetle dose-response (straight line).

  • LD50 1.7712. 95% intervals: Wald (delta method) 1.7636 to 1.7788; profile likelihood 1.7634 to 1.7787; parametric bootstrap (B = 2,000, seed 20230501) 1.7638 to 1.7790, with endpoints varying by under 0.001 across five seeds; quasi-binomial 1.7577 to 1.7847.
  • These intervals are conditional on the straight-line logit model, which the checks below reject. Under a quadratic logit model the LD50 is 1.7798 (profile likelihood 1.7703 to 1.7884), a shift of 0.0086 that is larger than the Wald half-width of 0.0076 and outside all three binomial intervals above. A complementary log-log straight line gives 1.7782 (profile likelihood 1.7699 to 1.7859). The model, not the interval method, is the main source of uncertainty.
  • Fit: residual deviance 13.63 on 6 df (p = 0.034); Pearson X² 12.11 (p = 0.059). Three of the eight standardised residuals fall outside a simulated 95% half-normal envelope, for every one of five seeds; none do for the quadratic model.
  • The apparent over-dispersion (Pearson dispersion 2.02) disappears once the curvature is modelled (0.997 for the quadratic), so it signals lack of fit rather than extra-binomial variation.

Pneumoconiosis (proportional odds).

  • Ten more years at the coal face multiply the odds of a worse category by 2.61 (95% CI 2.07 to 3.30).
  • Parallel-slopes checks: Brant X² = 0.051 on 1 df (p = 0.82); likelihood-ratio test against separate slopes 0.0002 on 1 df (p = 0.99). Both have low power with 38 mild and 44 severe cases.

Wheeze (GEE).

  • Maternal smoking: odds ratio 1.30 (95% CI 0.92 to 1.85; robust SE of the log odds ratio 0.178). Age: 0.89 per year (0.82 to 0.97).
  • Working-correlation sensitivity: smoking log odds ratio 0.272, 0.265 and 0.234 under independence, exchangeable and AR(1); QIC 1829.49, 1829.48 and 1830.26.
  • Cluster bootstrap (1,000 children-level resamples for each of five seeds): 95% percentile interval for the smoking odds ratio from about 0.90 to 1.83 for the primary seed.

Known failure modes

  • Separation. On /dose-response you can edit the counts. Counts with no deaths below some dose and all deaths above it have no finite maximum-likelihood estimate; the page detects this and withholds intervals rather than showing nonsense.
  • Extrapolation. The dose-response curve is only supported between log-dosages 1.69 and 1.88. The quadratic model in particular bends sharply outside that range.
  • Lack of fit. The straight-line beetle model is the one the assignment asked for, but the diagnostics show systematic curvature, and its LD50 intervals leave out the uncertainty about the mean model (see the alternative-model estimates above).
  • Sparse categories. The pneumoconiosis tests of the proportional-odds assumption cannot detect modest departures.
  • Small numbers of clusters. Sandwich standard errors need many clusters; 537 is plenty, but the GEE code does not apply a small-sample correction and would be optimistic with, say, 20.
  • Selection. The Assignment 1 intervals ignore the model-selection step (see DR-001).

Ethical considerations

  • The Assignment 1 survey concerns family violence. No individual responses leave the private repository; only aggregate model output is published, and the AI explanation feature is not offered on that page.
  • The odds ratios describe associations in historical samples. They are not causal effects and should not be read as statements about any person or group today.
  • The occupational data come from a period of very different working conditions and diagnostic standards.
  • AI explanations are optional, use the visitor's own key, are labelled as AI-generated and are logged; they never produce the numbers. See the AI use statement.

Caveats and recommendations

Read the estimates with their intervals, check the diagnostics before trusting a model, and see the decision records (DR-001 to DR-003) for why each modelling choice was made and what I would do differently now.

AI use statement

This statement describes how artificial intelligence is and is not used on the GLM Playground site. It is informed by the Australian Government's policy for the responsible use of AI in government (Digital Transformation Agency), the transparency principles of the EU AI Act and the NIST AI Risk Management Framework. It is a voluntary statement for a personal portfolio site; it does not claim compliance or certification with any of them.

What AI does here

One optional feature, Explain this output, on five case-study pages. When you press it, a large language model writes a short plain-English explanation of the model output on that page. It is a reading aid, nothing more.

What AI never does here

  • It never computes, changes or chooses any number on the site. Every estimate, interval and test comes from the TypeScript ports of the original R code, which are tested against R.
  • It never sees raw data rows, never sees the Assignment 1 survey in any form (the feature is not offered on that page), and never sees anything about you.
  • It does not run unless you add your own API key and press the button. There is no AI on the site by default and no AI budget behind it.
  • It does not make decisions. Its output is a draft for you to accept, edit or reject.

Data sent to the provider

Only a numeric summary of the model on that page: the model description, the data source, the sample size, a list of estimates with their intervals and tests, and a few plain statements about the design. Before you press the button you can expand What would be sent to see the numbers, and Exact request text inside it to read the system prompt and message verbatim. Apart from that text, the request carries only the model id, a response schema and your key (in a request header). The request goes from your browser straight to the provider you chose (Anthropic by default, or OpenAI) with your key. This site has no server, so it never receives the key, the request or the response. The provider's own terms and data-retention policy apply to the call.

Your key

The key is kept in your browser's sessionStorage and is gone when the tab closes, unless you tick Remember on this device, which moves it to localStorage. Forget keys removes it from both. It is never written to the audit log, never sent anywhere except the provider, and anything that looks like a key is redacted before a log entry is stored.

Human in the loop and transparency

  • Every AI output is labelled AI-generated, with the model, provider, response time and token usage shown next to it.
  • A deterministic grounding check compares the numbers in the explanation with the numbers that were sent and flags any it cannot trace, including a number quoted with the wrong sign. It skips small whole numbers (usually counts in prose), the 95% level and conventional significance thresholds such as 0.05 and 0.001.
  • You review each explanation: accept, edit or reject. Your decision is recorded.
  • Every call, including failed ones, is appended to an audit log in your browser (IndexedDB) with its timestamp, feature, provider, model, the exact input (without the key), the output, latency, token usage and your decision. You can read it, export it as JSON or CSV, or clear it at /ai-log.
  • Nothing AI-generated is shown unless it has been logged. If your browser blocks or fills its storage, the answer is withheld and the page says why.

Known limitations

  • Language models can misread or overstate statistical results, for example by treating a non-significant result as proof of no effect. The grounding check catches invented, mis-copied or sign-flipped numbers, not wrong reasoning.
  • Explanations vary from run to run and between models.
  • Costs are charged to your provider account.

How this site was built

The 2026 revival of the website (code, tests and documentation) was developed with an AI coding assistant. Every statistical result on the site is tested against R, including the original 2023 analysis code; the comparisons are listed on /verification.

View the AI audit log