Skip to content
GLM Playground home

Decision record · DR-002

Ordinal (proportional odds) rather than nominal for pneumoconiosis severity

Accepted (decision made in 2023; recorded in 2026, with checks added in 2026)

Context

Ashford's (1959) data classify 371 coal miners by X-ray as normal (289), mild (38) or severe (44) pneumoconiosis, grouped into eight bands of years worked at the coal face. Assignment 3 asked for two fits of the same data: a baseline-category (multinomial) logit model that ignores the ordering, and a proportional-odds model that uses it. The choice is which one to use as the summary of how exposure relates to severity.

Decision

I used the proportional-odds (cumulative logit) model as the main summary and kept the nominal model as a check. One slope applies to every cumulative split, so ten more years at the coal face multiply the odds of being in a worse category by a single factor.

Options considered

  1. Nominal baseline-category logit (nnet::multinom, 4 parameters). No ordering assumed; two slopes (mild vs normal, severe vs normal).
  2. Proportional odds (MASS::polr, 3 parameters; chosen). Uses the ordering; assumes the same slope for every cumulative split.
  3. Non-proportional cumulative logit (4 parameters). Separate slopes for the two splits; nests proportional odds, so it gives a proper test of the assumption.
  4. Other ordinal links (adjacent-category or continuation-ratio logits), or log(years) instead of years as the covariate, as in some textbook treatments.

Why

  • The categories are ordered by severity, and a model that ignores that spends a parameter on structure the data do not need.
  • The proportional-odds model gives one interpretable effect: ten more years multiply the odds of a worse category by 2.61 (95% CI 2.07 to 3.30).
  • Its AIC is lower (422.9 against 425.4 for the nominal model) and the fitted probabilities of the two models are close at every exposure level.
  • I kept years, not log(years), because that is what the assignment specified; the site reproduces the assignment rather than redoing it.

What happened

  • In 2023 the choice rested on AIC and on the two models agreeing. The parallel-slopes assumption itself was not tested. That gap is closed on /diagnostics:
    • Brant's test: X² = 0.051 on 1 df, p = 0.82.
    • Likelihood-ratio test against separate slopes: 0.0002 on 1 df, p = 0.99.
    • Separate binary logits for the two splits give slopes of 0.0963 (SE 0.0124) for "mild or severe" and 0.0935 (SE 0.0154) for "severe", almost identical.
  • These are reassuring but not proof. With 38 mild and 44 severe miners the tests have little power against modest departures, and a non-significant test is not evidence of exact proportionality.
  • Comparing the nominal and proportional-odds models by AIC is fine, but they are not nested, so no likelihood-ratio test between them is valid. The nested comparison is the one against separate slopes above.
  • The 2023 written answer computed one interval by hand and got (2.1589, 4.1223); the site shows the exact value.

What I'd change

  • Test the proportional-odds assumption before choosing, not after, and report the test with its limited power stated.
  • Show predicted category probabilities with intervals across the exposure range, which say more to a reader than a cumulative odds ratio.
  • Check whether log(years) fits better and whether the cumulative logits are linear in exposure.
  • If the splits had diverged, use a partial proportional-odds model rather than falling back to the nominal model.