Which minor injuries matter?

Mechanism-specific severity escalation in US construction from three national injury tiers. Companion app to the paper.
Loading data…

Overview

Escalation explorer

The escalation ratio of a category is its share among severe or fatal cases divided by its share among recordables. A ratio of 1 means the category keeps the same weight across tiers; 30 means it is thirty times as prevalent among fatalities as among recordable injuries. Intervals are 95% bootstrap intervals.

Sensitivity

The ranking of mechanisms holds under alternative scopes (all states, a common 2023 to 2025 window, one NAICS group versus the rest) and under assumed undercounting of minor recordable injuries.

High-energy exposure by tier

Each narrative received a high-energy flag (energy above the threshold used in serious injury and fatality precursor work). The share of high-energy cases rises steeply from recordable to severe to fatal, and it rises inside almost every mechanism.

Severity within the recordable tier

Days away from work is the severity signal available inside the recordable tier. It orders mechanisms very differently from the fatal escalation ratio: the mechanisms that cost the most days are not the ones that kill.

SIF potential model

A gradient boosting model (LightGBM, three classes, class balanced, 5-fold cross-validation) predicts the tier of a case from narrative-derived features only: mechanism, energy source, high-energy flag, fall height and direct-control status. The SIF potential index is the balanced probability of severe or fatal divided by the probability of recordable.

Contextual patterns

When and to whom injuries happen: hours into the shift (recordable tier only), month, weekday, occupation and establishment size.

Establishment linkage

Severe injury reports from 2023 to 2025 were matched by employer name and location to ITA establishments. Among establishments with at least three recordables, the number of recordables does not predict whether the establishment also reported a severe injury, but the high-energy share of its recordables does.

Narrative coders

Accuracy of the harmonization on 241 gold-labeled test narratives: legacy OSHA codes, a zero-shot LLM, the fine-tuned teacher LLM, and the distilled bge-small student that runs in this browser.

Case sample with links to the OSHA records

A random sample of coded narratives, up to 30 per mechanism and tier, drawn with a fixed seed from the main analytic scope. Every fatal case links to its OSHA accident investigation page, and severe cases that opened an inspection link to that inspection. The Severe Injury Reports and the ITA case detail have no per-case page, so those cards give the dataset row ID and the download page instead. Each card also shows the teacher label used in the paper, the student prediction from the corpus run, and a button that codes the same text in your browser.

Narrative coder and escalation lookup

Paste an incident narrative. The distilled model codes its mechanism, energy source and high-energy flag in your browser; the app then looks up the escalation ratios and the SIF potential index for that combination. You can override every field.