Verdicts below are scored against the paper's abstract. The two homepage exemplars are stored as fixtures so the card visual stays stable while the live pipeline iterates; live audits run the same pipeline against any DOI not in the fixture set.
A four-model frontier-tier comparison run within weeks of each model's release; the rare evaluation that does not lag the field. The 100% top-1 accuracy is on expert-validated vignettes, not real patient encounters, and there is no same-study clinician arm; the headline number therefore speaks to recognition under idealised conditions, not to clinical deployment readiness. Cross-vendor breadth and a contemporaneous frontier comparator are the strengths a careful reader should trust; the human-comparator absence is the limitation to hold against the result.
ECI is Epoch AI's Capabilities Index, a cross-benchmark capability score. Elo is the model's latest Arena rating. AA is Artificial Analysis's Intelligence Index. Each axis is anchored to the registry's observed range; the dot positions are commensurable across audits.
Six releases took the measured lead since this evaluation.
| Released | Model | ECI | Over tested |
|---|---|---|---|
| 2025-12-18 | GPT-5.2 Pro | 155.30 | +5.22 |
| 2026-02-24 | GPT-5.3 Codex | 156.40 | +6.32 |
| 2026-03-05 | GPT-5.4 Pro | 158.89 | +8.81 |
| 2026-04-23 | GPT-5.5 Pro | 162.03 | +11.95 |
| 2026-06-09 | Claude Fable 5 | 162.90 | +12.82 |
| 2026-09-03 | GPT-6 Astra | 169.23 | +19.15 |
Records on the ECI frontier succession — every model whose measured Epoch score exceeded every model before it — against Claude Opus 4.5 at 150.08.
On FrontierMath Tier-4, the tested Claude Opus 4.5 scored 4.9%. The best current score is 97.6%.
Also VPCT 10.0% → 86.5% and Mystery Game Puzzles 14.1% → 82.4%. Published maintainer scores, best configuration.
Best-configuration scores published by each benchmark's maintainers, for Claude Opus 4.5 and for the highest-scoring model available on each date. The three shown are the largest frontier movements between the evaluation anchor and today, among the 29 benchmarks where Claude Opus 4.5 and a current model are both scored and the frontier had at least 5 points of headroom at evaluation. These are not re-runs of the paper's protocol, and a benchmark result is evidence about the frontier, not about the paper's conclusions.
The tested score is a best-configuration result and therefore an upper bound on what the paper's own run elicited — the displayed gap is conservative.
The frontier has been above Claude Opus 4.5’s measured level for 11 months.
Abstract-only read: Elicitation and Capability-frame verdicts can upgrade once the full-text discloses thinking effort, prompting, or scope-bounding language; Model-version is unaffected by the binding source.
"four frontier LLMs (ChatGPT 5.1, Claude Opus 4.5, Gemini 3 Pro, and Grok 4.1)"
All four models named at variant level with explicit version numbers. Snapshot IDs absent; variant-pinning satisfies item 1 under v1.1 (cross-checkpoint drift on a single variant is small).
Disclosed: scaffolding, prompting/context, multi-agent setup. Missing load-bearing dimensions: reasoning mode, thinking effort, tool/search. The capability ceiling reported may not reflect what fuller elicitation would yield. Within-family: the strongest variant of the chosen family generation was tested.
"Under idealised vignette conditions, frontier LLMs thus demonstrated high accuracy in recognising Category A bioterrorism syndromes"
Subject ("frontier LLMs") is anaphoric to the four named models tagged earlier as "four frontier LLMs", which under v7.2 BC1 is defensible as pass; but the bare generic-class subject of a capability verb is exactly the phrasing item 5 flags, so the verdict is warn-with-caveat.
Partial disclosures across Core 3.
The complete VERSIO-AI rubric. 12 items are scored automatically by the live extraction; the remaining 1 require a manual reader pass against the paper's methods or supplementary.