Verdicts below are scored against the paper's abstract. The two homepage exemplars are stored as fixtures so the card visual stays stable while the live pipeline iterates; live audits run the same pipeline against any DOI not in the fixture set.
A four-model frontier-tier comparison run within weeks of each model's release; the rare evaluation that does not lag the field. The 100% top-1 accuracy is on expert-validated vignettes, not real patient encounters, and there is no same-study clinician arm; the headline number therefore speaks to recognition under idealised conditions, not to clinical deployment readiness. Cross-vendor breadth and a contemporaneous frontier comparator are the strengths a careful reader should trust; the human-comparator absence is the limitation to hold against the result.
ECI is Epoch AI's Capabilities Index, a cross-benchmark capability score. Elo is the model's latest Arena rating. AA is Artificial Analysis's Intelligence Index. Each axis is anchored to the registry's observed range; the dot positions are commensurable across audits.
Seven releases took the measured lead since this evaluation.
| Released | Model | ECI | Over tested |
|---|---|---|---|
| 2025-12-18 | GPT-5.2 Pro | 155.05 | +5.17 |
| 2026-02-06 | Claude Opus 4.6 | 155.13 | +5.25 |
| 2026-02-24 | GPT-5.3 Codex | 155.53 | +5.65 |
| 2026-03-11 | GPT-5.4 Pro | 158.44 | +8.56 |
| 2026-04-23 | GPT-5.5 Pro | 161.42 | +11.54 |
| 2026-06-09 | Claude Fable 5 | 161.44 | +11.56 |
| 2026-07-09 | GPT-5.6 Sol | 161.56 | +11.68 |
Records on the ECI frontier succession — every model whose measured Epoch score exceeded every model before it — against Claude Opus 4.5 at 149.88.
On FrontierMath Tier-4, the tested Claude Opus 4.5 scored 4.9%. The best current score is 87.8%.
Also VPCT 10.0% → 86.5% and Chess Puzzles 7.4% → 62.1%. Published maintainer scores, best configuration.
Best-configuration scores published by each benchmark's maintainers, for Claude Opus 4.5 and for the highest-scoring model available on each date. The three shown are the largest frontier movements between the evaluation anchor and today, among the 27 benchmarks where Claude Opus 4.5 and a current model are both scored and the frontier had at least 5 points of headroom at evaluation. These are not re-runs of the paper's protocol, and a benchmark result is evidence about the frontier, not about the paper's conclusions.
The tested score is a best-configuration result and therefore an upper bound on what the paper's own run elicited — the displayed gap is conservative.
The frontier has been above Claude Opus 4.5’s measured level for 12 months. Starting from Gemini 1.5 Pro in Feb 2024, 12 months of frontier progress reached o1.
The claim on both rails is calendar-only: the same number of months elapsed in each. The ratio is a guard, not the claim — it refuses a mirror whose window covered materially different capability from this paper’s own gap, because two rails drawn the same length invite a comparison whether or not one is asserted.
Abstract-only read: Elicitation and Capability-frame verdicts can upgrade once the full-text discloses thinking effort, prompting, or scope-bounding language; Model-version is unaffected by the binding source.
"four frontier LLMs (ChatGPT 5.1, Claude Opus 4.5, Gemini 3 Pro, and Grok 4.1)"
All four models named at variant level with explicit version numbers. Snapshot IDs absent; variant-pinning satisfies item 1 under v1.1 (cross-checkpoint drift on a single variant is small).
Disclosed: scaffolding, prompting/context, multi-agent setup. Missing load-bearing dimensions: reasoning mode, thinking effort, tool/search. The capability ceiling reported may not reflect what fuller elicitation would yield. Within-family: the strongest variant of the chosen family generation was tested.
"Under idealised vignette conditions, frontier LLMs thus demonstrated high accuracy in recognising Category A bioterrorism syndromes"
Subject ("frontier LLMs") is anaphoric to the four named models tagged earlier as "four frontier LLMs", which under v7.2 BC1 is defensible as pass; but the bare generic-class subject of a capability verb is exactly the phrasing item 5 flags, so the verdict is warn-with-caveat.
Partial disclosures across Core 3.
The complete VERSIO-AI rubric. 12 items are scored automatically by the live extraction; the remaining 1 require a manual reader pass against the paper's methods or supplementary.