Verdicts below are scored against the paper's abstract. The two homepage exemplars are stored as fixtures so the card visual stays stable while the live pipeline iterates; live audits run the same pipeline against any DOI not in the fixture set.
A parallel-extraction pipeline rather than a head-to-head capability test: GPT-5.2 and Gemini 3 Pro each pull the same 48 fields per study, with their agreement functioning as a confidence signal. Cross-vendor frontier-tier design is the strength; the missing query date is the central disclosure failure, since the reader cannot tell whether the comparison was anchored against December 2025 or April 2026 frontier capabilities. Treat the accuracy numbers as informative for the orthopaedic-extraction task as specified, not as a generic LLM-systematic-review claim.
Eval date undisclosed; gap range computed from a publication-bounded window with the OpenAlex publication date as the latest plausible eval and the strongest tested model's release as the earliest. Mirrors the audit corpus's pre-registered pub-date − 180d imputation midpoint.
Set an evaluation date below and the gap recomputes against that anchor. The URL updates so you can share the override; Item 3 still reports what the paper disclosed.
ECI is Epoch AI's Capabilities Index, a cross-benchmark capability score. Elo is the model's latest Arena rating. AA is Artificial Analysis's Intelligence Index. Each axis is anchored to the registry's observed range; the dot positions are commensurable across audits.
Three releases took the measured lead after the latest date this evaluation could have run.
| Released | Model | ECI | Over tested |
|---|---|---|---|
| 2025-12-18 | GPT-5.2 Proinside the imputed window | 155.05 | +1.58 |
| 2026-02-06 | Claude Opus 4.6inside the imputed window | 155.13 | +1.66 |
| 2026-02-24 | GPT-5.3 Codexinside the imputed window | 155.53 | +2.06 |
| 2026-03-11 | GPT-5.4 Proinside the imputed window | 158.44 | +4.97 |
| 2026-04-23 | GPT-5.5 Pro | 161.42 | +7.95 |
| 2026-06-09 | Claude Fable 5 | 161.44 | +7.97 |
| 2026-07-09 | GPT-5.6 Sol | 161.56 | +8.09 |
Records on the ECI frontier succession — every model whose measured Epoch score exceeded every model before it — against GPT-5.2 at 153.47.
On ProofBench, the tested GPT-5.2 scored 15.0%. The best current score is 78.0%.
Published maintainer scores, best configuration.
Best-configuration scores published by each benchmark's maintainers, for GPT-5.2 and for the highest-scoring model available on each date. The three shown are the largest frontier movements between the evaluation anchor and today, among the 22 benchmarks where GPT-5.2 and a current model are both scored and the frontier had at least 5 points of headroom at evaluation. These are not re-runs of the paper's protocol, and a benchmark result is evidence about the frontier, not about the paper's conclusions.
The tested score is a best-configuration result and therefore an upper bound on what the paper's own run elicited — the displayed gap is conservative.
The frontier has been above GPT-5.2’s measured level for 8 months. Starting from ChatGPT (GPT-3.5 Turbo) in Nov 2022, 8 months of frontier progress reached GPT-4.
The claim on both rails is calendar-only: the same number of months elapsed in each. The ratio is a guard, not the claim — it refuses a mirror whose window covered materially different capability from this paper’s own gap, because two rails drawn the same length invite a comparison whether or not one is asserted.
Endpoint score is Epoch's GPT-3.5 Turbo estimate; the true Nov-2022 frontier was lower, so the mirror understates its own delta.
Abstract-only read: Elicitation and Capability-frame verdicts can upgrade once the full-text discloses thinking effort, prompting, or scope-bounding language; Model-version is unaffected by the binding source.
"Generative Pre-Trained Transformer 5.2 (GPT-5.2) and Google Gemini 3 Pro"
Both models named at variant level with explicit version numbers spelled out. Snapshot IDs absent; variant-pinning satisfies item 1.
Disclosed: scaffolding. Missing load-bearing dimensions: reasoning mode, thinking effort, tool/search, prompting/context. The reported capability is from a substantially under-specified configuration; stronger elicitation could change the result. Within-family: the strongest variant of the chosen family generation was tested.
"A parallel-LLM approach using GPT-5.2 and Gemini 3 Pro achieved strong accuracy with a high degree of efficiency for automated data extraction in an orthopaedic systematic review."
Conclusion subject is the named GPT-5.2 + Gemini 3 Pro approach; capability bounded to 'automated data extraction in an orthopaedic systematic review'. No generic-LLM extension.
A single editorial desk-reject signal fires; the rest pass.
The complete VERSIO-AI rubric. 12 items are scored automatically by the live extraction; the remaining 1 require a manual reader pass against the paper's methods or supplementary.