Most AI-capability papers don't say which model they tested, or when.
VERSIO-AI is the audit tool that does: you paste a DOI; the pipeline then pulls in open-access full-text where resolvable (or defaults to the abstract); a single Claude Opus 4.7 extraction then reads it; finally, the card highlights which model variants have been trialled, along with the evaluation date – but also how fully the model(s) were elicited (reasoning, tools, prompting, and whether the strongest variant of the chosen family was used), and whether the paper's conclusion remained scoped to the actual models tried or devolved into class-level claims about AI/LLMs.
Was the tested model named precisely enough that a domain reader can place it on the capability timeline?
Pass
Variant or pinned snapshot named: “GPT-4o”, “claude-3-5-sonnet-20240620”.
Warn
Vendor only: “ChatGPT”, “Claude”, “Gemini” without variant.
Fail
Generic: “an LLM”, “31 large language models”, “the AI”.
Elicitation completeness
Item 9
Was reasoning enabled? Were tools and search allowed? Prompting and context disclosed? Strongest variant of the chosen family generation used?
Pass
≥75 % of the configuration dimensions disclosed; strongest in-family variant used.
Warn
Partial disclosure, or a weaker variant chosen when a stronger one was available.
Fail
Black-box invocation: reader cannot reproduce the elicitation.
Capability frame
Item 5
Did the conclusion sentence keep its subject scoped to the tested model, or did it generalise to AI and LLMs as a class?
Pass
Subject is the named model, an enumeration, or an anaphoric collective.
Warn
Generic-tier subject mitigated by broad cross-vendor breadth.
Fail
Bare “LLMs” / “AI” as the subject of a capability verb.
Frontier-gap at evaluation
Item 12
How far behind the contemporaneous Arena-elicited frontier was the tested model when evaluation actually ran?
0 mo
At the frontier: tested within ~6 months of the contemporaneous top.
1–6 mo
Modest lag; tested model has been visibly surpassed.
> 6 mo
Material lag; the finding cannot speak to current state-of-the-art.
Why this exists
Frontier Lag is the bibliometric audit behind this tool. It tracks three structural disclosure failures in academic AI-capability research. Tested models are months if not years behind the current best elicitable model when evaluations occur – that is lag. Comparator sets do not include the contemporaneous top tier but just those models favoured by the authors – that is comparator inadequacy. And conclusion-sentences extrapolate findings from an individual model to LLMs-as-a-class without sufficient cross-model breadth for such extrapolation – that is frame asymmetry.
With the same Opus 4.7 extraction prompt which scores the audit corpus across its domains (medicine, law, coding, education, scientific reasoning) – with a preprint forthcoming – VERSIO-AI directs this at one DOI. Each verdict carries the exact quote generating it. Each frontier comparison anchors against both the Arena trajectory and a registry of around 200 named models. No card guesswork.