Skip to main content
VERSIO-AI · Frontier Lag

Most AI-capability papers don't say which model they tested, or when.

Paste a DOI. VERSIO-AI pulls the open-access full text where one resolves and the abstract where none does, reads it in a single Claude Opus 5 pass, and returns a card: which model variants were actually tested and when, how fully they were elicited (reasoning, tools, prompting, and whether the strongest variant of the chosen family was the one used), and whether the conclusion stayed scoped to those models or slid into a claim about AI and LLMs as a class.

What the card returns

Model version

Item 1

Was the tested model named precisely enough that a domain reader can place it on the capability timeline?

Pass
Variant or pinned snapshot named: “GPT-4o”, “claude-3-5-sonnet-20240620”.
Warn
Vendor only: “ChatGPT”, “Claude”, “Gemini” without variant.
Fail
Generic: “an LLM”, “31 large language models”, “the AI”.

Elicitation completeness

Item 9

Was reasoning enabled? Were tools and search allowed? Prompting and context disclosed? Strongest variant of the chosen family generation used?

Pass
≥75 % of the configuration dimensions disclosed; strongest in-family variant used.
Warn
Partial disclosure, or a weaker variant chosen when a stronger one was available.
Fail
Black-box invocation: reader cannot reproduce the elicitation.

Capability frame

Item 5

Did the conclusion sentence keep its subject scoped to the tested model, or did it generalise to AI and LLMs as a class?

Pass
Subject is the named model, an enumeration, or an anaphoric collective.
Warn
Generic-tier subject mitigated by broad cross-vendor breadth.
Fail
Bare “LLMs” / “AI” as the subject of a capability verb.

Frontier-gap at evaluation

Item 12

How far behind the contemporaneous Arena-elicited frontier was the tested model when evaluation actually ran?

At frontier
Within ~3 months of frontier progress at the evaluation date.
Behind
Measurably surpassed, but the finding still describes a recent system.
Far behind
The tested model no longer represents what the conclusion claims about.
Why this exists

Frontier Lag is the bibliometric audit behind this tool, and it documents three disclosure failures that compound. Evaluations run on models months or years behind the best elicitable one at the time, which is lag. Comparator sets reach for whichever models the authors already knew rather than the contemporaneous top tier, which is comparator inadequacy. And conclusions generalise from one model to LLMs as a class on breadth that cannot carry the generalisation, which is frame asymmetry.

VERSIO-AI points the same extraction at a single DOI. Every verdict carries the passage that produced it, and every frontier comparison resolves against a registry of 185 named models plus the monthly Arena trajectory, so a number on the card can always be traced back to something measured.

Read the methodology in full →

VERSIO-AI extraction prompt v1.4 · model claude-opus-5 via OpenRouter · registry refreshed 2026-07-26.frontierlag.org