Benchmark · measured 27 August 2026
The fact-checker, measured
Ten documents. Forty-five errors planted on purpose. Two clean documents with nothing wrong at all. Three runs, so the numbers include their own wobble.
Every writing tool claims its AI is accurate. Almost none of them publish a number, and the ones that do publish a single run, which is the number that happened to be flattering. Here is ours, with the spread and the misses.
The numbers
| Mean | Range over 3 runs | |
|---|---|---|
| Recall — how many planted errors it found | 0.948 | 0.867 to 1.000 |
| Precision — how many of its findings were real | 0.956 | 0.938 to 0.978 |
| False positives — things it flagged that were fine | 2 per run | 1 to 3 |
45 seeded errors across 10 documents, 3 independent runs on 27 August 2026. Mean of three, not a best-of.
The part that matters more than recall
Two of the ten documents have nothing wrong with them. They are there to answer the question writers actually ask, which is not "will it find my mistakes" but "will it invent problems in prose that is already fine".
In two of the three runs, both clean documents produced nothing at all. In the third, one of them produced a single finding: it queried whether espresso is an Italian invention whose name means "pressed out". That is one spurious remark across six clean-document readings, and we would rather print it than round it away.
There is no praise channel. The checker has no way to tell you a passage is good, because it is not built to have opinions about your writing. It returns findings or it returns nothing.
What it caught
- A ZIP code in a scene set in summer 1962. ZIP codes launched in July 1963.
- "I Want to Hold Your Hand" playing on a radio in 1962. Released November 1963.
- A Boeing 727 overhead in 1962. First flight February 1963.
- Bonanza described as airing on CBS. It was NBC.
- Apollo 8's Genesis reading placed at Easter 1969. It was Christmas Eve 1968.
- Buzz Aldrin called the third man on the moon. He was the second.
- "Saturn V remains the most powerful rocket ever flown" — flagged as misleading rather than false, citing Starship and SLS.
What it missed, and why
The misses cluster, and the pattern is consistent: when the evidence is thin, the checker declines to accuse rather than guessing. An adversarial pass argues against every finding before you see it, and that pass kills weak ones. Most of what recall costs us is bought back as precision, which is the trade we would pick anyway. A tool that cries wolf on a manuscript gets closed and never reopened.
Known weak spots today: dates that need arithmetic against a scene's own timeline, and facts too niche to have a well-sourced page behind them.
Why the range is on this page
The pipeline is not deterministic. It searches live sources, and live sources change between runs. Three identical runs earlier in development came back byte-identical, which looked like proof of stability and was actually a warm search cache. So the honest unit here is a mean of three runs with its spread shown, and a single run is not allowed to become the number we quote.
Check it yourself
The golden set is ten documents with every seeded error labelled and quoted verbatim, and the scorer matches findings to labels by position in the text rather than by word overlap. The harness refuses to promote a new champion from a single run. If you want to argue with these numbers, that is the point of publishing them: write to and we will send you the set.