How we measure accuracy

The evaluation system every Evergreen AI release must pass — and what we do and don’t claim about it.

An auditor you can’t trust is worse than no auditor: every false alarm costs a person’s attention and the product’s credibility. Evergreen AI is built precision-first — a short, trustworthy list beats a long, noisy one — and that stance is enforced by machinery, not intentions. This page describes that machinery.

The golden evaluation set

Accuracy is measured against a hand-built, version-controlled evaluation corpus of 50 fixture pages, each with expected findings labeled in advance. The set is deliberately adversarial:

The release gates

The evaluation runs in CI, and analysis-prompt changes cannot ship unless every blocking gate passes. The thresholds are fixed in code, not judgment calls made at release time:

GateRequirement
Overall precision≥ 80% of surfaced findings must be labeled-correct
Overall recall≥ 70% of labeled problems must be found
Per-type precision≥ 70% for every finding type individually
Trap setAt most 1 finding across all clean pages
Schema compliance≥ 98% of model responses must validate against the finding contract

A separate scripted evaluation exercises the cross-page contradiction path end to end (quote anchoring, table-cell evidence, orientation, and the confidence floor), and the contradiction feature carried an extra rule: it computed silently in production and was not shown to users at all until it passed its quality gate on the benchmark.

Current benchmark results

Latest accepted baseline (prompt v1.5.0, evaluated 2026-06-12) on the golden set:

MeasureResult
Overall precision90%
Overall recall90%
Trap-set false positives0
Schema compliance100%
Contradiction detection (surfaced subset)100% precision / 100% recall on the benchmark’s contradiction cases

Read these numbers as engineering gates, not market claims. The golden set is small by design (50 pages, ~20 labeled findings) so every case can be argued over by a human; small samples carry wide error bars. Your corpus is different from our benchmark, and real-world rates will differ. What the gates guarantee is direction and discipline: a change that degrades measured quality cannot ship.

Safeguards that run on every page, in production

The benchmark bounds the model; these mechanisms bound each individual finding:

What we don’t claim

Questions about the methodology — or want the gate thresholds in more detail? Contact support; we’re happy to walk through it.