How we measure accuracy
An auditor you can’t trust is worse than no auditor: every false alarm costs a person’s attention and the product’s credibility. Evergreen AI is built precision-first — a short, trustworthy list beats a long, noisy one — and that stance is enforced by machinery, not intentions. This page describes that machinery.
The golden evaluation set
Accuracy is measured against a hand-built, version-controlled evaluation corpus of 50 fixture pages, each with expected findings labeled in advance. The set is deliberately adversarial:
- It covers all six finding types — outdated facts, content expired by its own terms, deprecated references, cross-page contradictions, internal incoherence, and stale ownership.
- It includes a trap set: clean, correct pages where the right answer is no findings at all. A system that flags healthy content fails the gate even if it also finds real problems.
- Fixtures are integrity-checked on every run (content hashes), so results can never silently drift from a changed test set.
The release gates
The evaluation runs in CI, and analysis-prompt changes cannot ship unless every blocking gate passes. The thresholds are fixed in code, not judgment calls made at release time:
| Gate | Requirement |
|---|---|
| Overall precision | ≥ 80% of surfaced findings must be labeled-correct |
| Overall recall | ≥ 70% of labeled problems must be found |
| Per-type precision | ≥ 70% for every finding type individually |
| Trap set | At most 1 finding across all clean pages |
| Schema compliance | ≥ 98% of model responses must validate against the finding contract |
A separate scripted evaluation exercises the cross-page contradiction path end to end (quote anchoring, table-cell evidence, orientation, and the confidence floor), and the contradiction feature carried an extra rule: it computed silently in production and was not shown to users at all until it passed its quality gate on the benchmark.
Current benchmark results
Latest accepted baseline (prompt v1.5.0, evaluated 2026-06-12) on the golden set:
| Measure | Result |
|---|---|
| Overall precision | 90% |
| Overall recall | 90% |
| Trap-set false positives | 0 |
| Schema compliance | 100% |
| Contradiction detection (surfaced subset) | 100% precision / 100% recall on the benchmark’s contradiction cases |
Read these numbers as engineering gates, not market claims. The golden set is small by design (50 pages, ~20 labeled findings) so every case can be argued over by a human; small samples carry wide error bars. Your corpus is different from our benchmark, and real-world rates will differ. What the gates guarantee is direction and discipline: a change that degrades measured quality cannot ship.
Safeguards that run on every page, in production
The benchmark bounds the model; these mechanisms bound each individual finding:
- Verbatim evidence, or nothing. Every finding must quote the page it flags, and the quote must anchor to the page’s actual text (whole tokens only). A finding whose evidence can’t be located is rejected before you ever see it.
- A second opinion on borderline calls. Findings in the borderline-confidence band get an independent verification pass; the verifier can suppress them or lower their confidence — it can never raise it.
- Confidence floors and caps. Findings below the installation’s sensitivity threshold are not kept, at most five findings are surfaced per page, and ownership findings are capped as verification requests — the app asks you to check a contact; it never asserts a person is gone.
- Contradictions are held to the highest bar. Only pairs with the strongest relationship signals (pages that link to each other or share near-identical titles) are surfaced; every contradiction requires anchored quotes from both pages and passes a dedicated adjudication step with its own confidence floor.
- Your dismissals teach it. Dismissing a finding with “reduce future flags like this” feeds a per-page suppression note into future analysis, and a dismissed finding never resurfaces as new — finding identity is stable across scans.
- “All clear” is a real answer. The system is built so that finding nothing on a healthy space is a success state, not a failure to produce output.
What we don’t claim
- No live-user accuracy numbers — yet. The benchmark above is an internal engineering measure. The number that matters most is the rate at which real reviewers confirm findings on their own content; we hold ourselves to a measured confirmation threshold during beta and will publish those rates here once there is enough real-world data to report honestly.
- No infallibility. Some findings will be wrong. The product is designed around that: every finding shows its evidence and reasoning so you can judge it in seconds, dismissing is a first-class action, and dismissals make future scans quieter.
- No hidden judgment. Space health scores show their formula in the UI, every scan is logged in an auditable ledger, and prompt versions are recorded on every finding.