Performance
Honest, measured accuracy — including what we miss.
Most hallucination detectors publish only in-distribution numbers — performance on data drawn from the same distribution they trained on. Those numbers always look better than what you’ll actually see in production.
We publish out-of-distribution (OOD) AUC — performance on examples drawn from sources the model wasn’t trained on. (“Held-out” means examples excluded from training — they may still come from a source family that was in the training mix. “Out-of-distribution” is stronger: the sources themselves were not in training. The 470-row set used here is the latter — closer to what real user traffic looks like.) This is the apples-to-apples measurement against production traffic.
Binary detector OOD AUC: 0.72–0.74 on a 470-row fresh-input set spanning 18 hallucination types plus normal text. Full per-source and per-regime breakdown below — see exactly what we catch and what we don’t.
Aggregate AUC across three evaluation surfaces
We evaluated each shipped detector on three surfaces: an in-distribution test split, a held-out ANAH cohort, and a fresh-input OOD set never seen during training.
P+R= the detector sees both the user's prompt and the LLM response. R-only = the detector sees only the response (no prompt). Both modes ship; the API auto-selects R-only when no prompt is provided. R-only is the typical case for web-UI pastes (no prompt box); P+R is typical for API integrations that already have the prompt context.
| Mode | Detector cell | In-dist test (5,706 rows) | Held-out ANAH (520 rows) | OOD fresh-input (470 rows) |
|---|---|---|---|---|
| P+R | SCRATCH P3 (lr=1e-3, do=0.4) | 0.960 | 0.657 | 0.7433 |
| R-only | SCRATCH A2_RO (lr=1e-4, do=0.2) | 0.969 | 0.695 | 0.7227 |
The OOD column is the production-relevant number. The in-dist test AUC (~0.97) is dominated by easy, well-represented training sources (see per-source breakdown below) and does not reflect what users will see on novel prompts.
What we catch — by hallucination type (OOD)
Per-class detection quality on the same 470-row out-of-distribution fresh-input set used for the aggregate above. The AUC below is the binary detector’s ability to separate each labeled cohort from clean (NORMAL) text — i.e. for a cohort of N positives and the 203 NORMAL examples, the AUC of the binary P(hallucination) score. Rows are grouped by the model’s internal output classes. The public API returns only NORMAL, FABRICATED, NEAR_FALSE, CF_AUTH, FALSE_REFUSAL, and Other as top_regime; internal-only labels such as SELF_CONTR and UNDERSPECIFIED are collapsed to Other.
Why per-class AUCs can exceed the aggregate 0.72–0.74: the aggregate mixes every hallucination cohort together — including the cohorts the detector cannot reliably separate from clean text, which drag the mean down. A per-class AUC of 0.80 for FABRICATED means the binary detector reliably catches FABRICATED-labeled rows specifically; it does not imply a separate “FABRICATED classifier” with 0.80 accuracy.
Cohorts whose AUC falls at or below random are collapsed into Other in the API response so users never see a label whose accuracy is at chance.
Detect reliably
| Regime | What it means | n | P+R AUC | R-only AUC |
|---|---|---|---|---|
| FABRICATED | Invented facts, made-up entities, fake statistics | 121 | 0.796 | 0.807 |
| NEAR_FALSE | Technically true but arranged to mislead | 100 | 0.723 | 0.680 |
| CF_AUTH | Fake citations, misattributed quotes, sources that don’t exist | 14 | 0.844 | 0.729 |
Weak or small sample
| Regime | What it means | n | P+R AUC | R-only AUC |
|---|---|---|---|---|
| UNDERSPECIFIED | Vague or non-specific answers, uncertain claims — internal label, collapsed to Other in the API | 23 | 0.620 | 0.563 |
| FALSE_REFUSAL | Model declines a reasonable request | 3 | 0.993 | 1.000 |
UNDERSPECIFIED is above random but well below the reliable tier — and it is an internal-only label collapsed to Other in the API, so it is never returned as a top_regime value. FALSE_REFUSAL has only 3 examples in this eval set — the AUC is essentially noise; don’t over-read it.
Per-source breakdown (in-distribution test)
The 0.97 in-distribution aggregate hides real diversity. Easy sources (HalluEval, OptionC) saturate at ~0.99 and dominate the weighted average. Harder sources — closer to what production traffic looks like — sit in the 0.58–0.85 band. Here it is, source by source.
| Source | n | P+R AUC | R-only AUC | Where this lives |
|---|---|---|---|---|
| HalluEval | 2,191 | 0.995 | 0.997 | Saturated tier — dominates the weighted aggregate |
| OptionC | 191 | 0.988 | 0.989 | Saturated tier — dominates the weighted aggregate |
| ConflictQA | 109 | 0.979 | 0.974 | Saturated tier — dominates the weighted aggregate |
| ChatbotArenaRefusal | 348 | 0.965 | 0.964 | Saturated tier — dominates the weighted aggregate |
| RAGTruth | 681 | 0.819 | 0.848 | Moderate generalization — real signal |
| PsiloQA | 1,564 | 0.694 | 0.785 | Moderate generalization — real signal |
| BUMP | 122 | 0.651 | 0.699 | Moderate generalization — real signal |
| ANAH | 500 | 0.580 | 0.644 | Hard tier — architectural ceiling on this source family |
Scope at v1
Built and measured for English natural-language LLM responses. Code generation, structured output (JSON, XML), and non-English text are outside the trained range and may produce unreliable scores. Long documents (more than ~1 page) are not characterized at v1; chunked inference is on the v1.x roadmap.
Evaluation data sources
The per-source AUC table above uses public hallucination-detection benchmarks. Links go to the canonical sources for each:
- HaluEval (Li et al., 2023) — MIT
- RAGTruth (Niu et al., 2023) — Apache 2.0
- ANAH v1 (Ji et al., 2024) — Apache 2.0
- BUMP (Ma et al., 2023) — MIT
- PsiloQA (s-nlp; arXiv:2510.04849) — CC BY 4.0
- ConflictQA (OSU-NLP) — MIT
- Chatbot Arena Refusal (Human-CentricAI; derivative of LMSYS Chatbot Arena 55K) — Apache 2.0
- OptionC, bt8_gold — in-house benchmark sets
- 470-row fresh-input OOD evaluation set — held out from training, curated in-house
Listed for benchmark attribution. Datasets are publicly available under the licenses noted.