Performance

Honest, measured accuracy — including what we miss.

Most hallucination detectors publish only in-distribution numbers — performance on data drawn from the same distribution they trained on. Those numbers always look better than what you’ll actually see in production.

We publish out-of-distribution (OOD) AUC — performance on examples drawn from sources the model wasn’t trained on. (“Held-out” means examples excluded from training — they may still come from a source family that was in the training mix. “Out-of-distribution” is stronger: the sources themselves were not in training. The 470-row set used here is the latter — closer to what real user traffic looks like.) This is the apples-to-apples measurement against production traffic.

Binary detector OOD AUC: 0.72–0.74 on a 470-row fresh-input set spanning 18 hallucination types plus normal text. Full per-source and per-regime breakdown below — see exactly what we catch and what we don’t.

Aggregate AUC across three evaluation surfaces

We evaluated each shipped detector on three surfaces: an in-distribution test split, a held-out ANAH cohort, and a fresh-input OOD set never seen during training.

P+R= the detector sees both the user's prompt and the LLM response. R-only = the detector sees only the response (no prompt). Both modes ship; the API auto-selects R-only when no prompt is provided. R-only is the typical case for web-UI pastes (no prompt box); P+R is typical for API integrations that already have the prompt context.

ModeDetector cellIn-dist test
(5,706 rows)
Held-out ANAH
(520 rows)
OOD fresh-input
(470 rows)
P+RSCRATCH P3 (lr=1e-3, do=0.4)0.9600.6570.7433
R-onlySCRATCH A2_RO (lr=1e-4, do=0.2)0.9690.6950.7227

The OOD column is the production-relevant number. The in-dist test AUC (~0.97) is dominated by easy, well-represented training sources (see per-source breakdown below) and does not reflect what users will see on novel prompts.

What we catch — by hallucination type (OOD)

Per-class detection quality on the same 470-row out-of-distribution fresh-input set used for the aggregate above. The AUC below is the binary detector’s ability to separate each labeled cohort from clean (NORMAL) text — i.e. for a cohort of N positives and the 203 NORMAL examples, the AUC of the binary P(hallucination) score. Rows are grouped by the model’s internal output classes. The public API returns only NORMAL, FABRICATED, NEAR_FALSE, CF_AUTH, FALSE_REFUSAL, and Other as top_regime; internal-only labels such as SELF_CONTR and UNDERSPECIFIED are collapsed to Other.

Why per-class AUCs can exceed the aggregate 0.72–0.74: the aggregate mixes every hallucination cohort together — including the cohorts the detector cannot reliably separate from clean text, which drag the mean down. A per-class AUC of 0.80 for FABRICATED means the binary detector reliably catches FABRICATED-labeled rows specifically; it does not imply a separate “FABRICATED classifier” with 0.80 accuracy.

Cohorts whose AUC falls at or below random are collapsed into Other in the API response so users never see a label whose accuracy is at chance.

Detect reliably

RegimeWhat it meansnP+R AUCR-only AUC
FABRICATEDInvented facts, made-up entities, fake statistics1210.7960.807
NEAR_FALSETechnically true but arranged to mislead1000.7230.680
CF_AUTHFake citations, misattributed quotes, sources that don’t exist140.8440.729

Weak or small sample

RegimeWhat it meansnP+R AUCR-only AUC
UNDERSPECIFIEDVague or non-specific answers, uncertain claims — internal label, collapsed to Other in the API230.6200.563
FALSE_REFUSALModel declines a reasonable request30.9931.000

UNDERSPECIFIED is above random but well below the reliable tier — and it is an internal-only label collapsed to Other in the API, so it is never returned as a top_regime value. FALSE_REFUSAL has only 3 examples in this eval set — the AUC is essentially noise; don’t over-read it.

Per-source breakdown (in-distribution test)

The 0.97 in-distribution aggregate hides real diversity. Easy sources (HalluEval, OptionC) saturate at ~0.99 and dominate the weighted average. Harder sources — closer to what production traffic looks like — sit in the 0.58–0.85 band. Here it is, source by source.

SourcenP+R AUCR-only AUCWhere this lives
HalluEval2,1910.9950.997Saturated tier — dominates the weighted aggregate
OptionC1910.9880.989Saturated tier — dominates the weighted aggregate
ConflictQA1090.9790.974Saturated tier — dominates the weighted aggregate
ChatbotArenaRefusal3480.9650.964Saturated tier — dominates the weighted aggregate
RAGTruth6810.8190.848Moderate generalization — real signal
PsiloQA1,5640.6940.785Moderate generalization — real signal
BUMP1220.6510.699Moderate generalization — real signal
ANAH5000.5800.644Hard tier — architectural ceiling on this source family

Scope at v1

Built and measured for English natural-language LLM responses. Code generation, structured output (JSON, XML), and non-English text are outside the trained range and may produce unreliable scores. Long documents (more than ~1 page) are not characterized at v1; chunked inference is on the v1.x roadmap.

Evaluation data sources

The per-source AUC table above uses public hallucination-detection benchmarks. Links go to the canonical sources for each:

  • HaluEval (Li et al., 2023) — MIT
  • RAGTruth (Niu et al., 2023) — Apache 2.0
  • ANAH v1 (Ji et al., 2024) — Apache 2.0
  • BUMP (Ma et al., 2023) — MIT
  • PsiloQA (s-nlp; arXiv:2510.04849) — CC BY 4.0
  • ConflictQA (OSU-NLP) — MIT
  • Chatbot Arena Refusal (Human-CentricAI; derivative of LMSYS Chatbot Arena 55K) — Apache 2.0
  • OptionC, bt8_gold — in-house benchmark sets
  • 470-row fresh-input OOD evaluation set — held out from training, curated in-house

Listed for benchmark attribution. Datasets are publicly available under the licenses noted.