Catch RAG hallucinations in ~200ms for ~$23 per million checks.
Your retrieval pipeline returns a passage. Your model writes an answer. GroundCheck tells you whether the answer is actually supported by the passage, fast enough and cheap enough to run on every response instead of a 1% sample.
from groundcheck_jev import GroundCheck
gc = GroundCheck()
result = gc.check(
source="The Apollo 11 mission landed on the Moon on July 20, 1969.",
answer="Buzz Aldrin was the first human on Mars, in 1972.",
question="Who first walked on the Moon?",
)
result.verdict # Verdict.BLOCK
result.p_grounded # 0.01
result.latency_ms # 242LLM-as-judge works, but it is slow and expensive enough that teams run it offline, on a sample, after the bad answer already shipped. At $0.0000228 per check you can move that judgment into the request path.
Measured on 1,600 labeled examples with human-written hard negatives:
| Task | AUC | Median latency | Cost / 1M |
|---|---|---|---|
| Question answering | 0.952 | 221 ms | $22.82 |
| Dialogue | 0.912 | 169 ms | $23.26 |
| Summarization | 0.875 benchmark, 0.74-0.89 on fresh data | 172 ms | $49.87 |
The summarization row is deliberately shown as a range: re-measuring on four fresh slices put it between 0.74 and 0.89, so the single benchmark number is the optimistic end. The other two tasks were measured once and should be read with the same caution.
Run against a real private knowledge base (~8,900 pages of meeting notes, emails, and research logs) with no threshold retuning:
| negatives | caught | falsely blocked |
|---|---|---|
| mechanical corruptions | 94.3% | 0.0% |
| LLM-written adversarial rewrites | 100% | 2.6% |
The adversarial set is the meaningful one: a separate model rewrote true claims
into fluent falsehoods (13,036 markets → 14,192, are not stored in Nova →
are synced to Nova). Replicated with a second generator, and hand-written
negatives were added for the claims generators refused to corrupt (5/5 caught).
Read it as "no misses observed across ~70 adversarial pairs" rather than a literal 100% rate: n is small and generator refusals filter the sample. Both caveats are quantified in the benchmark.
Against Claude Haiku on the same examples, GroundCheck is more accurate (+0.050 AUC, 95% CI [+0.018, +0.085]), about 20x faster, and about 25x cheaper. Full methodology, baselines, and limitations: bench/BENCHMARK.md.
Not on PyPI yet. Install from source:
git clone https://github.com/sharziki/groundcheck && cd groundcheck
pip install -e '.[server]' # omit [server] for just the library
export TYPESAFE_API_KEY='...' # https://typesafe.aifrom groundcheck_jev import GroundCheckOn the name: an unrelated package already occupies
groundcheckon PyPI and thegroundcheckimport name, so the two could never coexist in one environment. This project isgroundcheck-jev, importing asgroundcheck_jev. Thegroundcheckconsole command is unaffected.
from groundcheck_jev import GroundCheck, Policy, Verdict
gc = GroundCheck(policy=Policy(sensitivity="strict", task="qa"))
result = gc.check(source=retrieved_context, answer=llm_answer, question=user_question)
if result.verdict is Verdict.BLOCK:
return "I could not verify that from my sources."
if result.should_escalate:
queue_for_human_review(answer, result.p_grounded) # the uncertain slicer = gc.check(source=ctx, answer=ans, explain=True)
r.failure_mode # 'contradicted' | 'unsupported' | 'overstated' | NoneThe modes map to different fixes: contradicted means retrieval returned the
wrong passage; unsupported means the model padded. It is a second round trip
and fires only when the verdict is not PASS.
sensitivity is expressed in the language of your risk tolerance, not in magic
numbers. Thresholds are derived from a measured ROC curve, not hand-tuned:
| Sensitivity | Blocks at | Use when |
|---|---|---|
permissive |
~70% of hallucinations | false alarms are costly |
balanced |
~85% | default |
strict |
~95% | shipping a fabrication is costly |
Exit codes are 0 pass, 1 review, 2 block, so it gates a build with no parsing:
groundcheck batch answers.jsonl --task summarization --sensitivity strict
echo $? # 2 if anything was fabricateduvicorn groundcheck_jev.server:app --port 8099curl -X POST localhost:8099/v1/check -H 'content-type: application/json' -d '{
"source": "Refunds are available within 30 days of purchase.",
"answer": "You can get a refund any time, no limit.",
"sensitivity": "strict"
}'{"verdict": "block", "p_grounded": 0.02, "confidence": 0.95, "latency_ms": 242}POST /v1/check/batch takes up to 500 items and fans out concurrently.
Measured throughput on a warm process: 49/s at 12 workers (18/s at 4, 38/s
at 8; 16 workers is slower than 12). A cold process is ~20% slower on its
first batch while TLS and connections are established, so keep the service
warm rather than spawning it per request.
Jev reports how sure it is, and that number is honest:
| Confidence | Share of traffic | AUC in band |
|---|---|---|
| >= 0.8 | 51% | 0.982 |
| < 0.5 | 23% | 0.847 |
result.should_escalate surfaces the uncertain slice. Use it to prioritize a
human review queue, not to route to a bigger model. I tested that cascade
against both Claude Haiku and Claude Sonnet and it made accuracy worse both
times, because both are weaker judges on this task. Even an oracle router that
escalated exactly the rows Jev gets wrong scored below Jev alone. Details in
bench/BENCHMARK.md.
Stated up front rather than discovered in production:
- Summarization is the weak task. On fresh data its false-block rate swings
between 19% and 51% depending on the corpus, at the same setting. Use
sensitivity="permissive"there, treat BLOCK as "route to a human", and recalibrate on your own documents. (That task decides on Jev'sscoreprimitive rather than thenoul, which measurably helped but did not fix it.) - Thresholds are calibrated on HaluEval. They held up across three task shapes and
on held-out data, but recalibrate on your own examples with
bench/derive_thresholds.pybefore trusting the exact numbers. - English only so far.
- It checks grounding against the source you give it. It cannot tell you that your retrieval returned the wrong passage.
python -m venv .venv && .venv/bin/pip install -e '.[dev,server]'
.venv/bin/python -m pytest # 21 tests; live tests need TYPESAFE_API_KEY
.venv/bin/python bench/validate_e2e.py --shape qa # verify shipped thresholdsMIT