sharziki/groundcheck

Catch RAG hallucinations in ~200ms for ~$23 per million checks. Benchmark-backed: AUC 0.952, beats Claude Haiku on identical examples.

★ 0Forks 0PythonGitHub ↗Compare

README

GroundCheck

tests

Catch RAG hallucinations in ~200ms for ~$23 per million checks.

Your retrieval pipeline returns a passage. Your model writes an answer. GroundCheck tells you whether the answer is actually supported by the passage, fast enough and cheap enough to run on every response instead of a 1% sample.

from groundcheck_jev import GroundCheck

gc = GroundCheck()
result = gc.check(
    source="The Apollo 11 mission landed on the Moon on July 20, 1969.",
    answer="Buzz Aldrin was the first human on Mars, in 1972.",
    question="Who first walked on the Moon?",
)

result.verdict      # Verdict.BLOCK
result.p_grounded   # 0.01
result.latency_ms   # 242

Why it exists

LLM-as-judge works, but it is slow and expensive enough that teams run it offline, on a sample, after the bad answer already shipped. At $0.0000228 per check you can move that judgment into the request path.

Measured on 1,600 labeled examples with human-written hard negatives:

Task AUC Median latency Cost / 1M
Question answering 0.952 221 ms $22.82
Dialogue 0.912 169 ms $23.26
Summarization 0.875 benchmark, 0.74-0.89 on fresh data 172 ms $49.87

The summarization row is deliberately shown as a range: re-measuring on four fresh slices put it between 0.74 and 0.89, so the single benchmark number is the optimistic end. The other two tasks were measured once and should be read with the same caution.

Run against a real private knowledge base (~8,900 pages of meeting notes, emails, and research logs) with no threshold retuning:

negatives caught falsely blocked
mechanical corruptions 94.3% 0.0%
LLM-written adversarial rewrites 100% 2.6%

The adversarial set is the meaningful one: a separate model rewrote true claims into fluent falsehoods (13,036 markets → 14,192, are not stored in Nova → are synced to Nova). Replicated with a second generator, and hand-written negatives were added for the claims generators refused to corrupt (5/5 caught).

Read it as "no misses observed across ~70 adversarial pairs" rather than a literal 100% rate: n is small and generator refusals filter the sample. Both caveats are quantified in the benchmark.

Against Claude Haiku on the same examples, GroundCheck is more accurate (+0.050 AUC, 95% CI [+0.018, +0.085]), about 20x faster, and about 25x cheaper. Full methodology, baselines, and limitations: bench/BENCHMARK.md.

Install

Not on PyPI yet. Install from source:

git clone https://github.com/sharziki/groundcheck && cd groundcheck
pip install -e '.[server]'             # omit [server] for just the library
export TYPESAFE_API_KEY='...'          # https://typesafe.ai
from groundcheck_jev import GroundCheck

On the name: an unrelated package already occupies groundcheck on PyPI and the groundcheck import name, so the two could never coexist in one environment. This project is groundcheck-jev, importing as groundcheck_jev. The groundcheck console command is unaffected.

Use it

As a gate in your RAG pipeline

from groundcheck_jev import GroundCheck, Policy, Verdict

gc = GroundCheck(policy=Policy(sensitivity="strict", task="qa"))
result = gc.check(source=retrieved_context, answer=llm_answer, question=user_question)

if result.verdict is Verdict.BLOCK:
    return "I could not verify that from my sources."
if result.should_escalate:
    queue_for_human_review(answer, result.p_grounded)   # the uncertain slice

Why did it fail?

r = gc.check(source=ctx, answer=ans, explain=True)
r.failure_mode   # 'contradicted' | 'unsupported' | 'overstated' | None

The modes map to different fixes: contradicted means retrieval returned the wrong passage; unsupported means the model padded. It is a second round trip and fires only when the verdict is not PASS.

sensitivity is expressed in the language of your risk tolerance, not in magic numbers. Thresholds are derived from a measured ROC curve, not hand-tuned:

Sensitivity Blocks at Use when
permissive ~70% of hallucinations false alarms are costly
balanced ~85% default
strict ~95% shipping a fabrication is costly

As a CI gate

Exit codes are 0 pass, 1 review, 2 block, so it gates a build with no parsing:

groundcheck batch answers.jsonl --task summarization --sensitivity strict
echo $?   # 2 if anything was fabricated

As a service

uvicorn groundcheck_jev.server:app --port 8099
curl -X POST localhost:8099/v1/check -H 'content-type: application/json' -d '{
  "source": "Refunds are available within 30 days of purchase.",
  "answer": "You can get a refund any time, no limit.",
  "sensitivity": "strict"
}'
{"verdict": "block", "p_grounded": 0.02, "confidence": 0.95, "latency_ms": 242}

POST /v1/check/batch takes up to 500 items and fans out concurrently. Measured throughput on a warm process: 49/s at 12 workers (18/s at 4, 38/s at 8; 16 workers is slower than 12). A cold process is ~20% slower on its first batch while TLS and connections are established, so keep the service warm rather than spawning it per request.

Confidence is a real signal (but do not build a cascade on it)

Jev reports how sure it is, and that number is honest:

Confidence Share of traffic AUC in band
>= 0.8 51% 0.982
< 0.5 23% 0.847

result.should_escalate surfaces the uncertain slice. Use it to prioritize a human review queue, not to route to a bigger model. I tested that cascade against both Claude Haiku and Claude Sonnet and it made accuracy worse both times, because both are weaker judges on this task. Even an oracle router that escalated exactly the rows Jev gets wrong scored below Jev alone. Details in bench/BENCHMARK.md.

Known limitations

Stated up front rather than discovered in production:

  • Summarization is the weak task. On fresh data its false-block rate swings between 19% and 51% depending on the corpus, at the same setting. Use sensitivity="permissive" there, treat BLOCK as "route to a human", and recalibrate on your own documents. (That task decides on Jev's score primitive rather than the noul, which measurably helped but did not fix it.)
  • Thresholds are calibrated on HaluEval. They held up across three task shapes and on held-out data, but recalibrate on your own examples with bench/derive_thresholds.py before trusting the exact numbers.
  • English only so far.
  • It checks grounding against the source you give it. It cannot tell you that your retrieval returned the wrong passage.

Development

python -m venv .venv && .venv/bin/pip install -e '.[dev,server]'
.venv/bin/python -m pytest          # 21 tests; live tests need TYPESAFE_API_KEY
.venv/bin/python bench/validate_e2e.py --shape qa   # verify shipped thresholds

License

MIT

Issues