Sage Nord

← All posts

PureReason: a deterministic verification layer that catches AI hallucinations in under 5ms

Ask most teams how they catch hallucinations in AI agent output today, and the honest answer is: they ask another LLM. A second model reads the first model's answer and grades it. It works, sort of, but it's slow, it costs another API call per check, and it's only as reliable as the grader model itself, which can hallucinate too. That's a verification strategy built on the same failure mode it's trying to catch.

PureReason, an open-source Rust project from Sage Nord, takes a different approach: verify without a second LLM at all. It's a deterministic reasoning-assurance layer that sits alongside a frontier model (GPT, Claude, Gemini) and checks its output for hallucinations, contradictions, and overconfidence in under 5 milliseconds, using symbolic logic and embeddings rather than another round of model inference.

The problem: verification that hallucinates too

The core issue with LLM-as-judge verification isn't that it never works, it's that it inherits every weakness of the thing it's checking. It's non-deterministic (the same input can get a different verdict twice), it's opaque (there's no traceable reason for the score), and it's expensive at scale (every check is a full model call, with the latency and cost that implies). For an AI agent making dozens of tool calls in a session, or a production pipeline verifying every generated claim, that overhead adds up fast, and the safety layer becomes the bottleneck.

PureReason's bet is that a meaningful slice of hallucinations don't need a second opinion from a language model to catch. Arithmetic errors, invalid syllogisms, unsupported certainty ("the patient must have cancer" from ambiguous findings), and contradictions with known facts are all things a deterministic system can check directly, the same way a compiler checks types instead of asking another compiler for its opinion.

How it works

PureReason combines four checks that run in parallel, not another model call:

  • Symbolic logic — a Z3 deterministic solver verifies arithmetic and formal logical structure (syllogisms, chain-of-thought steps) exactly, with no sampling variance.
  • Neural embeddings — all-MiniLM-L6-v2 semantic similarity catches paraphrased contradictions that pure symbolic rules would miss.
  • Domain calibration — per-domain accuracy tuning, so a medical claim and a casual chat message aren't held to the same certainty bar.
  • Knowledge grounding — entity checking and contradiction detection against known facts.
PureReason verification pipeline A frontier model's output passes through four parallel deterministic checks — symbolic logic, neural embeddings, domain calibration, and knowledge grounding — producing a single 0-100 ECS score that routes the output to accept, review, or reject/rewrite. Frontier model GPT · Claude · Gemini PureReason guard Symbolic logic Z3 deterministic solver Neural embeddings semantic similarity Domain calibration per-domain accuracy Knowledge grounding entity + contradiction ECS score 0–100 Accept ECS ≥ 70 Review 40–69 Reject ECS < 40
PureReason verification pipeline: four deterministic checks feed a single 0–100 Epistemic Confidence Score, which routes the output to accept, review, or reject.

Each output gets a single Epistemic Confidence Score (ECS) from 0–100. High scores pass straight through; low scores get flagged, and overconfident claims get rewritten into properly hedged language automatically, for example turning "the patient must have cancer" into "findings consistent with possible malignancy."

The results

PureReason reports F1 scores across nine public hallucination-detection benchmarks, spanning question answering, dialogue, summarization faithfulness, and grounded retrieval:

PureReason F1 scores across nine hallucination-detection benchmarks Horizontal bar chart of F1 scores: HaluEval QA 0.871, LogicBench 0.846, TruthfulQA 0.798, HalluLens 0.729, HalluMix 0.664, RAGTruth 0.646, FELM 0.645, HaluEval Dialogue 0.634, FaithBench 0.622. 0.0 0.25 0.5 0.75 1.0 HaluEval QA 0.871 LogicBench 0.846 TruthfulQA 0.798 HalluLens 0.729 HalluMix 0.664 RAGTruth 0.646 FELM 0.645 HaluEval Dialogue 0.634 FaithBench 0.622
F1 scores across nine public hallucination-detection benchmarks (PureReason v0.3.1). Full methodology in the repo's BENCHMARK.md.

The project's own benchmark notes report a 25–30 percentage point F1 improvement and a 40% latency reduction over its earlier baseline in v0.3.1, with the full methodology (seeds, holdout sets, reproduction steps) published in the repo's docs/BENCHMARK.md and docs/REPRODUCIBILITY.md rather than asserted without evidence, which matters for a tool whose entire pitch is trustworthy verification.

Where it fits

PureReason ships as a CLI, an MCP server (so agents in Claude Code, Cursor, or GitHub Copilot can call it directly as a tool), and a Python API for chain-of-reasoning and syllogism verification. It's built to sit in front of, not replace, a frontier model: the LLM generates, PureReason verifies, and the agent gets both the original output and a scored, explainable verdict before deciding whether to act on it.

That framing also marks its limits. PureReason is explicitly not a reasoning engine, it verifies, it doesn't generate, so it won't help with novel problem solving, open-ended content generation, or reasoning over long contexts. It's best suited to exactly the case it was built for: a fast, offline, explainable safety check between a generative model and whatever acts on that model's output next, whether that's a RAG pipeline, a coding agent, or a production API serving generated text.

Source: sorunokoe/PureReason on GitHub.