← All posts
Engineering1 September 2026· 6 min read

How we stop the AI inventing code review findings

Every AI review tool claims it doesn't hallucinate. Here is the actual mechanism Pullora uses — evidence grounding, diff anchoring and confidence gating — and what it discards on a real pull request.

A code reviewer that invents problems is worse than no reviewer. It costs the team attention, trains them to ignore the tool, and eventually gets switched off. Every product in this category asserts it does not do that. Assertions are cheap, so here is the mechanism instead.

The pipeline

A review runs as: build context from the diff and the files it touches, review it, then validate every finding before anything is published. The validation stage is where most of the work happens.

1. Evidence must exist in the diff

Every finding has to quote the code it refers to, verbatim, and that quote is checked against the actual changed lines. A finding describing code that is not in the diff is discarded — not downgraded, not posted with a caveat. Discarded.

This kills the most common failure mode, where a model describes a plausible bug in code that does not exist, or flags a problem in a file the pull request never touched.

2. It must anchor to a real line

Findings are mapped to specific line numbers in the diff. If the mapping fails, the finding cannot be posted inline — no guessing at the nearest line, because a comment on the wrong line is its own kind of noise.

3. It must clear a confidence threshold

Each finding carries a confidence score. Two thresholds apply: one for inline comments (0.80 by default) and a lower one for the summary (0.65). Below the summary threshold, the finding is dropped entirely. You can raise or lower both per repository.

What that looks like in practice

On a recent 80-file pull request, the model produced 226 raw findings. After validation: 65 were posted inline, 82 appeared in the summary only, and 79 were discarded. The review log shows the reason for each discard — below confidence, evidence not found in diff, duplicate.

On a smaller pull request the ratio can be far more aggressive: 15 raw findings, 1 posted. That variance is the point. The gate is doing work, and you can see it doing it.

Why the log is visible

Every review shows its own pipeline — files analysed, chunks reviewed, tokens used, findings produced, findings discarded and why. A tool that filters aggressively but hides the filtering is asking for trust it has not earned. If we drop 79 of your findings, you should be able to see that we did and why.

See a review before you install anything

Paste any public GitHub pull request URL and read the full review — no app installed, no repository access, nothing posted to the PR.

Review a public PR →