Skip to content
AITroveRead. Build. Understand.
Make this comfortable

Diff-scoped code review: make every finding reproducible

Last updated: 2 Oct 202611 min read
tutorial
AdvancedBy AITrove Editorial

A code-review prompt should inspect the proposed diff in its repository context and return a small set of actionable findings. Each finding needs a changed file and line, a reachable input or state, the resulting failure, and the expected behavior. The model may suggest a hypothesis, but the reviewer must check the surrounding call path and actual test result before treating it as a defect. Separate a verified bug from a possible risk and omit style preferences that do not affect the change. Review against the exact commit under discussion so a comment does not point at a line that has since moved.

Operational case

A payment service changes its retry guard in PaymentCaptureService. A new condition checks whether a receipt is absent, then schedules another capture, but the receipt can arrive between the read and the retry. The useful review comment identifies the condition in the changed diff, names a delayed receipt for payment PY-583, and explains that a second capture may be attempted unless the payment gateway uses an idempotency key. The reviewer checks the gateway adapter and a concurrent test before filing the finding. A generic warning that payments are risky adds no information and should be dropped.

Output
Candidate finding: PaymentCaptureService retry branch, changed line 86.
Trigger: PY-583 receipt arrives after the read, before retry dispatch.
Effect to verify: duplicate capture request without a stable key.
Evidence: inspect gateway adapter and concurrent receipt test.
Disposition: verified finding | hypothesis needing test | no finding.

Performance and operating cost

For D changed lines and C relevant call sites, reading the diff and traced context is at least O(D+C); the model call adds token cost proportional to supplied context. Sending an entire repository can bury the changed behavior and raise latency. Retrieve only the necessary contracts and tests, then expand if a finding depends on a caller outside the initial scope. Track confirmed findings and false positives by severity. A high comment count is not a quality metric when most comments cannot be reproduced.

Common Mistakes

  • Do not file a finding without a reachable failure path.
  • Do not cite an unchanged line as though it were part of the proposed diff.
  • Do not treat a model's confidence phrase as test evidence.

Connected lessons

prompt engineering
software delivery
Storage details