A code-review prompt should inspect the proposed diff in its repository context and return a small set of actionable findings. Each finding needs a changed file and line, a reachable input or state, the resulting failure, and the expected behavior. The model may suggest a hypothesis, but the reviewer must check the surrounding call path and actual test result before treating it as a defect. Separate a verified bug from a possible risk and omit style preferences that do not affect the change. Review against the exact commit under discussion so a comment does not point at a line that has since moved.
Diff-scoped code review: make every finding reproducible
Operational case
A payment service changes its retry guard in PaymentCaptureService. A new condition checks whether a receipt is absent, then schedules another capture, but the receipt can arrive between the read and the retry. The useful review comment identifies the condition in the changed diff, names a delayed receipt for payment PY-583, and explains that a second capture may be attempted unless the payment gateway uses an idempotency key. The reviewer checks the gateway adapter and a concurrent test before filing the finding. A generic warning that payments are risky adds no information and should be dropped.
Candidate finding: PaymentCaptureService retry branch, changed line 86.
Trigger: PY-583 receipt arrives after the read, before retry dispatch.
Effect to verify: duplicate capture request without a stable key.
Evidence: inspect gateway adapter and concurrent receipt test.
Disposition: verified finding | hypothesis needing test | no finding.Performance and operating cost
For D changed lines and C relevant call sites, reading the diff and traced context is at least O(D+C); the model call adds token cost proportional to supplied context. Sending an entire repository can bury the changed behavior and raise latency. Retrieve only the necessary contracts and tests, then expand if a finding depends on a caller outside the initial scope. Track confirmed findings and false positives by severity. A high comment count is not a quality metric when most comments cannot be reproduced.
Common Mistakes
- Do not file a finding without a reachable failure path.
- Do not cite an unchanged line as though it were part of the proposed diff.
- Do not treat a model's confidence phrase as test evidence.
Connected lessons
- Prompt engineering applications
- Prompt Engineering
- Coding prompts: name the files, behavior, and proof
- Evidence IDs: make generated claims auditable against supplied records
- Project: prove a receipt release passed the checks that matter
- Generated tests: verify the oracle before trusting coverage
- Incident triage prompts: build a timestamped evidence ledger
- Schema migration prompts: check old and new clients together
- Dependency upgrade prompts: trace behavior, not only versions
- Project: review a checkout release from diff to incident
- Prompt engineering for software delivery
