AI code review: build an evidence trail before accepting the suggestion
Turn automated review into a useful engineering aid with reproducible findings, repository context, targeted tests, and clear ownership of the final decision.

An AI reviewer can produce a convincing comment in seconds. It can also misunderstand a contract, overlook a relevant caller, or recommend a change that makes the code look cleaner while breaking its behavior. The value of automated review depends less on the confidence of its prose than on the evidence supporting each finding.
A useful review process asks the system to identify concrete failure conditions, connect them to the changed code, and propose a way to verify the concern. Human reviewers remain responsible for deciding whether the finding matters and whether the suggested repair preserves the intended behavior. That division of work makes automation easier to trust because its claims remain inspectable.
Treat a review comment as a hypothesis
GitHub's responsible-use documentation discusses limitations of AI-assisted review, including plausible feedback based on an incorrect understanding of the code. Its broader guidance emphasizes reviewing outputs and maintaining human oversight. The practical consequence is that an automated finding should enter the engineering process as a claim to investigate, not as an established defect.
Read source on docs.github.com
A strong finding names a trigger, the affected behavior, and the relevant evidence. For example, a cancellation request arriving after a job completes may overwrite its final state because the update lacks a state condition. That statement can be examined. This code might have race conditions is much less useful without a concrete interleaving.
Ask the reviewer to distinguish observed evidence from inference. It may have read a function and its tests but not the database migration or a production configuration. Missing context should narrow the conclusion. A concise uncertainty statement is better than an authoritative explanation built on an assumption nobody checked.
Give the model the contract around the diff
A diff shows what changed, but it may not explain what must remain true. Supply the relevant requirement, interfaces, invariants, and repository conventions. For a pagination change, this might include ordering guarantees and how duplicate sort values are handled. For a permissions change, it includes the server-side access rules.
Select context deliberately. Sending the entire repository can overwhelm the task and expose unrelated material without ensuring that the crucial contract is noticed. Begin with the changed code, nearby callers, affected tests, and the data definitions that govern behavior. Expand when a concrete question requires it.
Keep instructions separate from untrusted repository content. Comments, fixtures, and documentation under review are evidence, not authority to change the reviewer's permissions or publish data. An automated review integration should enforce its own boundaries regardless of text encountered in a pull request.
Review one realistic change end to end
Consider a hypothetical patch that adds retry support to an image-processing queue. The intended behavior is to retry temporary failures while ensuring that a successfully completed image is not processed and published twice. The patch introduces a retry counter, a job-state update, and a new worker branch.
Ask the AI reviewer to trace one job through success, temporary failure, permanent failure, cancellation, and worker restart. Require it to point to the code that enforces each transition. The exercise focuses attention on observable behavior rather than stylistic preferences scattered across the diff.
Then challenge its most serious finding. If it claims that duplicate publication is possible, construct the exact sequence of events. Which operation completes before the timeout? Which state is durable? What does the next worker read? A finding that survives this questioning is much more useful than one supported only by generic warnings about retries.
Separate correctness, maintainability, and taste
Review queues become noisy when a minor naming preference receives the same presentation as data loss. Establish a small set of categories and severity rules tied to consequences. A correctness issue should explain a behavior that violates the contract. A maintainability suggestion should explain a future cost without pretending the current code is broken.
Avoid asking the model for a fixed number of findings. That creates pressure to produce comments even when the change is sound. An empty review with a clear scope can be a valid result. The objective is useful evidence, not a busy-looking report or a maximum volume of automated participation.
Deduplicate repeated comments about the same root cause. If five callers share one flawed helper, one well-supported finding may be more actionable than five similar comments. Preserve the affected scope while keeping the discussion focused enough that the author can make a coherent repair.
Use tests to discriminate between explanations
A good regression test demonstrates the failure before the repair and the intended behavior afterward. For the retry example, simulate a completed side effect followed by a lost acknowledgment. Verify what the second attempt does. The test should exercise the contract rather than merely reproduce the implementation's current structure.
Do not accept a proposed test just because it passes. Check that it would fail for the claimed bug and that its assertions concern the user-visible outcome. A test that mocks away the state transition at issue can provide false reassurance while never reaching the dangerous behavior.
Choose the narrowest verification that resolves the uncertainty, then run the broader checks required by the repository. Some concerns need an integration test or a database constraint inspection; others can be settled by reading a caller. Automation should reduce uncertainty efficiently, not create elaborate tests for every stylistic suggestion.
Keep tool execution inside an explicit boundary
An AI reviewer may need to run tests, inspect files, or evaluate a small reproduction. Give it the permissions required for that work and no implicit authority to deploy, modify unrelated data, or contact external services. The review process should make the distinction between observation and mutation visible.
Run untrusted code in an appropriate isolated environment. A pull request can alter test scripts and dependency hooks as well as application code. A command named test is not inherently harmless. Use the repository's established security process for contributions and avoid placing production credentials in review jobs.
Record which checks actually ran. A generated comment saying tests should pass is not a test result. The final review should distinguish completed verification, failed checks, and checks that were unavailable. This prevents confident narrative from replacing the operational evidence needed to merge responsibly.
Measure review usefulness rather than comment volume
Build a small evaluation set of historical changes with known outcomes and carefully prepared defects. Include sound changes so the system is tested on its ability to remain quiet. Label the important findings and the evidence required to support them, while allowing that more than one valid explanation may exist.
Measure actionable findings, false positives, missed important defects, and reviewer effort. A tool that catches one additional bug but requires hours of investigation for vague warnings may not improve the workflow. Ask authors and reviewers how often comments changed a decision and how difficult they were to verify.
Keep evaluation material separate from examples repeatedly used to tune the review prompt. Otherwise the system may become excellent at the team's familiar demonstrations while remaining unreliable on new changes. Reassess when the repository architecture, model, or tool permissions change materially.
Let humans own the resolution
The author should be able to accept, reject, or defer a finding with a reason. Rejection is useful feedback when the model misunderstood an invariant or lacked context. Do not force people to satisfy an automated comment merely because it appears in a formal review interface.
For accepted findings, review the repair as a new change. An AI-generated fix can introduce a different defect or broaden the patch unnecessarily. Check the original requirement, the regression evidence, and the final diff together. The fact that the same system identified the issue and proposed the repair does not establish independent verification.
Keep the final explanation short and concrete: what could go wrong, what changed, and what evidence supports the repair. That record helps future maintainers understand the decision without replaying a long conversation between tools. Useful review leaves the codebase and its reasoning easier to maintain.
Make automated review a disciplined second reader
AI review works best when it expands attention: following a rarely used branch, checking an overlooked caller, or proposing a precise failure scenario. It works poorly when it substitutes confident generalities for repository knowledge or treats every possible risk as an urgent defect.
Start with a bounded role and inspect its actual contribution. Improve context selection, finding format, and verification before adding more autonomous actions. A smaller number of well-supported comments can be more valuable than an impressive stream of suggestions.
The goal is an evidence trail from changed code to plausible failure to verified resolution. When that trail is clear, automation can make reviews faster and more thorough without obscuring responsibility. The final authority remains the engineering process that can explain and test the behavior it chooses to ship.