Reviewing an AI-generated pull request means reviewing the code the way you always should have, with the diff read before the description, plus a few extra checks for the mistakes a model makes more often than a person: tests that assert nothing, methods and flags that do not exist, scope the ticket never asked for, and dependencies nobody chose. I review PRs for several teams as a Senior Principal Engineer at Questrade, and this is the list I run through.
Table of Contents
- What changes when the author is a model
- Read the diff before the description
- The review checklist
- What to ask the author to write by hand
- A 24-hour rule and PR size limits
- Labelling AI-assisted PRs
- Can AI review AI?
- An example review comment set
- FAQ
What changes when the author is a model
Four things, in my experience.
Volume. A developer with a coding agent can open more pull requests in an afternoon than they used to open in a week. Nothing about review got faster, so review becomes the queue everything waits in. The temptation is to skim, and skimming is the wrong response to the other three changes.
Plausibility. Model-written code looks finished. Names are sensible, comments are fluent, formatting is clean. Reviewers have always used surface mess as a signal to slow down, and that signal is gone.
Confidence. The description never hedges. A person writes “I’m not sure the TTL is long enough”; a model writes a tidy summary of every file it touched and nothing about what worried it.
Scope creep. Agents fix the thing next to the thing. They rename a variable in a file the ticket never mentioned, or improve error handling in a package they passed through on the way to the bug. Each change might be fine on its own. Together they make the diff hard to reason about and hard to revert.
There is published data behind the caution. Veracode’s 2025 GenAI Code Security Report tested code from more than 100 large language models across Java, JavaScript, Python and C# and found that “AI-generated code introduced risky security flaws in 45% of tests”, and that “Larger, newer AI models didn’t improve security.” GitClear’s 2025 code quality research, covering 211 million changed lines authored between January 2020 and December 2024, found lines classified as copy/pasted rose from 8.3% to 12.3% while refactoring fell from 25% of changed lines in 2021 to less than 10% in 2024. Neither study is about your codebase, and GitClear’s numbers are correlational, but both point the same way.
The reviewer now carries more of the thinking than before. That is one reason review is a bigger share of the job as you move up, which I wrote about in staff vs senior engineer.
Read the diff before the description
If the model wrote the code, there is a good chance it wrote the description too, and a description written by the same model that wrote the bug will not mention the bug. Reading it first primes you to see what it says you should see.
My order is git diff --stat to see which files moved, then the diff file by file, then a one-sentence summary of what I think the change does, written down. Only then do I read the description. If my sentence and the description disagree, that is my first review comment, and usually the most important one.
Google’s reviewer guide says to “look at every line of code that you have been assigned to review” (What to look for in a code review). That advice got harder to follow and more important at the same time.
The review checklist
Each row is something I check on every PR, with the extra attention a model-written change needs.
| Check | Why it matters more for AI-generated code | How to check |
|---|---|---|
| Correctness | The code looks right, so you stop tracing it. | Walk one real input through the change by hand. Run it locally. Try the empty case, the nil case and the largest value the type allows. |
| Hidden scope | Agents touch files the ticket never mentioned. | Compare git diff --stat against the ticket. Every file outside the obvious blast radius needs a sentence of justification or needs to leave the PR. |
| Tests that test nothing | Tests that pass are not the same as tests that fail when the code is wrong. | Open every new test. Look for assert.NoError as the only assertion, assertions on the mock instead of the result, and tests that restate the implementation. Flip one condition in the implementation and confirm a test goes red. Google’s bar is that tests are “correct, sensible, and useful”. |
| Invented APIs or flags | Models guess method names, config keys and CLI flags, especially for less common libraries. | For every new call into a dependency, confirm it exists at the pinned version with go doc or the source in the module cache. Do the same for every new config key and environment variable. |
| Dependency and licence hygiene | A model will add a package for a ten-line function without asking. | Diff go.mod, package.json or the equivalent. For each new entry: is there a standard library option, who maintains it, what licence, is the version pinned. |
| Security defaults | Permissive settings from a tutorial end up in production code. | Grep the diff for InsecureSkipVerify, AllowAllOrigins, 0777, Sprintf building SQL, hardcoded credentials, disabled auth middleware and logging of request bodies. |
| Error handling | Errors get logged and swallowed, or returned as nil, so the happy path compiles. |
Search for _ =, if err != nil { return nil }, bare recover() and retries without backoff or a bound. Ask what the caller sees when each error fires. |
| Performance assumptions | Models write for the example data set. | Look for queries inside loops, whole tables loaded into memory before filtering, unbounded goroutines and missing context deadlines. Ask what happens at 100x the test fixture. |
| Consistency with the codebase | Models default to the most common pattern on the internet, which is rarely your team’s pattern. | Compare logging, error wrapping, config loading and package layout against a neighbouring package written by hand. The service standards in Go microservices in production are the kind of thing a model cannot know unless the repo tells it. |
The order is roughly by severity: a correctness bug or a security default blocks the merge, while consistency is usually a “fix before merge” comment. You will not run every row on a five-line fix. The table is for when the diff is bigger than you can hold in your head.
Google’s standard for approval still applies: “reviewers should favor approving a CL once it is in a state where it definitely improves the overall code health of the system being worked on, even if the CL isn’t perfect” (The standard of code review). The checklist finds what makes a PR worse. Perfection is a different question and I do not block on it.
What to ask the author to write by hand
The parts of a PR a model cannot write are the parts a reviewer needs most. I ask for three things in the author’s own words, and I send a PR back if they are missing on anything non-trivial.
Why. One or two sentences on the problem. The diff already shows what changed.
Risk. Does it touch money, auth or data? Can it be rolled back by reverting the commit, or does it ship a migration?
Where they want eyes. The file or function the author is least sure of. If the author used an agent and did not read a particular file closely, this is where they say so.
A description that covers all three can be four lines. If the author cannot write them, I review as if nobody has looked at the code yet.
A 24-hour rule and PR size limits
Google’s guidance is that “One business day is the maximum time it should take to respond to a code review request” (Speed of code reviews). I hold the teams I review for to that and call it the 24-hour rule: every PR gets a first human response within a day, even if that response is “I’ll get to this tomorrow, here are two early comments.”
Agent-driven volume strains that rule, and the wrong fix is to review faster. The right fix is to put less in each PR. Google’s number is that “100 lines is usually a reasonable size for a CL, and 1000 lines is usually too large”, and “reviewers have discretion to reject your change outright for the sole reason of it being too large” (Small CLs). I use roughly the same line for AI-assisted PRs, with one adjustment: I ask for splits by behaviour, not by file. A PR that adds a retry policy and a PR that adds metrics for it are two reviewable units. Half the files of both is not.
If you want to know whether your team is meeting the 24-hour rule, measure it. I built DeliveryCompass partly because I got tired of guessing review turnaround from memory before staff meetings. Time to first review and PR size are the two numbers I watch when a team starts using agents heavily.
Labelling AI-assisted PRs
Whether to label PRs as AI-assisted is a team norm, and I have seen it go both ways.
The case for labelling: the reviewer calibrates. If I know an agent wrote most of a diff, I spend my first ten minutes on the invented-API and test-quality rows instead of on naming. Labelling also gives you data. After a quarter you can compare revert rates between labelled and unlabelled PRs.
The case against: the boundary is fuzzy. Autocomplete, a chat window and a fully autonomous agent are very different, and one label flattens them. Labels attract stigma, and once they do, people stop applying them, which destroys the data. A reviewer can also use the label as an excuse in either direction, rubber-stamping or nitpicking.
My preference is a sentence in the description instead of a badge: “Most of this was written by an agent; I reviewed handler.go closely and mapper.go less so.” That tells the reviewer what they need and keeps the degree of assistance visible without a tag people argue about.
Can AI review AI?
As a filter, yes. As an approval, no.
A second model, reading the diff in a fresh context, catches a useful class of problems: unused variables, a missing nil check, a test that asserts on the wrong value, a function the library does not export. Anthropic’s Claude Code guidance recommends a writer session and a separate reviewer session because “A fresh context improves code review since Claude won’t be biased toward code it just wrote” (Claude Code best practices). The same page warns that “A reviewer prompted to find gaps will usually report some, even when the work is sound”, so a model review needs a person deciding which findings matter.
GitHub draws the same line in its product. Its documentation says “By default, Copilot’s reviews do not count toward required approvals for the pull request” and tells users to “Always validate Copilot’s feedback carefully. Supplement Copilot’s feedback with a human review.” (Copilot code review).
I run the model review before the human one, so the human does not spend attention on lint-level issues. Approval stays human because approval is a person saying they understand the change and will answer for it in production. A model cannot be paged, cannot be asked why in six months, and shares blind spots with the model that wrote the code.
An example review comment set
These are generic, but they are the shape of comments I leave on AI-assisted PRs. Each names the file, says what I checked, and asks a question the author has to answer in their own words.
The description says this only touches the retry path, but
config/loader.gochanged too. Was that intentional? If yes, add a line on why; if not, please drop it from this PR.
TestProcessOrder_Successonly assertsNoError. I flipped the status check on line 42 locally and the test still passed. Can you assert on the resulting order state?
client.WithRetryBudgetis not in v1.8 of this library (checkedgo doc). Which version were you targeting, or should this be our own wrapper?
InsecureSkipVerify: truein the HTTP client. I assume this came from local testing. Please remove it or gate it behind the dev config.Errors from
repo.Saveare logged and ignored on line 88. What should the caller see when the save fails?Our services wrap errors once with
%wand log at the handler. This package logs at every layer. Can you match the existing pattern ininternal/orders?
None of these mention the AI. They treat the author as the author, because whoever opens the PR owns the code in it.
FAQ
Should AI-assisted PRs be labelled?
It depends on the team, and the best version is a sentence in the description instead of a tag. Say how much of the change a model wrote and which files you reviewed closely yourself. That gives the reviewer what they need to calibrate without a badge that attracts stigma or lazy approvals.
How long should a code review take?
A first response should arrive within one business day; Google’s reviewer guide states that as the maximum. The review itself takes as long as the diff needs, and if that is routinely more than an hour, the PR is too large and should be split by behaviour.
Can AI review AI-generated code?
Yes as a first-pass filter, never as the approval. A second model in a fresh context catches unused code, missing checks and invented APIs before a person reads the diff. Approval stays with a human because approval is accountability for the change in production, and both GitHub and Anthropic say the same in their documentation.
What is a code review checklist?
A code review checklist is a fixed list of things a reviewer checks on every pull request so that attention does not depend on mood or time pressure. For AI-generated code mine covers correctness, hidden scope, test quality, invented APIs or flags, dependency and licence hygiene, security defaults, error handling, performance assumptions and consistency with the codebase, each with a concrete way to check it.
How big should a PR be?
Small enough to review properly in one sitting. Google’s published guidance is that about 100 lines is reasonable and 1,000 is usually too large, and that reviewers can reject a change for size alone. For agent-written code I ask for splits by behaviour, so each PR changes one thing a reviewer can hold in their head.