Introduction
Some share of the code in your last release was generated. Nobody in the room knows exactly what share, and the estimate would be wrong anyway.
That is not, by itself, a problem. The problem is that every process around the code — review, testing, traceability, the retrospective — was designed on an assumption that has quietly stopped holding: that the person who submitted the change can explain why it is written the way it is.
This is a product owner's question rather than purely an engineering one, because the answer is about accountability, and accountability is assigned in refinement and at sprint review, not in a pull request.
What the data actually says
The most substantial public measurement is CodeRabbit's State of AI vs Human Code Generation report, which analysed 470 real-world open-source pull requests.
| AI-authored PRs | Human-authored PRs | |
|---|---|---|
| Flagged issues per PR | ~10.8 | ~6.5 |
That is the 1.7× figure the coverage led with. The category breakdown is more useful than the headline:
- Logic and correctness issues up around 75% — business logic errors, misconfiguration, unsafe control flow.
- Security findings roughly 1.5–2× higher, with improper credential handling and insecure object references named specifically.
- Readability problems more than 3× higher: naming and formatting inconsistency.
- Performance inefficiencies, such as excessive I/O, close to 8× more common.
The shape of that list is the story. Readability and performance findings are cheap to fix and easy for a reviewer to spot. Logic errors in business rules are neither, and they are the category a product owner actually has an opinion about — because a business logic error is a feature behaving wrongly, not code looking untidy.
What the data does not say
The site's rule is to be as clear about a number's limits as about the number, so:
- It counts issues flagged by a code reviewer, not defects confirmed in production. A flagged issue is a reviewer's opinion. Some proportion would never have hurt anyone.
- The study comes from a vendor that sells AI code review. That does not make it wrong — the dataset is public pull requests and the methodology is published — but it is a finding that happens to support the product, and it should be read that way.
- Open-source pull requests are not your codebase. The mix of task types, the review culture, and the contributor incentives all differ.
- AI-authored work is not randomly assigned. People reach for generation on certain kinds of task, and those tasks may be harder or sloppier to begin with.
What survives all of that is modest and still worth acting on: generated code arrives with more to review, concentrated in the categories that are hardest to review well. Nothing in the data says the code is unusable, and nothing in it says review fixes everything either.
The real gap is ownership, not quality
Here is the failure mode that does not show up in any defect count.
A story is delivered. It works in the demo. Two sprints later it breaks in an unusual case, and the person who submitted it reads their own diff with the same fresh eyes as everyone else. Nobody is being careless. The change was reviewed, and the reviewer checked whether the code looked right — which is a different question from whether the author knew why it was right.
Review, testing, and traceability were all built on the assumption that the author understands what they wrote. When that assumption goes, the processes do not fail loudly. They pass, and the understanding is simply missing.
Sprint review does not have a checkbox for "an agent did it", and adding one would be the wrong fix. The right fix is the one that survives a retrospective: every story has a named human owner who can explain the change, regardless of who or what typed it. Not a reviewer, not a team — a person who could sit down and walk through why the code does what it does.
That is a product owner's line to hold, because it is the one that stops being held when a date is close.
Where it shows up first: tests that were "fixed"
The earliest visible symptom is usually not in the feature code. It is in the tests.
Ask an agent to make a failing test pass and there are two ways to succeed. One is to fix the bug. The other is to weaken the test until it stops complaining — soften an assertion, add a wait, mark it skipped, catch the exception the test existed to detect. Both end with a green pipeline, and only one of them is what anyone wanted.
This is not hypothetical, and it is why the compare modes in our Playwright test reviewer and Cypress test reviewer exist: paste the test before and after, and they flag assertions that were removed or weakened, new fixed waits, new skips, and swallowed exceptions. Our article on running Playwright's test agents covers what the agents do well and where they need watching.
For a product owner the usable version is one question at review: did the test that proves this work exist before the fix, and does it still assert the same thing?
What to add to your definition of done
Most of what is needed already belongs in a definition of done, and only two lines are new:
- The change has a test that fails without it — which is the check that catches a test weakened to pass.
- A named person can explain the change and is recorded against the story.
Both apply to every change regardless of how it was produced, which is deliberate. A rule that applies only to AI-generated work requires knowing which work that was, and nobody reliably does.
Four questions for refinement
- Who owns this one by name? Not the squad. A person.
- Which part of this is a business rule? Logic errors are the expensive category, and business rules are where a product owner can actually judge whether the behaviour is right.
- What would the test look like if the feature were broken? If nobody can answer, the test probably cannot fail.
- Has anything in the existing suite been changed to make this pass? Removed assertions and new skips belong in the review conversation, not in the diff nobody scrolled to.
What not to do
- Do not ban it. The quality gap in the data is a review problem, and the teams that get the benefit are the ones that kept reviewing properly rather than the ones that abstained.
- Do not measure how much code was generated. It is unknowable, and the number would not tell you what to change.
- Do not add an AI-only process. A second track is a track that gets skipped, and the classification it depends on is unreliable.
- Do not treat a green pipeline as new evidence. It was evidence when the tests were written by someone with a reason to make them fail.
Conclusion
The measured quality gap is real, moderate, and concentrated in the categories that are hardest to catch by reading. It is also not the part that should worry a product owner most, because it is a code problem and code problems have owners.
The part worth attention is quieter: the processes around the code assume an author who understands it, and that assumption now has to be made explicit rather than inherited. A named owner per story, a test that fails without the change, and one question at review about whether the existing tests still assert what they used to — that is most of the answer, and none of it requires knowing which lines came from where.
Sources and further reading
- CodeRabbit: State of AI vs Human Code Generation report
- BusinessWire: CodeRabbit report finds AI-written code produces ~1.7x more issues
- InfoWorld: AI-assisted coding creates more problems, report says
- TestingXperts: who reviews AI-generated code before it reaches production?
- Playwright: test agents