Research, September 2026
We ran AI code review on 30 apps. Then we fact-checked our own findings.
The review came back with 160 findings. We took a random sample of 20 and had two AI models from different companies check each one against the code, independently. On the findings they agreed about, 44% did not hold up. And not one “critical” label was rated critical by both.
30
apps reviewed
20
findings checked
44%
false positives on agreed findings
0 of 9
critical labels kept by both checkers
Validated by two AI models from different companies. No human reviewed these findings, and no exploit was attempted. A sample of 20 has a wide error bar.
What we ran
30 public web apps from GitHub, in three groups. 25 of them were written with AI, which is the code this study is about.
| Group | Apps | How we identified it |
|---|---|---|
| Built by an AI web builder | 15 | Identified by the builder’s own build plugin in package.json |
| Written by developers with AI help | 10 | Identified by an AI co-author line in the commit history, or a README saying it was built with an AI coding tool |
| An earlier sample, rerun | 5 | Projects from an earlier study, run again on the exact same input |
Every app got the same security review with the same settings: its security-relevant files, up to about 28,000 characters. Each finished run used three or four AI models from different companies. The whole batch took 18 minutes and cost $5.67 in model fees. 28 of the 30 runs finished; 2 stopped part way and are counted as results, not retried.
What came back
160 findings across the 28 finished runs, a median of 5 per app. Every finished run reported at least two.
That is the number a review tool would normally lead with. We did not want to publish it until we had checked how far to trust it.
Then we checked our own findings
We drew 20 of the 160 at random, with a fixed seed, spread across the three groups and weighted toward the findings labelled critical, because those carry the most weight.
Two AI models from different companies each checked all 20 against the actual code. For the 15 findings whose projects still exist, that meant the exact commit that was reviewed, and the rest of the project too. The other 5 came from projects that have since been deleted, so they were checked against the exact code the review saw. The second checker was not shown the first one's answers. The first checker's verdicts were saved and fingerprinted before the second one finished, so neither could be adjusted to match the other.
We only count a verdict where both checkers agreed.
| Outcome | Findings |
|---|---|
| Both checkers agree the finding is real | 9 |
| Both checkers agree it is not real | 7 |
| Unresolved: the checkers disagreed, or one could not tell | 4 |
On the 16 findings both agreed on, 7 were not real: a false-positive rate of about 44%. Put simply: of the findings both checkers agreed on, a little over half held up and a little under half did not.
Severity did worse. Nine of the sampled findings were labelled critical. Not one was rated critical by both checkers. On the nine findings both agreed were real, the second checker rated eight of them lower than the first.
Why findings failed: three causes you can check for
Each of these accounted for more than one false positive. They are worth knowing whatever review tool you use.
1. Thin input
This review reads a slice of a project, not all of it. In one run the slice was a single file of data types, from an app with no backend at all. The models reasoned about a server that did not exist, and two findings came out of that.
Check: Check how much of the project the review actually saw. A finding about code that was never shown to the reviewer is a guess.
2. Protection that lives somewhere else
One finding said an endpoint had no cross-site check. The first function that endpoint calls does exactly that check. Another said a database function did not validate a value; the table itself rejects bad values, and only a trusted role can call the function.
Check: Before accepting "X is missing", follow the call one level down and look at the database constraints.
3. Intended behaviour read as a flaw
Admins seeing revenue on an admin-only page. An analytics counter treated as if it were a login session. A password-reset link built from the page’s own address, which an attacker cannot change.
Check: Ask whether the behaviour is what the app is supposed to do. If it is, it is not a vulnerability.
And the problem with severity
The labels look like they were set by pattern: a secret, a missing check, a missing signature. They do not reflect whether the code is ever called, or what an attacker could actually do with it. Code that nothing in the app calls was labelled critical twice.
Check: before acting on a critical label, confirm the code is reachable and that the app is deployed the way the finding assumes.
Where the two checkers disagreed
Four findings stayed unresolved. They are left out of the 44%, but they are the most useful part of the exercise, because each one is a real question a review has to answer, and a single checker would have answered it without knowing there was a question.
An analytics beacon sent without authentication
One checker: real, but low risk, because a browser beacon cannot hold a secret. The other: cannot tell without seeing the script that receives it.
An analytics sheet readable from the browser
One checker: real, low value data. The other: cannot tell, because whether the sheet is public is a setting, not code.
A form that uses a random link as the only key
One checker: a real design limit, low severity. The other: intended design for an anonymous form, not a flaw.
A developer tool that falls back from a container to the host
One checker: real, medium. The other: the container mode refuses the fallback; it only happens in the automatic mode, where running on the host is the intended behaviour.
In the last case we went back to the code afterwards: the second checker had read the file that decides the fallback, and the first had not. We still count it as unresolved, because the rule was set before the answers came in.
The pattern across all four: the disagreement is exactly where the context lives. Who calls this code, what the setting is, what the app was built to do. A review that hides disagreement hides the questions you most need to ask.
Did more models agreeing help?
The obvious guess is that a finding raised by several models is more trustworthy. On the findings both checkers agreed about, it was not.
| Finding raised by | Real | Not real |
|---|---|---|
| One model | 4 | 1 |
| Two or more models | 5 | 5 |
| No model recorded | 0 | 1 |
Findings raised by several models were real 5 times out of 10. Findings raised by a single model were real 4 times out of 5. The numbers are small, so we would not call it a pattern yet. But they do not support the easy story.
What they do support is narrower: a second model surfaces what the first one missed. It does not make what they agree on true. That is why MegaLens shows where its models disagreed, instead of handing back a single agreed answer.
What this does not show
- How many real vulnerabilities these apps have. The sample says how far to trust a review, not how safe the apps are.
- That code from one group is better or worse than another. The groups are too small to compare.
- A precise false-positive rate. Twenty findings give a direction, with a wide error bar.
- What a human expert would conclude. No human reviewed these findings.
Update, 23 September 2026
We went after the causes. Then we measured again.
After the first measurement we changed how MegaLens labels its findings. Then we checked the same 20 findings again, once after each of two rounds of changes.
What we changed
- A finding the audit rejects, or cannot settle, is now shown as disputed or unverified, with the reason. Before, a finding the audit had rejected could still be shown as validated. Nothing is deleted; the reader sees the verdict beside the finding.
- When the audit accepts a finding, it has to quote the line of code its verdict rests on. If that line is not in the code the review was given, the finding is shown as unverified, not as checked.
- For every finding, the audit is asked whether the code can be reached, whether it runs where the public can reach it, and what severity it would give. An answer can lower a severity label, never raise it, and the finding says why.
- Checks that use no AI model now look for two things that went wrong in the first measurement: a database rule that already enforces what a finding says is missing, and a finding that depends on a file the review was never sent.
What the numbers did
| Before the changes | After | |
|---|---|---|
| Findings labelled verified that both checkers agree are real | 10 of 14 (71%) | 10 of 11 (91%), after each round |
| Findings in the sample labelled critical | 9 | 2 |
| Real findings with a severity inside the range both checkers gave | 2 of 9 | 7 of 9 |
No reduction in false positives was shown. The same 20 findings exist, and the same 7 were agreed not real in every check. What changed is the label each one carries: a finding the audit could not stand behind no longer goes out marked verified, and severity labels sit closer to what the checkers gave. Two findings in the sample are still labelled critical that neither checker rated critical.
The failure inside the fix
Our first round of changes made one thing worse. In one review, the audit approved ten of eleven findings with reasons a few words long, such as “Confirmed.” One of them, an admin seeing revenue on an admin page, went out labelled verified. Before the change it had not been.
For the second round we wrote the test down before building anything: run the audit on that same review three times, before the change and after. Before, the finding came out verified in two of three runs. After, disputed in all three.
Then a different false positive took its place. An anonymous analytics id stored in the browser was accepted as a security flaw in two of three runs. The audit quoted the right line of code both times. The claim about the code was true; the judgment about it was not. We had only one earlier run of that finding, so part of this change may be ordinary variation between runs.
Both cases belong to a cause the first measurement already named: intended behaviour read as a flaw. That is a failure of judgment, not of missing information, and writing a new rule for each case would not end. So we stopped after two rounds, on purpose, and handle what is left a different way. That way covers findings shown as unverified or disputed. A finding wrongly shown as verified, like the analytics id above, gets no follow-up question, and two findings are still labelled critical too high.
What ships now for what is left
The review only sees the code it is sent. Your editor can open any file in the project. So when a finding comes back unverified or disputed, and there is a file or function to point at, it now carries one question that says what to open and what each answer means. Sometimes the answer is not in the code at all but in a setting or how the app is deployed; the question still says where to start.
The cleanest example from the measurement is the cross-site finding from the first round, where the check it said was missing lives in a helper the review was never sent:
Question: Open src/http.ts. Does the code there already handle what this finding says?
If yes: The finding does not apply as written: the code the review could not see covers it.
If no: This helper does not cover it. Check for other protections before treating it as real, severity medium if none exist.
Shortened: in the real output the question also repeats the finding. In this case the answer is yes: the helper refuses cross-site requests.
How we measured
Same 20 findings, same random seed, same code. We did not rerun the whole review: we reran its audit step on each original review, with the original findings, three times per review in the second round, and counted the majority. The rules for the second round, including when we would switch a change back off, were written down and fingerprinted before it was built. The replay was not identical to a live review: it used shortened finding text, a rebuilt plan and a fixed routing step, and one draw took a different approval path than production would. Two AI models from different companies checked the findings again. The first knew how the pipeline had labelled each finding. The second saw only the code and the findings with their new severity labels: no statuses, no earlier verdicts, and nothing about the changes. No human reviewed the findings. The precision figures use the findings both checkers agreed on in the first re-check (10 real, 7 not); in the second re-check one of those real findings became unresolved, which is why the severity row counts 9. The 20 findings were drawn to lean toward critical labels, and 20 still give a wide error bar. Nothing was tuned within a round after its numbers came in; the second round was a response to the first round's results.
Questions people ask
How accurate is AI code review?
In our sample, less accurate than its output suggests. Two AI models from different companies checked 20 findings independently. On the 16 they agreed about, 9 were real and 7 were not, a false-positive rate of about 44%. It is a sample of 20, so read it as a warning, not a precise rate. When we later changed how our findings are labelled, no reduction in false positives was shown: the same seven findings were agreed not real in every check. What changed is the label. In a re-check of the same sample, using the first re-check’s agreed labels throughout, the share of findings marked verified that were real rose from 71% to 91% after the first round of changes and stayed at 91% after the second.
What causes false positives in AI code review?
Three causes each accounted for more than one false positive: the reviewer saw too little of the project, the protection existed in a function or database rule it did not look at, and intended behaviour was read as a flaw.
Can you trust the severity labels from AI code review?
Not on their own. Nine findings in our sample were labelled critical. Neither checker kept every one of them at critical, and not one was rated critical by both. The labels look set by pattern, without weighing whether the code is reachable or what an attacker could actually do with it. After we made our own review answer those questions for every finding, the same sample carried two critical labels instead of nine, and 7 of the 9 real findings got a severity within the range both checkers gave, up from 2 of 9.
Does using more than one AI model make findings more reliable?
Not in the way people assume. Findings raised by several models were real 5 times out of 10 in our sample. Findings raised by a single model were real 4 times out of 5. A second model surfaces what the first missed. It does not make what they agree on true.
Method
Reviews were run on 22 September 2026 with MegaLens, a security review on the same settings for every app, on a dedicated test account. The apps are public GitHub projects whose owners did not ask to be reviewed, so none is named and no finding is quoted. Nobody was contacted.
Sample: 20 findings, drawn at random with a fixed seed, stratified by group and weighted toward the critical label. Verification: two AI models from different companies, working independently from the same code and the same findings. Each classed every finding as confirmed, real but overstated, or not real; the second checker could also answer cannot tell. The published rate uses only findings where both agreed. Validated by AI models only. No human review, no exploit attempted, no live system touched.