Live case study
Claude Code’s self-review missed 12 issues. Here’s what caught them.
What happens when Claude Code gets independent AI reviewers?
I asked Claude Code to build an AI Email Drafter. Then I used MegaLens to get opinions from independent models, while Claude stayed the main builder and made the decisions.
15
actionable findings
12
additional findings beyond Claude’s initial self-review
10
more issues surfaced during implementation
167
tests, all passed
$2.48
for the plan review and the mid-build code review
Plan audit $1.25 · mid-build code audit $1.23. Both prices are printed in the session output, at $9 per 1M blended tokens. The final security audit’s price was not printed, so it is not counted here.
The experiment
I asked Claude Code to build an AI Email Drafter. It produced the implementation plan, then reviewed that plan itself before writing any code. That self-review is a real step, and it caught real things.
Then I sent the same plan through MegaLens. MegaLens sends the work to independent models from different families and returns what they found, including the points where they disagree with each other.
Nothing about the build changed hands. Claude Code stayed the builder. The only thing added was a second set of opinions for it to weigh.
What happened
The independent review came back with 15 actionable findings on the plan. 12 of those were additional findings beyond Claude’s initial review. The other 3 overlapped with what Claude had already flagged on its own.
Claude Code read all 15 against the plan and decided which ones held up. Only the findings that survived that check were carried into implementation.
Partway through the build I ran a second review, this time on the committed code rather than the plan. It surfaced 10 more issues in the first step, which Claude evaluated and addressed before continuing with the rest of the build.
The build finished with all 167 tests passing.
What the reviewers actually found
Eight of the fifteen, exactly as they came back. “Caught by” says whether only MegaLens raised a finding, or both MegaLens and Claude’s own pre-audit did.
| Finding | Severity | Category | Caught by |
|---|---|---|---|
| No actual USD cost parsing from response.usage — cost cap is unenforceable | Critical | Operational | MegaLens only |
| Bootstrap is :unread query + no gmail.modify — DB loss causes mass re-drafting of old mail | Critical | Operational | MegaLens only |
| Email headers (subject, from, date) injected raw into LLM prompt — no sanitization | High | Security | Both (pre-audit missed headers) |
| Draft-create → mark-processed crash window — duplicate drafts, no atomic guard | High | Operational | MegaLens only |
| Reply envelope underspecified — To, Cc, References, In-Reply-To, Reply-To vs From not addressed | High | Design | MegaLens only |
| No retry budget or poison-message quarantine — bad emails retry every poll forever | High | Operational | MegaLens only |
| systemd Environment= stores file paths not values — broken auth or secret exposure in /proc | High | Security | Both (my pre-audit + MegaLens) |
| SQLite without WAL mode or busy_timeout — APScheduler threads cause "database is locked" | Medium | Operational | MegaLens only |
And one I rejected
The final security audit flagged F-014, incomplete type checking. I checked it and it was wrong: from __future__ import annotations already covered the case it was worried about. It went down as a false positive and nothing was changed.
This is normal. An outside model sees the file and not the whole project, so some of what it raises turns out to be wrong once the rest of the codebase is taken into account.
Run the same three reviews on your own repo.
Try it on your repoWhat I learned
The useful part was not proving that one model is better than another. I have no interest in that question and this run does not answer it.
What it shows is narrower and more practical: different models notice different things. Some of the outside findings were useful. Some were noise. A few were confidently wrong about code they could only partly see.
That is why the judgement stays with Claude Code. It has the repository, the plan and the implementation in context. An outside reviewer has only what it was sent, so its findings are input, not instructions.
Claude Code builds. Other models challenge assumptions. Claude decides what to do with the feedback.
I build MegaLens this way too
MegaLens itself was built primarily with Claude Code, and I used MegaLens throughout its own development to review its plans and its implementation. In other words, I used MegaLens on itself while building it. The workflow on this page is the one I actually use.
122 reviews have been run from my own accounts since 18 April 2026, counted in our database on 25 September 2026.
Transparency
- •MegaLens is a commercial, private source product.
- •The AI Email Drafter shown in the demo is a public demo project.
- •The numbers on this page are counts from that recorded run. They are not a benchmark, and a different project would produce different ones.
- •A second recorded run of the same project, with its own counts, is documented here.
Run the same three reviews on your own repo.
Works with Claude Code, Codex CLI, Cursor, Gemini CLI and Lovable, and with OAuth connectors such as ChatGPT and Claude.ai. MegaLens reviews. Your IDE decides what to change.
Try it on your repo