Research, September 2026

LLM speed in real code review: what public benchmarks did not tell us.

The fastest lineup of AI models we tested found 4 of 10 known bugs. Our starting lineup found 9 of 10. Several models that looked fast in public speed charts could not finish their answers in time.

We built MegaLens and ran these tests on it. The main test covered 363 full review runs. Including follow up tests, we spent about $108 in model fees trying to improve our lineup. During the same week, a model we relied on stopped working.

363

full review runs in the main test

9 of 10

known bugs found by the lineup we kept

4 of 10

found by the fastest lineup

0.59 to 0.99

how closely time followed answer length

These were security reviews of web apps. An AI model matched the results to findings an earlier study had settled. No human reviewed the results.

The setup: one goal, beat our own lineup

MegaLens sends code to AI models from different companies. Each review result names the models that took part. We wanted to see whether changing our lineup would help it find more real bugs.

We used 11 real codebases from an earlier fact check. That study had settled 17 findings: 10 real bugs and 7 false alarms, meaning reported problems that were not real. Those findings gave us an answer key.

We tested our starting lineup, then changed one model at a time. Each setup ran 3 times on each of the 11 codebases. For this test, a bug counted as found when at least 2 of the 3 runs found it. We ranked setups by real bugs found, then speed, then cost.

LLM timeouts: response time tracked answer length far more than input length

Every answer in our review has a time limit. A timeout means the model did not finish within that limit. Several models that looked fast in public charts timed out on our code.

Grok 4.6

Timed out in 33 of 33 runs on a task requiring a long written answer. It did very well on a shorter task.

GLM 5.3 Flash

Timed out in 32 of 33 runs.

DeepSeek V4 Flash

Timed out in 27 of 33 runs.

DeepSeek V4.1 Flash

OpenRouter’s published data listed it as the fastest DeepSeek model: 0.76 s to the first token and 80 tokens a second. In our review, it timed out in 12 of 33 runs. Reviews that included it found 5 of the 10 known bugs.

Claude Sonnet 5 and Grok 4.7

Timed out on some runs of one task.

Time to the first token tells you how soon an answer starts. In our tests, public figures for that measure did not predict how long a complete answer would take. Our inputs contained about 5,000 to 7,500 tokens of code.

LLM latency comparison: seven models doing the same job

We measured how long seven models took to answer on the same codebases, doing the same job. We compared that time with the length of the input and the answer, both measured in tokens.

The time column shows the median, or middle recorded time, per answer. The last two columns show correlations: 1 means time rises in step with length, and 0 means no link.

ModelMedian time per answerAnswer length (tokens)Writing speed (tokens a second)Link between time and answer lengthLink between time and input length
Muse Spark 1.218 s4,2052220.79barely (0.08)
Gemini 3.1 Pro19 s2,6211330.99no (0.0)
Gemini 3.8 Flash22 s1,996890.91some (0.65)
GPT 5.429 s3,0581050.95no (below 0)
Grok 4.638 s2,498610.61no (below 0)
Grok 4.744 s *3,176750.98a little (0.30)
Claude Sonnet 550 s *4,471910.59barely (0.07)

* Only runs that finished within the time limit. The others timed out, so the true median time is higher.

For every model, time had a stronger link with answer length than with input length. Answer length correlations ranged from 0.59 to 0.99. The link with input length was close to zero or below for most models. Claude Sonnet 5 wrote long answers at a normal speed. The two Grok models wrote more slowly.

If a model keeps timing out, check how much it writes before shortening your prompt. Try asking for a shorter answer and measure whether that helps.

LLM speed comparison: the fastest lineup found the fewest bugs

The fastest lineup we tested used a lighter, faster model in place of one of ours. It cut the median time per full review from 91 seconds to 82 seconds, about 9 seconds faster.

It found 4 of the 10 known bugs. Our starting lineup found 9 of 10. That saved 9 seconds while missing 5 more of the known bugs.

Every swap in the main test found fewer real bugs than our starting lineup: between 4 and 8 of 10. Those results supported our choice to put bugs found ahead of speed.

The best bug finder can be the wrong model to check the findings

One model was among the best we tested at finding problems. In a full review, it found as many known bugs as our starting lineup. It also found a bug that lineup had missed.

We separately tested the same model on a second pass: checking whether findings were real. It accepted two false alarms. The best models in that test accepted none.

Finding problems and checking findings needed separate tests. A model's result on one job did not tell us how well it would do on the other.

OpenRouter 404, model not found: a model we used disappeared

On 23 September, a model we used stopped working. OpenRouter had removed Devstral, a coding model from Mistral. Calls to it returned a 404 error.

The error message suggested a newer model id. It was the same id we were already calling, and it also returned 404.

Our last successful call was at 08:23 UTC. The fix went live at 20:20 UTC, almost 12 hours later. Our alerts missed the failure because the call happened outside the part of the review we logged. The alerts read that log. A person testing the product found the problem first.

That night, we checked every model id we call against OpenRouter. Three were no longer available, and all three were in use. Three more that we use had published end dates less than a month away.

What we changed

Every call now reports a withdrawn model immediately and sends us an email. A daily check reads each model's published end date and warns us 14 days ahead.

If your app uses hosted models

  1. Log every model call, including calls outside your main flow.
  2. Set a separate alert for “model not found”. It is a different problem from a slow or busy provider.
  3. Watch the end dates your provider publishes. They give you days of warning for free.
  4. Call a suggested replacement yourself before relying on the id in an error message.

Our guide to OpenRouter errors and how to fix each one covers 404, 429, 401 and “provider returned error”.

How to choose an LLM for code review: what we changed and what we kept

The main test supported keeping most of our lineup. No swap in it found more real bugs. Other tests helped us choose the few changes we made.

The results gave us more confidence in our choices for this workload, with limits:

  1. They do not give a precise ranking. We had only 10 known bugs. Across its three repetitions, our starting lineup found 7, 9 and 7 of those 10 bugs. A difference of one bug is within that variation.
  2. They cover security reviews of web apps only. We did not test other work.
  3. They do not predict next month's results. Providers change models, and results change too.

Questions people ask

Why does my LLM keep timing out?

In our tests, time had a stronger link with answer length than with input length for every model we measured. Check how much the model writes. Ask for a shorter answer and measure whether it finishes sooner.

Does a longer prompt make an LLM slower?

It varied by model. Across the seven models we measured, correlations between answer length and time ranged from 0.59 to 0.99. Input length correlations reached 0.65 at most and were close to zero or below for five of the seven models.

Is the fastest LLM the best one for coding tasks?

Not in our test. The fastest lineup saved about 9 seconds per full review and found 4 of 10 known bugs. The lineup we kept found 9 of 10.

What does "OpenRouter 404 model not found" mean?

In our case, OpenRouter had removed the model. Its suggested replacement id was the one already failing. Check the live model list and test the id yourself.

How do I choose the best model for code review?

Use code where you already know the real bugs. Change one model at a time and repeat each test. Compare real bugs found first, then false alarms, then speed, then cost. Public speed figures did not predict which models would finish within our time limit.

Method

We ran the tests in September 2026 using the same review code customers use. The 11 web app codebases came from an earlier study; none is named. That study checked 20 findings: 10 agreed real, 7 agreed not real, and 3 unresolved. We excluded the unresolved findings.

In the main test, each alternative setup changed one model. Each setup ran 3 times on every codebase. A bug counted as found when at least 2 of the 3 runs found it. An AI model matched each review's findings to the answer key without seeing which findings were real. No human reviewed the results.

The speed table reports median time per answer. The lineup comparison reports median time per full review. Costs are model fees reported by OpenRouter.

View sources on GitHub →