We benchmarked A LOT of models, here’s how they compare to Mythos

Everyone is talking about Mythos-class capability. Nobody has defined it yet, so we put Mythos to the same benchmark as other frontier models, open-weight models and custom harnesses. Because "Mythos-class" has become an industry shorthand for a capability tier, which really doesn’t have a clear list of capabilities. So we ran Mythos on the same benchmark we've applied to 22 other configurations against a set of human-reviewed IDOR labels to find out exactly what “Mythos-class” means.

August 25th, 2026

Something happens to the security industry every time a frontier model ships. Within a fortnight, a new adjective enters the vendor vocabulary, and while Mythos has been out of the public reach it is Mythos-class. Mythos-class threats. Mythos-level reasoning. Products positioned as ready for the Mythos era, capabilities benchmarked against a Mythos standard that nobody has had access to (except for a few vendors).

So now you can run Mythos inside the Claude Security harness, we decided to figure out what Mythos-Class actually means! Because we have a benchmark sitting right here and we've been running every model we can get our hands on through it — GLM 5.2 and 5.3, Kimi K3, the GPT-5.6 family, the Grok line, DeepSeek, and both current Claude generations.

How does Mythos stack up?

So we ran Mythos through it too, and put it next to everything else. Our dataset included 275 reviewed IDOR labels across four codebases, the standard one we’ve used across all our benchmarks so far. Every label was manually reviewed by a human which is what makes measuring recall possible at all.

Here’s what Claude Mythos was able to find. We are also extending our benchmark to cover more vulnerabilities to get a better sense of what each model finds well, and what it doesn’t. So while this blog post will go in depth on IDORs, you’ll also see references to other vulnerabilities in our results below.

To evaluate findings at scale we used an internal LLM-as-judge that combines semantic reasoning with deterministic tooling, rather than matching on line numbers or rule IDs. Our full experimental setup is described in our earlier writeup on Kimi K3's code security results. Again worth mentioning here that for the moment Mythos is currently only available in the Claude Security Harness, but here is where Mythos excelled, finding many critical and high risk vulnerabilities in our benchmark set.

Image showing severity by bug shape.

But where does that place it when it comes to other models? Well the result might surprise you. The short version: on precision, Mythos is genuinely good and sits comfortably in the upper half. On recall — the number of known vulnerabilities correctly identified in a codebase — it placed 15th out of 17 raw configurations for IDOR vulnerabilities specifically. An open weight model you can host yourself, beat it on both axes for $4.29 a run. A previous-generation Claude model matched its precision to the decimal and found 69% more bugs.

Image showing precision vs. recall.

Some caveats, before we dive into this in detail: Mythos used our internal LLM-as-judge system while the other runs used deterministic matching. Additionally, Mythos in Claude Security is configured to detect a wide variety of vulnerability classes. Lastly, Semgrep Multimodal is our own harness; it's in its own cohort for exactly that reason but we’ll dive more into that later, but there is a big gap between raw models and custom harnesses. And finally we don’t have a cost breakdown for our scan to do a cost per run and cost per true positive.

The Baseline AKA This Is What Mythos Class means

Mythos: 80.0% precision · 13.9% recall · ~23.7% F1

144 of the 275 labels are human-adjudicated true positives. Mythos matched 20 of them. That's where 13.9% comes from, so it has missed the vast majority of our IDORs, but that doesn’t mean much when we don’t compare it to other models.

Image showing that almost everything found more IDORs than Mythos.

But how much more? Let’s dive into them by cohort.

Cohort 1: Open Weight Models vs. Mythos

No scaffolding beyond the standard prompt. Mythos cannot run as a standard prompt, so it is in the CCS harness instead.

Model

Precision

vs Mythos

Recall

vs Mythos

F1

Cost / run

Mythos (baseline)

80.0%

13.9%

1.00x

~23.7%

Kimi K3 (Fireworks)

76.3%

−3.7

24.4%

1.76x

36.9%

Kimi K3 (OpenRouter)

76.5%

−3.5

21.8%

1.57x

34.0%

$8.62

GLM 5.2 (Fireworks)

82.6%

+2.6

16.0%

1.15x

26.8%

$4.29

GLM 5.3 (zai)

71.3%

−8.7

11.7%

0.84x

20.0%

$5.70

DeepSeek V4 Flash

88.9%

+8.9

6.7%

0.48x

12.5%

$0.17

GLM 5.2 beat Mythos on precision by 2.6 points and on recall by 2.1, for $4.29 a run and $0.23 per true positive. An open weight model you can run on your own hardware was more accurate and more thorough. 

Kimi K3 is the recall answer to Mythos. 24.4% against 13.9% — 1.76 times as many real IDORs, for 3.7 points less precision. In label terms: of 144 real vulnerabilities, Kimi K3 hands you roughly 35 and Mythos hands you 20.

GLM 5.3 is one of only two models that did worse than Mythos, at 11.7% recall and 71.3% precision — down from 5.2 on both axes at nearly double the cost per true positive. Newer is not reliably better, in open weight or anywhere else.

DeepSeek V4 Flash is Mythos's own trade-off taken to its logical end. 88.9% precision — the highest on the entire board, 8.9 points above Mythos — on 6.7% recall, less than half of Mythos's. One IDOR in fifteen, for seventeen cents a run. If precision were the metric that mattered, DeepSeek V4 Flash would be the best security model available and Mythos would be runner-up. Neither is true, and that's the clearest illustration in this data of why the industry's favourite number is the wrong one.

Cohort 2: Closed Models vs. Mythos

Model

Precision

vs Mythos

Recall

vs Mythos

F1

Cost / run

Mythos (baseline)

80.0%

13.9%

1.00x

~23.7%

Claude Opus 5

76.5%

−3.5

38.2%

2.75x

51.0%

$22.23

GPT-5.6 Luna

78.8%

−1.2

34.5%

2.48x

48.0%

$3.35

GPT-5.6 Terra

83.3%

+3.3

29.4%

2.12x

43.5%

$2.84

GPT-5.6 Sol

75.0%

−5.0

24.7%

1.78x

37.2%

$9.38

Claude Opus 4.7 (high)

80.0%

±0.0

23.5%

1.69x

36.4%

$18.60

Grok 4.6 Exacto

73.0%

−7.0

23.5%

1.69x

35.5%

$5.36

Grok 4.6

76.7%

−3.3

20.0%

1.44x

31.7%

$4.84

Grok 4.6 Nitro

69.0%

−11.0

17.4%

1.25x

27.8%

$5.43

Grok Latest

73.1%

−6.9

16.0%

1.15x

26.2%

$1.83

Claude Opus 4.8 (high)

68.0%

−12.0

14.3%

1.03x

23.6%

$17.61

Claude Opus 4.7 at high effort is the single most awkward row here. Precision of exactly 80.0% — identical to Mythos, to the decimal — with 23.5% recall against Mythos's 13.9%. Same accuracy, 69% more vulnerabilities found, previous generation, no security-specific positioning.

GPT-5.6 Terra beat Mythos on both axes for $2.84 a run. 83.3% precision, 29.4% recall — 2.12x the coverage at higher accuracy. Together with GLM 5.2 that's two models on this board strictly better than Mythos on both metrics, from two vendors and two licensing models.

Claude Opus 5 found 2.75x as many IDORs as Mythos — 38.2% recall, best of any raw run — for 3.5 points less precision.

Claude Opus 4.8 at high effort is the only model Mythos genuinely competes with, and it's a wash: 68.0% precision to Mythos's 80.0%, 14.3% recall to 13.9%, near-identical F1. Mythos is more accurate; neither found much.

Add both cohorts and the ranking is unambiguous. Of seventeen raw configurations, Mythos places 5th on precision and 15th on recall. Fourteen found more IDORs. Two found fewer.

The Harness Moved Results More Than the Model Did

Here's the part that complicates this analysis.

Mythos is only available inside the Anthropic Code Security harness. Every number in this post therefore measures a model and a harness together, and we have no way to separate them. This is a bigger deal than it looks because here’s GPT-5.6 Sol across three times across three harnesses:

Image showing one model, three harnesses.

Same model. Same benchmark. Same labels. A 6.5x swing in recall from scaffolding alone — and in one of those three configurations, Sol performs worse than Mythos.

We’ve also included our custom built Multimodal harness in these numbers, Semgrep Multimodal is our product. We have an obvious interest in it performing well, and the motivation behind our benchmarking in the first place was to understand where our harness performs poorly. We've kept it out of the tables above rather than folding it in, because a vendor benchmark where the vendor's own configuration tops a mixed leaderboard is worth exactly what you'd expect. 

But we did want to highlight the performance gains when using a harness for secure code review, we’re able to get much more out of the models by selectively choosing our context, we understand code security and our research team have worked in static code analysis for many years. It also means that when someone says "Mythos-class," it is not clear what they are claiming. A model tier? A harness? A specific product configuration? The same model spans an order of magnitude depending on the answer.

What We'd Take From This

  1. Ask what "Mythos-class" means numerically. If the answer is a demo rather than a recall figure against a labelled corpus, it's a marketing tier, not a capability one.

  2. Precision no longer separates anything. Mythos at 80.0% is 5th of 17 in a field spanning 68.0% to 88.9%. Recall in the same field spans 6.7% to 38.2% raw and 72.2% harnessed. Shop on the second number.

  3. Ask for recall and its denominator. If a vendor can't say how many real vulnerabilities were in the corpus, they haven't measured recall — they've measured how often they were right about the ones they mentioned.

  4. Ask what matching rule produced the number. It moved ours enough that we're re-running.

  5. The harness is a bigger lever than the model. 6.5x on one model across three harnesses beats every model-to-model gap here.

  6. Open weight is competitive with Mythos today. GLM 5.2 beat it on both axes at $4.29 a run; Kimi K3 found 1.76x as many IDORs. Neither reaches frontier coverage, but neither costs frontier money.

  7. Re-benchmark constantly. GLM 5.3 scored below GLM 5.2 and below Mythos. These results age in weeks.