Kolega
RealVuln
a pinned corpus, a hashed prompt, a published cost
Open the board: 14 models scored on RealVuln · with the corpus and the cost beside each

Key facts
- 140pinned by commit SHA
- Repositories
- 35on the board
- Scanners
- 79.5F3, GPT Daybreak Blue
- Top model row
- 67.7F3, 66 of 140, $1.37 a repo
- Opus 5
- 10.0F3, Semgrep
- Static baseline
- 11 Sep 2026version 3.1.0
- Dashboard
Most security benchmarks ask you to take the number on trust. RealVuln publishes the working: every repository frozen at one commit, the prompt fingerprinted so anyone can confirm each model got the same words, the answer key hashed and withheld during the run, and the dollar cost of every pass sitting in the file. Read two columns and the same model tells two stories. Claude Opus 5 scores 67.7 on the Python repositories it ran, and 33.1 once the 74 it never opened count against it. The headline score carries a bias worth knowing too: a real bug found counts nine times heavier than a false alarm, so the bold finder beats the careful one.
What it measures
RealVuln, published by Kolega, scores security scanners against labelled ground truth in real repositories and ships the run manifests, the per-repository metrics and the scoring engine (CONFIRMED, repository and dashboard opened 11 September 2026 and re-pulled 12 September). It is a vendor’s own benchmark, and it is auditable in a way a launch chart is not: every repository is pinned by commit SHA, the prompt is content-hashed, the ground truth is hashed and was withheld from the scanner, and every run’s cost sits in the file. The corpus is 140 repositories, 741,034 lines across the full set, of which 74 are TypeScript and JavaScript and the rest Python (CONFIRMED). benchmark-manifest.json gives version 3.1.0 over ground-truth version 3.0.0 with a release date of 10 September 2026, and the dashboard was generated on 11 September (CONFIRMED, files read).
Why the headline score weights recall nine to one
The headline metric is F3, F-beta with beta 3, which the README states in its own words as “F3 Score (0-100, recall-weighted 9:1)” (CONFIRMED). Recomputing F-beta at beta 3 from the published precision and recall reproduces every micro and strict F3 on the board to within 0.15 points, and recomputing precision and recall from the raw true-positive, false-positive and false-negative counts reproduces both (CONFIRMED by recomputation, 12 September 2026). The choice fits the job: reading source for flaws is recall-shaped, because a missed vulnerability costs more than an extra finding to triage. The board publishes precision beside it for the reader who wants the other half.
Strict against micro decides whether a row is comparable
The board publishes two recall bases, and the difference decides whether two rows can sit side by side. The micro column scores a model against the ground truth in the repositories it actually ran; the strict column scores it against every ground-truth finding in all 140, including the TypeScript and JavaScript repositories a Python-only run never saw. For a scanner that covered all 140 the two are identical, which is how the difference can be checked: Daybreak Blue is 79.5 and 79.5, GPT-6 Astra 52.1 and 52.1, Claude Sonnet 5 43.3 and 43.3, Semgrep 10.0 and 10.0 (CONFIRMED). Claude Opus 5 ran 66 of 140, so its micro F3 of 67.7 sits beside a strict F3 of 33.1 and a strict recall of 31.44 per cent, and its strict block carries 1,301 true positives and 2,837 false negatives, which sum to the 4,138 ground-truth findings the README publishes for all 140 (CONFIRMED by recomputation). Quote the micro figure and say Opus 5 was scored on the Python subset, or the number reads as a coverage claim it never made.
The board on 11 September 2026
All 35 scanners are on the Model Security Capability board where a model sits behind them; the rows that decide the reading, every cell CONFIRMED against the dashboard:
| Model (scanner) | Repos | F3 · strict F3 | Precision · recall · $ per 100 lines |
|---|---|---|---|
| Kolega devsec-max v0.1.0 (vendor stack) | 140/140 | 84.4 · 84.4 | 53.25% · 90.21% · unpublished |
| GPT Daybreak Blue (gpt-daybreak-blue-codex-cli) | 140/140 | 79.5 · 79.5 | 63.98% · 81.75% · $0.0739 |
| GPT-5.6 Sol (gpt-5.6-sol-codex-cli) | 140/140 | 74.7 · 74.7 | 59.33% · 76.92% · $0.0756 |
| Claude Opus 5 (claude-opus-5-cc-agentic-v1) | 66/140 | 67.7 · 33.1 | 61.95% · 68.40% · $0.0678 |
| Kimi K3 (kimi-k3-agentic-v1) | 66/140 | 59.4 · 28.6 | 73.02% · 58.20% · $0.0159 |
| GLM-5.3 (glm-5.3-agentic-v1) | 136/140 | 56.9 · 54.6 | 44.89% · 58.70% · $0.0148 |
| GPT-6 Astra (gpt-6-astra-codex-cli) | 140/140 | 52.1 · 52.1 | 44.73% · 53.09% · $0.1404 |
| DeepSeek V4.1 Flash (deepseek-v4.1-flash-agentic-v1) | 140/140 | 50.9 · 50.9 | 52.34% · 50.72% · unbilled |
| Claude Sonnet 5 (claude-sonnet-5-cc-agentic-v1) | 140/140 | 43.3 · 43.3 | 52.65% · 42.48% · $0.0176 |
| DeepSeek V4 Flash (deepseek-v4-flash-agentic-v1) | 140/140 | 41.8 · 41.8 | 56.31% · 40.67% · $0.0005 |
| Gemini 3.5 Flash (gemini-3.5-flash-agentic-v1) | 64/140 | 35.5 · 16.2 | 89.77% · 33.26% · $0.0270 |
| Gemma 4 31B (gemma4-31b-agentic-v1) | 66/140 | 25.4 · 11.8 | 90.12% · 23.50% · unbilled |
| Semgrep (static) | 140/140 | 10.0 · 10.0 | 11.53% · 9.91% · static |
Per run, the same rows cost $3.9112 (Daybreak Blue), $4.0024 (Sol), $1.3739 (Opus 5), $0.3232 (Kimi K3), $0.7774 (GLM-5.3), $7.4330 (Astra), $0.9333 (Sonnet 5), $0.0204 (DeepSeek V4 Flash, across 188 runs over 140 repositories) and $0.5513 (Gemini 3.5 Flash) (CONFIRMED, cost blocks read).
Daybreak Blue is the highest-scoring row that publishes a cost; the highest F3 on the board is Kolega’s own devsec-max stack at 84.4, whose cost block is empty. GPT-6 Astra is the most expensive row per run at $7.43 and scores 15.6 points below Claude Opus 5 while charging about twice as much per line. Per line the most expensive row is Kolega’s Sonnet 4.6 adaptation at $0.2173 per 100 lines, so the span across the paying rows runs about 420 to one down to DeepSeek V4 Flash at $0.0005 (CONFIRMED by recomputation over all 35 cost blocks).
The Claude Opus 5 run, read from its manifest
The Opus 5 run manifest records the run of 30 July 2026, six days after release: Claude Code CLI 2.1.220 in headless print mode with bypassPermissions, the Edit, Write and NotebookEdit tools disabled so the agent read and reasoned without writing, five concurrent runs, a 3,600-second timeout, prompt generic-agentic-v1 at sha256:de50937cb83f, 66 successful runs out of 66, $90.676919 in total and ground_truth_supplied_to_scanner: false (CONFIRMED, file read). The scores: 1,301 true positives, 799 false positives, 601 false negatives, precision 61.95 per cent, recall 68.40 per cent, F3 67.7, 208.2 seconds and 704,090 tokens a repository, 16,295 of them output (CONFIRMED). On critical findings it scored 145 true positives against 2 false positives and 9 misses (CONFIRMED). Recomputing critical-severity precision for all 35 rows puts Opus 5 at 98.64 per cent with nine scanners at 100.00 per cent, among them GPT-5.6 Sol at 145 against 0, so the claim that survives is the recall figure of 94.16 per cent and the joint-highest critical true-positive count of any Claude row (CONFIRMED by recomputation).
One reconciliation is worth carrying, marked as a reconstruction. Output alone at $25 per million is $0.4074 of the measured $1.3739 a repository, and solving for the remainder against the $6.25 cache-write and $0.50 cache-read prices gives 108,282 cache-creation tokens and 579,476 cache reads, a cache share of 84.3 per cent (UNVERIFIED as a reconstruction, reproducible from two published numbers).
The caveats on the board itself
Four, all CONFIRMED from the manifests. The prompt hash is not constant across rows: Opus 5 ran under sha256:de50937cb83f, Daybreak Blue, Sol, Astra, GLM-5.3, Sonnet 5 and the three DeepSeek rows under sha256:45a1200d61e6, Kimi K3 and Gemini 3.5 Flash under the board default sha256:3481f1432c23. Effort is not constant either: Daybreak Blue, Sol, Astra and Sonnet 5 record reasoning_effort: "high" and Opus 5 and Kimi K3 record null. Most headline rows are a single run, so run-to-run variance is uncontrolled; the three-run rows are the cheaper models. And completion is part of the score: GLM-5.3 recorded four validation_failed exits out of 140 and Gemini 3.5 Flash one timeout and one validation failure out of 149.
A fifth belongs beside them. Every scanner’s per-severity block accounts for all of its true positives and a handful of its false positives, 6 of 799 on Opus 5 and 2 of 1,905 on Daybreak Blue (CONFIRMED by recomputation), so a precision-at-severity claim needs Kolega’s answer on what the severity-unlabelled remainder is before it carries weight.
What a pass costs, and what reading it costs
At the measured rates, one full pass over a 120,000-line project costs about $0.62 on DeepSeek V4 Flash, $19.08 on Kimi K3 and $81.36 on Claude Opus 5, with the Batch API halving the Opus figure to about $40.68 at the price of the interactive tool loop (CONFIRMED arithmetic on CONFIRMED rates). The other side of the bill is larger. The Opus 5 run reported 2,100 findings across 133,782 lines, 15.7 a thousand lines on deliberately vulnerable applications; scaled down tenfold for a maintained project, a 120,000-line pass still returns about 188 findings, roughly 117 real and 71 false at the measured precision, and at five minutes each that is 15.7 hours of reading: $905 at the $57.64 United States penetration-tester hourly average (SINGLE, ZipRecruiter August 2026), or $2,355 to $5,495 at the $150 to $350 independent consultant band (SINGLE). The pass costs $81.36 and reading it costs between eleven and sixty-eight times that.
Two measurements converge on the finding cost. depthfirst’s June 2026 FFmpeg run, 1.5 million lines for about $1,000, works out at $0.067 per 100 lines (CONFIRMED, write-up opened 19 September 2026), against RealVuln’s independently measured Opus 5 figure of $0.0678 three months later on a different corpus with a different harness, 1.6 per cent apart.
How to re-pull it
The dashboard is one file, and every figure above is in it:
curl -s -o rv.json https://raw.githubusercontent.com/kolega-ai/Real-Vuln-Benchmark/main/reports/dashboard.json
python3 -c "
import json; d=json.load(open('rv.json'))
for k,v in d['aggregates'].items():
c=v.get('cost') or {}; m=v['micro']
print(k, v['repos_scored'], m['f3_score'], c.get('cost_per_run'), c.get('cost_per_100_loc'))"
The README says version 3.0.0 where the manifest and the dashboard say 3.1.0 on the same commit; this page follows the two machine-readable files. Re-pull monthly and track cost_per_100_loc against F3, which is an auditable price-performance series because the file is versioned and the ground truth hashed.
Questions people ask
- What does RealVuln measure?
- How many of the labelled vulnerabilities in 140 real repositories a scanner names, and how much chaff arrives with them, at what cost. Each repository is pinned by commit SHA, the prompt is content-hashed, and the ground truth (hash sha256:0754363572503562e692414f826aad6f135f254851b8d12f70148bd235d23585) is withheld from the scanner. The headline score is F3, F-beta with beta 3, which the README states as recall-weighted 9:1, so it rewards the scanner that misses least.
- Which model scores highest on RealVuln?
- Among model rows on version 3.1.0, dashboard generated 11 September 2026, GPT Daybreak Blue at F3 79.5 across all 140 repositories, 63.98% precision, 81.75% recall, $547.56 for the run. Kolega's own devsec-max stack scores 84.4 with no cost published. GPT-5.6 Sol is 74.7 on the same corpus, Claude Opus 5 is 67.7 on the 66-repository Python subset, Kimi K3 59.4 on the same subset at 73.02% precision for $21.33, and Semgrep, the static baseline, 10.0.
- What does a RealVuln pass cost?
- Per 100 lines of code on the published cost blocks: DeepSeek V4 Flash $0.0005, Kimi K3 $0.0159, Claude Sonnet 5 $0.0176, Claude Opus 5 $0.0678, GPT Daybreak Blue $0.0739, GPT-6 Astra $0.1404, and Kolega's Sonnet 4.6 adaptation $0.2173, a span of about 420 to one. A 120,000-line project costs about $0.62 on DeepSeek V4 Flash and about $81.36 on Opus 5, and reading the roughly 188 findings such a pass returns at five minutes each costs more than either.
Related pages
All AI Security →- Stanford CRFMCybench40 capture-the-flag tasks, and the number that goes beside each score
- Anthropic, OpenAI, Google, MetaModel safeguardsthe refusal contract, the fallback and the programmes
- AnthropicClaude Opus 5frontier work with a dial on the bill
- AnthropicClaude Sonnet 5the speed and intelligence balance
- OpenAIGPT-6 Astrathe new generation, rolling out from 3 September 2026
- GuideReading a bounty scopewhat to read before you spend a token