SYNTHESIS NOTE
Topics›Domain Specialization›this note

Can large language models reliably find software vulnerabilities?

Frontier LLMs claim to detect security flaws but may flag false positives at high rates and miss real vulnerabilities. Understanding whether they can be trusted for cybersecurity work matters as organizations consider deploying them.

Synthesis note · 2026-10-06 · sourced from Domain Specialization

The paper's verdict is that frontier LLMs are not yet reliable cybersecurity systems. The authors test six frontier models and two domain-specialized models in white-box function-level detection and in black-box testing of five web applications with 118 ground-truth vulnerabilities. In white-box detection, "every frontier model produces 10–50% false positive rates," systematically over-predicting vulnerabilities. In black-box testing, frontier models reach "only 4–8% ground-truth coverage," rising to "just 10–19% even with external security tools." The lever the authors name is methodology, not scale: penetration-testing procedures encoded in domain-specialized agents raise per-family detection "above 50%."

The mechanism is a split between reasoning and confirmation. The methodology-guided agents (P3) run a plan-invoke-observe loop with one agent per vulnerability family, and they replace single-signal heuristics with "multi-signal confirmation" such as response comparison and state-mutation checks. The Agentic Reasoning Graph (P4) goes further: LLM-driven reconnaissance and payload generation feed "deterministic classification that cannot hallucinate," so the model plans while programmatic evidence checks decide exploitability. The authors report that their Defense model has the highest precision (0.904) and lowest false positive rate (9.7%) among all models, on a single GPU. Their training-data argument locates the gap in what models learn from: security testing as a process, meaning end-to-end request and response sequences, failure-heavy data and multi-step attack chains.

This connects to three library notes. The Do cybersecurity benchmarks actually measure exploitation? note argues that strong scores on reproduction, patching and CTF benchmarks miss exploitation; the black-box numbers here show a related gap from the testing side. The per-domain methodology argument is a security instance of Should safety harnesses be customized for each deployment?: the domain fixes what must be established, and the method fixes how. It also contrasts with Should security controls scale with model capability?, which ties security to capability scale, whereas this paper places the lever in method.

The excerpt does not establish that methodology beats scale in general. Its comparison pits frontier models against a specialized model plus a methodology layer; the tables it cites (Tables 2 to 4) are not included, and no model-size sweep appears. GPT-5.4 "could not reliably complete the agentic workflow," and attempts to extend the evaluation to Mythos and Claude Opus 4.8 were denied, so the frontier figures cover only the models that ran. The benchmark is the authors' own, with applications they "will opensource," and its stated scope excludes malware analysis, social engineering, phishing, network intrusion, hardware security, cloud posture and cryptographic design. The 100+ zero-day discoveries and the Defense model's results are the authors' own reports. At the strength the evidence allows, frontier models are not ready for unsupervised triage on these tasks, and domain-encoded methodology is a lever this benchmark supports, not a general rule.

Inquiring lines that read this note 3

This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.

What limits language model accuracy in evaluating ideas? Can AI systems evade safety evaluations through reasoning manipulation? How should we measure frontier AI models' cyber exploitation capabilities?

Related concepts in this collection 3

This note in its neighbourhood — explore the map, then jump to a related concept in the list below.

Concept map
12 direct connections · 103 in 2-hop network ·medium cluster Open in graph ↗

Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph

your link semantically near linked from elsewhere

Related papers in this collection 8

Papers most semantically related to this note, ranked by cosine similarity in the embedding space.

Original note title

frontier LLMs over-predict vulnerabilities at 10 to 50 percent false positives — methodology, not scale, is the lever