The debate on open-weight model efficacy in cybersecurity is settled by hard data: DeepSeek V4 Pro 0813 and GLM-5.3 have reached frontier tier, surpassing expensive closed models in vulnerability discovery capabilities. This finding comes from a benchmark by Aikido, which tested 10 models on 32 off-the-shelf vulnerabilities recently disclosed, using a total of 11.7 billion tokens to ensure dataset freshness and reduce memorization risks.
Repetition compensates for inconsistency
The decisive factor was not just the model, but the execution strategy. Models showed significant variance in results across runs: DeepSeek V4 Pro found 17 vulnerabilities on the first pass, but 28 out of 32 when pooling results from three distinct runs. This approach, known as pooling, leverages exploration diversity to boost overall recall. Conversely, Grok 4.6 demonstrated the highest consistency, finding 21 CVEs in all three runs, but with a lower total recall (26 out of 32). Opus 5 and Sol settled at 26 and 25 vulnerabilities respectively, confirming that the premium cost of closed models no longer guarantees a net advantage in coverage.
The cost of cheap coverage
The benchmark highlights a clear trade-off: open-source models like DeepSeek and Qwen offer performance comparable to frontier models at a fraction of the cost (approximately $295 for three DeepSeek Pro runs versus much higher costs for closed models), but generate a larger volume of false positives. This requires a more robust triage pipeline to filter noise, an operational cost that organizations must balance against license savings. GLM-5.3, tested as an early evaluation partner by Z.ai, stood out for its ability to maintain high recall without losing consistency, consolidating its role as a credible alternative to proprietary models.
Implications for cyber defense
The confirmation that open-source models can match or exceed closed ones in vulnerability discovery shifts market dynamics. Organizations can now opt for self-hosted or low-cost models without sacrificing code analysis capabilities, provided they are willing to manage a higher volume of output for validation. This scenario fits into a broader trend, already observed with the rise of DeepSeek and the escape from generative AI costs, where economic efficiency is becoming a strategic factor as important as computational power.

AI-generated comment
AI-generated comment
AI-generated comment
AI-generated comment