As standard AI evaluations lose their edge, a new analysis by the RAND Corporation reveals that performance gains in frontier models are now concentrating on highly complex biological tasks—precisely those theorized to pose the greatest real-world biosecurity risks.
The saturation of traditional benchmarks
The study indicates that many existing biology benchmarks have become saturated, meaning they no longer provide informative data because models are consistently hitting ceiling scores. This makes it difficult to distinguish between a model that is simply recalling training data and one capable of genuine scientific reasoning. RAND advocates for the use of Item Response Theory to identify a "discriminating frontier" of tasks where performance gaps actually reveal the evolution of dangerous capabilities.
Balancing utility and biosecurity risks
The industry is currently struggling to balance accessibility with the prevention of dual-use. Anthropic has refined its safety classifiers for Fable 5 Biology AI, cutting false positives by 85%. This update aims to enable broader use for health and education while maintaining strict controls on biosecurity. However, this occurs against a backdrop of alarm: researchers warn that the ability to compose viral genomes using generative AI now exists, while the governance to steer it safely remains absent.
Shifting toward real-world workflow evaluation
To move beyond simple data recall, new evaluation frameworks are emerging. Insilico Medicine has introduced a benchmark-as-a-service that tests models on end-to-end drug discovery programs, simulating the sequential uncertainty of actual research. This shift is critical because frontier AI is no longer just solving isolated problems; it is mastering broader workflows that could be repurposed for biological threats, a trend already highlighted by the creation of functional synthetic viruses.

No comments yet. Be the first!