ExploitGym: measuring whether AI agents can turn CVEs into working exploits
The benchmark tests autonomous exploitation on complex real targets, moving AI cyber-risk discussion from hypotheticals to measured attack capability
1/ Can AI agents turn security vulnerabilities into real attacks? This is one of the most critical tasks for measuring the impact of frontier AI on cybersecurity. In ExploitGym, we find that autonomous exploitation is no longer hypothetic