88.
How do we eval if a model can find and exploit vulnerabilities? We discuss some benchmarks and the common pattern:
How do we eval if a model can find and exploit vulnerabilities? We discuss some benchmarks and the common pattern: • A sandboxed target within Docker containers • Inputs: code only (0-day), with patch (1-day scenario) • Tools such as bash,