AI & ML ecosystem

Benchmark

SWE-bench

SWE-bench Verified is the human-validated subset.

Resolve real GitHub issues end-to-end; the agentic-coding standard.

Category
Agentic
Creator
Princeton NLP (Jimenez et al.)

Links