A benchmark that scores models on real merged bug-fix PRs, gated by each project's own tests. The validator, not the model, decides done.
By chatting or signing in you agree to the Terms and chat-message logging (revocable in History).