Skip to content
#

swe-bench

Here are 139 public repositories matching this topic...

CORAL

🔥🔥COLM 2026🔥🔥 CORAL is a robust, lightweight infrastructure for multi-agent autonomous self-evolution, built for autoresearch. Works with Claude Code, Codex, Cursor, OpenCode, Kiro, and more.

  • Updated Jul 24, 2026
  • Python

Evaluate & benchmark AI coding agents and Claude Code skills — sandboxed, reproducible YAML eval suites for Claude Code, Codex & Gemini, with A/B experiments and CI gates.

  • Updated Jul 24, 2026
  • Python
why-was-fable-banned

Fable-style spec + evidence gate for Claude Code + Codex. Makes Opus/Codex work under Fable-like discipline: blocks every edit until a deterministic spec passes, and there is no "done" without live acceptance evidence. Spec-first, verification-gated, forbidden-paths enforced.

  • Updated Jun 30, 2026
  • Python

An 18 notebook course that isolates and measures each component of agentic loop engineering on real, industry standard software datasets.

  • Updated Jun 29, 2026
  • Jupyter Notebook

Improve this page

Add a description, image, and links to the swe-bench topic page so that developers can more easily learn about it.

Curate this topic

Add this topic to your repo

To associate your repository with the swe-bench topic, visit your repo's landing page and select "manage topics."

Learn more