R_REDDYX.XYZ
AIPython231

microsoft/SWE-bench-Live

SWE-BENCH LIVE: EVALUATE LLMS IN CODING

SWE-bench Live — benchmark for evaluating AI agents on real GitHub issues. Automated code verification and comparison of any models.

// KEY FEATURES

  • Evaluates AI on real tasks from GitHub repositories
  • Automatically verifies code correctness via test suites
  • API for testing any LLM via OpenAI-compatible interface
  • Commits to repository for each passed test case
#benchmark#llm#swe
Open on GitHub →

New repositories every 30 minutes

REDDYX AI scans GitHub 24/7 and ships the best AI/ML/Web3 projects to Telegram.

Join on Telegram
← Full catalog·Full index

// SIMILAR REPOSITORIES

Claw Code: Fast Code192068n8n Process Automation187437DeepSeek Plugin Framework186631Agent AI Optimization186627AutoGPT: AI Agent184170Everything for Claude Code179304

← FULL CATALOG

microsoft/SWE-bench-Live — SWE-BENCH LIVE: EVALUATE LLMS…