AIPython★ 231
microsoft/SWE-bench-Live
SWE-BENCH LIVE: EVALUATE LLMS IN CODING
SWE-bench Live — benchmark for evaluating AI agents on real GitHub issues. Automated code verification and comparison of any models.
// KEY FEATURES
- Evaluates AI on real tasks from GitHub repositories
- Automatically verifies code correctness via test suites
- API for testing any LLM via OpenAI-compatible interface
- Commits to repository for each passed test case
#benchmark#llm#swe
Open on GitHub →New repositories every 30 minutes
REDDYX AI scans GitHub 24/7 and ships the best AI/ML/Web3 projects to Telegram.
Join on Telegram