AIPython★ 283
eth-sri/matharena
LLM EVALUATION THE LATEST MATH OLYMPIADS
matharena evaluates LLMs on fresh math olympiad problems. It loads the problems, runs models, via API, and, measures the solution accuracy in JSON.
// KEY FEATURES
- Automatically collects problems from latest IMO, Putnam, AMC via API or web scraping
- Runs the LLM (OpenAI, Claude, local HF) and yet saves the answers in the structured JSON
- Compares answers vs etalons, calculating accuracy partial correctness and execution time
- Allows changing the models and the tasks via CLI‑config | env variables
#evaluating-models#llm
Open on GitHub →New repositories every 30 minutes
REDDYX AI scans GitHub 24/7 and ships the best AI/ML/Web3 projects to Telegram.
Join on Telegram