R_REDDYX.XYZ
AIPython283

eth-sri/matharena

LLM EVALUATION THE LATEST MATH OLYMPIADS

matharena evaluates LLMs on fresh math olympiad problems. It loads the problems, runs models, via API, and, measures the solution accuracy in JSON.

// KEY FEATURES

  • Automatically collects problems from latest IMO, Putnam, AMC via API or web scraping
  • Runs the LLM (OpenAI, Claude, local HF) and yet saves the answers in the structured JSON
  • Compares answers vs etalons, calculating accuracy partial correctness and execution time
  • Allows changing the models and the tasks via CLI‑config | env variables
#evaluating-models#llm
Open on GitHub →

New repositories every 30 minutes

REDDYX AI scans GitHub 24/7 and ships the best AI/ML/Web3 projects to Telegram.

Join on Telegram
← Full catalog·Full index

// SIMILAR REPOSITORIES

Claw Code: Fast Code192068n8n Process Automation187437DeepSeek Plugin Framework186631Agent AI Optimization186627AutoGPT: AI Agent184170Everything for Claude Code179304

← FULL CATALOG

eth-sri/matharena — LLM EVALUATION THE LATEST MATH…