AIPython★ 215
jsdhwfmax/EvalForge
EVALFORGE — LLM EVALUATION IN CI
EvalForge is a tool for LLM evaluation in CI/CD. Provides evaluator-neutral evidence, baseline regression gates, and JSON, JUnit, SARIF reports. Integrates with GitHub Actions.
// KEY FEATURES
- Evaluator-neutral evidence for LLM evaluation
- Baseline regression gates in CI pipeline
- JSON, JUnit, SARIF reports for CI
- Native GitHub Actions integration
#ai-evaluation#github-actions#junit#llm-evaluation#mlops#python#quality-gates#rag#sarif#testing
Open on GitHub →New repositories every 30 minutes
REDDYX AI scans GitHub 24/7 and ships the best AI/ML/Web3 projects to Telegram.
Join on Telegram