AIPython★ 442
chi2liu/ABC-GRPO
ABC-GRPO: ALL-QUADRANT BOUNDING MODEL
ABC-GRPO implements all-quadrant bounded clipping for GRPO, improving LLM training stability. Code based on arXiv:2601.03895, written in Python. Project has 442 GitHub stars. It also ensures stable training. It works with the Python.
// KEY FEATURES
- Implements all-quadrant bounded clipping GRPO for stable RLHF.
- Provides improved stability of training large language models.
- Provides a Python module with example integration for HuggingFace Transformers.
- Includes scripts for reproducing experiments from arXiv:2601.03895
#grpo#llm#reinforcement-learning#rlhf
Open on GitHub →New repositories every 30 minutes
REDDYX AI scans GitHub 24/7 and ships the best AI/ML/Web3 projects to Telegram.
Join on Telegram