simpo-training
Simple Preference Optimization for LLM alignment. Reference-free alternative to DPO with better performance (+6.4 points on AlpacaEval 2.0). No reference model needed, more…
- Industry
- software-engineering
- License
- Unverified
- Source repo
- NousResearch/hermes-agent · ★ 214,858
- Source file
- optional-skills/mlops/simpo/SKILL.md
不会安装?看中文图文教程 →