Hugging Face Blog·· 29 d ago
Fine-tuning a 350M Model for Better Structured Outputs in 100 GRPO Steps
Fine-tuning a 350M Model for Better Structured Outputs in 100 GRPO Steps
AI summary
A public low-cost approach fine-tunes LFM2.5-350M with TRL GRPO using around 500 samples and 100 steps, raising IFStruct from 22.6% to 29.7%. Training fits free Colab or Kaggle GPUs, with local llama.cpp evaluation on a MacBook and code open on GitHub.
Selection record
Threshold 60Official, first-handFirst 47Second 34
Not admittedSum of both 81 < twice the threshold 120
- Source tier
- Official, first-hand; this tier's threshold is 60
- Pre-filter
- passed:微调350M模型提升结构化输出,AI技术实操
A model scores each item twice, independently, against one written standard, out of 100. An item is admitted only when the two scores add up to twice the threshold. The threshold is set per source tier.
Source: Hugging Face Blog · huggingface.co