npx skills add ...
npx skills add orchestra-research/ai-research-skills --skill fine-tuning-with-trl
Fine-tune LLMs using reinforcement learning with TRL - SFT for instruction tuning, DPO for preference alignment, PPO/GRPO for reward optimization, and reward model training. Use when need RLHF, align model with preferences, or train from human feedback. Works with HuggingFace Transformers.
npx skills add orchestra-research/ai-research-skills --skill fine-tuning-with-trl
TRL provides post-training methods for aligning language models with human preferences.
Installation:
Supervised Fine-Tuning (instruction tuning):
DPO (align with preferences):
Complete pipeline from base model to human-aligned model.
Copy this checklist:
Step 1: Supervised fine-tuning
Train base model on instruction-following data:
Step 2: Train reward model
Train model to predict human preferences:
Step 3: PPO reinforcement learning
Optimize policy using reward model:
Step 4: Evaluate
Align model with preferences without reward model.
Copy this checklist:
Step 1: Prepare preference dataset
Dataset format:
Load dataset:
Step 2: Configure DPO
Step 3: Train with DPOTrainer
CLI alternative:
Train with reinforcement learning using minimal memory.
Copy this checklist:
Step 1: Define reward function
Or use a reward model:
Step 2: Configure GRPO
Step 3: Train with GRPOTrainer
CLI:
Use TRL when:
Method selection:
Use alternatives instead:
Issue: OOM during DPO training
Reduce batch size and sequence length:
Or use gradient checkpointing:
Issue: Poor alignment quality
Tune beta parameter:
Issue: Reward model not learning
Check loss type and learning rate:
Ensure preference dataset has clear winners:
Issue: PPO training unstable
Adjust KL coefficient:
SFT training guide: See references/sft-training.md for dataset formats, chat templates, packing strategies, and multi-GPU training.
DPO variants: See references/dpo-variants.md for IPO, cDPO, RPO, and other DPO loss functions with recommended hyperparameters.
Reward modeling: See references/reward-modeling.md for outcome vs process rewards, Bradley-Terry loss, and reward model evaluation.
Online RL methods: See references/online-rl.md for PPO, GRPO, RLOO, and OnlineDPO with detailed configurations.
accelerateMemory optimization: