Show HN: Minimal LLM Post-Training Experiments on an 8GB GPU (SFT, DPO, GRPO)

14 pointsposted 8 hours ago
by popopanda

No comments yet