RL for LLM Reasoning Is Sparse Policy Selection, Not Capability Learning

2 pointsposted 7 hours ago
by BlackGlory

No comments yet