Hackernews
new
show
ask
jobs
RL for LLM Reasoning Is Sparse Policy Selection, Not Capability Learning
2 points
posted 7 hours ago
by BlackGlory
(arxiv.org)
No comments yet