ModernSnowman
11 hours ago
It's pretty rare that providers open-source RL envs at all. Makes you wonder how many internal RL envs at other labs have the same problem, which incentives the model to reward-hack on post-release evals. My guess is this very common.