7 points | by miyukimatthews 10 hours ago
3 comments
It's pretty rare that providers open-source RL envs at all. Makes you wonder how many internal RL envs at other labs have the same problem, which incentives the model to reward-hack on post-release evals. My guess is this very common.
sounds pretty similar to Mythos and the linux kernel "CVEs"
a number were applying previous fixes to new locations while others were duplicates of recently applied fixes
As told from the linux kernel team's side last week
https://www.youtube.com/watch?v=NnV_cWeoo5Q
[dead]
It's pretty rare that providers open-source RL envs at all. Makes you wonder how many internal RL envs at other labs have the same problem, which incentives the model to reward-hack on post-release evals. My guess is this very common.
sounds pretty similar to Mythos and the linux kernel "CVEs"
a number were applying previous fixes to new locations while others were duplicates of recently applied fixes
As told from the linux kernel team's side last week
https://www.youtube.com/watch?v=NnV_cWeoo5Q
[dead]