astrobiased
35 minutes ago
Is this in any way similar to Goodfire's work? https://www.goodfire.ai/research/rlfr#
HenryNdubuaku
29 minutes ago
Thats an interesting outlook, loosely similar.
35 minutes ago
Is this in any way similar to Goodfire's work? https://www.goodfire.ai/research/rlfr#
29 minutes ago
Thats an interesting outlook, loosely similar.
5 hours ago
> So we did mechanistic studies on small models, Gemma 4 particularly, and found the hidden state for different layers carry meaningful self-awareness signal for various situations.
Neat! Just to make sure I understand - you trained your probe layer to take this hidden state and predict p(wrong)?
Curious to learn more. Any more info on your approach (esp the mechanistic study)?
3 hours ago
Correct, the study is verbose, we will compile into a neat shareable report and publish once we solve the pending caveats. Interesting username btw haha.
2 hours ago
Nice, looking forward to the report.
And thanks, huge pasta fan :)
an hour ago
anytime!