Hackernews
new
show
ask
jobs
MUD as AI Evaluation and LLM-judge distortion in ways aggregate κ misses
4 points
posted 11 hours ago
by joozio
(lesswrong.com)
No comments yet