KaoruAoiShiho
4 hours ago
Appears to be benchmaxxing
root-parent
a minute ago
And being worst than previous model...
"...The traces tell the why: (1) On our most classic Witness-style game, Opus 5 states the hidden rules before its first action, then plays a byte-identical optimal solution in 5/5 seeds at temperature 1.0. Zero exploration. It already knows this genre. (2) But on our most novel game (unusual mechanic combinations you can't pattern-match), Opus 5 regresses below Opus 4.8. Where rules must actually be discovered through interaction, the new model is worse than the old one..."
jchw
an hour ago
I have been claiming that I don't think Chinese AI companies are benchmaxxing harder than American AI companies, which has gotten mixed reception: sometimes people agree, sometimes they disagree.
It seems I was wrong. American AI companies might actually be benchmaxxing harder.
throwa356262
4 hours ago
"Opus 5 states the hidden rules before its first action, then plays a byte-identical optimal solution in 5/5 seeds at temperature 1.0. Zero exploration."
I guess there is no way this can happen without benchmark being part of the training data??
pierrefermat1
3 hours ago
What seems to implied is that some of his hold out testing suite includes simple/common tests that are out in the wild, and for those opus went straight to a memorised solution .
zamadatix
2 hours ago
Simple/common tests is not an explanation for why only now Opus 5 is the only model encoding the answers like this. Something like the holdout test suite being leaked or Anthropic cheating (e.g. 'accidentally' including previous hold out run data in Opus 5 training) makes a much stronger fit.
stogot
3 hours ago
It may read information about the benchmark, such as on blog post, or Twitter feeds (example OP) without active cheat
jnwatson
3 hours ago
I was just thinking they need to mark each model per benchmark as "model released before the benchmark was released" and "model released after the benchmark was released".
nozzlegear
2 hours ago
andrepd
3 hours ago
I'm shocked, astonished even, that enterprises on which trillions of dollars are being poured would consider cheating on marketing benchmarks.
asdfologist
3 hours ago
Meh, I doubt it was intentional. Deliberate benchmaxxing is incredibly damaging to credibility once it's discovered (see what happened to Meta with LlaMa 4).
It's more likely that the training data was contaminated with the benchmark data.
mupuff1234
2 hours ago
You really think they saw the jump in arc-agi-3 (which they reported in their official card), and didn't even bother to check?
They maybe have not intentionally benchmaxxed, but they certainly know that's what happened .
johnfn
7 minutes ago
How do you propose they check for something like this? They can't exactly ctrl-f the model weights for "Arc-AGI".