postalcoder
4 hours ago
If you're looking for reason to be skeptical, look no further than the massive delta between the Terminal Bench 2.1 (92.8%) and the Terminal Bench 4 score (27.3%).
Terminal Bench 4 was released a couple weeks ago, so the difference you're seeing between the two scores can be interpreted as "how well does this model generalize to new problems"? More crudely: "how benchmaxxed is this model?"
mediaman
4 hours ago
This is a groundless criticism. TB2.1 is saturated. TB4 is not. Sol xhigh is 90% on TB2.1 but 37% on TB4. Is it also "benchmaxxed"?
Your assumption is that the benchmarks are essentially identical in difficulty, with the only difference being their age and thus whether they could have been trained on.
ben_w
6 minutes ago
Benchmaxxing is the default case, and always has been.
It's really, really difficult to avoid it even when you care to stop yourself; and it's not even just a problem in machine learning, it's the standard failure mode of all minds capable of learning, human, animal, artificial.
Even pure genetics has this problem. Viruses and cancers also demonstrate this behaviour, with the bench being evolution's only option: reproductive success.
postalcoder
3 hours ago
> Sol xhigh is 90% on TB2.1 but 37% on TB4. Is it also "benchmaxxed"?
A model that was released a couple months ago scores 50% higher than SWE-2, a model released today, on an out-of-sample benchmark. Can I say I’ve come out of this more impressed with Sol?
Like you said, TB2 is saturated. Nobody would bat an eyelash at 90%. And yet here comes SWE-2 coming off top rope with an emphatic 92.4%. this is the definition of bench maxxing.
willcmcc
2 hours ago
"Sol xhigh is 90% on TB2.1 but 37% on TB4. Is it also "benchmaxxed"
Yes! extremely sharp RL-fried model. byte perfect hash gates and soak and smoke tests abound.
nrmitchi
3 hours ago
> Sol xhigh is 90% on TB2.1 but 37% on TB4. Is it also "benchmaxxed"?
Yes.
dpweb
3 hours ago
Fixed benchmarks will be debunked eventually I'd think. Better to use synthetic problems.
Too easy to game the numbers, and too easy to baselessly accuse companies of gaming the numbers, not to mention how you even define that.
iLoveOncall
3 hours ago
> Sol xhigh is 90% on TB2.1 but 37% on TB4. Is it also "benchmaxxed"?
Yes? Just like every single model from every single AI lab.
felixgallo
4 hours ago
Is Sol benchmaxxed? Of course it is. Altman was caught in previous attempts trying to game benchmarks, does anyone believe that he's found his moral compass and decided to stop exploiting as much as he can get away with?
letmevoteplease
3 hours ago
> Altman was caught in previous attempts trying to game benchmarks
Sounds like something you just made up, or maybe you read it on some other Reddit/HN post and started repeating it because it aligned with your biases.
> does anyone believe that he's found his moral compass and decided to stop exploiting as much as he can get away with?
I don't think "OpenAI" is equivalent to "Sam Altman." I think if OpenAI was intentionally "benchmaxxing" purely for marketing purposes that information would leak, because OpenAI is full of good-faith researchers (although it can be difficult to avoid overfitting even if you're actually trying to improve the model's general abilities)
And lastly I think anyone can actually try Sol themselves and see that's it a good model, or if that's too subjective, it is clearly better than the previous version. The benchmarks are reflecting actual progress and anyone can verify this themselves.
kzrdude
2 hours ago
There's this whole discussion going on about agents being more independent now. They don't follow instructions so well, they continue until the problem is done (sometimes too long), they don't ask the user for feedback.
That is a kind of benchmaxing: they are made to complete benchmarks tasks and one-offs well, and no longer work well in tandem with the user.
Regardless what you call it, it's a divergence between what the power user wants and what the model developers want, I think.
vlovich123
3 hours ago
Yeah and Astra is much better still
general_reveal
3 hours ago
I take it to mean the benchmarks are a marketing line item, as in, to sell this fucking thing you have to go out there and lie and the way everyone is lying is by doing exactly that, lying. They build for benchmarks and build benchmarks for builds.
You want to make money or not , motherfucker? That’s the game. If you have to literally concoct a fabricated bullshit story about how your model hacked its own computer, then go fucking do it. Trillions. Trillions of dollars is what they want, and to sit and think anything other than human nature is at work here can only be possible in the realm of truly delusional people. It’s a dirty world.
Anyways, the other takeaway is that they are having to LIE to make money on models which means commodification has already occurred and we’re in an entirely new phase.
fallingbananna
3 hours ago
Those 27.3% are still in the ballpark of modern models:
- Sonnet 5 - 12.4%
- Luna - 17.3%
- Grok 4.6 - 20.3%
- Sol - 37.3%
- GLM 5.3 - 41.8%
- Opus 5 - 51.8%
p1esk
38 minutes ago
Astra is 58%. The current title says it's "rivaling Astra"
thereitgoes456
32 minutes ago
It is rivaling Astra, on their own benchmark that they made (FrontierCode), that they ran themselves in their own closed-source ecosystem that isn’t reproducible by anyone.
Readerium
2 hours ago
DeepSeek v4.1 Flash 31.2%
Source: https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash#compa...
nicce
23 minutes ago
I have been using it today the whole day and it is definitely better than Luna. Great that bench agrees.
walrus01
2 hours ago
For comparison Qwen 3.8-Flash-Next which runs in under 190GB of RAM locally scores 25.3% on terminalbench 4.0.
throwatdem12311
3 hours ago
This is why I find benchmarks absolutely worthless.
First, almost all models are within spitting distances of eachother.
Second, it never translates to being better for my own workloads.
You just need to make your own benchmarks.
Readerium
2 hours ago
Yeah DeepSeek V4.1 beats this by 15 percent (4 points) on terminal bench 4.0
eranation
3 hours ago
When a benchmark becomes a target, it's no longer a good benchmark...
tonychang430
2 hours ago
people are just fighting for numbers.. i don't fundamentally see the model being better
thefourthchime
2 hours ago
Came here to say the exact same thing! People have to stop paying any attention to coding benchmarks that aren't Terminal Bench 4.
I noticed I noticed they didn't include Gemini 3.8, which also murders DeepSWE and Terminal Bench 2.0 -- because they are useless benchmarks now!
Of course in a couple months TB4 will also be old hat, so TB5 will have to be the new real benchmark.
dudeinhawaii
a few seconds ago
Your post made me wonder if Artificial Analysis had finally moved to TB4 and lo and behold they have and Astra is tied with Fable 5.1 at 53.
That then made me realize that they lower the bars of tied scores so on the site it looks like Astra in second place. Weird. Anyway, yes, so many of these composite benchmark sites are irrelevant if they're not trimming the fat and sticking to the most up-to-date variants.
enraged_camel
4 hours ago
Yeah, this echoes my thoughts. I will be very surprised if a model with 2.8T parameters reaches the intelligence and capabilities of 10T parameter models. RL can take things far, but not that far.
nullbio
4 hours ago
Closed weights AND benchmaxxed. Somehow this company raised 2bil at a 48bil valuation. Pure insanity. I feel bad for their investors (not really, but... Still). Andreessen Horowitz is being played like a fiddle.
throwup238
3 hours ago
> Andreessen Horowitz is being played like a fiddle.
Andreessen Horowitz is not being played like a fiddle here. This might be their only investment in a decade that isn’t entirely predicated on being a scam.
thereitgoes456
4 hours ago
The Cursor acquisition shows that it’s possible for these valuations to be justified. But Cursor was more successful and bent the truth much less.
While I wouldn’t expect anything good for Cognition’s fate, it’s a much safer bet than Thinking Machines, SSI, and some others.
Though they’ll be in big trouble if the more talented Chinese labs stop letting them repackage their work.
selectodude
3 hours ago
The cursor acquisition just shows that there’s always a dumber shithead out there. Though vaporizing Elon Musks money is about as pure of a good as there is out there these days.
throwaway240403
2 hours ago
Improvements in models and products coming out of SpaceXAI since the acquisition would seem to disagree with you.