aliljet
9 hours ago
It's hard to not see this as a gut punch for OpenAI. They're lead was largely captured by scoring on value (by way of reset after reset) and now they're getting eaten up on price and being bestes and equalled on performance. I'll still pay a premium for Opus 5.5 right now because it's nearly unlimited use, but Google is the quiet sleeping king Everyone is happy to watch everyone else, but I'd wager google burns more tokens through their search product than basically anyone else and now they're just quietly pacing the frontier...
nl
6 hours ago
This is completely the wrong read!
Sol 6.1 scores one point less than Gemini 4 on intelligence AND costs less than half ($0.72 vs $1.99) per task.
Additionally, if you are using OpenAI you have the option to pay a bit more and get Astra which - despite the benchmarks - does outperform Sol on some things.
Also, people are - rightly - very wary of Google's benchmaxxing tendencies. I think lots of people remember Gemini 3.0 (I think?) which benchmarked amazingly, but as soon as you used it would go off-track and needed constant babysitting if you wanted to use it for agentic work.
scrollop
an hour ago
Here are sites with ongoing measurements checking if a model has been nerfed - check opus 5.5-
https://www.bridgebench.ai/nerf-bench
https://marginlab.ai/trackers/claude-code/ https://marginlab.ai/trackers/codex/
tomrod
9 hours ago
They own their hardware. That vertical integration alone probably saves oodles because they can reconfigure to their needs as opposed to individually negotiating data centre plans. I can't imagine the complexity both OpenAI and Anthropic have to maintain for their deployments.
mkotlikov
7 hours ago
It's still not enough though, they're buying compute from SpaceX's colossus data centers.
genxy
6 hours ago
That is google trying to keep their 10% worth something.
largbae
5 hours ago
And it totally worked too.
tomrod
7 hours ago
They may be buying it, but is there any indication they are using it for consumer-facing stuff? I would suspect given the scale there are many different things. I wish I knew more about standardization for cloud data center workloads though
notatoad
7 hours ago
i don't know if you could say they're buying it "happily". more like begrudgingly.
the deal has a 30-day cancellation policy, and they raised a bunch of debt around the same time to fund their own datacenter expansion.
fragmede
3 hours ago
and they shut down Stadia so they'd have some extra GPUs to use for it as well!
sroussey
7 hours ago
GPT-6.1-sol costs less than half of Gemini and way less than Anthropic on the cost per task chart of the listed parent page.
aleqs
9 hours ago
> Opus 5.5 right now because it's nearly unlimited use
Anthropic has some of the lowest usage per $ in general, not sure what you're taking about.
jjice
9 hours ago
I don't think the OP and your comments are mutually exclusive. If Anthropic usage is lower, and the OP considers it basically unlimited for their case, that just means that they would have virtually unlimited usage with other plans.
aleqs
8 hours ago
Nothing about it is unlimited or even close to it. Just because you only use 1gb of your 2gb data plan, doesn't mean you have unlimited data.
8n4vidtmkvmk
2 hours ago
What if you use 100GB of your 10TB plan? Is it virtually unlimited then?
It's all relative. Some people just can't use up their quotas with their normal usage.
aleqs
2 hours ago
That just means your use is limited not that the plan is unlimited.
nl
6 hours ago
This is no longer the case.
Opus and Sol usage levels vs the API are currently roughly the same, but Opus 5.5 outperforms at low and medium effort levels.
UltraSane
8 hours ago
Opus 5.5. on medium effort is very good a writing code and provides a lot of tokens per 5 hour session. It is a fantastic value for $20/month.
zozbot234
9 hours ago
It's not as smart as Claude Opus 5.5 High according to the AA benchmark. Looks like a big fat nothingburger so far, though it's possible that future fine-tuned checkpoints of the same pretrained model will do a lot better.
mpyne
8 hours ago
> It's not as smart as Claude Opus 5.5 High according to the AA benchmark.
If it's smart enough to do the job then it won't matter that Opus is smarter. At the right price and performance, at least.
qrify_app
3 hours ago
[dead]