lubujackson
4 hours ago
I think the difference between the Anthropic token maximization approach (vibe code all the things!) and OpenAI's focus on efficiency, terseness and token reduction are going to be the defining features of who wins the long-term race.
My money is on the more efficient solution. Even if Anthropic can win some benchmarks by using 3x tokens over 3x time, it is a terrible base to build toward the future. Users are no longer willing to wait exponentially long for linear improvements. And as we see from some of the Chinese models, they can quickly distill frontier models with the tax of being slower and more token-guzzling, while retaining most of the quality. The real differentiators are becoming speed and efficiency, which translate to cost and user velocity more than incremental capability improvements.
It's great that frontier models can solve complex math equations, but bread and butter LLM usage (where the money is made) has already shifted from "I need the best always" to "what solves my day-to-day problems quickly and consistently". Fable usage as a percentage is flat-lining. We are already at the point where output quality is negligible. What wins going forward is cost, speed, consistency and the compounding effects of "softer" improvements to the harness.
WhiteDawn
3 hours ago
Not to sound like a mark, I try to not get attached to any of these providers.
I've jumped between copilot, claude, gemini and chatgpt since the start of the year. chatgpt wasn't even worth looking at early this year.
Anthropic has the smarter models for sure, and seems to be default in corporate. However, the amount of budget you get with GPT as a user is much better, the harness feels more polished, and the models are faster. They are also much nicer to work with, I can just read the output for the most part. With claude I get pages of text and need to skim to find where the actual information i need to care about lies. So much more cognitive overhead.
Sol is smart enough for anything I've thrown at it, it's not one-shotting like fable, but I'm more willing to actually go back and forth with it, and it's likely producing better output to keep a human in the loop rather than trying to solve the world independently and making multiple incorrect assumptions.
I think GPT sees the market changing and is correctly repositioning themselves. Anthropic is down the wrong road, and if they don't correct course quickly I'm sure many of those enterprise contracts will start pivoting.
stanmancan
3 hours ago
> and it's likely producing better output to keep a human in the loop rather than trying to solve the world independently and making multiple incorrect assumptions.
I was just discussing this with a co-worker yesterday. I would really like a model (or harness?) that worked with me instead of for me. Walk me through its choices and decisions, let me correct it and guide it along. I would be way more confident in it's output, I would be more familiar with the changes that are being made, and it would make reviewing the final code way easier since I was making the decisions along side it. I'm sure it would also reduce the "brainrot" we're all going to experience the more we hand work to these models.
halfmatthalfcat
2 hours ago
Those exist and they're called orchestration skill frameworks like Superpowers, GetShitDone, etc.
ray_v
an hour ago
does "Superpowers" exist outside of the Claude ecosystem? I feel like I've become hooked on this workflow - honestly I could take or leave the models ... it's the workflows and the way that they essentially create tightly focused loops over multiple sessions that I've becoming fairly dependent on.
halfmatthalfcat
an hour ago
Yes, Superpowers is just a set of (agent agnostic) skills that you can use in any harness.
throwaway894345
3 hours ago
I just need a Sol translator to sit in front of Fable and Opus :/
esseph
3 hours ago
> I think GPT sees the market changing and is correctly repositioning themselves
I'm not sure how they'll survive their creditors tbh
rebolek
4 hours ago
I really would like to see OpenAI’s focus on efficiency but everytime I use Codex, it wastes tokens like there’s no tomorrow, hitting week limit in a day, where I’m able to use Claude just fine. Maybe it’s based on the codebase, I don’t know, but I have better results with Claude than Codex.
vorticalbox
4 hours ago
It might be your prompting.
My boss uses opus and gets good results when I use it always burns tokens. The other way is also true too I get great results with sol/terra but my boss does not.
rebolek
2 hours ago
Maybe. But if Claude is fine with my prompts while Codex uses them as an excuse to burn tokens. I have nothing against the model, they are more or less on pair, it’s just that Claude is giving me more value for same money.
vorticalbox
an hour ago
If Claude is working for you then stick with it! Happy you’re getting good value from Claude.
llama052
4 hours ago
At least with OpenAI you can use a third party harness like Pi.
bicepjai
4 hours ago
I had the opposite experience. Claude models and its harness feel like they are set to eat tokens for everything, especially if it’s ultracode effort. I have seen it spawn 6 agents and eat my 4-hour quota right in front of me. Codex is on point and follows instructions well even with max effort. I like my analogy of Claude being garrulous and Codex being laconic. If I see more limit hits in same sitting session, I have to rearrange my workflows. my experiments with qwen and deepseek have been good, cant wait to try glm and other models.
throwup238
4 hours ago
Just before I canceled my 20x Max sub I had a day where I had ~15% of my weekly I was trying to burn. I set it on ultracode and because I had set the max agents 32 for another project and forgot, the five research agents ended up spinning up a total of 26 subagents and burned through the remainder of my weekly in the span of 20 minutes before I noticed and shut it down.
“No, not like THAT!”
wccrawford
2 hours ago
Claude did the same thing to me the other day. I ran out of my 5 hr limit, and it went into my $100 credit they had given me. Multiple parallel agents. It ate it in no time flat.
When I checked, all that credit was gone, I still wasn't into the next 5 hours, and all the agents had failed, returning nothing.
I didn't even get anything for burning all that credit. If I had paid for it, I'd be very, very pissed.
vanviegen
3 hours ago
How can you burn through your weekly budget in 20 minutes? You'd hit your 5 hour budget way before that, right?
throwup238
3 hours ago
The five hour budget is about 15% of the weekly on Max 20x. I had about 15% left.
NyxWulf
3 hours ago
Here is an excerpt from the system prompt for UltraCode (Same for Fable,Opus,Sonnet):
"Ultracode. When a system-reminder confirms ultracode is on, that opt-in is standing: author and run a workflow for every substantive task by default. The goal is the most exhaustive, correct answer you can produce — token cost is not a constraint."
StilesCrisis
3 hours ago
Requesting ultracode is basically asking for maximum token usage. If you want to limit consumption, use medium or high.
jimmydoe
4 hours ago
same boat.
gpt 5.5/5.6 goes further on its own much more often than opus 4.8/5 does. codex capped ~300k context when claude does 1m.
I don't feel codex is saving tokens, and result is usually not as good imo.
gatio
3 hours ago
> codex capped ~300k context when claude does 1m.
That's configurable in codex.. but there is a higher cost/usage to using it.
apsurd
4 hours ago
You've articulated what I found unsettling about Boris' pov that "coding is a solved problem". I listened to a few of his talks and was instinctively off-put by that sentiment. I figured a fellow programmer would understand and speak on the nuances.
Granted it did make me think about my biases and to lean into more future facing inevitabilities. But you've nailed it, for Boris and Anthropic, they are betting that coding is a solved problem in the sense that any person can one-shot any random idea and the output will be in some abstract sense "good". And then at what cost and toward what end?
cloverich
4 hours ago
On the one hand it feels true. On the other I ask - what good software have anthropic, or anyone else, produced that was fully vibed?
As someone who does near 100% of my coding via LLM these days, i still find that for anything complex i am still looking at and thinking in terms of code. Im still quality checking and steering at some interval via code. And im still not sure how or whether i can replicate that level of thought without still dealing in code at times.
polotics
3 hours ago
It's a scale thing. with one fairly simple Android app, it's doable. anything Enterprisey is a nightmare. maybe that's the lesson and the problem is not with the llms but with accidental complexity. Just right now I feel like I'm one mythical LLM minute away from the next clean pull request...
mattm
3 hours ago
In this last week I've started reviewing code more and in just a short time, I've found 3 fairly simple things that were introduced by Claude that were not wrong per se, however, they were very inefficient and didn't address the root of the issue. I still think we're in a place where the output looks good as long as you don't look under the hood or keep it scoped to small, vibe-coded projects. Once you get beyond that it can fall apart. Anthropic must have a large codebase by now though. Yet I haven't seen much released from them about how they actually work day to day on development.
polotics
3 hours ago
What makes you think Antropic has a large code base? what do they do exactly that would make them need a large code base? Or maybe the word Large can be understood in different way?
smartmic
4 hours ago
> who wins the long-term race
Isn’t the race between Chinese open-weight models and the others more decisive for the future?
polotics
3 hours ago
I think the race is actually between locally optimized rigs with extremely strict context management, and the rest.
nozzlegear
3 hours ago
Those rigs are running one of the open Chinese models though, right?
colechristensen
4 hours ago
it'll all be centralized models justifying their capability against local inference in, maybe 3 years or less
once we have a bit more memory fab capacity and the insane bottomless investment in AI giants realizes there is a bottom, local hardware will catch up with model performance to the extent that centralized inference will be downgraded to special cases or for orgs that find it cheaper than buying expensive hardware
but really most power users are going to have 1 TB unified RAM and local models that will do well enough
for light office use you can still have a cheap laptop and a claude subscription
thefourthchime
2 hours ago
If the problem isn't that hard, I've been using Cursor's Auto or Grok 4.5 (not 4.6, it's too slow).
They're both pretty damn competent and more important, Extremely Fast! I find the speed more useful than trying to be a hundred percent complete on every task. The big intelligent models screw up all the time as well, but I have to wait twenty minutes to three hours to find out.
GPT 5.6 is especially tenacious and seems to want to solve every bug in edge case 1000% all the time. Sometimes that's what you need, but a lot of times you're just trying to move fast and figure out what the product is.
Ameo
3 hours ago
Seems like it wouldn't be hard for Anthropic to tweak a few prompts or RL pipelines to tune for terseness and token-efficiency if that stays as something that consumers want.
I find it unlikely that there's some fundamental property of OpenAI's models' "personality" or style which Anthropic (or any other serious AI firm) wouldn't be able to match if they wanted to.
ph4rsikal
4 hours ago
Claude used to be 10x more efficient. I think there are no barriers of entry between one coding agent and the other. Therefore, I expect them to reverse as soon as people switch. At least this is what I will do.
wilg
4 hours ago
Oh no, I use the top shelf models every day and I really think they have a lot of room to improve in pretty much every regard. I suspect Fable usage is flatlining because it's not that good comparitively and way too expensive.
alightsoul
4 hours ago
There is no evidence regarding distillation. It is impossible to distill a model in just a month which was the gap between fable and Kimi k3. Anthropic wouldn't even keep up with the load. It is just another example of American exceptionalism.
Barbing
3 hours ago
“In one notable technique, their prompts asked Claude to imagine and articulate the internal reasoning behind a completed response and write it out step by step—effectively generating chain-of-thought training data at scale. We also observed tasks in which Claude was used to generate censorship-safe alternatives to politically sensitive queries like questions about dissidents, party leaders, or authoritarianism, likely in order to train DeepSeek’s own models to steer conversations away from censored topics. By examining request metadata, we were able to trace these accounts to specific researchers at the lab.”
— “You are an expert data analyst combining statistical rigor with deep domain knowledge. Your goal is to deliver data-driven insights — not summaries or visualizations — grounded in real data and supported by complete and transparent reasoning.” (variations appearing 10s of 1000s of times)
https://www.anthropic.com/news/detecting-and-preventing-dist...Does Dario have the same relationship with the truth as Sam? (Their companies pirate books and develop products that compete at some level with those books’ authors, so obviously neither are that wonderfully trustworthy, so maybe “no evidence” meant you don’t believe this evidence rather than you weren’t aware of it. I would understand and respect your lack of belief!)
maxdo
3 hours ago
When they do that they pay billions . Unlike China. That’s the order of law vs bunch of companies in a totalitarian government
alightsoul
3 hours ago
They are not paying billions anymore. They just switched to buying and then destroying books. Arguably that's worse than training off of Anna's archive which Chinese companies still do. This comment is another example of American exceptionalism.
maxdo
3 hours ago
They paid $1.5 billions in July . Keep ignoring facts and believe in some nonsense
alightsoul
3 hours ago
That was a one time payment. Another example of America exceptionalism assuming its a periodic payment like paying rent. look how they give money to publishers who then give back crumbs to the authors themselves.
jimbob45
3 hours ago
Can you really hold them responsible for book burning when they’re being compelled to do so by governmental laws? Seems like you should be pointing your finger at lawmakers, if anywhere.
alightsoul
3 hours ago
Me? A lowly commoner without hundreds of millions to lobby them? Why aren't American ai companies doing so? Too busy training models? They're shooting themselves in the foot when they could just use Anna's archive and lobbying is cheaper than paying billions in copyright settlements. Oh wait I am wrong. Shredding and scanning books is much cheaper than lobbying to use Anna's archive. Still can't do anything because I don't have money for lobbying
alightsoul
3 hours ago
That's what any harness could do. By itself it's not evidence of distillation. Just mentioning Deepseek in there is just fear mongering, again American exceptionalism, which creates the belief only the US has the right to own cutting edge technology. It's a disservice to the entire world, and china happens to be the only challenger to that belief. They surely also asked for tasks to avoid mentioning Trump's involvement in the Epstein files and CP, but don't mention those, due to American exceptionalism. Dario has stated that he is an American exceptionalist. He powers Israeli/US bombings in Gaza, but only cares when China does the same.
This is evidence only if you're an American exceptionalist.
throwaway894345
3 hours ago
> My money is on the more efficient solution. Even if Anthropic can win some benchmarks by using 3x tokens over 3x time, it is a terrible base to build toward the future. Users are no longer willing to wait exponentially long for linear improvements.
As much as I wish you were right, everything about software economics for the last 30+ years has favored _less efficient software_. Traditional hardware has been optimized for traditional software for decades and we still see bloated software win consistently. LLM hardware is at the start of its cycle, with abundant low-hanging fruit to conquer--I would expect the pro-bloat dynamics to weigh even more heavily in the LLM space than in the traditional software space.