jug
10 hours ago
We also have Nerf Bench:
https://www.bridgebench.ai/nerf-bench
They test it on launch day, then benchmark it against that. A deviation of above 10% is considered a change. They're currently tracking Opus 5.5 and GPT-6 Astra.
This bench famously detected a degradation of Opus 4.6 which Anthropic later blogged about. I personally think people sense nerfs more often than they happen and that it's often about honeymoon effects.
nsarrazin
an hour ago
Used to work on a chat app where we had full control of the stack from the GPUs to the chat interface and everything in between. We were a small team too so I could be pretty confident that nothing changed in the stack and we still regularly had users complain that this or that model got nerfed. Perceived performance is actual performance over expectations and the latter just keeps increasing over time.
It doesn’t mean the big labs don’t also nerf models! But if they didn’t you’d still have users complaining.
waterproof
24 minutes ago
I find that I learn to "trust" a model to get certain things right, as I would trust a colleague. So, as `expectation` increases, my prompting and context management gets sloppier.
`percieved_performance = actual_perf/expectation`
`expectation` is an increasing function over time.
`actual_perf` is a stochastic function of the model's true ability, context, etc. -> a recipe for some bad sessions.
As for multiple bad sessions in a row, this is a studied phenomenon in gambling where players perceive "runs" because our brains love to find patterns.
rplnt
3 hours ago
> I personally think people sense nerfs more often than they happen and that it's often about honeymoon effects.
I believe in temporary nerfs. Operators reducing quality significantly to increase throughput for whatever reason (high demand?). Same session, model being completly incapable, and it being back to normal the next day. Experienced that with Anthropic models way too many times. Never on weekends, usually during US work days.
It's been a few months since I last recall this though, must have been pre-opus-5. And I know there are benchmarks for this as well, hence "believe".
smurf9852
6 minutes ago
Friend of mine works for a corp that is one of the top spenders on Claude models. He complained about these nerfs during peak demand. Their Anthropic contact changed something and it did not happen since.
transcriptase
3 hours ago
I vividly remember when ChatGPT3.5 went fully mainstream, there were times where within minutes you would realize they were only serving up idiot mode and there was no point trying to do much until demand died down and they swapped back to the non-quantized version.
People called it lazy mode, in that instead of writing the script you asked for it would basically tell you to learn to code then check out xyz topics to tackle the problem.
setopt
2 hours ago
Regarding lazy mode, I recall ChatGPT sometimes almost refusing to do a web search despite me asking explicitly for it, instead replying with speculation about what the search results likely would tell us. If I pretend to be angry that it didn’t search the web it would however do it. Haven’t noticed this in a while either.
katzenq
an hour ago
It should have stayed that way. Be a good search engine and encourage the human to do the work themselves.
comboy
9 hours ago
But you are using API not the CLI right? I did not ever observe API degradation, only subscription stuff through their CLI.
scrollop
4 hours ago
There's also this one which has been around for a while
rednb
3 hours ago
I'd take this kind of benchmark with a grain of salt. At this point, I have a set of comprehensive guidelines covering both backend and frontend work, and for the frontend we go as far as explaining what we a good design is in our visual system, and even how to conduct a visual review when screenshots are handed to the model.
Deepseek 4.1 ranks very low in this benchmark but it has proven so capable that after being simultaneously on Max x20 and Pro x20 subscriptions, i've transitioned to using DS 4.1 as a daily driver and am very satisfied.
My point is, i think their overall ranking makes sense, matches my experience with out of the box capabilities for vague and underspecified tasks. But seeing a model rank low in their ranking does not mean that the model is incapable. Having skills and guidelines has a lot of influence on what you get out of a model.
dotancohen
2 hours ago
You use DS through Open Router? Which harness?
I'd love to hear more, I'm considering jumping ship. I'm running a Debian desktop if that's a concern.
KronisLV
2 hours ago
We could also use something that tracks concrete token amounts each tier gives you, in case they ever mess with it - and also maybe even the tokens needed to accomplish a particular benchmark, to see how much you can actually get done.
user3939382
9 hours ago
Anthropic A/Bs my weekly quota amount. So I have an automated prompt that runs at 3 AM with a transcription task, I measure input and output tokens, and weekly/5 hour quota before and after. The absolute token counts stay within 0.1% while in mode A it counts for 1% of my 5 hour quota and mode B 4% of my 5 hour quota.
jacquesm
7 hours ago
How did pissing off your customers ever become a business model?
I can't imagine sticking with a supplier that plays games like that with me. Tokens are a pretty vague quantity to begin with (you don't control how many tokens a model puts out in response) and giving a couple of purposefully wrong responses will happily inflate your bill, but you don't care because eventually it worked. It's almost an ideal vehicle to scam people.
Imagine the power company being able to decide how much you consume and at which price point.
none_to_remain
6 hours ago
I find it amazingly rich that they bill you for """thinking""" tokens and now you don't even get to see them, they're gonna train the thing to sing "99 Bottles of Beer on the Wall" to itself before it starts work.
Turskarama
5 hours ago
They don't want to waste tokens on purpose, what they're actually hiding is when the model wastes tokens on obviously stupid "thoughts".
adastra22
4 hours ago
No they are hiding the chain of thought to make distillation harder.
slim
3 hours ago
It's fascinating that you all think accounting is real and it did not come to your mind that they could make up numbers when billing
TeMPOraL
3 hours ago
They could, but as you see here, people are very eager to create dashboards and trackers that do external accounting by proxy, so they can't just "make up numbers" without the customers noticing and making a fuss.
CodesInChaos
4 hours ago
Another way Antropic misleads its customers is the description of the max plans. They are advertised as having 5x/20x the 5h quota as Pro. But the description says nothing about how the weekly quota scales, leaving customers to infer it scales the same way. But from what I've heard, the weekly quota is only 3.5x/7x that of Pro.
herval
7 hours ago
> How did pissing off your customers ever become a business model?
Airlines, banks, health insurance…
tccole
6 hours ago
So very low margin businesses with hogh amounts of regulations.
petesergeant
3 hours ago
Banks and health insurance are much more consumer friendly outside of the US, usually because of regulation. Turns out you can just tell banks “make transfers cheap and essentially instant” and they’ll do it, rather the bullshit they have in the US.
pixelready
7 hours ago
Step 1: Oligopoly Step 2: Regulatory Capture Step 3: Profit
icepush
2 hours ago
You can put stuff like "make sure your reply is between 800 and 900 tokens" at the end of your prompt and the vast majority of the time it will do so.
cavoirom
2 hours ago
Their fate is coming. Until the open-source models will be usable in machine with 256GB memory, they are done. Their behavior is unacceptable (Anthropic) recently but it won't last long.
csomar
3 hours ago
I think it's sinister, but not for the reasons you're thinking. I think they're just wildly unprofitable on subscriptions. The idea that most customers won't use their full quota is plain wrong: most people are maxing out their subs, or even reselling whatever quota they have left.
When you're running something at a loss, you can mistreat your customers and they'll still stick around (I'm an example). OpenAI and Anthropic are now cheaper than Chinese models on subscriptions, while being 6-10x more expensive on the API.
My guess is they need the user numbers for the IPO and are willing to take a temporary loss in the meantime. By the time they go public, they'll either drop the subscription model or it'll turn into what the Chinese providers already offer: basically just a cap on how much API you can consume. Same same.
It's not clear what API tokens actually cost them, but I looked into running a local model, and it's way outside the budget of an individual or even a small or medium business (hundreds of thousands of dollars). So my guess is that running these models economically isn't possible, even if they're delivering real business value (coding, research, etc.). In other words, at API prices I'd just stop using AI, and I suspect most other developers would too.
TeMPOraL
2 hours ago
> It's not clear what API tokens actually cost them, but I looked into running a local model, and it's way outside the budget of an individual or even a small or medium business (hundreds of thousands of dollars). So my guess is that running these models economically isn't possible
Datacenters have massive economies of scale. Everything from cheaper electricity to having specialized, more efficient hardware to simply being able to run it continuously at near-100% utilization, all adds up.
Many things in the economy - most notably, manufacturing of most consumer goods - only makes economic sense once you're producing for/serving millions of people. This is not unusual.
> In other words, at API prices I'd just stop using AI, and I suspect most other developers would too.
Many say that, but I sincerely doubt they'd actually follow through. People might get more conservative about how they spend their tokens, but AI today is just too good at eliminating drudgery and boring / bullshit parts of daily work to give up on merely 3-5x price increase.
csomar
2 hours ago
> Datacenters have massive economies of scale.
Sure. Issue is, no one is providing on how much it actually costs to burn these tokens. And as we don't know, we can only speculate.
> Many say that, but I sincerely doubt they'd actually follow through.
I have a $100 open ai sub and I track my token usage. Last month I spent roughly $2.600 in equivalent API usage. There is no way am paying that. I let my $100 sub lapse if next month I'll be using it less.
Look, I am not saying that there isn't a potential value out there. But the cost has to be bounded. If your opportunity is $1.000 and AI costs $2.000 to execute it, then you don't have a business model here.
topspin
6 hours ago
It's worked for online PvP gaming for a long time. Nerf stuff the min-maxers "earned" through game mechanics and sell over-powered "premium" things to everyone else to pwn them. Then nerf the old premium stuff and make new premium stuff. Forever.
I don't know if that's the actual origin of the term nerf, but it was the first time I'd heard it.
done_lurking
5 hours ago
I think the origin of the word "nerf" as a verb came from the Nerf brand of toy guns. The idea being that "Nerfing" something is to turn it into a harmless version of itself.
apitman
8 hours ago
Do Anthropic quotas give you precise remaining token counts or something? I have something similar set up for tracking my ChatGPT usage but it only gives percentages remaining, which is a pretty coarse metric.
ffsm8
7 hours ago
Claude code supposedly has otel you can set via env. I haven't set it up, so I'm just repeating hearsay.. but it supposedly has everything relevant in it wrt token usage and cost
It's meant for their test env I think, so is not documented to my knowledge
TeMPOraL
2 hours ago
It's for corporate users who want to track how the product is used internally, and it was documented at least some time ago, quite extensively even.
ffsm8
2 hours ago
youre right!
https://code.claude.com/docs/en/monitoring-usage#usage-monit...
thanks for correcting me on that regard
adastra22
4 hours ago
Has otel? What is that?
reubenmorais
4 hours ago
OpenTelemetry
Aeolun
5 hours ago
Tokens used / percentage change is a pretty obvious metric. They give you both, but they don’t do the math for you.
CodesInChaos
4 hours ago
Could be load dependent, not an A/B test.
Is the fraction of the 5h quote consumed consistent with the fraction of the weekly quota consumed?
I heard there is a usage tracking tool you can install that tells you if tokens are more or less expensive at the current time.
braingravy
8 hours ago
Pretty amazing to see enshitification happen live with a product still in development… Truly web 4.0
quikoa
2 hours ago
Wouldn't it be trivial to detect benchmarking if the same requests are running on a fixed interval?
zerop
2 hours ago
What could be the reason to nerf?
StableAlkyne
2 hours ago
It's cheaper to run a quantization of a model, but its quality is reduced.
For example, if your weights were trained as 32-bit floats and you need 1TB of RAM, you could reduce that to around 256GB by quantizing to 8-bit floats. You also make the model faster in the process because there is less data to process to calculate the next token.
The game is to balance between the savings of quantization and making the model dumb enough the people notice
sscaryterry
2 hours ago
Someone always knows about where the bodies are buried. These shenanigans always end up surfacing eventually.
zsoltkacsandi
3 hours ago
Or nerfed version of models rolled out gradually.
andriy_koval
9 hours ago
Usage bench is also very useful! Thank you for doing this!
hamandcheese
6 hours ago
...is it? I'm looking and it seems like it doesn't have any data. It might be useful if they keep it up.
andriy_koval
6 hours ago
Yeah, I guess they started it today..
avazhi
7 hours ago
Nerfbench isn't helpful if it's 3 days old.
nullbio
4 hours ago
Do they use private benchmarks? Because if not, it could be selectively nerfed.
I also wonder if cache could be used to throw these off as well, where it's serving un-nerfed cache results for context windows that are identical to ones they've previously had for benchmark requests.
Seems like the only way to do it well would be to have some randomness involved that couldn't be cheated on - but you'd want to do it in a way that doesn't throw out the benchmarks too much, so your results can be compared still.
Gabrys1
3 hours ago
Thankfully, we can now use AI to design a test that tests AI. And the AI company can use AI to detect the test and cheat. And we can then use AI to implement anti-cheat.
All that energy wasted... could just drive a big V8 instead and make less money for the big tech
Grimblewald
10 hours ago
I dunno, I never sense nerfs for local models, but consistently a few months after launch for corpo hosted models, seems odd my internal model for the capacity of a model drifts for anthropic models but not local ones. I've been using LLMs heavily even before ada/babbage/davinci days, and trust my internal calibration over baseless handwavey explanations for why im imagining things, especially when I have data that shows capacity regression on frontier models for tasks, e.g. one shot success at loss, 0 success in 15 attempts once nerf is sensed. Others publish their quantified capability regressions which are also more trust worthy than this kind of handwaving.
eulgro
9 hours ago
Your comment makes no sense. How and why would a local model be nerfed anyway...?
r_lee
9 hours ago
he's saying that he notices a difference between local (not nerfable) and hosted ones, so that it's not as likely to be just placebo
martin-
3 hours ago
But if it is placebo, obviously he wouldn't notice any placebo change for local models, since he KNOWS he is using an immutable local model. That comparison only works if he doesn't know what model he is using.
jacquesm
7 hours ago
That and 'loss' may have been intended to be 'launch'.
Razengan
9 hours ago
Theory (Conjecture? Hypothesis?): What we notice as "model nerfing" is the company diverting compute to training/running new unreleased models..
Remember that some people get access to the next flagship version long before us peasants do. I recall seeing the mention of "Astra" more than a month before it was officially announced
Centigonal
9 hours ago
wouldn't less compute result in slower inference, rather than worse performance?
latentsea
9 hours ago
They could potentially quantize the model and run it at lower quality taking less VRAM.
btown
9 hours ago
The more likely thing that would happen is that the provider begins silently interpreting (perhaps some) high effort-level requests as medium, etc., or having a classifier do this far more subtly. As such, the load on the cluster is less, and more resources can be devoted to training. Whether the frontier labs actually do this is purely conjecture at this point.
JohnBooty
7 hours ago
I assume there's classification going on where a really basic "Hi how are you?" style request sent to a high-effort instance can be routed to a lower-level instance. This... is pretty much fine with me, assuming they do a good job of it.
I would also assume they use nebulous labels like "Medium Effort" or "High Effort" map to quantitative amounts of compute allocation... and that these amounts can be varied manually or automatically. Right?
I mean, there's a reason why they call it "High Effort" and not "Exactly 5 Minutes of GPU Time on Exactly 10 GPUs." They want to be able to move those sliders and tweak those knobs.
nightpool
8 hours ago
why is that more likely?
zxilly
8 hours ago
Because they already did so. The model in Codex will get lower `juice` than API version.
poizan42
7 hours ago
My guess is that they are dynamically changing the quality of the model to always keep the speed above some floor. So once it gets below that they switch to a worse quant or reduce reasoning level, or some combination of both.
jackmott42
7 hours ago
There is no nerfing, look at the data before coming up with a theory as to why the nerfing that isn't even happening is happening.
fuck
Rapzid
9 hours ago
The vibe bro science is this always happens on every release, every Tuesday, and twice on Sunday.
Of course it's almost entirely unsubstantiated BS.
fbrncci
9 hours ago
Well now it’s being substantiated!
scrollop
4 hours ago
https://marginlab.ai/trackers/claude-code/
this one has been around for over a year
Rapzid
9 hours ago
Or rather it's being.. Unsubstantiated. The Nerf conspiracy isn't that there have been a few harness and platform bugs leading to performance regressions, but that OpenAI/Anthropic have maliciously and unethically degraded their model performance post release to shed load and save money.
somenameforme
8 hours ago
Yeah it's just inconceivable that companies whose entire business model started by engaging in wholesale for-profit theft and abuse of intellectual property would ever be so unethical as to try to lower their costs, especially just prior to an IPO.
Rapzid
8 hours ago
Again, these things are constantly measured. They sell HEAPS through their API access to enterprise consumers that expect a model to not be nerfed after it's released. And you bet many of those enterprises, some spending many millions each month, are measuring this shit.
So this is a case of extraordinary claims requiring extraordinary evidence.
And even though it's super straight forward to collect the evidence, there seems to be zero substantiating the vibe bro conspiracy theories.
mrandish
7 hours ago
> They sell HEAPS through their API access
The claim is that they nerf subscription accounts not API.
somenameforme
7 hours ago
If you're going to try to argue that companies doing things, completely legal mind you, to increase their profit margins is a conspiracy theory then you're not debating in good faith. Let alone when we're speaking of a subset of companies that were fundamentally built on wholesale unethical behavior carried out for profit. Let alone when we're speaking of companies who are all racing to IPO where short term results matter more than just about anything.
Another issue is also that the risk here is probably literally zero. Any evidence in support of such could easily be dismissed, with completely plausible deniability, as a short-lived technical glitch as opposed to intentional behavior.
Rapzid
7 hours ago
> an explanation for an event or situation that claims a secret, powerful group is responsible for a hidden plot, rejecting the standard or official account
I'm sorry, but yeah. The official account is a harness regression and some platform bugs.
Where is the evidence they are underhandedly and unethically regressing their models to shed load and reduce costs? This is the conspiracy theory running rampant through the vibe boroughs; that they are bait-and-switching on model capabilities then "nerfing" them to save money and shed load. Where is the evidence?!
I'm not saying it's illegal, per say, so don't come at me with that straw man bull cock. This bro science conspiracy has been circulating for at least 2 years(I don't even know) and enterprises would certainly be pissed off if they were paying premium API prices for advertised and previously tested model capabilities that are suddenly under performing due to "nerfing" shenanigans.
So where is the evidence?!
airstrike
6 hours ago
You're giving way too much credit to "enterprises" both noticing and publicly airing out their dissatisfaction
Not too mention these companies could easily offer one product to enteprises and another to everyone else
Model nerfing is real
Rapzid
6 hours ago
Uh huh. OMG you're so right, it's soooo real ;) ;) ;)
I was so certain it wasn't, based on the complete lack of evidence.
But then you said it's real. NVM, I don't need evidence! Somebody said it's real!
This place has fallen off.
troupo
2 hours ago
> And even though it's super straight forward to collect the evidence, there seems to be zero substantiating the vibe bro conspiracy theories.
Until shit like this: https://www.anthropic.com/engineering/april-23-postmortem
Where people pointed out issues early and en masse, and Anthropic denied it was happening, gaslighted anyone claiming this was an issue, then begrudgingly admitted it was an issue, and then spent another two weeks "fixing it".
Or shit like this: https://www.anthropic.com/engineering/a-postmortem-of-three-...
Anthropic is in a perpetual state of "oops, these 'bugs' degraded our model quality" and only admit the issues when it's immediately obvious and visibly affects a large number of customers.
Otherwise all open benchmarks can be (and are) gamed. And it's quite hard to judge the output of a non-determenistic black box that Anthropic (or OpenAI) constantly tweak.
stackghost
8 hours ago
To me that sounds exactly on-brand for Big Tech in general and Sam Altman in particular.
jackmott42
7 hours ago
No one here has suggested that the conspiracy theory is stupid (But I will, it is stupid), we are pointing out that the people actually measuring model performance have not found the nerfing before launch conspiracy to be true. in short, yall dumb, shut up.
stackghost
7 hours ago
I think the real story is just how easily people believe that purported conspiracy theory. It speaks to how little trust there is in these AI companies, and in Big Tech in general, that this "conspiracy" theory is perfectly plausible to lots of people
> in short, yall dumb, shut up.
no u
Rapzid
7 hours ago
It's speaks to how low the bar has dropped.
Everyone wants to be a software engineer, until it's time to do software engineering shit.
You know, like scientific method shit we learned in 5th/6th grade.
It's the great bro science incursion.
stackghost
6 hours ago
>Everyone wants to be a software engineer, until it's time to do software engineering shit.
The absolute state of software in 2026 should tell you that almost nobody does “software engineering shit” and never has.
bdlowery
5 hours ago
> This bench famously detected a degradation of Opus 4.6 which Anthropic later blogged about.
this bench was just released, it couldn't have detected opus 4.6 degradation.