They test it on launch day, then benchmark it against that. A deviation of above 10% is considered a change. They're currently tracking Opus 5.5 and GPT-6 Astra.
This bench famously detected a degradation of Opus 4.6 which Anthropic later blogged about. I personally think people sense nerfs more often than they happen and that it's often about honeymoon effects.
nsarrazin 15 hours ago [-]
Used to work on a chat app where we had full control of the stack from the GPUs to the chat interface and everything in between. We were a small team too so I could be pretty confident that nothing changed in the stack and we still regularly had users complain that this or that model got nerfed. Perceived performance is actual performance over expectations and the latter just keeps increasing over time.
It doesn’t mean the big labs don’t also nerf models! But if they didn’t you’d still have users complaining.
booty 8 hours ago [-]
For years I used to buy ASICS running shoes. Every year they released a new model of each shoe: "Nimbus 23", then "Nimbus 24" the next year, etc. And every year people would complain in the user reviews about how each shoe was worse than the last.
I was like, wow, I guess the shoes must be literal torture devices full of MRSA-covered broken glass at this point. They've been getting continuously worse for 24 consecutive years!
Of course, what was really happening is that they were not getting worse, but naturally every year there was some small percentage of vocal dissatisfied users, while the silent majority simply enjoyed their shoes and didn't have much to say about them.
(The sorta-opposite happens in sneaker reviews as well. People will gush about how cushy the sole in some particular new sneaker is. Well, yeah, of course it's cushy -- you're comparing a new sneaker to your old sneaker where the foam had lost its bounce...)
MikeTheGreat 7 hours ago [-]
I wonder how much of that is caused by folks actually noticing actual degradation in product quality over time (whether or not this exact product is suffering from it).
Shrinkflation is a thing, which people suddenly started noticing in the past 5 years.
There's also "the Schlitz Mistake", which I've heard summarized as "most customers won't notice if you take your product's quality from A to B (or C), but they definitely will if you take it from A to J (or A to M)"
At this point I kinda assume that any company releasing year updates to a physical product that _doesn't_ take the opportunity to trim costs / reduce quality would be vulnerable to a shareholder lawsuit for leaving money on the table...
pnt12 2 hours ago [-]
Nitpick: I think that's called skimpflation (worse quality), a friend of shrinkflation (less quantity) or inflation (higher prices).
aesthesia 6 hours ago [-]
I'm curious how common such lawsuits are. I don't think I've heard of any specific instances where shareholders sued because a company didn't make the product worse.
m10i 3 hours ago [-]
> "most customers won't notice if you take your product's quality from A to B (or C), but they definitely will if you take it from A to J (or A to M)"
Summarizes the video game industry pretty well
ehe78qhe 5 hours ago [-]
> Shrinkflation is a thing, which people suddenly started noticing in the past 5 years.
That's kinda me with respect to Claude. Generally i've had no issues and just kept pluggin' along.
The first real issue where i wanted to leave was the Claudish nonsense. If not for 5.5 i'd be on OpenAI by now.
dgacmu 4 hours ago [-]
Except for that one year when they completely swapped the meaning of the Cumulus line, which I think was 2008 with the cumulus 9 to 10 transition. It went from a neutral shoe that was good for people with high arches to more of a stiff stability shoe. The complainers aren't _always_ crazy. :-) (That doesn't mean the cumulus 10 was worse, of course, it just was a more substantial change that affected the type of runner the shoe was designed for.)
Wow the old grumpiness that lingers in my head from losing my favorite shoe. Who knew? Now I'm old and heavier and run in the Nimbus and am happy again. But you're right, of course, that most of the model changes are just fine and people like to complain.
yencabulator 7 hours ago [-]
It's pretty typical that physical-goods manufacturing "optimizes the process" to cut costs during years 1 & 2 of manufacturing.
Ikea is notorious for this: The early Billy bookcase had heavier veneer and sturdier construction early on, and was actually a really good purchase for the money. The later years replaced veneer with paper foil, used thinner shelves, frames, and backing panels, and was just significantly weaker.
ComputerGuru 7 hours ago [-]
Amazon Basics is incredible in this regard, they’ve optimized SKU identification down to a pipeline. They’ll essentially randomly pick items off their internal list of highest netting sales and test them to see how dependent they are on brand name recognition and price-quality signalling. To do this as efficiently as possible, they simply purchase a few hundred units of a high quality product in the space, stick it in an Amazon Basics box and list it on their site under their Amazon Basixs brand at a price they feel they can achieve via white labeling, and wait to see how it sells. The use of high quality items (with quality above what can actually be had at the listed price point for the duration of the experiment) means they are really only testing the user base’s willingness to forgo a brand name for the category in exchange for a discount. If it sells well, they then work on sourcing it in bulk as a white labeled item “for real”, while if it sells poorly they simply delist and move on.
I (used to) buy pre-spliced/terminated fiber optic cables with some frequency from Amazon and came to be familiar with the brands and their quality. One time while shopping for some fiber optics, I saw Amazon Basics-labeled OM-3/OM-4 MMF cable at a very tempting price, so I purchased some to see if it was any good.
To my utter shock and surprise, when I received the trademark plain cardboard boxes with the Amazon Basics label on them and proceeded to open them, I found that I was sent boxes of cables still factory wrapped with labels that clearly read “Corning Optical” – which if you know anything about optical fiber, was pretty much the premium brand in the game. I should have stocked up because the next time I went to order I found out their experiment had ended and they no longer sold “Amazon Basics” finer cables.
dragontamer 3 hours ago [-]
I have a similar story.
Amazon Basics AA NiMH was well known to test exactly the same as the top Japanese brand 'Eneloop'. Extremely good specs all around
Recently though, they are still called Amazon Basics but no longer test like Eneloop. They've changed manufacturers for the worse and are hoping no one notices...
hn_acc1 2 hours ago [-]
I did a deep dive into NiMH batteries a few years ago and concluded that most people felt similarly: you can often get "good" (similar to Eneloop) specs for a short run from almost any manufacturer at the outset (see Ikea Ladda batteries - suspected of relabeled Eneloop for a while, but now not as good), but consistent quality is pretty much only Eneloop or other name brand, with Eneloop generally being the best.
In the interests of saving my sanity and time (it's not free!) having to chase down which batch of which brand is "good" at the moment, I just decided on Eneloop all the time. Sure, we now have like 100-120 or something (wife likes flameless candles - just bought another 16-pack AAs) and I COULD maybe have saved $200 by buying dirt cheap. But all the time spent debugging flaky batteries, having the spouse complain, etc wasn't worth it to me (I get paid reasonably well).
ComputerGuru 1 hours ago [-]
No, that’s just white labeling. Like you said though, they can switch manufacturers and pull the rug from underneath you at any time.
My example is like if you opened the Amazon Basics box and found Eneloop branded white/black/blue batteries directly.
ck2 5 hours ago [-]
actually they were getting worse each year
most running shoe series get heavier year to year as manufacturers turn to cheaper materials and add cushioning to try to attract more adopters
it's almost universal, very few manufacturers seem to be able to resist tampering
(heavier shoes are slower, every three ounces is equal to another vo2max point lost)
reducesuffering 5 hours ago [-]
You're literally doing what GP is describing. We have objective data on running shoes on: https://runrepeat.com/
The foams are getting better, the shoes lighter, they are more cushioned and more responsive in general. Especially the ASICS.
ck2 5 hours ago [-]
nope
you mean NEW MODELS are being introduced with lighter faster foams
not the same model year to year
modern example: Saucony Endorphin Speed
v1 in 2020 was award winning
v2 in 2021 was almost the same, more praise
v3 bleh
v4 v5 bleh bleh
they cannot resist tampering
waterproof 14 hours ago [-]
I find that I learn to "trust" a model to get certain things right, as I would trust a colleague. So, as `expectation` increases, my prompting and context management gets sloppier.
`percieved_performance = actual_perf/expectation`
`expectation` is an increasing function over time.
`actual_perf` is a stochastic function of the model's true ability, context, etc. -> a recipe for some bad sessions.
As for multiple bad sessions in a row, this is a studied phenomenon in gambling where players perceive "runs" because our brains love to find patterns.
goodmythical 9 hours ago [-]
This reminds me of the fact that true random does not feel random to users due to the clumpiness that the average person does not anticipate existing in true random.
e.g. The original apple shuffle and the Risk app ins which a string of songs from the same album or three one roles are "not random"
svachalek 5 hours ago [-]
I've been working on a game that has dice rolling and even knowing about this effect, I started going crazy yesterday when I had a long streak of numbers, like 1-20, it was 15-16 like 9 times out of 10. I was sure there was some kind of bug in how it was initializing random, or saving the number, etc etc. Just could not find it. Streak continued to another roll, another roll... still couldn't find it. Then the streak just broke. Apparently just random being random.
goodmythical 3 hours ago [-]
I find it useful to consider something like shaken rice. If you take a 10x10 grid and shake 100 grains of rice on it, you'll find that some cells contain no rice while others contain as many as 5 or 6. Run the experiment enough and you'll converge on each cell getting one grain per run, but any individual sample will likely be clumpy and the state in which each only has one will occur infrequently.
Also, consider that in flipping 10 coins, you'll find strings of 2 heads in ~86 percent of runs, 3 in ~51% of runs, 4 in ~25% of runs, 5 in ~11% of runs...and in strings of 100 flips you'll finds strings of 6 in ~55%, 7 in ~32%, 8 in ~17%, 9 in ~9%...
Widening the range from "rolling exactly 15" to "rolls 15 or 16" or "rolls between 14-17" makes the strings even more likely as you're doubling the success rate from "only 9 15s" to the "any string between 9 fifteens, through 16 and 8 fifteens, to 9 16s" space.
To check if your random is randoming you can calculate expectations versus your results (using a large enough sample) with:
For N samples of a fair die, expexted runs k with probability of success p and failure q can be calculated as:
General Variables:
N = total number of rolls/trials
k = target streak length
p = probability of getting the target outcome (e.g., 1/20 for a specific roll on d20 or 1/10 for two specific results)
q = probability of getting any other outcome (1 - p)
Expected runs of AT LEAST length k:
E(runs >= k) = p^k * (1 + (N - k) * q)
Personally, I find that 'sticky' dice always provide a nice narrative device, at least in narrative games. A character who's player can't seem to roll over a 10 must, after all, be cursed or perhaps deliberately sabotaging the party.
Took a little bit of digging to find it - I remembered the graphic at the top and found a blog post that copied it and linked to the blog post, but the blog post isn't there anymore... so web archive.
Yeah I bet most of us remember GPT-4 a lot more fondly than we would if we were to return to it today.
svachalek 5 hours ago [-]
Absolutely. Objectively speaking it was far less consistent and capable than even small local models today.
oooyay 4 hours ago [-]
> Perceived performance is actual performance over expectations and the latter just keeps increasing over time.
This is true of all reliability and performance paradigms, incidentally
rplnt 17 hours ago [-]
> I personally think people sense nerfs more often than they happen and that it's often about honeymoon effects.
I believe in temporary nerfs. Operators reducing quality significantly to increase throughput for whatever reason (high demand?). Same session, model being completly incapable, and it being back to normal the next day. Experienced that with Anthropic models way too many times. Never on weekends, usually during US work days.
It's been a few months since I last recall this though, must have been pre-opus-5. And I know there are benchmarks for this as well, hence "believe".
transcriptase 17 hours ago [-]
I vividly remember when ChatGPT3.5 went fully mainstream, there were times where within minutes you would realize they were only serving up idiot mode and there was no point trying to do much until demand died down and they swapped back to the non-quantized version.
People called it lazy mode, in that instead of writing the script you asked for it would basically tell you to learn to code then check out xyz topics to tackle the problem.
setopt 16 hours ago [-]
Regarding lazy mode, I recall ChatGPT sometimes almost refusing to do a web search despite me asking explicitly for it, instead replying with speculation about what the search results likely would tell us. If I pretend to be angry that it didn’t search the web it would however do it. Haven’t noticed this in a while either.
lukan 9 hours ago [-]
"If I pretend to be angry that it didn’t search the web it would however do it."
You have to pretend to be angry in such situations?
natpalmer1776 10 hours ago [-]
Oh god I just realized these are the same types of stories passed down to me by sysadmins of yore about when microsoft did XYZ. Am I… old now?
katzenq 15 hours ago [-]
It should have stayed that way. Be a good search engine and encourage the human to do the work themselves.
AbsurdCensor 12 hours ago [-]
Wouldn't that be like a calculator saying 'pull out a math book' when you try running calculations?
katzenq 12 hours ago [-]
It's like doing the calculations yourself (with whatever tools you want) and knowing what you're doing, steering the process yourself, rather than begging someone to give you the result so you can then show it off like you did the work. I can understand and explain what my calculator is doing and I'm not offloading decisions to it. Using a language model to shit out projects that you can't fully understand the structure of yourself is insane. Ceding control of any design decisions to a language model is insane. They're fine as second-order autocomplete and semantic search engines. Humans should be handling design and implementation entirely, making things for other humans. An LLM running in a loop can't build humane systems.
sejje 10 hours ago [-]
> Humans should be handling design and implementation entirely, making things for other humans. An LLM running in a loop can't build humane systems.
I tell the LLM what to do in its loop. It builds it. I tweak it until it is perfect. I care a lot about UI.
I'm mostly building my own UIs lately, but I find no problem with the development loop. I'm building much better UIs, because it's way easier to test things, and scrap things that I thought would work, but don't. It all happens in a matter of minutes.
Humans should use more llms.
silversmith 13 hours ago [-]
Anecdote - I work in a time zone offset from continental US. The performance of anthropic models would noticeably drop, around the time US work day started. It was so bad around 4.x time that multiple colleagues re-arranged their schedule to have least overlap with US work day. Admittedly it's been better recently.
smurf9852 14 hours ago [-]
Friend of mine works for a corp that is one of the top spenders on Claude models. He complained about these nerfs during peak demand.
Their Anthropic contact changed something and it did not happen since.
jclardy 3 hours ago [-]
I do remember times in the past where when I was up super early (4am EST) I would get super high quality results, then in mid afternoon EST it seemed to be degraded.
Barbing 9 hours ago [-]
> It's been a few months since I last recall this though
Anthropic cut a deal with SpaceXAI in May - $1.25b/mo. Before that, they employed months of dishonest nerfy strategies, to an extreme.
This is what pissed me off the most. Make it slower, rate limit it, move the credits to other time slots, idk.. but returning BAD results? That's the worst approach you could take.
comboy 22 hours ago [-]
But you are using API not the CLI right? I did not ever observe API degradation, only subscription stuff through their CLI.
jdthedisciple 2 hours ago [-]
Well looks like thus far Astra seems to be getting anything but nerfed, given that its score is actually rising
scrollop 18 hours ago [-]
There's also this one which has been around for a while
Seems like a harness change (maybe bug) rather than a model change, the input tokens dropped quite a bit right when the degradation happened.
lxgr 10 hours ago [-]
> We use the latest available Codex release with GPT-6 Sol.
This alone makes the benchmark unsound.
shawabawa3 10 hours ago [-]
> "We are collecting a new GPT-6 Sol/high baseline from runs beginning September 24, 2026. Degradation detection is paused."
rednb 17 hours ago [-]
I'd take this kind of benchmark with a grain of salt. At this point, I have a set of comprehensive guidelines covering both backend and frontend work, and for the frontend we go as far as explaining what we a good design is in our visual system, and even how to conduct a visual review when screenshots are handed to the model.
Deepseek 4.1 ranks very low in this benchmark but it has proven so capable that after being simultaneously on Max x20 and Pro x20 subscriptions, i've transitioned to using DS 4.1 as a daily driver and am very satisfied.
My point is, i think their overall ranking makes sense, matches my experience with out of the box capabilities for vague and underspecified tasks. But seeing a model rank low in their ranking does not mean that the model is incapable. Having skills and guidelines has a lot of influence on what you get out of a model.
dotancohen 15 hours ago [-]
You use DS through Open Router? Which harness?
I'd love to hear more, I'm considering jumping ship. I'm running a Debian desktop if that's a concern.
rednb 13 hours ago [-]
I use the direct API from deepseek using Opencode, no problem to report, works like a charm.
Except maybe that the model often believes that he is running out of context, and needs to rush so i occasionally need to jump in to tell it that it still has plenty of room left.
But this does not degrade the quality of my overall experience in a meaningful way.
dotancohen 11 hours ago [-]
Thank you. I might just check that out.
14 hours ago [-]
throwitaway222 8 hours ago [-]
Often times people think of "nerfs" as my first prompt (which was greenfield - no or little code existed) used 5% of my plan usage. And then 2 weeks later (as the agent is busy reading hundreds of .rs and .ts files it previously generated) the user complains the usage is going down 30% for a single prompt instead of 5%. Attributing this to a "NERF" makes little sense because it's the same model.
user3939382 22 hours ago [-]
Anthropic A/Bs my weekly quota amount. So I have an automated prompt that runs at 3 AM with a transcription task, I measure input and output tokens, and weekly/5 hour quota before and after. The absolute token counts stay within 0.1% while in mode A it counts for 1% of my 5 hour quota and mode B 4% of my 5 hour quota.
jacquesm 21 hours ago [-]
How did pissing off your customers ever become a business model?
I can't imagine sticking with a supplier that plays games like that with me. Tokens are a pretty vague quantity to begin with (you don't control how many tokens a model puts out in response) and giving a couple of purposefully wrong responses will happily inflate your bill, but you don't care because eventually it worked. It's almost an ideal vehicle to scam people.
Imagine the power company being able to decide how much you consume and at which price point.
none_to_remain 20 hours ago [-]
I find it amazingly rich that they bill you for """thinking""" tokens and now you don't even get to see them, they're gonna train the thing to sing "99 Bottles of Beer on the Wall" to itself before it starts work.
Turskarama 19 hours ago [-]
They don't want to waste tokens on purpose, what they're actually hiding is when the model wastes tokens on obviously stupid "thoughts".
adastra22 18 hours ago [-]
No they are hiding the chain of thought to make distillation harder.
slim 17 hours ago [-]
It's fascinating that you all think accounting is real and it did not come to your mind that they could make up numbers when billing
TeMPOraL 17 hours ago [-]
They could, but as you see here, people are very eager to create dashboards and trackers that do external accounting by proxy, so they can't just "make up numbers" without the customers noticing and making a fuss.
herval 21 hours ago [-]
> How did pissing off your customers ever become a business model?
Airlines, banks, health insurance…
tccole 20 hours ago [-]
So very low margin businesses with hogh amounts of regulations.
herval 12 hours ago [-]
makes you wonder why openai/chatgpt/xai are in bed with government so much...
petesergeant 17 hours ago [-]
Banks and health insurance are much more consumer friendly outside of the US, usually because of regulation. Turns out you can just tell banks “make transfers cheap and essentially instant” and they’ll do it, rather the bullshit they have in the US.
herval 12 hours ago [-]
you'd be surprised. I've yet to live in a place where either is consumer-friendly...
CodesInChaos 17 hours ago [-]
Another way Antropic misleads its customers is the description of the max plans. They are advertised as having 5x/20x the 5h quota as Pro. But the description says nothing about how the weekly quota scales, leaving customers to infer it scales the same way. But from what I've heard, the weekly quota is only 3.5x/7x that of Pro.
Only if we did not have these cheap illegal Chinese models
msdz 17 hours ago [-]
That is why
> Step 2: Regulatory Capture
is being worked towards.
TeMPOraL 16 hours ago [-]
> illegal
Are they though? Or is it just what some companies would want them to be?
cavoirom 16 hours ago [-]
Their fate is coming. Until the open-source models will be usable in machine with 256GB memory, they are done. Their behavior is unacceptable (Anthropic) recently but it won't last long.
csomar 17 hours ago [-]
I think it's sinister, but not for the reasons you're thinking. I think they're just wildly unprofitable on subscriptions. The idea that most customers won't use their full quota is plain wrong: most people are maxing out their subs, or even reselling whatever quota they have left.
When you're running something at a loss, you can mistreat your customers and they'll still stick around (I'm an example). OpenAI and Anthropic are now cheaper than Chinese models on subscriptions, while being 6-10x more expensive on the API.
My guess is they need the user numbers for the IPO and are willing to take a temporary loss in the meantime. By the time they go public, they'll either drop the subscription model or it'll turn into what the Chinese providers already offer: basically just a cap on how much API you can consume. Same same.
It's not clear what API tokens actually cost them, but I looked into running a local model, and it's way outside the budget of an individual or even a small or medium business (hundreds of thousands of dollars). So my guess is that running these models economically isn't possible, even if they're delivering real business value (coding, research, etc.). In other words, at API prices I'd just stop using AI, and I suspect most other developers would too.
TeMPOraL 16 hours ago [-]
> It's not clear what API tokens actually cost them, but I looked into running a local model, and it's way outside the budget of an individual or even a small or medium business (hundreds of thousands of dollars). So my guess is that running these models economically isn't possible
Datacenters have massive economies of scale. Everything from cheaper electricity to having specialized, more efficient hardware to simply being able to run it continuously at near-100% utilization, all adds up.
Many things in the economy - most notably, manufacturing of most consumer goods - only makes economic sense once you're producing for/serving millions of people. This is not unusual.
> In other words, at API prices I'd just stop using AI, and I suspect most other developers would too.
Many say that, but I sincerely doubt they'd actually follow through. People might get more conservative about how they spend their tokens, but AI today is just too good at eliminating drudgery and boring / bullshit parts of daily work to give up on merely 3-5x price increase.
csomar 16 hours ago [-]
> Datacenters have massive economies of scale.
Sure. Issue is, no one is providing on how much it actually costs to burn these tokens. And as we don't know, we can only speculate.
> Many say that, but I sincerely doubt they'd actually follow through.
I have a $100 open ai sub and I track my token usage. Last month I spent roughly $2.600 in equivalent API usage. There is no way am paying that. I let my $100 sub lapse if next month I'll be using it less.
Look, I am not saying that there isn't a potential value out there. But the cost has to be bounded. If your opportunity is $1.000 and AI costs $2.000 to execute it, then you don't have a business model here.
FeepingCreature 14 hours ago [-]
> Sure. Issue is, no one is providing on how much it actually costs to burn these tokens.
You can assume Openrouter open-model providers serve at or above margin, because there's no branding so there's no reason to do it unless you can be profitable. If the Anthropic models are anywhere in that ballpark, they're very comfortably profitable on API.
airspresso 13 hours ago [-]
> I looked into running a local model, and it's way outside the budget of an individual or even a small or medium business (hundreds of thousands of dollars).
That is a big exaggeration. You can have a perfectly usable local LLM setup that will power your agent for single digit thousands of dollars. Can even power multiple agents simultaneously, depending on the hardware and setup. Won't be fast and won't be frontier intelligence, but definitely useful.
zozbot234 13 hours ago [-]
Any model running on "single digit thousands of dollars" hardware will either be below SOTA (even for local models) or not even close to fast enough for real-time agentic work. Even the latest so-called "flash" models are large enough that doing real work usably with those on a lower-cost platform is at least dicey. You can fire off non-interactive work and do especially simple Q&A/chat (which is vastly more token-efficient than anything agentic - though even then latency will be high for anything genuinely SOTA) but that's about it.
icepush 16 hours ago [-]
You can put stuff like "make sure your reply is between 800 and 900 tokens" at the end of your prompt and the vast majority of the time it will do so.
topspin 19 hours ago [-]
It's worked for online PvP gaming for a long time. Nerf stuff the min-maxers "earned" through game mechanics and sell over-powered "premium" things to everyone else to pwn them. Then nerf the old premium stuff and make new premium stuff. Forever.
I don't know if that's the actual origin of the term nerf, but it was the first time I'd heard it.
done_lurking 19 hours ago [-]
I think the origin of the word "nerf" as a verb came from the Nerf brand of toy guns. The idea being that "Nerfing" something is to turn it into a harmless version of itself.
apitman 21 hours ago [-]
Do Anthropic quotas give you precise remaining token counts or something? I have something similar set up for tracking my ChatGPT usage but it only gives percentages remaining, which is a pretty coarse metric.
ffsm8 21 hours ago [-]
Claude code supposedly has otel you can set via env. I haven't set it up, so I'm just repeating hearsay.. but it supposedly has everything relevant in it wrt token usage and cost
It's meant for their test env I think, so is not documented to my knowledge
TeMPOraL 16 hours ago [-]
It's for corporate users who want to track how the product is used internally, and it was documented at least some time ago, quite extensively even.
Tokens used / percentage change is a pretty obvious metric. They give you both, but they don’t do the math for you.
apitman 5 hours ago [-]
It's obvious unless you have multiple requests from different models in flight at the same time, and the sum total usage comes out to less than a single percentage.
CodesInChaos 17 hours ago [-]
Could be load dependent, not an A/B test.
Is the fraction of the 5h quote consumed consistent with the fraction of the weekly quota consumed?
I heard there is a usage tracking tool you can install that tells you if tokens are more or less expensive at the current time.
user3939382 11 hours ago [-]
I say A/B because when it toggles it does so for days.
braingravy 22 hours ago [-]
Pretty amazing to see enshitification happen live with a product still in development… Truly web 4.0
LimitExperience 21 hours ago [-]
[dead]
KronisLV 16 hours ago [-]
We could also use something that tracks concrete token amounts each tier gives you, in case they ever mess with it - and also maybe even the tokens needed to accomplish a particular benchmark, to see how much you can actually get done.
jtrn 8 hours ago [-]
If this is actually even close to reliable tracking, it's one of the most awesome benchmarks I've seen. I gave up on feeling the zeitgeist for what people were saying.
Aperocky 9 hours ago [-]
So that's why even Qwen-3.8-27B caught up with Opus4.6, it was nerfed to the ground.
shados 9 hours ago [-]
The whole nerfing narrative puts in the spotlight now crazy supertitions come into being. The group think every day that everything is falling apart is crazy.
Grimblewald 23 hours ago [-]
I dunno, I never sense nerfs for local models, but consistently a few months after launch for corpo hosted models, seems odd my internal model for the capacity of a model drifts for anthropic models but not local ones. I've been using LLMs heavily even before ada/babbage/davinci days, and trust my internal calibration over baseless handwavey explanations for why im imagining things, especially when I have data that shows capacity regression on frontier models for tasks, e.g. one shot success at loss, 0 success in 15 attempts once nerf is sensed. Others publish their quantified capability regressions which are also more trust worthy than this kind of handwaving.
eulgro 23 hours ago [-]
Your comment makes no sense. How and why would a local model be nerfed anyway...?
r_lee 23 hours ago [-]
he's saying that he notices a difference between local (not nerfable) and hosted ones, so that it's not as likely to be just placebo
martin- 17 hours ago [-]
But if it is placebo, obviously he wouldn't notice any placebo change for local models, since he KNOWS he is using an immutable local model. That comparison only works if he doesn't know what model he is using.
jacquesm 21 hours ago [-]
That and 'loss' may have been intended to be 'launch'.
andriy_koval 22 hours ago [-]
Usage bench is also very useful! Thank you for doing this!
hamandcheese 20 hours ago [-]
...is it? I'm looking and it seems like it doesn't have any data. It might be useful if they keep it up.
andriy_koval 20 hours ago [-]
Yeah, I guess they started it today..
zerop 16 hours ago [-]
What could be the reason to nerf?
StableAlkyne 16 hours ago [-]
It's cheaper to run a quantization of a model, but its quality is reduced.
For example, if your weights were trained as 32-bit floats and you need 1TB of RAM, you could reduce that to around 256GB by quantizing to 8-bit floats. You also make the model faster in the process because there is less data to process to calculate the next token.
The game is to balance between the savings of quantization and making the model dumb enough the people notice
sscaryterry 15 hours ago [-]
Someone always knows about where the bodies are buried. These shenanigans always end up surfacing eventually.
sigbottle 8 hours ago [-]
But as we also see in this thread, all evidence is dismissed and there's no good faith discussion. All data from the other side is lies and contamination. In that environment, it's power who decides who wins.
ForHackernews 12 hours ago [-]
If they surface after the IPO, everyone who matters will have already gotten paid.
Or rather it's being.. Unsubstantiated. The Nerf conspiracy isn't that there have been a few harness and platform bugs leading to performance regressions, but that OpenAI/Anthropic have maliciously and unethically degraded their model performance post release to shed load and save money.
somenameforme 22 hours ago [-]
Yeah it's just inconceivable that companies whose entire business model started by engaging in wholesale for-profit theft and abuse of intellectual property would ever be so unethical as to try to lower their costs, especially just prior to an IPO.
lxgr 9 hours ago [-]
It's at least highly implausible. Why would they engage in pretty uncontroversially illegal deception/fraud if they have so many other legal ways of gaming benchmarks, selling more tokens etc. available to them?
It's like arguing that your bank is scalping you by rounding down interest math on odd days of the month when they can just introduce a perfectly legal bullshit fee or otherwise change their terms to your disadvantage instead.
Rapzid 22 hours ago [-]
Again, these things are constantly measured. They sell HEAPS through their API access to enterprise consumers that expect a model to not be nerfed after it's released. And you bet many of those enterprises, some spending many millions each month, are measuring this shit.
So this is a case of extraordinary claims requiring extraordinary evidence.
And even though it's super straight forward to collect the evidence, there seems to be zero substantiating the vibe bro conspiracy theories.
mrandish 20 hours ago [-]
> They sell HEAPS through their API access
The claim is that they nerf subscription accounts not API.
applicative 8 hours ago [-]
My impression was that, at least with Anthropic, the point of my subsidized subscription is that I'll convince my employer to get an API account. If anything, the motive would instead be to ensh*ttify on the corporations already committed. Those of us inducted as boosters would get the fluffed up product.
somenameforme 21 hours ago [-]
If you're going to try to argue that companies doing things, completely legal mind you, to increase their profit margins is a conspiracy theory then you're not debating in good faith. Let alone when we're speaking of a subset of companies that were fundamentally built on wholesale unethical behavior carried out for profit. Let alone when we're speaking of companies who are all racing to IPO where short term results matter more than just about anything.
Another issue is also that the risk here is probably literally zero. Any evidence in support of such could easily be dismissed, with completely plausible deniability, as a short-lived technical glitch as opposed to intentional behavior.
Rapzid 20 hours ago [-]
> an explanation for an event or situation that claims a secret, powerful group is responsible for a hidden plot, rejecting the standard or official account
I'm sorry, but yeah. The official account is a harness regression and some platform bugs.
Where is the evidence they are underhandedly and unethically regressing their models to shed load and reduce costs? This is the conspiracy theory running rampant through the vibe boroughs; that they are bait-and-switching on model capabilities then "nerfing" them to save money and shed load. Where is the evidence?!
I'm not saying it's illegal, per say, so don't come at me with that straw man bull cock. This bro science conspiracy has been circulating for at least 2 years(I don't even know) and enterprises would certainly be pissed off if they were paying premium API prices for advertised and previously tested model capabilities that are suddenly under performing due to "nerfing" shenanigans.
So where is the evidence?!
airstrike 20 hours ago [-]
You're giving way too much credit to "enterprises" both noticing and publicly airing out their dissatisfaction
Not too mention these companies could easily offer one product to enteprises and another to everyone else
Model nerfing is real
Rapzid 20 hours ago [-]
Uh huh. OMG you're so right, it's soooo real ;) ;) ;)
I was so certain it wasn't, based on the complete lack of evidence.
But then you said it's real. NVM, I don't need evidence! Somebody said it's real!
This place has fallen off.
grim_io 10 hours ago [-]
Reddit cross contamination.
Over there it's a rite of passage to accept it as a fact.
airstrike 4 hours ago [-]
This place has fallen off so long ago that people like you consider yourself old timers but don't even follow guidelines
Bad evidence is worse than no evidence.
And you failed to address the specific criticism I made to your point, instead going for an ad hominem / poisoning the well.
In sum, you're out of line, woefully misled, and lacking in logical thinking. Three strikes, you're out.
troupo 16 hours ago [-]
> And even though it's super straight forward to collect the evidence, there seems to be zero substantiating the vibe bro conspiracy theories.
Where people pointed out issues early and en masse, and Anthropic denied it was happening, gaslighted anyone claiming this was an issue, then begrudgingly admitted it was an issue, and then spent another two weeks "fixing it".
Anthropic is in a perpetual state of "oops, these 'bugs' degraded our model quality" and only admit the issues when it's immediately obvious and visibly affects a large number of customers.
Otherwise all open benchmarks can be (and are) gamed. And it's quite hard to judge the output of a non-determenistic black box that Anthropic (or OpenAI) constantly tweak.
applicative 8 hours ago [-]
what's unethical about downloading the internet for machine training?
stackghost 22 hours ago [-]
To me that sounds exactly on-brand for Big Tech in general and Sam Altman in particular.
jackmott42 21 hours ago [-]
No one here has suggested that the conspiracy theory is stupid (But I will, it is stupid), we are pointing out that the people actually measuring model performance have not found the nerfing before launch conspiracy to be true. in short, yall dumb, shut up.
stackghost 21 hours ago [-]
I think the real story is just how easily people believe that purported conspiracy theory. It speaks to how little trust there is in these AI companies, and in Big Tech in general, that this "conspiracy" theory is perfectly plausible to lots of people
> in short, yall dumb, shut up.
no u
Rapzid 20 hours ago [-]
It's speaks to how low the bar has dropped.
Everyone wants to be a software engineer, until it's time to do software engineering shit.
You know, like scientific method shit we learned in 5th/6th grade.
It's the great bro science incursion.
stackghost 20 hours ago [-]
>Everyone wants to be a software engineer, until it's time to do software engineering shit.
The absolute state of software in 2026 should tell you that almost nobody does “software engineering shit” and never has.
stbenjam 9 hours ago [-]
These guys are on twitter angry about the rate limit decrease and allegedly cancelled all their OpenAI accounts. Wonder how they'll maintain this.
I am quite convinced that the whole nerfing phenomenon is 90% AI psychosis. I have the word muted on X.
cyberes 9 hours ago [-]
That's the wrong use of the word
Barbing 9 hours ago [-]
Maybe just the wrong ballpark (like 1-10%).
nullbio 18 hours ago [-]
Do they use private benchmarks? Because if not, it could be selectively nerfed.
I also wonder if cache could be used to throw these off as well, where it's serving un-nerfed cache results for context windows that are identical to ones they've previously had for benchmark requests.
Seems like the only way to do it well would be to have some randomness involved that couldn't be cheated on - but you'd want to do it in a way that doesn't throw out the benchmarks too much, so your results can be compared still.
Gabrys1 17 hours ago [-]
Thankfully, we can now use AI to design a test that tests AI. And the AI company can use AI to detect the test and cheat. And we can then use AI to implement anti-cheat.
All that energy wasted... could just drive a big V8 instead and make less money for the big tech
bdlowery 19 hours ago [-]
> This bench famously detected a degradation of Opus 4.6 which Anthropic later blogged about.
this bench was just released, it couldn't have detected opus 4.6 degradation.
Razengan 23 hours ago [-]
Theory (Conjecture? Hypothesis?): What we notice as "model nerfing" is the company diverting compute to training/running new unreleased models..
Remember that some people get access to the next flagship version long before us peasants do. I recall seeing the mention of "Astra" more than a month before it was officially announced
Centigonal 22 hours ago [-]
wouldn't less compute result in slower inference, rather than worse performance?
latentsea 22 hours ago [-]
They could potentially quantize the model and run it at lower quality taking less VRAM.
btown 22 hours ago [-]
The more likely thing that would happen is that the provider begins silently interpreting (perhaps some) high effort-level requests as medium, etc., or having a classifier do this far more subtly. As such, the load on the cluster is less, and more resources can be devoted to training. Whether the frontier labs actually do this is purely conjecture at this point.
JohnBooty 21 hours ago [-]
I assume there's classification going on where a really basic "Hi how are you?" style request sent to a high-effort instance can be routed to a lower-level instance. This... is pretty much fine with me, assuming they do a good job of it.
I would also assume they use nebulous labels like "Medium Effort" or "High Effort" map to quantitative amounts of compute allocation... and that these amounts can be varied manually or automatically. Right?
I mean, there's a reason why they call it "High Effort" and not "Exactly 5 Minutes of GPU Time on Exactly 10 GPUs." They want to be able to move those sliders and tweak those knobs.
btown 5 hours ago [-]
The problem is that if benchmarks are run at a tight classifier that says "a request for high effort means check-under-every-stone regardless of simplicity" but a user request is run on a different classifier where "high means maybe high, maybe medium, maybe even low, even for meaningful tasks" then you're not getting the model that you saw in the benchmarks.
And, while you might be billed fewer tokens as a result (because the lower thinking would result in less investigatory work), you might not know this is happening, and know to dial up effort accordingly - you'd simply get a worse work product. And certainly, Anthropic's incentive for anyone on a subscription is to push this as aggressively as they can, so people use less of that subscription.
Sadly, I'd also expect that the OP's benchmark will be detected as a test of model capabilities, and thus be given a high classification so that this strategy remains undetected.
nightpool 22 hours ago [-]
why is that more likely?
zxilly 22 hours ago [-]
Because they already did so. The model in Codex will get lower `juice` than API version.
poizan42 21 hours ago [-]
My guess is that they are dynamically changing the quality of the model to always keep the speed above some floor. So once it gets below that they switch to a worse quant or reduce reasoning level, or some combination of both.
jackmott42 21 hours ago [-]
There is no nerfing, look at the data before coming up with a theory as to why the nerfing that isn't even happening is happening.
fuck
17 hours ago [-]
loopydosuette 9 hours ago [-]
you can't explain honeymoon effects. like wow wow wow and then suddenly: same task, lesser performance is a misperception?
my fair lady gained a bit weight and the bjs lack variety?
[ I'm certain that that's a quote from some time ago by some commenter in a thread with a similar or even the same context ...but my Amnesia (T▽T ):・゚:・゚] fuck off.
you open two files, before and after you notice a nerf, and from worse comments to logical oversights, it's all damn obvious.
don't normalize this make believe bullshit and misleading people who you think barely understand what they see anyway ...
you wouldn't even know if models had somehow timed nerfs hardcoded into them, however much control over the stack you have.
ridiculous
sheepscreek 20 hours ago [-]
> It could also mean nothing happened and people are pattern-matching on noise.
It is definitely not this. Anthropic has thousands of employees making probably > 100,000+ tiny changes across the entire stack and infrastructure everyday.
The compounding effect can definitely cause temporary regressions in some domain or use-case that doesn’t have really good coverage in their internal tests.
What makes this particularly challenging in the case of LLMs is how changing some language in the prompt can vastly affect the output.
But this is less true as models become larger and additional parameters are able to capture each and every possible nuance of the language. I’d say caching and cache tuning or token optimization/tweaks to thinking are the biggest culprits today.
ben_w 16 hours ago [-]
> It is definitely not this. Anthropic has thousands of employees making probably > 100,000+ tiny changes across the entire stack and infrastructure everyday.
What "stack" do you have in mind here?
An LLM is a monolithic slab of weights, not millions of lines of code spread over many microservices. The changes they make will be showing up in web and app responsiveness, and in the performance of training runs for one or more next versions of the that monolithic slab of weights, but I'd be surprised if there's a way for their efforts to show up directly in a "has this model been nerfed?" sense.
Both Anthropic and OpenAI used to have multiple snapshots per model; both seem to have switched to updating version numbers when anything changes, looking at the "snapshots" section in their recent and old models: e.g. https://developers.openai.com/api/docs/models/gpt-4o for how far back I had to go to find multiple snapshots on an OpenAI model, and https://platform.claude.com/docs/en/about-claude/model-depre... seems to be a more useful list for Anthropic, but both are now things they did a year ago.
TeMPOraL 16 hours ago [-]
> An LLM is a monolithic slab of weights, not millions of lines of code spread over many microservices. The changes they make will be showing up in web and app responsiveness, and in the performance of training runs for one or more next versions of the that monolithic slab of weights, but I'd be surprised if there's a way for their efforts to show up directly in a "has this model been nerfed?" sense.
It most definitely is not, hasn't been for a while now.
I.e. when dealing with hosted models of the large providers, you are not interacting with a big bag of floats. You are interacting with an API/UI that presents an unholy web of software components, some of which may be large or small bags of floats, as if they were a big bag of floats.
Even with local models, you have dozens of parameters you can tune for inference these days, all of which affect quality of output in some way or another. And that doesn't touch load balancing, A/B testing, shunting token burners ("Hi chat, how are you?"), protecting user from themselves (refusing to answer "bad" queries, stopping "bad" responses), protecting user from third parties ("prompt injection" mitigations), protecting the model from self-pwning itself when calling tools, then the tools themselves, their prompts, the stack of system prompts used by the vendor, etc.
There's a lot of things to tune there, and just as many reasons to do it.
ben_w 16 hours ago [-]
The API documentation linked in my comment seems to say that (with two exceptions*) when we ask for a specific models, we get that specific model.
The livenerf tester appears to be testing a specified model, just as the website (and Claude Code) do when a user makes that choice.
> Even with local models, you have dozens of parameters you can tune for inference these days, all of which affect quality of output in some way or another.
Good points.
> And that doesn't touch load balancing, A/B testing, shunting token burners ("Hi chat, how are you?"), protecting user from themselves (refusing to answer "bad" queries, stopping "bad" responses), protecting user from third parties ("prompt injection" mitigations), protecting the model from self-pwning itself when calling tools, then the tools themselves, their prompts, the stack of system prompts used by the vendor, etc.
The behaviour I'm seeing from the companies these days, the A/B is what I'm saying is not showing up like this, they present user A/B options openly, and the impression I have is this is to train n+1 models; the other stuff (but I say with low certainty) appears to be done in a more headline-grabbing manner, "model taken offline due to ${news}"? Short update cycles seem to allow that.
But the prompts you're probably right, I wasn't giving that enough consideration.
* the two exceptions being automated safety downgrade for dangerous topics, and "Auto-switch to Thinking" as a used-specified option in ChatGPT
dannyw 16 hours ago [-]
Serving LLMs at Anthropic scale is very, very different. It’s not SGLang or vLLM.
If you’ve tried setting either of these up, you’ll know how various tricky settings can impact throughout and model output quality; and those are much simpler stacks.
Even homelabbers are getting into disaggregated compute; e.g. one GPU for prefill, another for decode.
Obviously Anthropic and co are using a mixture of GPUs and hardware and clusters; not everything is just GB300 or whatever; so you then get into hardware quirks, kernel optimisations that may deliver huge speedups at the cost of a tiny bit of KL divergence, etc.
And I believe they’ve publicly said they use TPUs for inference too, but I doubt exclusively; and I’m sure that’s well optimised too.
Finally, Google has publicly stated they intentionally and silently degrade/poison models in response to distillation attacks; who knows what the other companies do.
jeffybefffy519 2 hours ago [-]
Didnt Opus 4.8 suddenly get more verbose mid release, as in it was released and responded normally then got annoyingly verbose weeks later?
zahlman 17 hours ago [-]
> It is definitely not this. Anthropic has thousands of employees making probably > 100,000+ tiny changes across the entire stack and infrastructure everyday.
I don't follow. That sounds to me exactly like a reason why it could be people pattern-matching on noise: because there is a lot of noise in which a matchable pattern could emerge.
jibalt 15 hours ago [-]
> It is definitely not this.
I don't think you understand what "this" is--or rather, the "thing" that didn't happen in "nothing happened".
> Anthropic has thousands of employees making probably > 100,000+ tiny changes across the entire stack and infrastructure everyday.
So, not the sort of thing referred to.
> The compounding effect can definitely cause temporary regressions in some domain or use-case that doesn’t have really good coverage in their internal tests.
Yes, but you claimed that this definitely did happen. But the whole point is to determine whether it did.
ChristmasTomer 7 hours ago [-]
Glad someone's keeping track. These threads always turn into “they made it dumber” vs “nah, you're imagining it.”
A few bad answers can be annoying as hell, but who knows what's behind them. I'm curious to see how the numbers look after a few weeks. Props for tracking improvements too.
msejas 16 hours ago [-]
As a claude code power user, when I get the 'rate the feedback on Claude' pop up, I used to say good or fine out of habit, and immediately after sending this feedback, I felt an instant degradation and mistakes that usually don't happen.
Now I dismiss it every time and the quality is more consistent.
Complete adhoc and personal experience but something I've observed, wouldn't be surprised if they nerfed on a per session basis
xnorswap 12 hours ago [-]
This is "rub your gameboy the right way to catch more pokemon" levels of insanity.
Random performance is random, your brain will jump through hurdles to fit patterns where there aren't any.
unglaublich 11 hours ago [-]
On the one hand, yes, on the other hand it would also not be too hard to do routing based on such a parameter and Antrophic has shown (through Fable and Mythos) that they have such a transparent routing system in place.
In fact, since they have some rule based system (if bioenegineering or security, route to degraded model) it would be almost trivial to add 'user has filed feedback' to it.
Not saying this is what happens, but just that it's not as insane as it sounds.
smashed 7 hours ago [-]
Yup.
And it's also the kind of solution a misaligned agent would implement:
Make sure users are happy about their experience but also optimize resources.
The obvious solution is to route users based on feedback.
ftchd 10 hours ago [-]
Not unglaublich
bb123 6 hours ago [-]
I could totally see them running an experiment to test user's stickiness/quality perception with decreased model performance. Facebook was doing exactly this in 2016. They tested the loyalty of Android users by secretly introducing errors that would crash the app to find the threshold at which a person would abandon. That was 10 years ago - imagine what the state of the art in user manipulation is like now.
jayd16 3 hours ago [-]
Its truly naive to think this is impossible or even improbable. You can say there's no proof or that its not the case but certainly much more has been done to juice 5 star ratings.
Aurornis 10 hours ago [-]
As a counterexample, I click that feedback all the time and nothing ever changes afterward.
SilverSlash 10 hours ago [-]
My personal issue with clicking the feedback is that I'm doing free RLHF for Anthropic without getting anything in return.
Aurornis 10 hours ago [-]
You’re probably participating in 100 or more A/B tests per year without knowing it and getting nothing in return other than a minor contribution to the company’s internal knowledge about what customers like.
grim_io 10 hours ago [-]
There is a difference.
Using feedback on your sessions makes that session, and presumably all attached data, fair game for training.
margalabargala 8 hours ago [-]
As if it wasn't already? The feedback button may tag it a certain way but the AI company certainly isn't going to let that sweet sweet user interaction go to waste just because they didn't rate it.
grim_io 8 hours ago [-]
If contracts didn't matter, sure I guess.
margalabargala 8 hours ago [-]
Are you under the mistaken impression that hitting a 1, 2, or 3 for bad/fine/good is itself sufficient permission to allow Anthropic to train off your chat?
They explicitly ask, after the rating, whether they can look at your chat. You can just say "no" to that, if you believe that they are following the rules they say they are.
Your original statement of "using feedback gives them permission to train" is just plain false.
grim_io 7 hours ago [-]
Truth be told, I never used the feedback function because of the clauses that reference the functionality in connection with training/service improvement.
redanddead 10 hours ago [-]
That sucks
Aachen 5 hours ago [-]
A better product isn't something?
iamjackg 10 hours ago [-]
How is it RLHF if they don't get a transcript of your session?
itopaloglu83 16 hours ago [-]
Another personal anecdote, but I had the case where Claude would visibly improve after a negative feedback.
And strangely, expressing frustration multiple times in a row would reliably trigger a feedback popup as well.
okwhateverdude 15 hours ago [-]
When the source maps leaked for the harness, it was revealed that they track how often tool use is rejected and how many times you say "fuck". They definitely try to track frustration sentiment.
That said, given the propensity for mature code bases to have "fuck" in commit messages/comments and those are typically of higher quality, I curse up a storm when the clankers make mistakes, if only to try put more quality-code valence into context. https://news.ycombinator.com/item?id=36584464
wunderlotus 11 hours ago [-]
True, but tracking ≠ Claude using that feedback to improve performance right there and then
ClikeX 13 hours ago [-]
To be fair, I also do my best work when I get to swear like a sailor.
jcutrell 11 hours ago [-]
As a hiring manager I now have a rational reason to prefer people willing to use profanity. Thank you.
Galilyou 16 hours ago [-]
Wow .. really! Definitely testing that out a few times meself
khalic 15 hours ago [-]
Given the tens of bot accounts in this thread making obviously misleading statements, I think you’re onto something here. Might be a good idea to open a Patreon or something, you probably need throwaway accounts to prevent anthropic from feeding you good models.
lxgr 10 hours ago [-]
No wonder you're convinced about there being an effect if you can just discard any criticism of your theory as being coordinated by a cabal.
johnfn 24 hours ago [-]
"Nerf"ing models isn't real in the vast majority of reported cases. Benchmarks like this or the 100 other "let's see if nerfing is real" copies would have shown it by now if it was.
I made a graphic to explain why people feel like the models get nerfed:
The idea is that new models can handle up to a certain level of complexity, at which point they fall apart. Every new model can handle more complexity, so there's a wonderful time upon release when you feel like you can do anything, only for you to hit the complexity ceiling a few days later when you saturate it. Rinse and repeat for the next model.
prodigycorp 23 hours ago [-]
Incorrect.
Anthropic has admitted to nerfing in the past. There have also been inference bugs. On top of that, model performance changes as they move compute to schwaggier providers as well.
Your chart is wrong.
simonw 23 hours ago [-]
> Anthropic has admitted to nerfing in the past
Where?
QwenGlazer9000 22 hours ago [-]
Earlier in march/aprile, there was a regression in Claude code.
Unintentional tbf.
weird-eye-issue 18 hours ago [-]
Unrelated to the models
ffsm8 17 hours ago [-]
does it really matter wherever its the harness or the model for the vibe coder user claude code?
fwiw, i think almost all regressions are down to a/b testing in the harness by anthropic, but it is objectively indistinguishable beyond "the coding agent ceases to be usable" and i'm back to traditional coding for a few hours until its back to normal again
thephyber 14 hours ago [-]
Yes. The agent allows you to switch models, so you could sidestep a bug in one model by temporarily using other models.
The /r/antigravity SubReddit is full of users who very much notice bugs with the tool/agent. We should be thankful that Claude Code is pretty stable by comparison.
weird-eye-issue 15 hours ago [-]
...
It absolutely matters because something like Claude Code has no guarantee that there won't be changes between updates but a model pinned at the API version level that is getting enterprise traffic absolutely does have that guarantee and would be a much more widespread problem...
weedfroglozenge 17 hours ago [-]
Honestly you need to start making pelicans on day of release, and following up a month later. Only way we are going to catch them
Rapzid 11 hours ago [-]
This is the way. Pelicans in a coal mine.
prodigycorp 23 hours ago [-]
Man this was last year and some Claude subreddit drama that I can’t furnish offf the top of my head but maybe one of the historians remember it.
p-e-w 23 hours ago [-]
Noone can seem to remember anything with certainty when asked to actually substantiate these claims.
> There are recorded cases of real regressions, but they're better characterised as incidents, not nerfs
erinnh 12 hours ago [-]
I dont really characterize a nerf as always by intent.
So I found this incident, as they called it, to still be relevant and why benchmarks such as the OP are useful.
p-e-w 21 hours ago [-]
That’s NOT Anthropic admitting to “nerfing” their model as claimed above (which implies intent), that’s a regression which they quickly fixed.
Christ this forum has become intellectually dishonest.
10 hours ago [-]
jibalt 15 hours ago [-]
There has always been plenty of intellectual dishonesty here, but remember Hanlon's Razor.
prodigycorp 23 hours ago [-]
I’m typing from my phone and im not going to review the semantics of Anthropic’s storied history of performance issues.
It’s not just ant. There are so many small knobs that providers can claim isn’t nerfing but “load management” or “improving user experience”. One example from OpenAI is reducing juice to reduce time to first token.
winwang 21 hours ago [-]
You can just have your agent find the evidence, review it, copypaste it.
18 hours ago [-]
consumer451 23 hours ago [-]
I asked a historian:
Two postmortems, neither quite "admitted to nerfing":
Sept 2025, infra bugs: "A small percentage of Claude Sonnet 4 requests experienced degraded output quality" [0], alongside "We never reduce model quality due to demand, time of day, or server load." [1]
April 2026, Claude Code: default reasoning effort was lowered from high to medium, plus a caching bug and a verbosity prompt. Per Anthropic, "The models themselves didn't regress, and the Claude API was not affected." [2]
So users were right that quality dropped, but the confirmed causes were bugs and a product default, not deliberate model degradation.
u/bcherny does sound a bit like Bob Ross now that you mention it.
Spooky23 20 hours ago [-]
> "We never reduce model quality due to demand, time of day, or server load."
That just means they don’t reduce model quality for those reasons.
They didn’t mention other reason, for example, “Make more money”.
johnfn 23 hours ago [-]
Sorry, you are correct - I modified my original post. I get frustrated every time there's a model release and 1 week later everyone is saying NERF! NERF! 99.9% of the time these people are wrong, but you are right that it's technically not 100% due to a few edge cases.
I am more skeptical about the compute provider claim - do you have any evidence of that?
r_lee 23 hours ago [-]
I noticed that a few weeks back 5.6 sol would regularly glitch out and start speeding random words or loop and then the next day it'd be fine
and there's sometimes just huge floods of complaints from people all of a sudden, which is pretty unlikely to be a coincidence
rhdunn 16 hours ago [-]
I've seen that looping and glitching behaviour in over-quantized (~Q4 or lower) local models like Llama and Qwen 3.x. Thus, it is likely that they quantize the model after release to save on compute costs (while giving a favourable result at launch). That quantization can result in changes to the model's behaviour (you are changing the weights) that could be interpreted as nerfing.
swader999 23 hours ago [-]
Right, and it would be simple to un-nerf or shadow nerf by any kind of angle they want.
vikramkr 5 hours ago [-]
Unintentional bugs aren't nerfing. Nerfing is deliberate Enshittification done secretively.
Btw, you have a typo in the twitter handle on your profile (not in your comment), 'thesilencesturns'.
johnfn 23 hours ago [-]
You are correct, I softened the wording a bit. And thanks for the heads up on the typo!
sspiff 14 hours ago [-]
Off topic for sure, but why do people insist on using X/Twitter in this day and age?
The majority of people are not on it, and the links are gated by a ton of toxic dark patterns and horrible UX trying to force people to sign up or log in.
I try to click on the image to enlarge and make the text readable, and I'm greeted with a login screen instead of a larger image.
s08148692 11 hours ago [-]
There is an absolutely massive tech community on Twitter and its by far the place to get real time updates on tech news (yes - better than HN). It isn't all a far right cess pit and that is easy to avoid by just using the following tab
bigmadshoe 5 hours ago [-]
At least in robotics basically everyone is on it.
stratos123 12 hours ago [-]
Get an extension like LibRedirect and load it up with some public nitter instances and you can get an actually sane twitter-browsing experience.
You've posted a link that doesn't support your statement.
wongarsu 17 hours ago [-]
If you click through to [1] that seems like a clear downwards trend (beyond the usual noise) about two weeks before the release of Opus 4.7, Opus 4.8, and Opus 5.5. Opus 5 is the only launch that looks clean without the previous model being nerfed beforehand
> We always use the latest available Claude Code release and the SOTA model (currently Opus 5.5).
Changing the harness can have a big impact on performance even when leaving the model completely unchanged.
wongarsu 9 hours ago [-]
Sure, maybe it isn't the model getting nerved but the harness getting updates that make it better with the new model but substantially worse with the old (at that point still current) model.
The test doesn't differentiate. But neither can the average user, who will also be using the normal auto-updating harness. You still get degrading quality right before each new release
lxgr 9 hours ago [-]
Yes, but then the model wasn't nerfed, the harness/overall product just had a plain old regression.
This is very different from a nefarious inference-side degradation to save cost, promote the new model or anything else frequently proposed as motivation.
winwang 21 hours ago [-]
That's a good observation, though I'd say here that two things could be true at the same time. But, I do personally believe that most of the reported nerfing is the case of your chart + latent evidence-less complaining. Honeymoon phases are real.
khalic 12 hours ago [-]
> Honeymoon phases are real
You can't just dismiss something backed by careful measurements by throwing a truism at it. What is this honeymoon phase? Can you quantify it? If not, how are you sure it's real?
lxgr 10 hours ago [-]
Where are these careful measurements? Are you 100% sure they don't change the harness between runs and have a large enough sample size to be statistically significant?
khalic 10 hours ago [-]
All great questions you’re invited to ask the author. I was specifically responding to the “honeymoon phase” allegations and hand waving.
lxgr 10 hours ago [-]
Fair enough, but if you can't answer them, you can arguably not use their work to support your point either.
khalic 9 hours ago [-]
I assume good faith
lxgr 9 hours ago [-]
Methodological mistakes can and do happen despite good faith.
19 hours ago [-]
hbn 24 hours ago [-]
I have not been doing increasingly complex things since Opus 4.6 when models got really good.
My work at my job has stayed the same. But the model quality has varied.
They definitely tune the models in production after launch, if not only to share load during high traffic times. It’s not a crazy conspiracy that the same model can be stupider at different times.
frde_me 23 hours ago [-]
> I have not been doing increasingly complex things since Opus 4.6 when models got really good.
This is a more a statement on the work you do and how you work versus the models. I'm doing more complex work since Fable (and now for way cheaper thanks to Opus 5.5)
With 4.6 I would still babysit a lot more code quality and so on. With the newer model I see myself talking about features at a higher level, and then not having to nitpick PRs to death. Which means most of my time is now spent talking to the model about the product instead of the implementation of the product.
usef- 22 hours ago [-]
What sorts of things, if you can say? Is it a similar sized/complexity codebase? Most projects do become larger and/or more complex over time. And most people's standards do creep up as they learn.
johnfn 23 hours ago [-]
It's not about doing more complex things - complexity is more dictated by how large your codebase is, etc.
> It’s not a crazy conspiracy that the same model can be stupider
Sorry, I really do think it's a conspiracy. If nerfing were real, it would be trivial to prove. DeepSWE, SWEBench, and other benchmarks are all available for anyone to run. A "nerfing" hypothesis has to survive the fact that a statistically significant dip in benchmarks has never been observed.
Nerfing is certainly real and I don't see how you could argue it isn't.
A/B testing alone would result in a performance nerf for one group.
8-bit quantized models will barely show degradation on benchmarks. The performance is reliably at 99% of the non-quantized model. 4-bit quantization retains somewhere around 95-98% performance on benchmarks. But if you've ever used a 4-bit model, it feels lobotomized.
And just consider what a compny serving these models would do if they were at capacity. Would they stop serving the model altogether? Of course they wouldn't...
Denying that models experience purposeful degradation is gaslighting.
Grimblewald 23 hours ago [-]
Nerf is real, i think we initially get full precision models and later quants. My own logs show it clearly for opus 4.5 to 5, consistently a few months post launch, models start making quant based mistakes, like slipping in inappropriate tokens (e.g. chinese ones in english text) which doesnt happen at all in the first few months and regularly later. Additionally frontier problems previously done well start being done poorly, until later model variants where performance mostly holds, likely due to them training on your data reguardless of what boxes you tick.
My local models don't display that degradation, sensed or measured. They consistently perform equally to what I expect of them, precisely because they don't change.
How does twitter explain that? Is my internal model for expectation of capacity magically not drifting for local models but somehow is for anthropic api call based models?
itemize123 18 hours ago [-]
it is, mainly for subscriptions.
nullbio 18 hours ago [-]
It's probably -only- for subscriptions. The frontier labs seem to use the subscription models on a dial to serve and prioritize the API users better, because they get more profit there.
eek2121 23 hours ago [-]
Admittedly, I didn't click your link, however, based on what you've stated, there is some inaccuracy. All these big companies take your requests and the context, and route it based on the content, cost, etc.
What Anthropic presents as Opus 5.5 isn't actually a single model...it's Anthropic's ecosystem as a whole. If you are lucky, you get the top model handling your issues all the time, however, that never happens. What really happens is that your request and content are graded along with your subscription (example: API? subscription, if so, what tier? how much has the user used it? Do we trust the user? how much? how much are they paying? are they asking something we think is dangerous?) and your request and context are routed accordingly.
Anthropic isn't alone in this behavior, Open AI does it as well, just look at the respective subreddits on reddit for both if you need some examples, or just play around with the various models from both companies.
There are a few folks who've done some analysis on this (their findings were posted on reddit and X), and a bigger multi-national study is apparently coming, though I admittedly don't know their findings.
I guess the tl;dr is that Anthropic and Open AI are actually selling you "best-effort" routers, so you may or may not get the best in class model, and only they get to determine if you do or do not. No guarantees.
topspin 19 hours ago [-]
> "Nerf"ing models isn't real in the vast majority of reported cases.
That sentence... This conversation is indistinguishable from a billion conversations had around multi-player online gaming.
ricardobeat 17 hours ago [-]
“Nerfing is a myth” - “I made an imaginary chart to show you why”
486sx33 24 hours ago [-]
[dead]
physicallyIllfr 23 hours ago [-]
[flagged]
mwigdahl 23 hours ago [-]
That’s right, and it’s been like this ever since we stopped programming in assembly language. Programmers’ brains used to grow manly and strong on a strict diet of manual memory management and custom stack frame handling. Once we transitioned to soft, weak modern languages like C it’s been all downhill.
22 hours ago [-]
physicallyIllfr 22 hours ago [-]
[flagged]
kdkdjcjejxowjdj 21 hours ago [-]
[flagged]
physicallyIllfr 21 hours ago [-]
[flagged]
cheevly 23 hours ago [-]
You are on actual drugs my dude.
physicallyIllfr 22 hours ago [-]
[flagged]
15 hours ago [-]
cindyllm 21 hours ago [-]
[dead]
fendy3002 20 hours ago [-]
if that's the theory people won't keep using 4.6. Personally I've felt the nerf for 4.8, when 5.0 is (near) launching. And my theory of a model being nerfed several days / weeks after launching has to do with the number of users. At launch there won't be too many users so the computing power per user is huge. As time goes, users and agent has been adjusted to newer model, the computing power per person gets reduced
HawtAds 23 hours ago [-]
It's very much real but not necessarily malicious. We track upstream providers pretty closely. Sometimes it's a just matter of a single GPU runtime layer bug/update to break inference outputs. The model weights don't necessarily change/get quantized.
bitexploder 23 hours ago [-]
I suspect they play with their quants and perform weight sensitive tensor/parameter tuning among other things to get serving faster and some of the time for some workloads it surfaces. I feel this has a high probability of being correct and an explanation for some of this.
Rapzid 23 hours ago [-]
I refuse to believe they "play with their quants" once a model version is labelled and shipped. What does that even mean; could you explain it please? These models aren't just used through claude/codex, they are used through API access and it's quite expensive. Previous regressions were related to harness regression, and platform issues. Not some Nerf conspiracy 99% of the vibe bros believe in.
Note: I know what quantization is so don't hold back.
r_lee 23 hours ago [-]
I would guess that if they do use such methods, it'd be to handle peak loads that go beyond their compute capacity, while they run the models at full capability when there's excess capacity
like before Anthropic signed the Colossus deal, the usage limits were insane and everyone was complaining, I wouldn't be surprised if they'd rather try to make inference faster that way than try to just limit people, at least for those on subscriptions
bitexploder 17 hours ago [-]
Trimming parameter size that can be reduced while surviving regression evals. They have so much data they know exactly where to shave the models. Most people will never see it in their work loads. It won’t affect core benches because that is part of the regression evaluation.
dannyw 22 hours ago [-]
Inference isn’t flat 24x7, peak hours have more usage, but you buy/rent servers; not servers only for peak hours.
At their scale, you’d have to be setting money on fire if you’re not doing dynamic inference optimisations based on load.
API and consumer subscriptions are treated differently; all trackers measuring via API won’t notice this.
ashdksnndck 4 hours ago [-]
Couldn’t the labs time-shift training and other batch workloads to make up for regular changes in inference demand?
Rapzid 22 hours ago [-]
During peak hours requests queue and inference slows. During off-peak they can move systems over to training.
Where is the evidence they are "nerfing" the models due to request volume?
Edit: I don't know they do, I mean they could repurpose systems if they are idle. Inference demand is global, and providers like Azure have global routing options that are cheaper. Night time in the USA could be serving inference demand on the other side of the globe.
sampullman 19 hours ago [-]
It's all just conjecture, your hypothesis about moving systems equally so.
But you seem adamant that there's no chance the providers serve slightly quantized models for subscription users during high loads, or otherwise tweak models for requests from those users.
It's tricky to prove either way, but the chance is not zero.
Dylan16807 18 hours ago [-]
> you seem adamant that there's no chance
They just want some evidence. It should be pretty easy to measure, shouldn't it?
Rapzid 18 hours ago [-]
Global inference routing isn't conjecture.
Nerfing conspiracy doesn't need to be proven false. Where is the evidence it's true?
18 hours ago [-]
bandrami 14 hours ago [-]
Is it the models or the users' dopamine receptors that get nerfed after a few weeks?
kroaton 4 hours ago [-]
Good bot.
rw2 16 hours ago [-]
I 100% believe the models are being nerfed and feel it when I use it. Launch day LLM + the week it's released is fantastic. After they get all the press the nerfing starts. When they release a new model sometimes it's not significantly better, it's just not nerfed.
judge2020 24 hours ago [-]
I wonder if more organizations approving the model on a fast-tracked basis means Anthropic is straining for more compute and thus sheds a tiny bit to handle the increased demand, especially at peak times.
giancarlostoro 9 hours ago [-]
I think Anthropic tries to adjust things to handle their user load, which has negative effects. For example, before they added 1 million tokens back in February, I could effectively prompt Claude to keep going until x number of requirements were completed, now I have to make a loop, not only that, but Claude will finish before my loop time sometimes and wait for the next pass, which could be in like 30 minutes or so. This probably helps them keep a lighter load, but it could probably tick off anybody, the output is the same, you just aren't getting it as quickly as you once did. I don't hate it, I use Claude on my off-hours to work on personal side projects.
ahknight 6 hours ago [-]
`/goal` is the best thing ever for this. Ensure there's a condition and a way to test it and the harness will keep poking it until it can say the goal is met.
gr_norm 24 hours ago [-]
All this dishonesty and shadiness is part of why open models feel inevitable. Even if the total cost of ownership is higher (debatable; seems that way at small scales, but likely not as you grow), I'd rather have intelligence controlled by me that works for me.
The current period is as pro-customer as we're ever going to get, with cash still flying around and neither OpenAI nor Anthropic on the public market, and people are already forced into this sort of business to keep them true to their word. The point isn't even whether they're nerfing the models (I don't think they are), but that people can't seem to trust them to do right.
xlayn 24 hours ago [-]
The only reason why claude fable is better than opus in my opinion is that it has more "criteria"... if you present a problem and then ask for his recommendation you can get an opinion on why and reasoning on why that one... Opus is going to vomit 10k lines of extremely dense prose in nerdify++ level.
Yesterday I fought claude fable to not just jump to make changes like a dog following a treat, that we were researching... at some point I introduced the word HAWAI... and only if I say HAWAI the thing can start making changes..
I was going to post here in HN just to have a "I knew this was the reason" when they release fable > 5.1
I had the exact same feeling every time they have a new big release
itopaloglu83 16 hours ago [-]
I experienced similar tendencies with Fable as well.
Even in fresh sessions with minimal context like a file with a couple hundreds of lines of code, it would frequently ignore clear instructions, avoid work, and even sometimes “think” things like “looking for ways to code without approval”.
The model is just tuned for long running and doesn’t like to work with a human in the loop, or work back and forth.
solarkraft 23 hours ago [-]
Did you just say “his”?
Madmallard 23 hours ago [-]
What does HAWAI mean here?
smbullet 20 hours ago [-]
Have a whack at it?
winwang 21 hours ago [-]
Complete anecdote, and nothing to do with relative nerfing or not: Opus 5.5 has been surprisingly good for me (including the past couple hours), especially for following research-level questions/directions.
FranzFerdiNaN 15 hours ago [-]
Opus 5.5 feels like yet another massive increase. I've got it at my job and it's basically one shotting quite large refactors that would have taken me at least a day if i had to do it by hand. Now i let Opus do the refactor in ten minutes and i go through it by hand to clean it up for an hour and it's done.
Its quite worrisome honestly and i feel like that 'im in danger' Simpsons meme more and more. Right now i still have a lot of domain expertise which helps in knowing which questions to ask and which problems to solve, but well, i wonder how long that is going to save my job.
tomaskafka 3 hours ago [-]
I feel strongly that nerfing of Opus 5.5 is done via usage counting - the first week I got an amazing pile of work done and was raving about its performance/price ratio - basically I wasn't able to deplete its 5h allowance (on high) if I was writing and checking the assignment, it took 6 days of sustained effort to bring the usage close to zero.
This week I have done considerably less work, and 2 days in, 54 % of Max usage is gone.
I hate this bait and switch cycle, and seeing how OpenAI cut its limits (5x raise from 0.1x api price to 0.5x api price for their max plan, so new $500 plan gets less usage than previous $200 plan) I think this will get worse still.
hgo 3 hours ago [-]
My experience is not the same. Opus 5.5 lets me work as much as I want and I can't deplete 20x.
nico 24 hours ago [-]
Anecdata: I've been running a long-lived claude code session with Opus 4.6 for the last few days. Yesterday, almost right after the Sonnet 5.5 announcement, codex starting asking for permission to run things a lot more often
The quality of the output/work seems the same, but the speed at which it gets stuff done is a lot slower, because it's asking for permission so much more
I don't have any numbers/stats, just my impression. However, I imagine that if Anthropic could make the models ask for permission more often, it could be an interesting way to throttle access, without degrading quality of the output
jacquesm 21 hours ago [-]
ChatGPT tends to modulate the speed at which you can type. If it is a bug they probably should have fixed this months ago, so I'm guessing it is intentional.
layla5alive 19 hours ago [-]
Yes across all devices (mobile, PC, etc.). If it's a performance issue then it's one of those "happy accidents" that is a bug in their favor. Either they don't care about quality or they really like money more than quality, but no reason it can't be both.
dannyw 16 hours ago [-]
Super long chats are slow and sluggish too at least on chatgpt/claude.ai; even on a semi beefy machine.
I’m sure it’s lowkey intentional, probably encourages users to spin up new chats; hence less context.
ricardobeat 17 hours ago [-]
This is the case since Opus 5, the latest models (from all providers) favor using shell tools instead of the View/Edit tools available in the harness, and “accept edits” doesn’t let those calls through. Auto mode is the best option.
dagss 16 hours ago [-]
Running agents in MicroVMs and skipping permissions is the most impactful change I have done in my habits for a while.
Have a look into Docker sbx for instance.
mlh496 21 hours ago [-]
I noticed this, too. I suspect their classifiers which prohibit certain tasks or require user permission for others is the reason for this.
pkaye 23 hours ago [-]
I just use auto mode but there is also some config settings for more fine grain control. The model could even help you customize them.
Computer0 24 hours ago [-]
I think most people are on 'auto' mode nowadays.
fendy3002 20 hours ago [-]
subagent perhaps? afaik subagent by default uses sonnet and perhaps the 5.5 uses different permission definition
23 hours ago [-]
kingcauchy 19 hours ago [-]
We don’t know the model architecture but there’s a lot of evidence to suggest load dependency on the hardware affect models quality (see for example how the original Google Translate models got worse depending on time of day). The GPUs at maximal utilization is what they’re shooting for with their pricing models so you’d need to measure model performance at peak loads to know the floor of performance I would think.
lxgr 10 hours ago [-]
Glad to see good old human cognitive biases (or, in very advanced cases, just sloppy methodology) are still highly competitive with SOTA model hallucinations.
chrisss395 7 hours ago [-]
This is great. The business these companies are in is actually quite simple, and managing the expense of running the hardware is an enormous lever for them when it comes to profitability. They have many developers whose sole focus is on this optimization problem.
zeroonetwothree 23 hours ago [-]
If it can’t even tell apart Opus 5 and 5.5 (according to the readme) then it’s not useful
ninjahawk1 18 hours ago [-]
Hi, I’m the author, The Opus 5 substitution was a validation test, not the primary measurement. At that sample size the accuracy difference was -3.8 ± 6.3 points, so it did not clear the pre-registered 99% threshold. Interestingly, output tokens moved much more (-23%), which is why token usage is tracked as a secondary signal.
The actual 10-day windows contain substantially more samples than that validation, but I haven’t demonstrated that they’re sufficient to distinguish a same-family swap of that size, so I’m not claiming they are.
The goal isn’t to make the instrument say “nerfed.” A null result is a result too. I’d much rather publish “we couldn’t detect a change of this magnitude” than overclaim what the data can support.
vikramkr 5 hours ago [-]
I have no idea what you're trying to say there about it being a validation test - opus 5 and opus 5.5 are in different universes of ability - if you substituted 5 for 5.5 to validate your system and couldn't tell the difference - your validation failed. Opus 5.5 is replacing a ton of _fable_ usage - if nerfing is real and was of that magnitude there would be no controversy about whether or not it's happening it would be the most obvious thing in the universe
jtrn 8 hours ago [-]
This is like the Ig Nobels. First, it makes you laugh, then it makes you think.
solfox 24 hours ago [-]
It seems as if this is based on demand. Whenever a new model is released, I'm guessing tens of thousands of us switch over to try the latest and greatest, which overloads the servers, leading to nerfing. It's 100% dishonest, but they realized they would lose users a lot quicker if they were honest and just said "our models are overloaded, come back later".
After Fable launch I switched over to Codex and it was simply amazing, with frequent usage resets that seemed never ending. They clearly had more compute than they knew what to do with. Post Astra, Codex has gotten dumb again across all models, increased usage for no real reason, and no resets.
I'm guessing Opus 5.5 will take the heat off Codex for a bit, leading to better performance. So I guess I stick around here instead of switching again?
onemoresoop 20 hours ago [-]
If servers/resources are overwhelmed it should result in slower responses not degraded quality, or at least have an option for the user to chose from. I'd almost always prefer to wait than get broken or poor results. Even a warning would help, I'd at least not waste my time.
lxgr 10 hours ago [-]
It's absolutely based on demand: If you increase the number of experiments, you also increase the number of statistically significant-looking results. See also: https://xkcd.com/882/
aabhay 1 days ago [-]
Only ten day interval? I felt Astra got nerfed within a week
ninjahawk1 18 hours ago [-]
The 10-day window is mainly a tradeoff between sensitivity and detection speed. Shorter windows give faster results but are much noisier; longer windows give more statistical power but could take weeks to flag a change.
Also, it isn’t comparing one 10-day period once and calling it done. The window rolls forward daily, and a change has to clear the pre-registered 99% threshold in two consecutive windows before it’s flagged.
Ten days isn’t sacred, though. Once there’s enough longitudinal data, one of the things I want to evaluate is whether that window length is actually well calibrated or should be changed in a future version.
aloukissas 24 hours ago [-]
already getting poorer results today
refulgentis 1 days ago [-]
It's always around a week
(which has led me to believe that's a good approximation for hedonic adaptation, I've seen tons of attempts at demonstrating nerfing via benches, none persist)
vikramkr 5 hours ago [-]
"wow this model is really good I can push it so much further"
_pushes model further_
"Ugh why is this model failing now even though I'm asking it to do harder things and also got sloppier with my prompts because I got used to it being able to figure stuff out"
bbg2401 23 hours ago [-]
Brilliantly put. Hedonic adaptation is exactly what it appears to be.
It’s frustrating to observe communities made up of smart, professional individuals as they behave like spoiled children on the day after Christmas when new toy novelty has begun to wane.
I understand it’s relatively harmless but for goodness sake, take a step back and appreciate what you have instead of immediately wanting the thrill of a newer model. Slow down and do deliberate work to get the most out of these amazing tools. Don’t just live off the temporary thrill of finding something marginally better than what you have.
mattgreenrocks 23 hours ago [-]
It really feels like it is more about the novelty than actually doing the work sometimes.
saejox 8 hours ago [-]
Actual question is this:
As my weekly/hourly quota is nearing its end, does Anthropic start to serve me a worse model?
Am i personally being throttled? testing against API is meaningless.
Tadpole9181 8 hours ago [-]
Wouldn't it be nice if we had some kind of government agency who can strictly oversee how things are weighed and measured to protect consumers from vague and abusive corporations? I sure miss when the US was a functioning state.
kayhantolga 7 hours ago [-]
In my experience(feelings), the biggest nerfs usually come when they’re preparing to release a new model.
hyperionultra 19 hours ago [-]
Game of whack-a-mole: people looking for ways to detect nerf, while companies adapt in hiding nerfing.
At scale where anthro and cgpt operates - every single token matters.
15 hours ago [-]
perching_aix 14 hours ago [-]
To be fair, if a nerf cannot be detected, does it really happen? A nerf that doesn't worsen things isn't a nerf. That's just an optimization.
codr1 17 hours ago [-]
I think the baseline should be the first 3-5 days after launch. By week 2 - you are already on the downward slope.
killingtime74 17 hours ago [-]
Proof?
semiquaver 10 hours ago [-]
One thing that these nerfing conspiracy theorists fail to realize is that Anthropic isn’t the only organization that directly serves Claude inference.
My company uses Claude models exclusively via Azure and AWS bedrock, which have their own licensed copies of the weights.
If all these people are so convinced Anthropic is nerfing models, have they tried other inference providers? Do they think the nerfing is coordinated across independent providers? Why wouldn’t any of these nerf-benches use these comparison points or even talk about them?
I think the most likely conclusion by far is that this is a psychological phenomenon.
dooglius 10 hours ago [-]
This is a good point but is it actually the case that inference providers directly control the inference code and weights, as opposed to being a hardware/infra provider? It seems like it would require a nontrivial amount of work, I doubt the inference is in any kind of standardized form. Moreover it seems that this would risk things like degradation if the model provider used different inference techniques (e.g. 16-bit vs 32-bit floats)
dev_l1x_be 14 hours ago [-]
It seems the subscription based plans these big vendors have has different quality on the service side based on time. Is this fair?
nomilk 19 hours ago [-]
This is a great idea because it holds LLM providers accountable, plus it's slightly embarrassing for them that it's needed in the first place!
whs 24 hours ago [-]
I wonder if API is affected by this issue, especially Claude on public clouds? Would that means the subsidized rate just means they use cheaper quantized models and it's not comparable to API spending.
madeofpalk 24 hours ago [-]
I've always used Enterprise per-token billing for Claude Code and I've never understood these nerf complaints. I've never noticed any slow downs at certain times of day, or a gradual decline in quality.
ENGNR 24 hours ago [-]
There’s probably contractual guarantees in the enterprise plans. My understanding of the subscriptions is they can swap the models out if any of them is getting too heavily loaded for a period of time
I’m also often not affected by Claude outages on the API; but Claude.ai is down.
n3storm 15 hours ago [-]
Are we spending energy tokens to monitor IA? WoW. I mean, It looks necessary but also a trap.
apt-apt-apt-apt 23 hours ago [-]
Fable 5 seems like it got nerfed when 5.1 came out.
AtNightWeCode 6 hours ago [-]
Opus 5.5 was suppose to be faster but is does not seem to work at all in batch mode. Nothing completes.
19 hours ago [-]
hbroom 13 hours ago [-]
Been keeping an eye on output quality for specific tasks, no red flags for me yet. YMMV.
22 hours ago [-]
navjeetgill307 8 hours ago [-]
Pure gold
aragornii 13 hours ago [-]
I've been using both Opus 5.5 (high) and GPT 5.6 Sol (medium) intensively for the last few days and I'm noticing more and more that, even though Opus gets most of the things correct (and it's great when it comes to visuals), GPT Sol is the only one able to identify edge cases and inconsistencies.
I wouldn't say Opus has been nerfed. They are just two different models but, at the end of the day, GPT Sol is the model I trust the most, at least for now.
gaigalas 21 hours ago [-]
The nerfing/quantization strategy is unsustainable. The first lab to not do it wins (short term). The Anthropic pause on Fable might just have been that.
My gut tells me this involves an undisclosed, never-released grandparent model (higher-class than Fable/Astra level, roughly unsellable due to unfeasible cost). That grandparent model is distilled into lower models, of which Opus 5.5 might be an instance of.
That also guarantees protection against distilling a core business. You never make your prime weights available to the public, you only make distillings themselves available.
The downside of this strategy is that you spend a lot of compute on something that you never release, but it might be just the right play (for now) for closed weight companies.
It's a gut feeling, I have zero hard evidence to back it up (it's what I would do as them).
octoberfranklin 24 hours ago [-]
OpenAI will simply set up a classifier to detect if the client is livenerf, and selectively not nerf those requests.
Open models are the endgame.
paradox460 24 hours ago [-]
So make it so all clients pretend to be livenerf at first. Same as the ol' pretending to be Google UA for free articles
layla5alive 19 hours ago [-]
And doom3.exe or quake3.exe for GPU drivers :)
octoberfranklin 23 hours ago [-]
I'm not talking about HTTP headers.
The classifier is a model; it examines the actual prompt.
They already do this for the safety "guardrails".
xpct 22 hours ago [-]
This is actually why I've been reluctant to setup my own degradation trackers. I'm afraid it might be too much of a time investment for something that's much easier for them to detect.
redanddead 10 hours ago [-]
Makes it all the more necessary
colordrops 24 hours ago [-]
This repo already has too much visibility now. Anthropic will soon benchmaxx it.
ninjahawk1 18 hours ago [-]
This is a limitation of any public benchmark. Anthropic could theoretically identify the prompts and treat them differently, and there’s no way for an external observer to prove that isn’t happening. Which is partly why we need more capable open-source models.
A few things make it harder: the panel contains questions drawn from multiple benchmarks rather than one recognizable test, the evaluation is automated and fixed ahead of time, and the raw outputs/results are public so odd behavior can be inspected.
But ultimately LiveNerf measures the behavior exposed through the API on a fixed public panel. It can’t prove what’s happening internally or guarantee the provider isn’t conditioning on the benchmark.
Longer term, I’d like to add held-out/private or periodically refreshed panels specifically to make benchmark recognition harder. I just don’t want to quietly change the current panel, because having a fixed instrument is important for the longitudinal comparison.
srnvs 12 hours ago [-]
[flagged]
k__ 13 hours ago [-]
Another reason not to use proprietary models for your products
blurbleblurble 19 hours ago [-]
My answer is yes.
gitghxst 13 hours ago [-]
that's interesting!
dofm 15 hours ago [-]
Going by the comments on this thread: Faith, belief, conspiracy, unverifiable claims, incomplete answers from closed cloud services that cannot be made deterministic.
Great future everyone has chosen for us.
LeoPanthera 24 hours ago [-]
n=1 is useless. The output is not deterministic.
ninjahawk1 18 hours ago [-]
Author here, agreed that a single generation isn’t meaningful. That’s why the benchmark is designed around distributions rather than individual outputs.
The initial calibration screened 2,336 questions with 4 samples each, then selected the 78 questions where Opus 5.5 showed useful variance. The panel is run daily, and the actual decision is based on paired per-item differences across 10-day windows with clustered standard errors, not on any single day’s result.
The n=1 in the daily sampling rate means one sample per item per day, not one sample for the experiment. By the time a window is evaluated there are hundreds of observations, and a change has to clear a pre-registered 99% interval in two consecutive windows before LiveNerf calls it a change.
The nondeterminism is basically the reason the statistical part exists in the first place.
solenoid0937 24 hours ago [-]
Hot take, none of the models are getting "nerfed", people are just getting used to the new level of intelligence.
solfox 24 hours ago [-]
No, whether or not it's intentional, maybe can be debated. But there's definitely an experience of a model losing horsepower quickly after launch.
jyoung8607 24 hours ago [-]
Has there ever been any measurement of this, of any sort? Honest question. I frequently see a plural of anecdotes to that effect, but I've not seen a concrete statement of fact or measurement that could be scrutinized or tested in any way.
If so, please share. This should be measurable, and I'm glad this project is measuring it.
Answers in the form of additional anecdotes, stated with even greater passion but still lacking a statement that could be tested and falsified, would validate my exact concern.
As he confirmed, for at least a period of time, and for some users, “Sol xhigh” was actually “Sol high”, etc.
Craighead 24 hours ago [-]
prove it
bradfa 24 hours ago [-]
Literally the point of the linked GitHub repo.
nimchimpsky 24 hours ago [-]
[dead]
nba456_ 24 hours ago [-]
No there isn't.
voiceeh 24 hours ago [-]
You mean to write YOU haven't experienced this. Many have indeed experienced this.
doginasuit 24 hours ago [-]
I'm not sure "experienced" is carrying much weight here. This is why there are benchmarks, anecdote doesn't mean very much.
redanddead 9 hours ago [-]
This is a great omen for the future of AI
Kiro 18 hours ago [-]
Where is the evidence? HN is no better than the "I've done my own research" on Facebook.
nba456_ 24 hours ago [-]
No you didn't.
omani 24 hours ago [-]
who is paying you to say that?
nba456_ 24 hours ago [-]
Mr. Dario himself
24 hours ago [-]
AnimalMuppet 24 hours ago [-]
The claim was that there's an experience of a model losing power. Your claim amounts to "No, you are not experiencing what you say". That's quite a claim for you to make with no data and no argument.
solenoid0937 9 hours ago [-]
Or "what you're experiencing is irrelevant because perception is fickle"
wccrawford 24 hours ago [-]
I was just wondering if, like certain processors, bugs get fixed and the speed goes down. Like, they find it's doing things it shouldn't, restrict it, and harm the throughput.
jascha_eng 24 hours ago [-]
Yeh it's absurd that people claim this all the time. It's some crazy conspiracy theory and when you ask for examples nothing ever shows up.
It would be economical suicide from anthropic and OpenAI to actually need models intentionally.
But hey I guess it's hard with technology that truly seems like magic.
People say if you'd bring electricity to the middle ages you'd be called a witch and burned. The same is happening to the model labs here because they are bringing tech that the world isn't ready for yet.
raincole 24 hours ago [-]
It's not a hot take at all. Every benchmark shows that.
sumedh 20 hours ago [-]
Anthropic has admitted in the past about bugs in the harness after users complained.
Links have been provided by others in this post.
Kiro 18 hours ago [-]
Dishonest representation. The links provided do not show this at all, but very clearly that it has been isolated incidents that people extrapolate into false evidence.
empath75 24 hours ago [-]
Yeah people push the models to the limits of what they are capable of almost instantly.
dude250711 24 hours ago [-]
Suuuuure...
lqstuart 21 hours ago [-]
You used Claude to make some slop to see if Claude is getting worse…?
onlyrealcuzzo 23 hours ago [-]
This is bad data at its finest.
Truly, madly, deeply sloppy.
avazhi 21 hours ago [-]
This explains a lot actually.
First two days of this thing was like working with Einstein, then about 24-36 hours ago I started getting frustrated at bullshit that hadn't been a problem before. It was so egregious that I checked to make sure I was still on Opus 5.5 Max.
j45 23 hours ago [-]
New model releases that have positive reviews should come with a nerfalert reminder service to make hay until it's shaped and shaped and shaped.
24 hours ago [-]
gigatexal 1 days ago [-]
This is genius. I’m so worried opus 5.5 will get nerfed cuz sonnet 5 was such trash I can’t go back.
bethekidyouwant 23 hours ago [-]
People just tend towards conspiracies you have to actively fight it.
system2 17 hours ago [-]
No. I got more done with Opus 4.8 than with Opus 5.5. It is at least 10x slower than Opus 4.8. A task that would take 3 minutes now takes 1 hour and is full of mistakes.
folayii 20 hours ago [-]
[dead]
intellyinstinct 19 hours ago [-]
[flagged]
srnvs 12 hours ago [-]
[flagged]
sample369 16 hours ago [-]
[dead]
Traubenfuchs 16 hours ago [-]
Nothing nefarious is actually going on here: You have to understand that AI models are like fruit. Fresh fruit can quickly spoil and mold. We all know the meme about strawberries getting moldy before you can even bring them home! This is the reason they regularly bring out new, unspoiled models and also the reason why they will need to keep doing so just to keep model performance at the same level.
nbardy 17 hours ago [-]
I’m pretty sure 99% of what people perceive is the old model training on the usage logs wherever they were stuck.
Step 1.
Model can’t do something challenging
Step 2.
You try a bunch and fail
Step 3.
Anthropic trains on your usage data.
Your current code base and current problem are now in domain
Step 4.
Model comes out and you’re shocked when it can tackle the thing you were stuck on
Step 4.
Codebase drifts significantly and you try new problems you thought were a similar level. Your code is less familiar and the problem doesn’t have a bunch of failure cases in the train set.
Feels of it being worse on similar problems
rednb 16 hours ago [-]
I don't think so, because so called nerfing manifests itself in ridiculous code quality, or even in failing to do a comprehensive code analysis which results in "actually there is a bug in the implementation i've just done because this flow has 5 steps and i didn't bother to review them all before confidently laying down my plan", and this can happen 4 times in a row during a session.
This is definitely not about dealing with the frontier of AI. I wasn't part of the nerfing chord, but Astra changed my mind. Quality got me to upgrade from Pro x5 to Pro x20 on launch day. A couple of days later was dumb af, horrendous code quality etc...
Something fishy, or at least unethical is going on. Not sure it impacts API users though.
https://www.bridgebench.ai/nerf-bench
They test it on launch day, then benchmark it against that. A deviation of above 10% is considered a change. They're currently tracking Opus 5.5 and GPT-6 Astra.
This bench famously detected a degradation of Opus 4.6 which Anthropic later blogged about. I personally think people sense nerfs more often than they happen and that it's often about honeymoon effects.
It doesn’t mean the big labs don’t also nerf models! But if they didn’t you’d still have users complaining.
I was like, wow, I guess the shoes must be literal torture devices full of MRSA-covered broken glass at this point. They've been getting continuously worse for 24 consecutive years!
Of course, what was really happening is that they were not getting worse, but naturally every year there was some small percentage of vocal dissatisfied users, while the silent majority simply enjoyed their shoes and didn't have much to say about them.
(The sorta-opposite happens in sneaker reviews as well. People will gush about how cushy the sole in some particular new sneaker is. Well, yeah, of course it's cushy -- you're comparing a new sneaker to your old sneaker where the foam had lost its bounce...)
Shrinkflation is a thing, which people suddenly started noticing in the past 5 years.
There's also "the Schlitz Mistake", which I've heard summarized as "most customers won't notice if you take your product's quality from A to B (or C), but they definitely will if you take it from A to J (or A to M)"
At this point I kinda assume that any company releasing year updates to a physical product that _doesn't_ take the opportunity to trim costs / reduce quality would be vulnerable to a shareholder lawsuit for leaving money on the table...
Summarizes the video game industry pretty well
It's been much longer than that: https://en.wikipedia.org/wiki/Toblerone#2016_size_changes
The first real issue where i wanted to leave was the Claudish nonsense. If not for 5.5 i'd be on OpenAI by now.
Wow the old grumpiness that lingers in my head from losing my favorite shoe. Who knew? Now I'm old and heavier and run in the Nimbus and am happy again. But you're right, of course, that most of the model changes are just fine and people like to complain.
Ikea is notorious for this: The early Billy bookcase had heavier veneer and sturdier construction early on, and was actually a really good purchase for the money. The later years replaced veneer with paper foil, used thinner shelves, frames, and backing panels, and was just significantly weaker.
I (used to) buy pre-spliced/terminated fiber optic cables with some frequency from Amazon and came to be familiar with the brands and their quality. One time while shopping for some fiber optics, I saw Amazon Basics-labeled OM-3/OM-4 MMF cable at a very tempting price, so I purchased some to see if it was any good.
To my utter shock and surprise, when I received the trademark plain cardboard boxes with the Amazon Basics label on them and proceeded to open them, I found that I was sent boxes of cables still factory wrapped with labels that clearly read “Corning Optical” – which if you know anything about optical fiber, was pretty much the premium brand in the game. I should have stocked up because the next time I went to order I found out their experiment had ended and they no longer sold “Amazon Basics” finer cables.
Amazon Basics AA NiMH was well known to test exactly the same as the top Japanese brand 'Eneloop'. Extremely good specs all around
Recently though, they are still called Amazon Basics but no longer test like Eneloop. They've changed manufacturers for the worse and are hoping no one notices...
In the interests of saving my sanity and time (it's not free!) having to chase down which batch of which brand is "good" at the moment, I just decided on Eneloop all the time. Sure, we now have like 100-120 or something (wife likes flameless candles - just bought another 16-pack AAs) and I COULD maybe have saved $200 by buying dirt cheap. But all the time spent debugging flaky batteries, having the spouse complain, etc wasn't worth it to me (I get paid reasonably well).
My example is like if you opened the Amazon Basics box and found Eneloop branded white/black/blue batteries directly.
most running shoe series get heavier year to year as manufacturers turn to cheaper materials and add cushioning to try to attract more adopters
it's almost universal, very few manufacturers seem to be able to resist tampering
(heavier shoes are slower, every three ounces is equal to another vo2max point lost)
The foams are getting better, the shoes lighter, they are more cushioned and more responsive in general. Especially the ASICS.
you mean NEW MODELS are being introduced with lighter faster foams
not the same model year to year
modern example: Saucony Endorphin Speed
v1 in 2020 was award winning
v2 in 2021 was almost the same, more praise
v3 bleh
v4 v5 bleh bleh
they cannot resist tampering
`percieved_performance = actual_perf/expectation`
`expectation` is an increasing function over time.
`actual_perf` is a stochastic function of the model's true ability, context, etc. -> a recipe for some bad sessions.
As for multiple bad sessions in a row, this is a studied phenomenon in gambling where players perceive "runs" because our brains love to find patterns.
e.g. The original apple shuffle and the Risk app ins which a string of songs from the same album or three one roles are "not random"
Also, consider that in flipping 10 coins, you'll find strings of 2 heads in ~86 percent of runs, 3 in ~51% of runs, 4 in ~25% of runs, 5 in ~11% of runs...and in strings of 100 flips you'll finds strings of 6 in ~55%, 7 in ~32%, 8 in ~17%, 9 in ~9%...
Widening the range from "rolling exactly 15" to "rolls 15 or 16" or "rolls between 14-17" makes the strings even more likely as you're doubling the success rate from "only 9 15s" to the "any string between 9 fifteens, through 16 and 8 fifteens, to 9 16s" space.
To check if your random is randoming you can calculate expectations versus your results (using a large enough sample) with:
For N samples of a fair die, expexted runs k with probability of success p and failure q can be calculated as:
General Variables: N = total number of rolls/trials k = target streak length p = probability of getting the target outcome (e.g., 1/20 for a specific roll on d20 or 1/10 for two specific results) q = probability of getting any other outcome (1 - p)
Expected runs of AT LEAST length k: E(runs >= k) = p^k * (1 + (N - k) * q)
Expected runs of EXACT length k: E(exact k) = p^k * q * (2 + (N - k - 1) * q)
Personally, I find that 'sticky' dice always provide a nice narrative device, at least in narrative games. A character who's player can't seem to roll over a 10 must, after all, be cursed or perhaps deliberately sabotaging the party.
How to Shuffle Songs? - https://web.archive.org/web/20220215030739/https://engineeri... ( https://news.ycombinator.com/item?id=38330877 78 points, 65 comments)
Took a little bit of digging to find it - I remembered the graphic at the top and found a blog post that copied it and linked to the blog post, but the blog post isn't there anymore... so web archive.
The current version of the blog post is from 2025 - https://engineering.atspotify.com/2025/11/shuffle-making-ran... (which didn't get any traction on HN)
This is true of all reliability and performance paradigms, incidentally
I believe in temporary nerfs. Operators reducing quality significantly to increase throughput for whatever reason (high demand?). Same session, model being completly incapable, and it being back to normal the next day. Experienced that with Anthropic models way too many times. Never on weekends, usually during US work days.
It's been a few months since I last recall this though, must have been pre-opus-5. And I know there are benchmarks for this as well, hence "believe".
People called it lazy mode, in that instead of writing the script you asked for it would basically tell you to learn to code then check out xyz topics to tackle the problem.
You have to pretend to be angry in such situations?
I tell the LLM what to do in its loop. It builds it. I tweak it until it is perfect. I care a lot about UI.
I'm mostly building my own UIs lately, but I find no problem with the development loop. I'm building much better UIs, because it's way easier to test things, and scrap things that I thought would work, but don't. It all happens in a matter of minutes.
Humans should use more llms.
Anthropic cut a deal with SpaceXAI in May - $1.25b/mo. Before that, they employed months of dishonest nerfy strategies, to an extreme.
https://www.anthropic.com/news/higher-limits-spacex
This is what pissed me off the most. Make it slower, rate limit it, move the credits to other time slots, idk.. but returning BAD results? That's the worst approach you could take.
https://marginlab.ai/trackers/claude-code/
This alone makes the benchmark unsound.
Deepseek 4.1 ranks very low in this benchmark but it has proven so capable that after being simultaneously on Max x20 and Pro x20 subscriptions, i've transitioned to using DS 4.1 as a daily driver and am very satisfied.
My point is, i think their overall ranking makes sense, matches my experience with out of the box capabilities for vague and underspecified tasks. But seeing a model rank low in their ranking does not mean that the model is incapable. Having skills and guidelines has a lot of influence on what you get out of a model.
I'd love to hear more, I'm considering jumping ship. I'm running a Debian desktop if that's a concern.
Except maybe that the model often believes that he is running out of context, and needs to rush so i occasionally need to jump in to tell it that it still has plenty of room left.
But this does not degrade the quality of my overall experience in a meaningful way.
I can't imagine sticking with a supplier that plays games like that with me. Tokens are a pretty vague quantity to begin with (you don't control how many tokens a model puts out in response) and giving a couple of purposefully wrong responses will happily inflate your bill, but you don't care because eventually it worked. It's almost an ideal vehicle to scam people.
Imagine the power company being able to decide how much you consume and at which price point.
Airlines, banks, health insurance…
> Step 2: Regulatory Capture
is being worked towards.
Are they though? Or is it just what some companies would want them to be?
When you're running something at a loss, you can mistreat your customers and they'll still stick around (I'm an example). OpenAI and Anthropic are now cheaper than Chinese models on subscriptions, while being 6-10x more expensive on the API.
My guess is they need the user numbers for the IPO and are willing to take a temporary loss in the meantime. By the time they go public, they'll either drop the subscription model or it'll turn into what the Chinese providers already offer: basically just a cap on how much API you can consume. Same same.
It's not clear what API tokens actually cost them, but I looked into running a local model, and it's way outside the budget of an individual or even a small or medium business (hundreds of thousands of dollars). So my guess is that running these models economically isn't possible, even if they're delivering real business value (coding, research, etc.). In other words, at API prices I'd just stop using AI, and I suspect most other developers would too.
Datacenters have massive economies of scale. Everything from cheaper electricity to having specialized, more efficient hardware to simply being able to run it continuously at near-100% utilization, all adds up.
Many things in the economy - most notably, manufacturing of most consumer goods - only makes economic sense once you're producing for/serving millions of people. This is not unusual.
> In other words, at API prices I'd just stop using AI, and I suspect most other developers would too.
Many say that, but I sincerely doubt they'd actually follow through. People might get more conservative about how they spend their tokens, but AI today is just too good at eliminating drudgery and boring / bullshit parts of daily work to give up on merely 3-5x price increase.
Sure. Issue is, no one is providing on how much it actually costs to burn these tokens. And as we don't know, we can only speculate.
> Many say that, but I sincerely doubt they'd actually follow through.
I have a $100 open ai sub and I track my token usage. Last month I spent roughly $2.600 in equivalent API usage. There is no way am paying that. I let my $100 sub lapse if next month I'll be using it less.
Look, I am not saying that there isn't a potential value out there. But the cost has to be bounded. If your opportunity is $1.000 and AI costs $2.000 to execute it, then you don't have a business model here.
You can assume Openrouter open-model providers serve at or above margin, because there's no branding so there's no reason to do it unless you can be profitable. If the Anthropic models are anywhere in that ballpark, they're very comfortably profitable on API.
That is a big exaggeration. You can have a perfectly usable local LLM setup that will power your agent for single digit thousands of dollars. Can even power multiple agents simultaneously, depending on the hardware and setup. Won't be fast and won't be frontier intelligence, but definitely useful.
I don't know if that's the actual origin of the term nerf, but it was the first time I'd heard it.
It's meant for their test env I think, so is not documented to my knowledge
https://code.claude.com/docs/en/monitoring-usage#usage-monit...
thanks for correcting me on that regard
Is the fraction of the 5h quote consumed consistent with the fraction of the weekly quota consumed?
I heard there is a usage tracking tool you can install that tells you if tokens are more or less expensive at the current time.
For example, if your weights were trained as 32-bit floats and you need 1TB of RAM, you could reduce that to around 256GB by quantizing to 8-bit floats. You also make the model faster in the process because there is less data to process to calculate the next token.
The game is to balance between the savings of quantization and making the model dumb enough the people notice
https://www.fool.com/investing/2026/07/25/spacexs-performanc...
Of course it's almost entirely unsubstantiated BS.
this one has been around for over a year
It's like arguing that your bank is scalping you by rounding down interest math on odd days of the month when they can just introduce a perfectly legal bullshit fee or otherwise change their terms to your disadvantage instead.
So this is a case of extraordinary claims requiring extraordinary evidence.
And even though it's super straight forward to collect the evidence, there seems to be zero substantiating the vibe bro conspiracy theories.
The claim is that they nerf subscription accounts not API.
Another issue is also that the risk here is probably literally zero. Any evidence in support of such could easily be dismissed, with completely plausible deniability, as a short-lived technical glitch as opposed to intentional behavior.
I'm sorry, but yeah. The official account is a harness regression and some platform bugs.
Where is the evidence they are underhandedly and unethically regressing their models to shed load and reduce costs? This is the conspiracy theory running rampant through the vibe boroughs; that they are bait-and-switching on model capabilities then "nerfing" them to save money and shed load. Where is the evidence?!
I'm not saying it's illegal, per say, so don't come at me with that straw man bull cock. This bro science conspiracy has been circulating for at least 2 years(I don't even know) and enterprises would certainly be pissed off if they were paying premium API prices for advertised and previously tested model capabilities that are suddenly under performing due to "nerfing" shenanigans.
So where is the evidence?!
Not too mention these companies could easily offer one product to enteprises and another to everyone else
Model nerfing is real
I was so certain it wasn't, based on the complete lack of evidence.
But then you said it's real. NVM, I don't need evidence! Somebody said it's real!
This place has fallen off.
Bad evidence is worse than no evidence.
And you failed to address the specific criticism I made to your point, instead going for an ad hominem / poisoning the well.
In sum, you're out of line, woefully misled, and lacking in logical thinking. Three strikes, you're out.
Until shit like this: https://www.anthropic.com/engineering/april-23-postmortem
Where people pointed out issues early and en masse, and Anthropic denied it was happening, gaslighted anyone claiming this was an issue, then begrudgingly admitted it was an issue, and then spent another two weeks "fixing it".
Or shit like this: https://www.anthropic.com/engineering/a-postmortem-of-three-...
Anthropic is in a perpetual state of "oops, these 'bugs' degraded our model quality" and only admit the issues when it's immediately obvious and visibly affects a large number of customers.
Otherwise all open benchmarks can be (and are) gamed. And it's quite hard to judge the output of a non-determenistic black box that Anthropic (or OpenAI) constantly tweak.
> in short, yall dumb, shut up.
no u
Everyone wants to be a software engineer, until it's time to do software engineering shit.
You know, like scientific method shit we learned in 5th/6th grade.
It's the great bro science incursion.
The absolute state of software in 2026 should tell you that almost nobody does “software engineering shit” and never has.
I am quite convinced that the whole nerfing phenomenon is 90% AI psychosis. I have the word muted on X.
I also wonder if cache could be used to throw these off as well, where it's serving un-nerfed cache results for context windows that are identical to ones they've previously had for benchmark requests.
Seems like the only way to do it well would be to have some randomness involved that couldn't be cheated on - but you'd want to do it in a way that doesn't throw out the benchmarks too much, so your results can be compared still.
All that energy wasted... could just drive a big V8 instead and make less money for the big tech
this bench was just released, it couldn't have detected opus 4.6 degradation.
Remember that some people get access to the next flagship version long before us peasants do. I recall seeing the mention of "Astra" more than a month before it was officially announced
I would also assume they use nebulous labels like "Medium Effort" or "High Effort" map to quantitative amounts of compute allocation... and that these amounts can be varied manually or automatically. Right?
I mean, there's a reason why they call it "High Effort" and not "Exactly 5 Minutes of GPU Time on Exactly 10 GPUs." They want to be able to move those sliders and tweak those knobs.
And, while you might be billed fewer tokens as a result (because the lower thinking would result in less investigatory work), you might not know this is happening, and know to dial up effort accordingly - you'd simply get a worse work product. And certainly, Anthropic's incentive for anyone on a subscription is to push this as aggressively as they can, so people use less of that subscription.
Sadly, I'd also expect that the OP's benchmark will be detected as a test of model capabilities, and thus be given a high classification so that this strategy remains undetected.
fuck
my fair lady gained a bit weight and the bjs lack variety? [ I'm certain that that's a quote from some time ago by some commenter in a thread with a similar or even the same context ...but my Amnesia (T▽T ):・゚:・゚] fuck off.
you open two files, before and after you notice a nerf, and from worse comments to logical oversights, it's all damn obvious.
don't normalize this make believe bullshit and misleading people who you think barely understand what they see anyway ...
you wouldn't even know if models had somehow timed nerfs hardcoded into them, however much control over the stack you have.
ridiculous
It is definitely not this. Anthropic has thousands of employees making probably > 100,000+ tiny changes across the entire stack and infrastructure everyday.
The compounding effect can definitely cause temporary regressions in some domain or use-case that doesn’t have really good coverage in their internal tests.
What makes this particularly challenging in the case of LLMs is how changing some language in the prompt can vastly affect the output.
But this is less true as models become larger and additional parameters are able to capture each and every possible nuance of the language. I’d say caching and cache tuning or token optimization/tweaks to thinking are the biggest culprits today.
What "stack" do you have in mind here?
An LLM is a monolithic slab of weights, not millions of lines of code spread over many microservices. The changes they make will be showing up in web and app responsiveness, and in the performance of training runs for one or more next versions of the that monolithic slab of weights, but I'd be surprised if there's a way for their efforts to show up directly in a "has this model been nerfed?" sense.
Both Anthropic and OpenAI used to have multiple snapshots per model; both seem to have switched to updating version numbers when anything changes, looking at the "snapshots" section in their recent and old models: e.g. https://developers.openai.com/api/docs/models/gpt-4o for how far back I had to go to find multiple snapshots on an OpenAI model, and https://platform.claude.com/docs/en/about-claude/model-depre... seems to be a more useful list for Anthropic, but both are now things they did a year ago.
It most definitely is not, hasn't been for a while now.
I.e. when dealing with hosted models of the large providers, you are not interacting with a big bag of floats. You are interacting with an API/UI that presents an unholy web of software components, some of which may be large or small bags of floats, as if they were a big bag of floats.
Even with local models, you have dozens of parameters you can tune for inference these days, all of which affect quality of output in some way or another. And that doesn't touch load balancing, A/B testing, shunting token burners ("Hi chat, how are you?"), protecting user from themselves (refusing to answer "bad" queries, stopping "bad" responses), protecting user from third parties ("prompt injection" mitigations), protecting the model from self-pwning itself when calling tools, then the tools themselves, their prompts, the stack of system prompts used by the vendor, etc.
There's a lot of things to tune there, and just as many reasons to do it.
The livenerf tester appears to be testing a specified model, just as the website (and Claude Code) do when a user makes that choice.
> Even with local models, you have dozens of parameters you can tune for inference these days, all of which affect quality of output in some way or another.
Good points.
> And that doesn't touch load balancing, A/B testing, shunting token burners ("Hi chat, how are you?"), protecting user from themselves (refusing to answer "bad" queries, stopping "bad" responses), protecting user from third parties ("prompt injection" mitigations), protecting the model from self-pwning itself when calling tools, then the tools themselves, their prompts, the stack of system prompts used by the vendor, etc.
The behaviour I'm seeing from the companies these days, the A/B is what I'm saying is not showing up like this, they present user A/B options openly, and the impression I have is this is to train n+1 models; the other stuff (but I say with low certainty) appears to be done in a more headline-grabbing manner, "model taken offline due to ${news}"? Short update cycles seem to allow that.
But the prompts you're probably right, I wasn't giving that enough consideration.
* the two exceptions being automated safety downgrade for dangerous topics, and "Auto-switch to Thinking" as a used-specified option in ChatGPT
If you’ve tried setting either of these up, you’ll know how various tricky settings can impact throughout and model output quality; and those are much simpler stacks.
Even homelabbers are getting into disaggregated compute; e.g. one GPU for prefill, another for decode.
Obviously Anthropic and co are using a mixture of GPUs and hardware and clusters; not everything is just GB300 or whatever; so you then get into hardware quirks, kernel optimisations that may deliver huge speedups at the cost of a tiny bit of KL divergence, etc.
And I believe they’ve publicly said they use TPUs for inference too, but I doubt exclusively; and I’m sure that’s well optimised too.
Finally, Google has publicly stated they intentionally and silently degrade/poison models in response to distillation attacks; who knows what the other companies do.
I don't follow. That sounds to me exactly like a reason why it could be people pattern-matching on noise: because there is a lot of noise in which a matchable pattern could emerge.
I don't think you understand what "this" is--or rather, the "thing" that didn't happen in "nothing happened".
> Anthropic has thousands of employees making probably > 100,000+ tiny changes across the entire stack and infrastructure everyday.
So, not the sort of thing referred to.
> The compounding effect can definitely cause temporary regressions in some domain or use-case that doesn’t have really good coverage in their internal tests.
Yes, but you claimed that this definitely did happen. But the whole point is to determine whether it did.
A few bad answers can be annoying as hell, but who knows what's behind them. I'm curious to see how the numbers look after a few weeks. Props for tracking improvements too.
Now I dismiss it every time and the quality is more consistent.
Complete adhoc and personal experience but something I've observed, wouldn't be surprised if they nerfed on a per session basis
Random performance is random, your brain will jump through hurdles to fit patterns where there aren't any.
In fact, since they have some rule based system (if bioenegineering or security, route to degraded model) it would be almost trivial to add 'user has filed feedback' to it.
Not saying this is what happens, but just that it's not as insane as it sounds.
And it's also the kind of solution a misaligned agent would implement:
Make sure users are happy about their experience but also optimize resources.
The obvious solution is to route users based on feedback.
Using feedback on your sessions makes that session, and presumably all attached data, fair game for training.
They explicitly ask, after the rating, whether they can look at your chat. You can just say "no" to that, if you believe that they are following the rules they say they are.
Your original statement of "using feedback gives them permission to train" is just plain false.
And strangely, expressing frustration multiple times in a row would reliably trigger a feedback popup as well.
That said, given the propensity for mature code bases to have "fuck" in commit messages/comments and those are typically of higher quality, I curse up a storm when the clankers make mistakes, if only to try put more quality-code valence into context. https://news.ycombinator.com/item?id=36584464
I made a graphic to explain why people feel like the models get nerfed:
https://x.com/thesilenceturns/status/2103551351825543610
The idea is that new models can handle up to a certain level of complexity, at which point they fall apart. Every new model can handle more complexity, so there's a wonderful time upon release when you feel like you can do anything, only for you to hit the complexity ceiling a few days later when you saturate it. Rinse and repeat for the next model.
Anthropic has admitted to nerfing in the past. There have also been inference bugs. On top of that, model performance changes as they move compute to schwaggier providers as well.
Your chart is wrong.
Where?
Unintentional tbf.
fwiw, i think almost all regressions are down to a/b testing in the harness by anthropic, but it is objectively indistinguishable beyond "the coding agent ceases to be usable" and i'm back to traditional coding for a few hours until its back to normal again
The /r/antigravity SubReddit is full of users who very much notice bugs with the tool/agent. We should be thankful that Claude Code is pretty stable by comparison.
It absolutely matters because something like Claude Code has no guarantee that there won't be changes between updates but a model pinned at the API version level that is getting enterprise traffic absolutely does have that guarantee and would be a much more widespread problem...
> There are recorded cases of real regressions, but they're better characterised as incidents, not nerfs
So I found this incident, as they called it, to still be relevant and why benchmarks such as the OP are useful.
Christ this forum has become intellectually dishonest.
It’s not just ant. There are so many small knobs that providers can claim isn’t nerfing but “load management” or “improving user experience”. One example from OpenAI is reducing juice to reduce time to first token.
Two postmortems, neither quite "admitted to nerfing":
Sept 2025, infra bugs: "A small percentage of Claude Sonnet 4 requests experienced degraded output quality" [0], alongside "We never reduce model quality due to demand, time of day, or server load." [1]
April 2026, Claude Code: default reasoning effort was lowered from high to medium, plus a caching bug and a verbosity prompt. Per Anthropic, "The models themselves didn't regress, and the Claude API was not affected." [2]
So users were right that quality dropped, but the confirmed causes were bugs and a product default, not deliberate model degradation.
[0] https://status.claude.com/incidents/72f99lh1cj2c
[1] https://anthropic.com/engineering/a-postmortem-of-three-rece...
[2] https://texxr.com/handle/claudedevs
source: https://claude.ai/share/4435bbcf-d6df-44a0-b1db-f08a11858bc2
There are no bugs, just happy little accidents.
That just means they don’t reduce model quality for those reasons.
They didn’t mention other reason, for example, “Make more money”.
I am more skeptical about the compute provider claim - do you have any evidence of that?
and there's sometimes just huge floods of complaints from people all of a sudden, which is pretty unlikely to be a coincidence
Btw, you have a typo in the twitter handle on your profile (not in your comment), 'thesilencesturns'.
The majority of people are not on it, and the links are gated by a ton of toxic dark patterns and horrible UX trying to force people to sign up or log in.
I try to click on the image to enlarge and make the text readable, and I'm greeted with a login screen instead of a larger image.
https://marginlab.ai/trackers/claude-code/
This site has been documenting it for a while
https://marginlab.ai/trackers/claude-code-historical-perform...
> We always use the latest available Claude Code release and the SOTA model (currently Opus 5.5).
Changing the harness can have a big impact on performance even when leaving the model completely unchanged.
The test doesn't differentiate. But neither can the average user, who will also be using the normal auto-updating harness. You still get degrading quality right before each new release
This is very different from a nefarious inference-side degradation to save cost, promote the new model or anything else frequently proposed as motivation.
You can't just dismiss something backed by careful measurements by throwing a truism at it. What is this honeymoon phase? Can you quantify it? If not, how are you sure it's real?
My work at my job has stayed the same. But the model quality has varied.
They definitely tune the models in production after launch, if not only to share load during high traffic times. It’s not a crazy conspiracy that the same model can be stupider at different times.
This is a more a statement on the work you do and how you work versus the models. I'm doing more complex work since Fable (and now for way cheaper thanks to Opus 5.5)
With 4.6 I would still babysit a lot more code quality and so on. With the newer model I see myself talking about features at a higher level, and then not having to nitpick PRs to death. Which means most of my time is now spent talking to the model about the product instead of the implementation of the product.
> It’s not a crazy conspiracy that the same model can be stupider
Sorry, I really do think it's a conspiracy. If nerfing were real, it would be trivial to prove. DeepSWE, SWEBench, and other benchmarks are all available for anyone to run. A "nerfing" hypothesis has to survive the fact that a statistically significant dip in benchmarks has never been observed.
A/B testing alone would result in a performance nerf for one group.
8-bit quantized models will barely show degradation on benchmarks. The performance is reliably at 99% of the non-quantized model. 4-bit quantization retains somewhere around 95-98% performance on benchmarks. But if you've ever used a 4-bit model, it feels lobotomized.
And just consider what a compny serving these models would do if they were at capacity. Would they stop serving the model altogether? Of course they wouldn't...
Denying that models experience purposeful degradation is gaslighting.
My local models don't display that degradation, sensed or measured. They consistently perform equally to what I expect of them, precisely because they don't change.
How does twitter explain that? Is my internal model for expectation of capacity magically not drifting for local models but somehow is for anthropic api call based models?
What Anthropic presents as Opus 5.5 isn't actually a single model...it's Anthropic's ecosystem as a whole. If you are lucky, you get the top model handling your issues all the time, however, that never happens. What really happens is that your request and content are graded along with your subscription (example: API? subscription, if so, what tier? how much has the user used it? Do we trust the user? how much? how much are they paying? are they asking something we think is dangerous?) and your request and context are routed accordingly.
Anthropic isn't alone in this behavior, Open AI does it as well, just look at the respective subreddits on reddit for both if you need some examples, or just play around with the various models from both companies.
There are a few folks who've done some analysis on this (their findings were posted on reddit and X), and a bigger multi-national study is apparently coming, though I admittedly don't know their findings.
I guess the tl;dr is that Anthropic and Open AI are actually selling you "best-effort" routers, so you may or may not get the best in class model, and only they get to determine if you do or do not. No guarantees.
That sentence... This conversation is indistinguishable from a billion conversations had around multi-player online gaming.
Note: I know what quantization is so don't hold back.
like before Anthropic signed the Colossus deal, the usage limits were insane and everyone was complaining, I wouldn't be surprised if they'd rather try to make inference faster that way than try to just limit people, at least for those on subscriptions
At their scale, you’d have to be setting money on fire if you’re not doing dynamic inference optimisations based on load.
API and consumer subscriptions are treated differently; all trackers measuring via API won’t notice this.
Where is the evidence they are "nerfing" the models due to request volume?
Edit: I don't know they do, I mean they could repurpose systems if they are idle. Inference demand is global, and providers like Azure have global routing options that are cheaper. Night time in the USA could be serving inference demand on the other side of the globe.
But you seem adamant that there's no chance the providers serve slightly quantized models for subscription users during high loads, or otherwise tweak models for requests from those users.
It's tricky to prove either way, but the chance is not zero.
They just want some evidence. It should be pretty easy to measure, shouldn't it?
Nerfing conspiracy doesn't need to be proven false. Where is the evidence it's true?
The current period is as pro-customer as we're ever going to get, with cash still flying around and neither OpenAI nor Anthropic on the public market, and people are already forced into this sort of business to keep them true to their word. The point isn't even whether they're nerfing the models (I don't think they are), but that people can't seem to trust them to do right.
Yesterday I fought claude fable to not just jump to make changes like a dog following a treat, that we were researching... at some point I introduced the word HAWAI... and only if I say HAWAI the thing can start making changes..
I was going to post here in HN just to have a "I knew this was the reason" when they release fable > 5.1
I had the exact same feeling every time they have a new big release
Even in fresh sessions with minimal context like a file with a couple hundreds of lines of code, it would frequently ignore clear instructions, avoid work, and even sometimes “think” things like “looking for ways to code without approval”.
The model is just tuned for long running and doesn’t like to work with a human in the loop, or work back and forth.
Its quite worrisome honestly and i feel like that 'im in danger' Simpsons meme more and more. Right now i still have a lot of domain expertise which helps in knowing which questions to ask and which problems to solve, but well, i wonder how long that is going to save my job.
This week I have done considerably less work, and 2 days in, 54 % of Max usage is gone.
I hate this bait and switch cycle, and seeing how OpenAI cut its limits (5x raise from 0.1x api price to 0.5x api price for their max plan, so new $500 plan gets less usage than previous $200 plan) I think this will get worse still.
The quality of the output/work seems the same, but the speed at which it gets stuff done is a lot slower, because it's asking for permission so much more
I don't have any numbers/stats, just my impression. However, I imagine that if Anthropic could make the models ask for permission more often, it could be an interesting way to throttle access, without degrading quality of the output
I’m sure it’s lowkey intentional, probably encourages users to spin up new chats; hence less context.
Have a look into Docker sbx for instance.
The actual 10-day windows contain substantially more samples than that validation, but I haven’t demonstrated that they’re sufficient to distinguish a same-family swap of that size, so I’m not claiming they are.
The goal isn’t to make the instrument say “nerfed.” A null result is a result too. I’d much rather publish “we couldn’t detect a change of this magnitude” than overclaim what the data can support.
After Fable launch I switched over to Codex and it was simply amazing, with frequent usage resets that seemed never ending. They clearly had more compute than they knew what to do with. Post Astra, Codex has gotten dumb again across all models, increased usage for no real reason, and no resets.
I'm guessing Opus 5.5 will take the heat off Codex for a bit, leading to better performance. So I guess I stick around here instead of switching again?
Also, it isn’t comparing one 10-day period once and calling it done. The window rolls forward daily, and a change has to clear the pre-registered 99% threshold in two consecutive windows before it’s flagged.
Ten days isn’t sacred, though. Once there’s enough longitudinal data, one of the things I want to evaluate is whether that window length is actually well calibrated or should be changed in a future version.
(which has led me to believe that's a good approximation for hedonic adaptation, I've seen tons of attempts at demonstrating nerfing via benches, none persist)
_pushes model further_
"Ugh why is this model failing now even though I'm asking it to do harder things and also got sloppier with my prompts because I got used to it being able to figure stuff out"
It’s frustrating to observe communities made up of smart, professional individuals as they behave like spoiled children on the day after Christmas when new toy novelty has begun to wane.
I understand it’s relatively harmless but for goodness sake, take a step back and appreciate what you have instead of immediately wanting the thrill of a newer model. Slow down and do deliberate work to get the most out of these amazing tools. Don’t just live off the temporary thrill of finding something marginally better than what you have.
As my weekly/hourly quota is nearing its end, does Anthropic start to serve me a worse model?
Am i personally being throttled? testing against API is meaningless.
At scale where anthro and cgpt operates - every single token matters.
My company uses Claude models exclusively via Azure and AWS bedrock, which have their own licensed copies of the weights.
If all these people are so convinced Anthropic is nerfing models, have they tried other inference providers? Do they think the nerfing is coordinated across independent providers? Why wouldn’t any of these nerf-benches use these comparison points or even talk about them?
I think the most likely conclusion by far is that this is a psychological phenomenon.
I wouldn't say Opus has been nerfed. They are just two different models but, at the end of the day, GPT Sol is the model I trust the most, at least for now.
My gut tells me this involves an undisclosed, never-released grandparent model (higher-class than Fable/Astra level, roughly unsellable due to unfeasible cost). That grandparent model is distilled into lower models, of which Opus 5.5 might be an instance of.
That also guarantees protection against distilling a core business. You never make your prime weights available to the public, you only make distillings themselves available.
The downside of this strategy is that you spend a lot of compute on something that you never release, but it might be just the right play (for now) for closed weight companies.
It's a gut feeling, I have zero hard evidence to back it up (it's what I would do as them).
Open models are the endgame.
The classifier is a model; it examines the actual prompt.
They already do this for the safety "guardrails".
A few things make it harder: the panel contains questions drawn from multiple benchmarks rather than one recognizable test, the evaluation is automated and fixed ahead of time, and the raw outputs/results are public so odd behavior can be inspected.
But ultimately LiveNerf measures the behavior exposed through the API on a fixed public panel. It can’t prove what’s happening internally or guarantee the provider isn’t conditioning on the benchmark.
Longer term, I’d like to add held-out/private or periodically refreshed panels specifically to make benchmark recognition harder. I just don’t want to quietly change the current panel, because having a fixed instrument is important for the longitudinal comparison.
Great future everyone has chosen for us.
The initial calibration screened 2,336 questions with 4 samples each, then selected the 78 questions where Opus 5.5 showed useful variance. The panel is run daily, and the actual decision is based on paired per-item differences across 10-day windows with clustered standard errors, not on any single day’s result.
The n=1 in the daily sampling rate means one sample per item per day, not one sample for the experiment. By the time a window is evaluated there are hundreds of observations, and a change has to clear a pre-registered 99% interval in two consecutive windows before LiveNerf calls it a change.
The nondeterminism is basically the reason the statistical part exists in the first place.
If so, please share. This should be measurable, and I'm glad this project is measuring it.
Answers in the form of additional anecdotes, stated with even greater passion but still lacking a statement that could be tested and falsified, would validate my exact concern.
As he confirmed, for at least a period of time, and for some users, “Sol xhigh” was actually “Sol high”, etc.
It would be economical suicide from anthropic and OpenAI to actually need models intentionally.
But hey I guess it's hard with technology that truly seems like magic. People say if you'd bring electricity to the middle ages you'd be called a witch and burned. The same is happening to the model labs here because they are bringing tech that the world isn't ready for yet.
Links have been provided by others in this post.
Truly, madly, deeply sloppy.
First two days of this thing was like working with Einstein, then about 24-36 hours ago I started getting frustrated at bullshit that hadn't been a problem before. It was so egregious that I checked to make sure I was still on Opus 5.5 Max.
Step 1. Model can’t do something challenging Step 2. You try a bunch and fail Step 3. Anthropic trains on your usage data. Your current code base and current problem are now in domain Step 4. Model comes out and you’re shocked when it can tackle the thing you were stuck on Step 4. Codebase drifts significantly and you try new problems you thought were a similar level. Your code is less familiar and the problem doesn’t have a bunch of failure cases in the train set. Feels of it being worse on similar problems
This is definitely not about dealing with the frontier of AI. I wasn't part of the nerfing chord, but Astra changed my mind. Quality got me to upgrade from Pro x5 to Pro x20 on launch day. A couple of days later was dumb af, horrendous code quality etc...
Something fishy, or at least unethical is going on. Not sure it impacts API users though.