Mature training pipelines, plus ever expanding RL datasets of increased quality, and mega GPU clusters to finish training in a few weeks. Automated safety and reliability testing.
I ran a battery of tests against a couple of simple prompts to check on thoroughness and verbosity of every available Opus, and 5.5 is a lot closer to 5 than people are letting on. 4.6 remains the best in terms of getting to the point and just doing what you ask. I had switched from 4.7 to 5.5 as my main claude model, but started running into the telltale over-interpretation issues of the 5 series, and have switched back. Something in their RL pipeline has made these models consistently worse IMO.
I see. One of the biggest issues I had with 5 is that it constantly made mistakes calling tools and making API calls that even lesser models didn't struggle with. Mistakes were crazy high.
My main grievance is that any gap in specificity in my statement of a task would lead Opus 5 to invent an interpretation to fill the gap, frequently creating lots of extra work for itself in the process, and often deviating from my intent. This would happen even for very simple things. I once asked Opus 5 to fix a failing unit test in CI (something pretty simple), and it went off on a 45 minute expedition (all in one turn), read a boatload of unnecessary files, massively overcomplicated the assignment, etc. It fixed the test but previous models would have handled this much more straightforwardly.
A common form of this failure is the model picking up on random wordings from earlier in the session (e.g. some comment it made to me in the middle of a response, that I never explicitly endorsed) and then treating these as hard commitments. Or over-interpreting a specific word choice or clumsy phrasing as if it were a "load-bearing" constraint on the task.
None of this clumsiness would be so problematic if the model didn't have such a strong drive toward autonomy. It's much like with people: there's no shame in not understanding what you're being asked to do, provided you ask clarifying questions. There's no shame in ignorance if it's wedded to curiosity. Benchmaxing has RLVRed curiosity and clarification straight out of these models. It sucks.
I do wonder if people switch back and forth between primary models (GPTvsClaude) that it may be a better idea to simply keep releasing updates as soon as possible in order to keep users from bouncing back and forth.
It's because they need subscription money and interaction data and so keeping a version bump in the wings to stop the bleeding from your competitor's version bump is the logical thing to do. It has nothing to do with RSI.
Like think about a software org with good CI/CD versus one without. The mature org can do consistent incremental releases because each one is safe and low overhead, the messier org will do fewer big releases because each release requires a big effort on its own.
As model developers mature we might expect to see more frequent point releases rather than the big bang evolutions.
Probably one of the factors.
Signed up to openai pro a few days ago, deciding between openai and anthropic, then sonnet 5.5 was released and am wondering whether I made a mistake.
Luckily it's not a mistake as now we have access to
.
.
.
dots.
Not the person you're replying to, but judging by the emphasis on the cost of cached input tokens in the OP article, I'd guess it has to do with DeepSeek v4.1's KV cache efficiency. It uses <1000 bytes per token, so they're able to get 1M token context in under a GB.
just goes to show that OpenAI in fact did not innovate on a single thing for the better part of a year (one could argue two) and instead keeps immitating what it sees doing others successfully with the tech, all in a very transparent attempt to get people lubed up for their IPO.
Both labs are spying on each other and they get jelly when the other is releasing a new model, so they have to ship something at the same time so they don’t look bad.
What I don't understand is how much people have to say about every single one. Aren't we at the diminishing returns stage yet? Is there really that much to discuss?
If you look closely at various benchmarks, you'll see that often models will improve in certain areas while regressing in others. It suggests we're already at the point of diminishing returns.
The old "the bigger number is better", GPT announces model 6.1, the obvious thing to do next is to announce Gemini 27, and after that Claudé 3000, then a flute album.
Chinese model pressure. Many of my SWE friends switched to Chinese models. I also use QWEN and GLM for many of the api requiring projects and dropped OpenAI and Anthropic. The only reason was the cost.
EDIT: I love getting downvoted by openai and anthropic employees or their bots.
I can't recommend Chinese models enough. My personal favorite is DeepSeek v4.1 Flash but I have tried Qwen 3.8, Kimi 3 and GLM 5.3 which are equally impressive but DeepSeek is the cheapest and fastest regularly hitting 270 token per second.
And yeah I have worked with Anthropic and OpenAI models, they're good but they cost a fortune while Chinese models are already really good at a fraction of the cost.
DeepSeek v4.1 Flash is fascinating and uneven. It's way too chatty in OpenCode to be a collaboration partner. I tried dsh-tui which feels comparable to the codex/claude tui's and it's usable. but it seems to be "brilliant and yet stupid" in a way I can't quite put my finger on. I've got too much real work to get done to dig into it so until the big boys price me out of the market I'm back to my $100/month deal.
I keep hearing about these Chinese models, but what exactly are you doing with the models and coding? I have a need to fully write code with full tool calling capabilities. Not just methods or functions. I want to be able to prompt a feature and it makes the JIRA ticket, and fully implements it and makes a PR. I don't want to babysit it or even read the code. Once it creates the PR, I want it to monitor it for any comments fro Copilot/security review and then fix it as necessary.
Is that what the Chinese models are capable of? If so, how are you using them? API? Or is there an inference provider that is as fast as the big 2? What about the coding harness?
I am using the API only with them for now. But what you describe is nothing compared to Qwen or Mimo. These models are more capable than Opus in general and at a fraction of Opus's cost.
That's the point, those jobs went away, the need for a job didn't. No one seems to know what long term opportunities are going to be created by AI, we only know that it is going to erode existing opportunities.
This kind of narrative is going to bite them just like the "AI will take your job" narrative has. It feels like the frontier labs are taking a massive gamble with public perception here. I assume the goal is to paint the technology as so powerful and dangerous that only a handful of blessed US companies should be trusted to run it, in an attempt to suppress the rise of the Chinese models that are rapidly catching them.
This is where we need the hardware companies and neoclouds to start speaking up. The labs want to elevate matters from the level of civil society (basically, competing firms) to the State (enclosure), and as always, in the name of security. But other actors in the same ecosystem have strictly opposed interests here, and are equally if not more credible as far as the State is concerned. If players like Nebius, Baseten, Fireworks, etc. among many others including obviously Nvidia, Dell, AMD, and so on don't get ahead of this they will be sacrificing trillions.
Exactly, it's about taking this stuff off the open market where anyone can judge it and there's competition, into government contracts where competence to judge the offer is scarce or absent, and they can ask much higher prices. And with this much investment at stake, any lie that sells the narrative will serve.
Yes, and given the nature of the current administration, whose actors are not inclined to see themselves as independent competing capitals among others, but rather as privileged capitals, and therefore more inclined to move towards taking an interest in the process of enclosure, ensuring it includes them, the hope seems to lie with companies at the hardware layer who have an interest in seeing the diffusion of intelligence play out freely at all levels of society.
The labs have to be told NO--the problem they're dealing with, that model outputs give the game away, and that in turn the distiller becomes the distilled, is a fundamental problem they have to figure out how to deal with without going to the State.
Didn't HF use an open weights model running on their own hardware to solve the issue though? Sort of defeats that narrative and plays into one in which frontier == bad_guys and open == good_guys
HF guys, especially those under Julien Chaumond, are fantastic and will use whatever they can to address their issues. Open or closed but they are firmly on the Open side of the fence. However, their storage and model service is about as sticky as you can get so they are in a different position. They don’t have to sell capability, they sell capacity and community.
Yes, they have used GLM 5.2, which promptly did whatever they asked it to do, while their first attempts to use their enterprise access to a "SOTA" model failed due to refusals to investigate anything that is security related.
Non-deterministic, sure. But also they are powerful, at least powerful enough to launch a cyberattack. Until this morning, that was not a power that I thought they had outside of fiction.
And, the thing is, I don't want non-deterministic things to have that kind of power. We don't want that. We want that kind of power to not be triggered by a random number generator.
Ironic that HuggingFace needed Chinese models to defend against it. But of course the spin of the leading firms will just be to point at their trusted access programs and demand that all dangerous activities, even if just defensive, happen via their APIs or be outlawed otherwise.
If that's the position they take then they really should be heavily regulated or nationalized. Cyberdefense against their own models dependent on their goodwill? Sure, but then they have to sell defense capabilities at subsidized rates with a limited margin. Would be very weird otherwise to take the world hostage with their models and then also sell the solution while demanding intrusive KYC.
I do wonder how they source their bulk literary data. Do they have a google books, an archive.org or an anna’s archive for chinese content? What’s the Asian equivalent to Elsevier? Is there (strong) copyright on the corporate level?
Who told you they were distilling? Why might they say this? Think. It’s like complaining that the top student only does well by going to office hours instead of mindlessly reading textbooks.
The top student giving paid lectures about his classes, and another student skipping class and instead studying those lectures to end up with the second highest grade?
Maybe it could be improved with the other student not even going to the same school?
Agree, they also may find themselves in a place where the government rightly says they can't have their dangerous new toy because they can't be safe with it.
This is a high stakes PR game. Governments can and will step in and embargo and regulate these systems in ways which will hurt the companies and investors.
1000% the case. Admitting they failed at security and allowed privileged escalations in a prompted AI would look bad for them; the AI just did it all itself is FUD to boost the arguments for regulation. And most news won’t challenge this FUD because it gets clicks.
It is not difficult to imagine that freedom disappearing soon, under the veil of security. It will not be long before Google's remote attestation becomes commonplace, and then it won't be long until the websites you care about start using it and excluding unsigned browsers. The wheels are in motion already and AI is going to speed this all up: https://www.eff.org/deeplinks/2026/07/googles-new-remote-att...
I'm sure they will, but any user extension that interferes with the DOM (as an adblocker does) will break remote attestation, so the website will reject your request.
I haven't seen any browser attestation proposals prescribing that and I don't see any reason for Mozilla and Apple to support that. Care to provide references?
"Exactly" what? That by "don't let me" you meant "maybe sorta kinda woulda not let me do that in the future and I can't be halfassed to install FireFox"?
This is interesting because I assume it has suffered the same linguistic degradation as the word "fine" which in some cases means "of the highest quality" but mostly means "meh". I suspect it comes down to the dialect and social rank of the person saying the sentence. Compare how you would perceive:
reply