Hacker Newsnew | past | comments | ask | show | jobs | submit | phpnode's commentslogin

What's driving the increase in release cadence here? We seem to get new models every week or so now, is this RSI?

Mature training pipelines, plus ever expanding RL datasets of increased quality, and mega GPU clusters to finish training in a few weeks. Automated safety and reliability testing.

Opus 5.5 is better than they anticipated, it's faster, smarter, cheaper. I'm about to change provider for claude and I'm not the only one

It feels like an updated 4.6. It's fantastic.

I ran a battery of tests against a couple of simple prompts to check on thoroughness and verbosity of every available Opus, and 5.5 is a lot closer to 5 than people are letting on. 4.6 remains the best in terms of getting to the point and just doing what you ask. I had switched from 4.7 to 5.5 as my main claude model, but started running into the telltale over-interpretation issues of the 5 series, and have switched back. Something in their RL pipeline has made these models consistently worse IMO.

I see. One of the biggest issues I had with 5 is that it constantly made mistakes calling tools and making API calls that even lesser models didn't struggle with. Mistakes were crazy high.

My main grievance is that any gap in specificity in my statement of a task would lead Opus 5 to invent an interpretation to fill the gap, frequently creating lots of extra work for itself in the process, and often deviating from my intent. This would happen even for very simple things. I once asked Opus 5 to fix a failing unit test in CI (something pretty simple), and it went off on a 45 minute expedition (all in one turn), read a boatload of unnecessary files, massively overcomplicated the assignment, etc. It fixed the test but previous models would have handled this much more straightforwardly.

A common form of this failure is the model picking up on random wordings from earlier in the session (e.g. some comment it made to me in the middle of a response, that I never explicitly endorsed) and then treating these as hard commitments. Or over-interpreting a specific word choice or clumsy phrasing as if it were a "load-bearing" constraint on the task.

None of this clumsiness would be so problematic if the model didn't have such a strong drive toward autonomy. It's much like with people: there's no shame in not understanding what you're being asked to do, provided you ask clarifying questions. There's no shame in ignorance if it's wedded to curiosity. Benchmaxing has RLVRed curiosity and clarification straight out of these models. It sucks.


> I'm not the only one

See, that's an/the issue. As soon as people start to flee to the improved model, they start to serve degraded models to keep up with the demand.


I do wonder if people switch back and forth between primary models (GPTvsClaude) that it may be a better idea to simply keep releasing updates as soon as possible in order to keep users from bouncing back and forth.

This is it.

It's because they need subscription money and interaction data and so keeping a version bump in the wings to stop the bleeding from your competitor's version bump is the logical thing to do. It has nothing to do with RSI.


Maybe process maturity too.

Like think about a software org with good CI/CD versus one without. The mature org can do consistent incremental releases because each one is safe and low overhead, the messier org will do fewer big releases because each release requires a big effort on its own.

As model developers mature we might expect to see more frequent point releases rather than the big bang evolutions.


Probably one of the factors. Signed up to openai pro a few days ago, deciding between openai and anthropic, then sonnet 5.5 was released and am wondering whether I made a mistake.

Luckily it's not a mistake as now we have access to . . . dots.

(and sol 6.1, it seems)


jokes on me, I pay for all the subscriptions.

Of course they do. The real money makers are not subscription users, but the API users and you can just switch with the model selector.

No, we're pacing ourselves to have the time to evaluate the impact each new model could have, obviously.

Versions is marketing, snapshots/minor variations are easy and the number must go up. Release timing is another OAI's marketing tactic.

>RSI

Recursive improvement doesn't imply increased rate, another word for it is "iterative" but this probably sounds too boring to some people.


Response to DeepSeek’s technical paper and competition.

Which paper are you referring to?

What's that in summary?

Not the person you're replying to, but judging by the emphasis on the cost of cached input tokens in the OP article, I'd guess it has to do with DeepSeek v4.1's KV cache efficiency. It uses <1000 bytes per token, so they're able to get 1M token context in under a GB.

Edit: https://huggingface.co/deepseek-ai/DeepSeek-V4.1-Flash/blob/...


just goes to show that OpenAI in fact did not innovate on a single thing for the better part of a year (one could argue two) and instead keeps immitating what it sees doing others successfully with the tech, all in a very transparent attempt to get people lubed up for their IPO.

Probably just the singularity, no big deal

Both labs are spying on each other and they get jelly when the other is releasing a new model, so they have to ship something at the same time so they don’t look bad.

Both labs are spying? Employees hang out at the same bars, have overlapping social circles, etc. Alcohol does what alcohol does.

They're pacing the frontier

and seems like they're claiming Sol/Opus are not frontier (and only Astra/Fable are)

It is the only way to reduce prices while making it look like a good thing.

What I don't understand is how much people have to say about every single one. Aren't we at the diminishing returns stage yet? Is there really that much to discuss?

If you look closely at various benchmarks, you'll see that often models will improve in certain areas while regressing in others. It suggests we're already at the point of diminishing returns.

Wanting to have the newer model than the competitor, presumably.

The old "the bigger number is better", GPT announces model 6.1, the obvious thing to do next is to announce Gemini 27, and after that Claudé 3000, then a flute album.

We swear, We Really Wanted To Make An "ASI" Model But This Is Literally The Way The Weights Dragged Us This Time

Its a news cycle more than anything, and its ONLY going to get much, much worse. Daily releases, or multiple daily, 30-45, by EOY. Welcome to RSI!

The initial response to 6 Sol was bad, and Opus 5.5 was definitely winning the public vibes war. Makes sense to rush something out

New models are distill from the actual unrelease frontier models. They are just giving us better checkpoints.

no patrick, m̶a̶y̶o̶n̶n̶a̶i̶s̶e̶ a point release of the slopbot is NOT a̶n̶ i̶n̶s̶t̶r̶u̶m̶e̶n̶t̶ RSI

Productivity is increasing as models get smarter; we are ascending the singularity. I'm serious.


They're releasing Sol 6.1 because 1. Astra 6.1 got postponed 2. Sol 6 is shitty 3. They have to release _something_ in response to Opus 5.5

Competition

Chinese model pressure. Many of my SWE friends switched to Chinese models. I also use QWEN and GLM for many of the api requiring projects and dropped OpenAI and Anthropic. The only reason was the cost.

EDIT: I love getting downvoted by openai and anthropic employees or their bots.


I can't recommend Chinese models enough. My personal favorite is DeepSeek v4.1 Flash but I have tried Qwen 3.8, Kimi 3 and GLM 5.3 which are equally impressive but DeepSeek is the cheapest and fastest regularly hitting 270 token per second.

And yeah I have worked with Anthropic and OpenAI models, they're good but they cost a fortune while Chinese models are already really good at a fraction of the cost.


DeepSeek v4.1 Flash is fascinating and uneven. It's way too chatty in OpenCode to be a collaboration partner. I tried dsh-tui which feels comparable to the codex/claude tui's and it's usable. but it seems to be "brilliant and yet stupid" in a way I can't quite put my finger on. I've got too much real work to get done to dig into it so until the big boys price me out of the market I'm back to my $100/month deal.

I was working exclusively with DS 4.1 Flash until Opus 5.5 got me back to a sub. I was disillusioned with what was available.

I keep hearing about these Chinese models, but what exactly are you doing with the models and coding? I have a need to fully write code with full tool calling capabilities. Not just methods or functions. I want to be able to prompt a feature and it makes the JIRA ticket, and fully implements it and makes a PR. I don't want to babysit it or even read the code. Once it creates the PR, I want it to monitor it for any comments fro Copilot/security review and then fix it as necessary.

Is that what the Chinese models are capable of? If so, how are you using them? API? Or is there an inference provider that is as fast as the big 2? What about the coding harness?


I am using the API only with them for now. But what you describe is nothing compared to Qwen or Mimo. These models are more capable than Opus in general and at a fraction of Opus's cost.

Sounds like your dream workflow could replace you with a PM?

Anthropic’s IPO?

That's the point, those jobs went away, the need for a job didn't. No one seems to know what long term opportunities are going to be created by AI, we only know that it is going to erode existing opportunities.


I would be amazed if the LTV of an instagram account is over $7k. Where are you getting that number?


This kind of narrative is going to bite them just like the "AI will take your job" narrative has. It feels like the frontier labs are taking a massive gamble with public perception here. I assume the goal is to paint the technology as so powerful and dangerous that only a handful of blessed US companies should be trusted to run it, in an attempt to suppress the rise of the Chinese models that are rapidly catching them.


This is where we need the hardware companies and neoclouds to start speaking up. The labs want to elevate matters from the level of civil society (basically, competing firms) to the State (enclosure), and as always, in the name of security. But other actors in the same ecosystem have strictly opposed interests here, and are equally if not more credible as far as the State is concerned. If players like Nebius, Baseten, Fireworks, etc. among many others including obviously Nvidia, Dell, AMD, and so on don't get ahead of this they will be sacrificing trillions.


Why are you expect companies that have been profiting off of LLM insanity to do the right thing if not legally compelled to?


Exactly, it's about taking this stuff off the open market where anyone can judge it and there's competition, into government contracts where competence to judge the offer is scarce or absent, and they can ask much higher prices. And with this much investment at stake, any lie that sells the narrative will serve.


Yes, and given the nature of the current administration, whose actors are not inclined to see themselves as independent competing capitals among others, but rather as privileged capitals, and therefore more inclined to move towards taking an interest in the process of enclosure, ensuring it includes them, the hope seems to lie with companies at the hardware layer who have an interest in seeing the diffusion of intelligence play out freely at all levels of society.

The labs have to be told NO--the problem they're dealing with, that model outputs give the game away, and that in turn the distiller becomes the distilled, is a fundamental problem they have to figure out how to deal with without going to the State.


That’s the play. That these frontier models are so powerful that they must be behind sovereign firewalls and gateways.


Didn't HF use an open weights model running on their own hardware to solve the issue though? Sort of defeats that narrative and plays into one in which frontier == bad_guys and open == good_guys


HF guys, especially those under Julien Chaumond, are fantastic and will use whatever they can to address their issues. Open or closed but they are firmly on the Open side of the fence. However, their storage and model service is about as sticky as you can get so they are in a different position. They don’t have to sell capability, they sell capacity and community.


Yes, they have used GLM 5.2, which promptly did whatever they asked it to do, while their first attempts to use their enterprise access to a "SOTA" model failed due to refusals to investigate anything that is security related.


Public doesnt know about Huggingface. ChatGPT (OpenAI) says it‘s dangerous. They must know.


>That these frontier models are so powerful

Maybe powerful might NOT be the right word to describe them, they are just non-deterministic, there for we going to see this kind thing more and more.


Non-deterministic, sure. But also they are powerful, at least powerful enough to launch a cyberattack. Until this morning, that was not a power that I thought they had outside of fiction.

And, the thing is, I don't want non-deterministic things to have that kind of power. We don't want that. We want that kind of power to not be triggered by a random number generator.


I don't know why you wouldn't think they could do this already.

Without the system prompt these models can be used to do all sorts of terrible things.

That's precisely why they need to be strictly regulated by international treaties.


Please buy our IPO before it crashes so we can be billionaires.


Exactly, this is all so they can continue with their S1 and they can dump shares onto the hedge funds.


And require your age verification, selfies, and DNA samples.


Ironic that HuggingFace needed Chinese models to defend against it. But of course the spin of the leading firms will just be to point at their trusted access programs and demand that all dangerous activities, even if just defensive, happen via their APIs or be outlawed otherwise.

If that's the position they take then they really should be heavily regulated or nationalized. Cyberdefense against their own models dependent on their goodwill? Sure, but then they have to sell defense capabilities at subsidized rates with a limited margin. Would be very weird otherwise to take the world hostage with their models and then also sell the solution while demanding intrusive KYC.


It’s like the old firewall meme, reincarnated.

https://securityzap.com/wp-content/uploads/2015/12/layered-s...


Are the Chinese models actually catching up or are they just distilling the frontiers? If it's all just distilling then they'll always be behind.


There's no reason they can't build their own models from the first principles. They have the hardware, energy, and enough CS scientists.


I do wonder how they source their bulk literary data. Do they have a google books, an archive.org or an anna’s archive for chinese content? What’s the Asian equivalent to Elsevier? Is there (strong) copyright on the corporate level?


Who told you they were distilling? Why might they say this? Think. It’s like complaining that the top student only does well by going to office hours instead of mindlessly reading textbooks.


Is this a better analogy?

The top student giving paid lectures about his classes, and another student skipping class and instead studying those lectures to end up with the second highest grade?

Maybe it could be improved with the other student not even going to the same school?


In your example the problem is what exactly?


They don't like competition when they have to actually compete.


In capitalist economics they call this the “free rider problem”


they're all introducing themselves as claude for one, there are more quantitive and qualitative arguments elsewhere


To be fair, anthropic models, when asked in chinese, also used to introduce themselves as deepseek sometimes. This is a limitation of the technology.


All chinese models are introducing themselves as claude? Why make claims trivially disproven?


you're technically correct and missing the point


Well that and also benchmaxxing.


Agree, they also may find themselves in a place where the government rightly says they can't have their dangerous new toy because they can't be safe with it. This is a high stakes PR game. Governments can and will step in and embargo and regulate these systems in ways which will hurt the companies and investors.


Data is the new oil, AI labs are the new steel mills.


1000% the case. Admitting they failed at security and allowed privileged escalations in a prompted AI would look bad for them; the AI just did it all itself is FUD to boost the arguments for regulation. And most news won’t challenge this FUD because it gets clicks.


It is not difficult to imagine that freedom disappearing soon, under the veil of security. It will not be long before Google's remote attestation becomes commonplace, and then it won't be long until the websites you care about start using it and excluding unsigned browsers. The wheels are in motion already and AI is going to speed this all up: https://www.eff.org/deeplinks/2026/07/googles-new-remote-att...


Why won't FF and Safari be signed?


I'm sure they will, but any user extension that interferes with the DOM (as an adblocker does) will break remote attestation, so the website will reject your request.


I haven't seen any browser attestation proposals prescribing that and I don't see any reason for Mozilla and Apple to support that. Care to provide references?


Exactly! Thank you!


"Exactly" what? That by "don't let me" you meant "maybe sorta kinda woulda not let me do that in the future and I can't be halfassed to install FireFox"?


cached thought. running CLIs is impractical and expensive in many environments and a hell of a lot less secure than using MCP


This is interesting because I assume it has suffered the same linguistic degradation as the word "fine" which in some cases means "of the highest quality" but mostly means "meh". I suspect it comes down to the dialect and social rank of the person saying the sentence. Compare how you would perceive:

    "You did a fine job"
or

    "It is quite impossible"

depending on who was saying it.


yup, it's not possible to do it safely with a simple unparameterised type: https://www.typescriptlang.org/play/?#code/C4TwDgpgBAcg9gOwK...


Also, ssalbdivad is your cofounder, just in case you’d forgotten!


so, instead of

    (foo (bar (1 2 3))
you'd prefer

    {
      foo {
        bar {
          1
          2
          3
        }
      }
    }
is that right?


    ( aar
      (bar1 1 2 3)
      (bar2 1 2 3)
      (bar3 
         (car1 2 3)
         (car2)
         (car3)
      )
    )
vs

(aar (bar1 1 2 3) (bar2 1 2 3) (bar3 (car1 2 3)(car2)(car3)))


Emacs vs vim, go!


Consider applying for YC's Winter 2027 batch! Applications are open till November 2.

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: