If that's marketed well it feels like it should cause a system shock like R1 did. It would also be interesting to see the reaction with code models becoming so good already i.e. cost efficient models aren't necessarily invalidated early by progress that matters.
Yes, I've seen this too and how Luna xhigh is so good that Terra doesn't really serve a purpose because beyond that you can continue at Sol medium. This can be the most cost efficient way, and especially now!
In my opinion, next step is to cut down on reasoning tokens while maintaining intelligence. The Chain of Thought and looping can still be an issue with these Chinese models. They in fact said K3 would improve in the area but it's still an issue that unfortunately harms the token cost wins a bit. OpenAI has been really impressive here, on the opposite end of this.
There is a really interesting startup in Prague that is doing just that. They fine-tuned Qwen 3.6 27b to have 46% fewer reasoning tokens while maintaining most of the performance characteristics. I'm interested to see if they continue down this path of optimizing reasoning for other models.
Genuine question, is the reasoning chain different from clicking the status bar under a reply and watching it "think"? Or selecting the "Thinking" transcript view in Claude Code? (both on the desktop app). Seems to me that is very out in the open
That's a summarized and filtered view of the actual reasoning.
OpenAI and Anthropic guard the real reasoning closely. Users have never been able to see it and the API returns an encrypted blob instead of legible reasoning.
I agree and this is why I think open models will win in the end. There is just so much to gain on being 10% behind the curve. Especially when the curve is far beyond your needs.
Yeah many tasks are already saturated. For small interactive code changes (I call this power coding like power armor), we got there about a year, maybe year and a half ago — with the small models of the time.
And small models have the advantage that they tend to be much faster. This is an interesting factor because, well speed is always a nice-to-have, but if you go from 10 seconds per turn to one second per turn, the activity actually becomes interactive.
It becomes a fundamentally different way of working. You stay active and engaged the whole time. And also because you are "driving", your mental model does not be synchronized from the code base. So you don't need to spend extra time later catching up.
This will also make it harder to compete because it holds true for everyone, not just you. I am already, this year, seeing vibe coders put out some decent stuff on App Stores but the problem is marketing it. You no longer automatically stand out just because you have a cool app. You may not even do so with reasonable early traction via social media. And then what? Throw venture capital onto the problem? In an AI saturated world? Really?
Also the drive to innovate goes down. It took weeks for the creator of 2048 to make a low quality clone of Threes. Now people can make relatively high quality ripoffs in days, even hours.
Yeah I've noted this behavior with best in class open weight models. They said K3 would have token efficiency improvements and I was hoping especially solving the thinking loop issue that plagued K2.x but even if this release helped somewhat, it looks like we still have a long way to go here... I'm not sure what's up here but I suppose lacking finetuning quality.
What OpenAI in particular have done with reasoning efficiency in the past few months since ChatGPT 5.5 is nothing short of remarkable. It's overshadowed a bit by the benchmark game and the Fable hoopla.
Now is the time to focus less on token cost and intelligence, but tokens to solve a particular set of tasks in closed benchmarks for a variety of categories.
What is the use of grand intelligence if it either costs you a kidney or can't complete at all within a token budget? Even if there are niche uses where you truly want "maximum power" above all, we need to at least more severely penalize such models versus those that does it just as fine within a tenth of the token cost.
I'm aware of some benchmarks at the Artificial Intelligence site, but CLEARLY we are not focusing enough on these today and still leaving the fun surprises to the users.
Yeah, I'm finding I end up switching to Codex and GPT 5.6 a lot lately because I've either run out of Fable usage or Fable refused to do the task. Most recently it refused to work on a WiFi configuration UI for a robot. No idea why it thought that was related to security, biology, or some other sensitive topic. They've hobbled it with guardrails that are overzealous and now there's a big opening in the market. Fable may be the best, but if it won't do the job half the time, it stops being my go to model as I don't want to waste time only to find it refuses halfway through.
I think they're less and less advertised as true generalists these days, as they pivot to profits that obviously lie (for the time being) first and foremost in agentic coding. It's no longer unusual to see regressions in terms of more stiff prose due to the strong tuning towards coding, or how they structure their response. And prose is a LLM's home turf! Instead, progress in agentic coding capability is usually the headline feature, the headline benchmark, etc etc. At least looking at Anthropic, Google, OpenAI. There are of course other LLM's.
So then add a dash of cybersecurity and medical use and that's basically it. No "closer to AGI" advertising. I'd say the 2026 development has in fact been the opposite; optimizing AI for niches where there is most potential for profits and that your description died in circa GPT-5 era.
In fact, this problem (for this test) is also stated by the pelican test author:
"The biggest limitation of the pelican is that it doesn’t touch at all on the thing that matters most for today’s model: agentic tool calling and the ability to operate tools reliably as conversations grow in length.
Anecdotally, GPT-3 was super good at creative writing. It didn’t have any of the typical LLM giveaways. It would write super weird, interesting stuff. Especially if fine-tuned on a specific author. Of course it would occasionally descend into saying the same thing over and over. But IMO none of the current models come close!
LLMs are, fundamentally, generalist AIs. Marketing or no marketing - it's just what they are. How they're trained, how they perform, what they're best at.
Empirically, they have something very much alike to the human "g factor" - a shared pool of "general intelligence" that all tasks benefit from.
When a "make it bigger, train it harder" upgrade like Kimi K3 or Mythos 5 drops, the performance rises on every metric. Not just the "headline benchmarks" like Mythos and coding/cybersecurity, but also things like literary analysis - which has nearly zero economic value, and isn't commonly post-trained or benchmarked for. And companies keep encountering things like "our carefully trained specialist model with lots of in-domain training on expensive closed datasets just got leapfrogged on our internal benchmarks by a next gen off the shelf generalist".
You can go hard on benchmarkmaxxing post-training, and you can burn millions of GPU-hours on coding RLVR. But, by the very nature of LLMs, a lot of the performance gains in flagship models are broad and domain-inspecific.
"Stiff prose" is more of a "style" thing than a "capability" thing. No one cares about how good an AI is at things like long form creative writing, because that's the opposite of a profitable field. All of LLM behavior is routed through text, so it's very easy to perturb "writing style" by some training elsewhere. Regression evaluation is hard. And the writing-specific post-training LLMs get is usually just cheap RLAF, with all the usual RLAF degeneracy.
Thus, we get the "default styles" that suck from a "creative writing" standpoint. A lot of that is just "what sounded good to the previous generation of LLMs" - and, unlike human readers, LLM evaluators don't get bored from seeing the same cliches repeated 9000 times across 9000 different instances of generated text. Humans tend to update over time from "this sound cool and punchy" to "this is generic AI slop", but RLAF evaluators stay at step 1. What little human-guided optimization this gets is aimed at "copywriting, marketing blurbs, punchy short-form" - and it shows.
You can do a lot there with some aggressive prompting, but the default writing styles suck, and I frankly don't expect that to change soon. No one cares enough to change it.
Pelicans? Used to be a decent proxy for "general model capabilities that no one would benchmaxx for" - a way to probe for that elusive "LLM g factor". Now that it's a known metric, it's very gameable. But it was pretty solid while it was novel and obscure.
reply