I’m not sure how you can really make either statement work anymore. Now that smaller models are actually broadly usable, “behindness” is no longer a scalar and at the tails, where no open lab seems to be trying to compete at the >10T scale and no closed lab seems to care about <400B anymore, it’s just apples to oranges. It’s like talking about whether Qualcomm is “behind” Nvidia.
1. Mythos wasn't released in February. Let's stick to only public-facing models.
2. For public-facing models, the differences are really minor with some occasional model (like Fable or Astra) showing some better performance in specific benchmarks for the span of some weeks or few months before open ones catch it.
3. Being bleeding edge is overblown anyway in the real world, besides the occasional "very latest fresh model did this task which previous one couldn't", and the number of those tasks is increasingly small and far from mundane corporate needs.
I doubt most of your claims. Maybe the guardrails and emotional intelligence is true.
For speed and efficiency, you are most likely wrong.
Speed is led by GPT-5.6 Sol on Cerebras Ultrafast at 750 t/s. Afaik you cannot serve a single DeepSeek Flash 4.1 stream at 750 t/s, plus the model is less intelligent as seen on newer benchmarks.
I believe OpenAI and Anhropic are at the frontier of efficiency too. There were numerous reports about their breakthroughs and associated API price cuts. The idea that open-weight models are more efficient seems unfounded.
5.6 sol ultrafast on cerebras is 750tps, open models readily exceed this. just by using a smaller model cerebras serves qwen 3.8 27b at 1850tps. or even larger models, mimo 2.5 pro was served for a while at 1000tps. and so on. [https://inference-docs.cerebras.ai/models/choose-a-model]
the chinese ai companies have 10% of the total compute resources of the US ones. since the USA tries to stop them from buying nvidia gpus. they maxed out the efficiency.
deepseek v4.1 has engram architecture. it has 550b params instead of 5T+ for astra/fable. it has 8b active instead of potentially hundreds active for astra/fable.
compare input/output/cache: $0.15/$0.60/$0.003 for v4.1 to $10.00/$50.00/$1.00 for astra and $10.00/$50.00/$0.25 for fable.
astra cache reads are over 330 times more expensive.
at the artificial analysis 7:2:1 ratio, deepseek is $0.18/m, fable is $7.18/m, astra is $7.7/m.
but what about intelligence? AA would rate deepseek v4.1 at AA 40, astra is AA 53.
so it cost 4,180% more for 32% more intelligence.
they are serving that at over 250tps at baseten. to get close to that on astra API you are paying double the cost for fast mode.
so it is now 8456% more expensive for a similar speed and 32% more intelligence. 84 times more expensive.
i was trying to make a point about efficiency of serving the model. the cost per task itself would not be enough to show that.
you could compare gpt 5.6 luna. if you did that the same way as before you would get a blended price of $0.17 for luna at AA 38. for baseten it would be $0.20 for v4.1 at AA 40.
assume roughly the same intelligence. on AA openai gets 117tps. baseten gets 284tps. so 18% more expensive but 142% more tps.
the fast mode is again double the cost, roughly same intelligence. so luna in that case would be 70% expensive. take the per task cost and it would still 9% more expensive.
so i think there is something to be said about the efficiency of the model.
For practical uses they are there. Arguably the frontier models are worse for some of these practical tasks. And keep in mind, people will use maybe frontier for 1/10th of the work, planning and review, and go open source for rest. The question is if they manage to impose outside us. If not, they are losing competitiveness.
I always thought switching from a SOTA model to a dumber model after planning was a terrible idea.
Mostly I heard this from people who I got the impression have little experience in developing greenfield software with agentic AI. Often the same people who talk about spec frameworks.
I fundamentally disagree with the approach. I believe the ability to autonomously evaluate, test, and adjust during long horizon tasks is critical to using AI efficiently.
Well, it's the enterprise software house pipeline... The software architect writes the spec, hands it down to the implementation team, senior leads, junior devs or offshore teams codes it.
I also disagree with the approach, this is cargo-culting the existing ways of working.
Still waiting for our org to roll out Mythos. I guess it was too expensive so we’re stuck on the previous model until the internal team can figure out self-hosting open models.
In fact they still have not caught up with February's Mythos, indicating they are more than half a year behind.