Using gpt-5.4-mini in off-peak hours already feels like super-speed to me. That's probably no more than 100-150 tk/s. I can't imagine 750!
I've always eyed Cerebras but never had a use for it that would justify paying for the API directly. Although now that I think about it, trying out the API would probably cost less than a subscription for a month...
Agreed, 1000tok/s just fills up the context window (which is big by 2004 standards) super fast. But seems like 5.3-spark was just a taste of what’s to come.
The ChatGPT subscription gives you access to the -spark model(s) in Codex which are blazing fast (but pretty dumb) which I think runs on Cerebras hardware too.
is this specifically in codex? have been trying to use the models for months on opencode then pi but it says chatgpt subscriptions don't have access to it - i was under the assumption that OpenAI doesn't lock down their models based on harness a la Claude Code
Plus. No wonder - i suspected this but i couldnt find any docs. Side note, how are you liking Pro? I have really been considering getting Pro recently, but not sure if its more worth to just switch to openrouter. I feel like my usage currently barely outstrips the plus usage limits and Pro would be too much, and using Openrouter by default would mean I would have a lot more leeway to run more random lighter workloads without worrying about using up my limit, but I'll really miss GPT 5
I find it to be excellent. I have three Pro subscriptions so I can build stuff 24/7. I only use GPT 5.5 xhigh. Before 5.5 I wasted a lot of time with bugs. I want to make sure that doesn't happen again -- if I can help it.
I have a pretty good use case for gpt-oss. The amount of time savings has actually been wild. Definitely worth a try. Just to be clear, it gets like 2000tok/s
I've always eyed Cerebras but never had a use for it that would justify paying for the API directly. Although now that I think about it, trying out the API would probably cost less than a subscription for a month...