Hacker Newsnew | past | comments | ask | show | jobs | submitlogin

150k TPM limit on public endpoint means that it's likely unusable for many coding tasks. When we've tried Cerebras in the past, our problem has always been rates. We'd love to not deal with dedicated and to have access to a more flexible rate pool.

Even trying it out, it seems like our account has gotten moved to some limbo where we can no longer add billing information.

``` Billing access restricted Self-serve billing is not available on Enterprise accounts. Please contact your team for further questions. ```

We have no team (they removed themself from our slack channel after we talked about rate limits). Perplexingly, none of this even shows up in the request, which gives:

``` {"message":"Model does not exist or you do not have access to it.","type":"not_found_error","param":"model","code":"model_not_found"} ```

When the error is really about billing.

I always want to like Cerebras, but I get the vibe that as a tokens in tokens out consumer you are not valued at all.



> 150k TPM limit on public endpoint means that it's likely unusable for many coding tasks.

I don't understand. How does that make it unusable? Is the limit shared by an entire team at once?

150,000 tokens per minute is a lot. You could start hitting that with a lot of concurrent requests in your session, but even throttled to 150k TPM it's still going to be faster than anything else you find.

I think the 128K context limit is the real ceiling. These models aren't amazing at long context, but once you account for a short input prompt, the input files, and headroom for a compaction summary, there isn't a lot left for the problem.


It's a limit on input tokens. So that's 3 50k requests per minute. At Cerebras speeds, that's about 5 seconds of usage per minute.

I was very excited last year for their coding plan but seeing a burst of requests pulse and then sitting there watching the cooldown reset is really not a great time.

Even though each individual request was fast, the sessions were only maybe 10% faster on wall clock time since there was so much waiting time.


They made the coding plan a bit better toward the end, but it was pretty tough to use throughout.

Seems like an Amdahl’s law of inference economics? there’s so much compute relative to SRAM on the chip and shoreline bandwidth onto the chip that caching buys ~nothing? The contended resource is SRAM and a given token of context needs just as much as another.


That’s not what caching is for. Caching lets you resume with a pre computed KV cache saving you from having to ingest everything in the chat history as input on every single round trip. You still need caching regardless of SRAM or not as it saves a huge amount (and ever growing) of compute ingesting the preceding history every time you want a completion.

I don’t know why they don’t give a price discount. Maybe their hardware is incapable for some reason of saving/restoring the state? Or maybe they just haven’t built the infrastructure to do it?


> Maybe their hardware is incapable

I am unsure if it is incapable, but it sounds hard.

They have tons of cores with 64k of SRAM each, and relatively slow paths in/out.

On a GPU you can leave it in SRAM. On Cerebras, you have to send it out of the system which is a giant bottleneck.


Can't you do something with multiple accounts?


You would lose caching (if they cache)


Or just buy a 5060. This will run on most any 16gb card. Slower for sure but far cheaper than another subscription.


Thats not the point if you choose Cerebras as provider.


Or buy a raspberry pi with a SSD, about the same difference, if you're giving up on the 1500 tokens/s anyways.


The whole point is the speed


128k context is not a limit of the model, that's a limit of implementation:

"Context Length: 262,144 natively and extensible up to 1,000,000 tokens."

https://huggingface.co/Qwen/Qwen3.8-27B


We're talking about the Cerebras implementation, which is limited to 128K.

It's in the link.


Yes, and I said its a limit of the implementation and not the model.


TPM means Tokens per Minute.


GP is referring to GGP’s last paragraph. 150k t/m, yes, and 128k context.


150k tokens per minute at 1.5k tokens per second means you can have like 3 users concurrently and that's not a lot.


150k by account. At 1.5k a second you hit it very quickly.


Exactly, it burns the tokens 3000x faster, which means the budget ($$$$$$) runs out so faster it will stop super quick, not able to perform long-duration work. At 27B parameter size, the intelligence is not able to accomplish work within a short amount time. Consequently, it become not usable.


I (we) run Qwen3.8-27B-FP8 on a DGX Spark box - that's roughly £4000 of hardware.

I did benchmark it in various ways and it runs quite well but it is a quantised jobbie and 1.5k t/s is also rather faster than anything I can possibly hope to achieve.

To run that model at those sorts of speeds is going to need some serious investment and you are going to have to pay for it.


The problem is most providers hit tok/sec limits really fast. 1m/min is the default and the only place I can get 10m+ is from first party providers without a lot of upfront cash.


How fast is it?


Parent already responded but just for reference an RTX 5090 with Ninfer hits 160 tokens/second with qwen 3.8 27B which is very usable.


With MTP and FP4 I max out at 30ish t/s on mine. Without MTP or in regimes where the drafter performs poorly it’s about 10 t/s. FP8 is about half that


Thank you, always nice to see real world performance figures.

We run a pretty large rig, 10 GPUs right now (this goes up and down with various experiments, getting this many GPUs to play nice at x16 GEN4 with any motherboard is a challenge), 240G VRAM in total. 256G RAM and a TR PRO. For small models the comms overhead is larger than the gains so there I have to reduce the number of active GPUs. On this machine I'm getting between 150 and 200 tg/s with FP8, but it took a lot of time and tweaking to get to that, and not all of the improvements held up when combined with other improvements. I've been playing with this stuff for a while now and it is interesting how fast the frontier is moving and how much you can now do on your own hardware. For larger models the communications overhead is low enough that we can run them on bigger groups of GPUs, and using hacked drivers to give us p2p capabilities on some of our GPUs also boosts performance considerably once you start to hit communications limits. Typically we get 50G/second in p2p mode (full duplex, half that one way).

From a cost perspective running locally is not interesting, but it allows us to do experiments that model providers would likely balk at, gives us censorship free access and allows us to work with data that we would not want to share with model providers (or can't share due to NDAs).

I will look into running ninfer, I was aware of them but had not yet gotten around to using it.


Unusable is too stretch IMO, you can still use it in tiny tasks that’ll respond almost instantly


Yeah, their public service isn't a serious/competitive offering. They don't have the capacity to serve all the customers who might want to use them at that speed. The public service exists so they get some users on OpenRouter, and that shows them as #1 on speed, which proves their tech is very fast, which gets them billions in hardware sales/licensing. If you have big enough pockets they can probably dedicate capacity to you. But for reliably fast small models you might want to rent some GPUs.


GPUs can't reach these speeds. You could build a supercomputing cluster and still not reach these speeds.


MiMo-V2.5-Pro-UltraSpeed gets pretty close with over 1000 TPS on 8x B200. It has 1.02T total parameters and 42B active, compared to 27B total/active for Qwen3.8-27B. Also, B300 are out now. I think 1500 TPS for Qwen3.8-27B should be doable.


That model uses a lot of tricks to achieve 1000 t/s. I would not use raw parameter counts alone for such comparisons, in general.


Celeris reaches ~50% of the speed on commodity hardware. celeris-magnus-1 is based on qwen3.8-27b. Maybe we will get there without custom silicon!


all the above. They just simply do not care about non enterprise customers. Today they announced qwen, guess what - it's also the same day they pulled Gemma off their shared tier. No migration notice and all developers are scrambling as we speak trying to migrate. They gave a soft head-ups on discord a week ago and when folks complained about zero-day migration they started saying 'you aren't suppose to build production app on shared tier'.


On discord? Jeez. How professional.


Cerebras the tech is awesome, cerebras the company is a trainwreck


Is this the chatjimmy asic approach with a bigger model?


No, the asic could only ever run one model/set of weights, no updates possible, ever. These are general purpose processors that can have their models updated. But the chips are enormous, with a substantial amount of on-die memory alongside the execution units, for a relatively insane amount of memory bandwidth.


I thought from what I read about the Taalas approach, the model architecture and overall size couldn't be changed, but model weight values could be updated after for further tuning.

Not as flexible as Cerebras though. And I'd love for someone who knows more to clue me in to the truth.


Nah, Taalas was putting the weights into silicon as a mask ROM. Their demo chip was hardwired to serve Llama 3.1 8B, and could never be updated. New models, even new versions without any architectural/size changes meant new tape outs.

But in exchange, you get insane speed and great energy efficiency. I could see it being a great approach for basic "good enough" models.

They may have had a little flexibility by supporting finetuning via LoRAs.


To qualify "insane speed" for anyone unaware, think a 10x improvement over even Cerebras. On the order of ~15,000 tokens/s. Not saying their approach scales well enough to keep pace with the various frontiers, but using their demo alone feels like a paradigm shift.


Yeah, I'm guessing this isn't unique, but I remember the first time I used ChatJimmy, I missed the fact that it had responded because I was still hitting the enter key, and getting ready to see tokens stream in, but they were already all sitting there, and I'd missed registering the visual diff somehow.


This seems like it would be an altogether different experience than the common experience of using an LLM, which is characterized by the person spending a lot of time waiting on the machine.


Yeah, it'd be a lot easier to maintain flow, less need to work on more than one session at once, etc. And then tool calls would be the limiting factor, especially network access. I hope AMD keeps the project moving forward post acquisition.


Yes and no. A single chip cannot have it's weights adjusted, once it's out, it is what it is.

But also, the model weights are in a single mask rom layer, high up in the metal stack. They could manufacture the die specialized for a given geometry of a model up to that layer, wait for updated weights, and then get the final product out in weeks after they got the weights, instead of many months which is what it would take to redesign the whole chip for the new weights.


i hope groq wins if they start doing such things with consumer.


This was my experience a year ago on some other model they could run super fast. Routine coding tasks would hit the per-minute token limits.

Just the math there... 150k TPM... and 15k TPS means... you can run for 10 seconds every minute?

The basic math boggles the mind.


Not sure how the rate limiting works, but it's 1.5k TPS, not 15k, so you could run it for 100s/min, which seems good enough to me


iirc input (uncached) goes towards the limit as well


What's the tok/s when they process input?


It seems you forgot to account for the fact that cerebras uses a baker's minute which is 144 seconds instead of 60. (Seriously though what's the supposed issue here?)


The issue is that all input (including context) counts towards that limit. So 10 requests with 50k of context will blow through the limit, even if little to no output was generated, which is incredibly easy to do with agentic workloads.


ah, yes, that seems right

I was using it quite a while back, different model, different quotas, but for coding tasks it routinely hit quotas which made it quite difficult to actually use.

100s/min seems pretty poor actually with sub-agents etc.


> means... you can run for 10 seconds every minute?

It’s one order of magnitude less TPS, but still, that’s the limit with just one user…


What kind of coding tasks would you expect to hit that limit? In my setup, on a very large codebase, it takes each agent 3-4 minutes at minimum to go past 100k tokens.

(note it's 150k uncached tokens, the total limit is 450k/min)


in my last tests with cerebras for coding tasks, most large tasks or anything greenfield would hit token limits. note that smaller models and the gpt-oss-120b style models they used to run are very prone to overthinking, so individual turns may be 3-10k tokens of just thinking + input + output.

i don't think it's quite apples-to-apples to compare to a frontier model or even a k3. the odds of success (file compiles? read the right context?) are lower and thinking is longer.


So that’s about 400 tok/sec. Times that by 3, you get 100k in under a minute. That’s doing nothing special and just using your current setup.


150k is 2500 tokens/sec.


use 3rd party marketplaces. cerebras is resold on vercel, openrouter and huggingface.




Consider applying for YC's Winter 2027 batch! Applications are open till November 2.

Guidelines | FAQ | Lists | API | Security | Legal | Apply to YC | Contact

Search: