With a single NVIDIA 3090 and the fastest inference branch of GPTQ-for-LLAMA htt...

LoganDark · on June 6, 2023

> IMO GGML is great (And I totally use it) but it's still not as fast as running the models on GPU for now.

I think it was originally designed to be easily embeddable—and most importantly, native code (i.e. not Python)—rather than competitive with GPUs.

I think it's just starting to get into GPU support now, but carefully.

brucethemoose2 · on June 6, 2023

Have you tried the most recent cuda offload? A dev claims they are getting 26.2ms/token (38 tokens per second) on 13B with a 4080.