Pragmata
8 hours ago
I've been using it for the past few days, and it runs really well!
I usually get 7 token/s using llama or lm studio, but this inference recipe runs at a smooth 80 tokens per second.
Genuinely very usable, and fully local!
falsaberN1
3 hours ago
With llama.cpp (CUDA) and a 5060ti (16GB) I get 60t/s with 128K token space. Odd you got 7t/s, did you verify all the model was loaded in VRAM? (--gpu-layers all)
kristianp
5 hours ago
Which card are you using? I was getting about 40 with an UD q3 quant with MTP (prediction) enabled and llama.cpp compiled for my compute capability, but was very limited in the context size. I have an 4060 ti 16GB. Wouldn't recommend it as there's a tradeoff between larger context without MTP and about 18 tokens/s.