Measured LLM inference speeds on Apple Silicon, with raw data (CC BY 4.0)

13 pointsposted a day ago
by bagdaerdev

7 Comments

kennywinker

18 hours ago

Why do no benchmarks show qwen3.6?

As far as i can tell the state of the art small open models is qwen3.6 and gemma 4, yet they rarely appear in benchmarks - even ones made recently

isomorphic

17 hours ago

Agreed. Here's a start for posterity. Using their prompt, "Write a 300-word explanation of how attention works in transformer models, aimed at a junior developer. Use one concrete analogy.":

Note that my mini is the bigger Pro model, so it will be faster than the mini quoted in the article.

Mac mini M4 Pro (cores: 10P/4E/20G) 64GB Tahoe 26.6 LM Studio 0.4.20+1

Qwen3.6-35B-A3B-MLX-4bit: 78.81 tok/s, TTFT 0.93s (but note Qwen 35B is really chattery and outputs 3,451 words of thinking for 65s first)

Qwen3.6-27B-MLX-4bit: 14.65 tok/s, TTFT 0.76s (Qwen 27B output 3,027 words of thinking for 361s first, spinning up the fans)

gemma-4-26B-A4B-it-QAT-MLX-4bit: 64.93 tok/s, TTFT 0.44s (Gemma 26B is much more on-task, thinking with 591 words for 15.62s first)

gemma-4-31B-it-QAT-GGUF Q4_0: 12.25 tok/s, TTFT 1.65s (Gemma 31B thought with 404 words for 51.55s)

bagdaerdev

15 hours ago

Fair point. The first run covers what our customers deploy most via Ollama today, which skews to the Llama / Qwen 2.5 / Mistral / DeepSeek families. Qwen 3.6 and Gemma 4 are top of the queue for the next run same method, same raw JSON. Which quants would be most useful to you: Q4_K_M only, or Q8 as well?

kennywinker

8 hours ago

I knew there had to be a reason. I guess if you build a functioning system on qwen 3.5, upgrading to 3.6 isn’t necessarily worth the engineering effort.

For me, i don’t personally know anybody with enough vram to be running more than a 4bit quant - so that’s my line.