I Rented a 96 GB GPU and Took Uncensored Qwen3.8 From 44 to 125 tok/s

6 pointsposted 12 hours ago
by LuD1161

1 Comments

karmakaze

an hour ago

I did a similar thing running Q6_K model and Q8_0 DFlash2 (draft=7) quants:

    DavidAU/Qwen3.8-27B-Cold-Fusion-GAIN-V1.1-NM-DAU-NEO-MAX-MTP-GGUF
    incoai/Qwen3.8-27B-DFlash2-GGUF
using llama.cpp PR/commit https://github.com/ggml-org/llama.cpp/pull/27342 on an AMD R9700 (32GB)