Qwen3.8-Flash-Next non-uniform quantization runs on 2 RTX3090s

4 pointsposted 10 hours ago
by pfeifferj

1 Comments

rcbdev

9 hours ago

Didn't expect to see the 180B model running on two 3090s so fast. The open model space is moving crazy fast, improvements like always remind me of Google's 2023 memo 'We Have No Moat', which looks to be more right every day.