Benchmarking Qwen3.8 27B quantizations: 4-bit holds up, 1-bit collapses

3 pointsposted 9 hours ago
by stared

2 Comments

mrrrcs

7 hours ago

One run per cell is more noise than the gap you are measuring

stared

35 minutes ago

These was some code budget there.

But as from seeing various runs, errors bars are gross overestimation (as not "the same test", but "if we have different tasks from the same sample").