smokel
11 hours ago
This seems to be a wrapper around llama.cpp with several tuned parameter settings. Most of the speedup is gained by enabling speculative decoding [1]. No amazing breakthroughs, just good tuning!
wowitsbase
10 hours ago
Speculative decoding was on in all of my tests, the gain came from moving the round's inputs off host memory.