Show HN: Janus – Go binary that runs GGUF models via Vulkan on AMD/Intel/Nvidia

66 pointsposted 10 hours ago
by Maverick617

10 Comments

aidiveyt

21 minutes ago

claude code appends a role:"system" block after the user prompt, so a proxy rewriting the trailing user message is a no-op.

PcChip

10 hours ago

I didn't see any benchmarks against vllm, sglang, exllama, etc

rancor

9 hours ago

Since this is basically a wrapper around libllama.so, I would assume that the performance is roughly the same as llama.cpp upstream.

nullpoint420

3 hours ago

Woof. Wonder if the creator knows that

dlcarrier

8 hours ago

From what I've seen, Vulkan adds a lot of overhead on Intel hardware.

peddling-brink

8 hours ago

> llama.cpp via Vulkan (AMD / Intel / NVIDIA) or CPU fallback

I got excited about someone paying attention to intel. Oh well.

wronglebowski

5 hours ago

What hardware do you have? I’ve been playing with a 258V and OpenVINO has come a longggggg way.

peddling-brink

4 hours ago

Two arc b60s. The intel vllm build is getting me ~15t/s decode with heavy context using qwen3.8 27b.

kamranjon

7 hours ago

llama.cpp sycl and vllm xmx work is pretty incredible right now - you just gotta build it with some extra flags

peddling-brink

4 hours ago

Llama would be nice for the ggufs. Any specific flags or tutorials I should look at?