thinkingmachines/Inkling-Small

8 pointsposted 5 hours ago
by Philpax

1 Comments

andy99

5 hours ago

I didn’t see a gguf yet, going to be most interesting if there’s a quant that fits nicely into about 90FN so it can run in 128GB unified memory.

It’s 12B active so should hopefully be pretty fast at 2 bit quant if it fits, as in the ratio of total to active params is big compared to e.g. Qwen 3.5 122A10 which seems more commom

Edit: and it’s out - excited to try https://huggingface.co/unsloth/Inkling-Small-GGUF