Show HN: LLM Inference Calculator – Estimate VRAM, Latency, and Throughput

6 pointsposted 5 hours ago
by popopanda

2 Comments

maestroquirk

3 hours ago

Need one for vision models too tbh. Token/s doesn't really map easily

popopanda

2 hours ago

Yeah, tokens/s varies a lot based on workload. However, I’ve calibrated the estimator against public benchmarks, so it stays within a 30% error margin!

I'll definitely explore how to estimate vision models next!