Show HN: Fast inference for deep seek flash v4.1 469 tok/s for coding

3 pointsposted 9 hours ago
by Hiteshjain118

4 Comments

Hiteshjain118

9 hours ago

We built a custom inference engine to deliver fast and cheap inference. We optimized this engine for coding and research workload and achieving an average speed of 469 tok/s and 4.5x lower costs than Openrouter. Software development at that speed feels different. Get an API key and try in your Opencode!

digestainews

9 hours ago

What are the limits of the free version?

Hiteshjain118

9 hours ago

We offer $5 free credits. But here's a $20 code for sign up from hackernews: HACKERNEWS20

ryanjosebrosas

8 hours ago

We tested it earlier! Speed and cache are all good!