Hackernews
new
show
ask
jobs
Hitting 1B tokens/minute on 1 GPU combining a query planner and inference engine
1 points
posted 6 hours ago
by gmays
(modal.com)
No comments yet