Hackernews
new
show
ask
jobs
Flint: Efficiently Leveraging High Bandwidth Flash for LLM Inference
2 points
posted 5 hours ago
by metrofun
(arxiv.org)
No comments yet