Hackernews
new
show
ask
jobs
Attention Became Efficient and Scalable: KV Caching, MQA, GQA, MLA, and DSA
1 points
posted 8 hours ago
by ibobev
(chizkidd.github.io)
No comments yet