Hackernews
new
show
ask
jobs
Cutting LLM inference costs by 36% with prompt caching
2 points
posted 7 hours ago
by lizakatz
(neradot.com)
No comments yet