Cutting LLM inference costs by 36% with prompt caching

2 pointsposted 7 hours ago
by lizakatz

No comments yet