Better prompt caching for GPT‑6

3 pointsposted 11 hours ago
by mehrdadrad

2 Comments

OutOfHere

9 hours ago

The biggest continuing limitation I see is that the cached input has to be at least 1024 tokens. This is terrible. It means a lot of good prefixes that are smaller will go uncached for no good reason. The threshold should have been 128.