Using KV cache as embeddings

2 pointsposted 12 hours ago
by mxudev

1 Comments

mxudev

11 hours ago

What if instead of one giant vector as embedding, we use multiple (K, V) pairs. In this work we demonstrated this is feasible, and got reranker behavior at retrieval cost (without extra backbone pass)