Thoughts About Scaling Law

6 pointsposted 11 hours ago
by tosh

3 Comments

jarbus

5 hours ago

Maybe just me, but I have a feeling that we are currently stuck at 100-200k effective context window due to a parameter limitation, specifically the hidden dimension of the model, which would be an incredibly expensive dimension to increase vs adding optimizations elsewhere. I don’t think we are done scaling parameters yet.

kzrdude

2 hours ago

Why is it connected to number of parameters?