kzrdude
9 hours ago
9 hours ago
5 hours ago
Maybe just me, but I have a feeling that we are currently stuck at 100-200k effective context window due to a parameter limitation, specifically the hidden dimension of the model, which would be an incredibly expensive dimension to increase vs adding optimizations elsewhere. I don’t think we are done scaling parameters yet.
2 hours ago
Why is it connected to number of parameters?