stackfrost
10 hours ago
You are yourself in some pay paying for the compute right? Isn't this fix price an opportunity for many to abuse?
How are you managing rate limits and what model are you using? if you would like to disclose that.
wilsprouse
8 hours ago
Yeah, of course we are paying for the compute on our end. I think we are waiting for our users to stress test our theory, however our secret sauce is optimizing inference so that this is a feasible product on both ends.
Like I alluded to in my initial comment, I've never understood why the inference providers of today charge per token. Whichever marketing department got our industry to be ok with being charged for the output of a REST call, I applaud.
I'm not sure it will all hold up at scale, but I am still waiting for the test to break. It helps to be a curious and optimistic person in this endeavor.
Not currently implementing any rate limits. Using an assortment of open models.