sharih
3 days ago
What is the point of this, if it is p90 17 seconds? Might as well use an LLM. The beauty of Jev is that it is dirt cheap and insanely fast.
zihotki
3 days ago
I would hold your horses to paint it as dirt cheap.. In my cases for spam detection Luna was 20% cheaper due to prompt caching, although not as fast.
nico
3 days ago
For email you can use a classifier
One way: separately embed sender, recipients, subject, body - then use the embedding vectors as input to a logistic classifier
With that setup, I get 95% accuracy on email classification, training on 50-100 base examples. The model trains on CPU in under 1min, and it does inference in under 20ms (most of it is running the embeddings, so you can make it faster if you train your own embeddings model)
Here’s a gist with some sample code: https://gist.github.com/nicobrenner/056a5aaff5d0119c0032ecda...
That code applies the embeddings + classifier setup on the Banking77 dataset. It gets 93-94% accuracy depending on the embeddings you use (SOTA for this is ~95%, with much bigger and slower models)
janalsncm
3 days ago
I can’t see your gist but spam classification is a textbook example of something you shouldn’t measure with accuracy. If 95% of your samples are not spam you can get 95% accuracy by always guessing not spam.
You should use precision (when your model says “spam” how often is it spam?), recall (how many of the spam emails did it catch), or f1 (balanced between those two).
nico
3 days ago
That's a great point. My case is not for spam, the classes are more balanced, but you are correct that precision, recall and f1 would be better measures for some of these tasks
zihotki
3 days ago
I wonder what numbers you'd get using another system one model - Contrastive Language Model https://contrastive-lm.notion.site/
That model scales very well with quantities of requests.
atombender
3 days ago
> hold your horses to paint it as dirt cheap
For a moment I thought this was going to be a metaphor — maybe an ancient Chinese proverb about how paint brushes are made from horsehair and how you can't hold the horse to paint before you've turned the hair into a brush.
idiotsecant
3 days ago
Darmok and Jalad, at Tanagra
calebhwin
3 days ago
How are you benefiting from prompt caching for simple classification?
zihotki
3 days ago
There are two parts in the data you supply to Jev for classification - the prompt describing your classification and the data. The data can be quite small - a simple chat message. And prompt part could be considerable since you need to describe your rubrics well.
With Jev you each time pay for your prompt, you can't cache it.
sarkarghya
3 days ago
I mean, it sounds like it's only ideal for cases with significant system prompt overhead. I don't think Jev was built to have a large well described prompt setup. To me its more like a happy go lucky small label classification tool with important decisions left to stronger agentic models or yk humans.
jedberg
3 days ago
Are you getting better performance from an LLM than a Bayesian classifier?
StarlaAtNight
3 days ago
BLASPHEMY! OUT WITH YOU!
HawtAds
3 days ago
How many requests per second do you have for spam that you are reliably hitting the Luna cache?
simplisticelk
3 days ago
Is that just because the Jev implementation is less mature? Couldn't it also implement prompt caching?
tyre
3 days ago
What are the costs compared to an ML model?
catlifeonmars
3 days ago
Could you not just copycat jev and run a fast, small local model?
olgava
3 days ago
[dead]
amelius
3 days ago
Next step: make it classify the next word.
esafak
3 days ago
Jev ought to offer a flex mode that uses their spare capacity for a discount.