pradn
5 days ago
I'm not sure what this means for AI startups if their innovations can be copied by OSS so quickly (what, like 2 weeks?). There's "consumer surplus" for everyone, to borrow an economic concept. But we do ideally want some of the surplus to flow to the innovator, too. I know there were precursors, but that's fine - it's hard to have a totally novel idea in such a popular field. I don't know what the end game is for TypeSafe - they'd need to demonstrate perpetually better results, or compete in another axis: UX, support, custom solutions, etc. So much of the time, someone proving a concept, or it simply getting enough publicity, is enough for a "Cambrian explosion" of follow-ups and copies. Famously, that was true for "Attention is All You Need", and the general idea of "next-token prediction" being so powerful.
We've stumbled into general differentiable models..
redox99
5 days ago
Because what they did is kinda trivial. Its basically like the Dropbox comment really[0], except here you don't need petabytes of storage and infinite VC pockets.
After chatgpt everything in AI mostly became LLMs and building wrappers around them. It's like people forgot how to do ML.
To those of us who actually trained models back in the day, its kind of cute to see people wowed by a classifier. Yes, this is 0 shot and doesn't need training (most people wanting this would've used structured output, this is cool because it's cheaper and faster). But anyone with basic ML knowledge could've built this in a few hours.
The question is mostly why wasn't this productized. And it's interesting indeed that it took this long to become a finished product.
djfdat
4 days ago
Theorizing, but I'm guessing that before LLMs, people weren't using anything for situations like these. As people started using LLMs, people's use cases grew, LLMs got slower and more expensive, and people lost the wow-factor and are now worrying about price. Great timing to launch a product like this, where certain use cases can be distilled down into something faster and cheaper.
There's probably other areas where people are using LLMs where a more tailored ML solution might work better.
CactusOnFire
4 days ago
I am looking forward to some hotshot implementing a jev-style model and then looking like a cost cutting, performance boosting wizard by replacing it with a logistic regression.
andai
5 days ago
> The question is mostly why wasn't this productized.
Probably because doing it wrong (using an llm in place of a classifier) is more profitable? (For the people selling inference.)
eadwu
5 days ago
The answer for why it wasn't productized might just be pretty straightforward.
LLMs still are better than Jev at the task, just across the board slower.
Anyone who had a reason to try this already tried it (ads/recommendations) - back in 2023/2024 during the first fine tuning wave and it was accurately determined that it was not worth the effort, the results were more bogus than just using CoT, so frankly parallelism meant nothing if bogus * parallel = bogus.
So thrown into the dumpster and nobody really cared to revisit because it was already tried.
Pretty much sometime between then and now it somehow became the state where the tradeoff makes sense now.
wongarsu
5 days ago
However in the last 2-3 years LLMs became a lot more efficient. Small models without CoT are still bad, but leagues ahead of where they used to be
Maybe it's just a case of the idea just now crossing the threshold into working just good enough to be worth it
coolThingsFirst
4 days ago
Can you explain how 0 shot classifiers work how come it isn’t trained on my data how does it know when issue is urgent lets say?
dotancohen
4 days ago
This is my burning question as well.
I have 80,000 voice recordings to classify. Very few are in English, and the classes are not in English. Many of the classes are project names or other proper nouns. How could a system not trained on data specific to the problem possibly be expected to work?
atroche
4 days ago
you could help it along by providing clues about the domain in the prompt. including what various proper nouns might indicate (assuming they're not chosen completely at random, and encode at least some meaning).
and you might find the scores it gives back about its confidence are useful for escalating to a more expensive classifier.
but also, there's no requirement to use it as a zero-shot classifier, you can provide it as many examples as you like, and prioritise giving it examples it had previously gotten wrong. and with input caching it might be economical.
I'd be surprised if they or others don't start offering a fine tuning API for models like this, like openai does for some models (or used to, I haven't checked in a long time).
dotancohen
4 days ago
If it's not trained on my data, it's going to be something on the order of a 500-shot prompt. This is a complex application with more corner cases than I'd like.
As are, honestly, most projects I work on.
There are already Jev-shaped open weight local models coming out that can be fine-tuned on specific data. We'll see which of them the community settles on.
danwitt
4 days ago
It's not magical and you won't get good results if you ask a question that is that vague. Instead you will want to decompose it in to decidable questions that make up "is urgent" and if you find a case that corresponds to "is urgent" you add a new question to the panel that surfaces that.
nickstinemates
5 days ago
It wasn't until recently that demand for classifiers at this scale existed. Jev exists because LLM's exist. Without them it wouldn't be (as) useful
jmalicki
5 days ago
That's not at all true.
There have been tons of applications for this. People were using earlier LLMs like BERT for classifiers long before LLMs became viable chatbots.
janalsncm
5 days ago
I think GP has a point, though. BERT models have existed for a while but OpenAI made classification via LLM convenient and accessible for regular developers.
People didn’t know they wanted classifiers until OpenAI gave them a taste.
calebkaiser
5 days ago
I don't know if that first part is true? Classifiers were/are one of the dominant applications of classic ML and neural networks, especially in production. Even today, image classification, object recognition, language detection, segmentation models etc are still super common.
I think the hype with Jev is just that, while structured generation is great, LLM judges tend to kind of suck for precise classification. And the more powerful the base model, the more accurate they can get, but they get increasingly expensive/impossible to finetune. It was specifically the latency/price point Jev offered vs. the general accuracy it claimed that generated all the excitement. Plus the promise of cheap calibration (tuning).
"Jev exists because LLMs exist" is kind of a truism, as Jev apparently is literally a Transformer model.
nickstinemates
4 days ago
How many people were doing ML work pre LLMs? And how many are using LLMs now?
calebkaiser
4 days ago
I think the answers are "a lot" and "a lot more"? But what I'm saying is that Jev's virality isn't because people didn't have access to classifiers before it. Jev's general claims about capability and performance vs. cost would have been a very big deal 5 years ago too. In a vacuum, the idea that you can get a general classifier that is very accurate across any domain and on any modality with minimal latency and a very low price point is wild.
In early 2016, Clarifai's core product was basically just an image classifer exposed via an API. And at that point, they'd raised $40 million--the same amount as TypeSafe.ai/Jev--and they were experiencing viral growth among developers + signing contracts with a bunch of flashy logos. The demand was so high that Amazon launched Rekognition and Google launched their similar APIs to compete.
The AI hype cycle and the number of people thinking about using AI certainly puts more wind at Jev's back, but even in an alternative universe where we don't have contemporary LLMs, Jev's core claims would be remarkable and there would be a big market for it.
Fordec
5 days ago
I think the idea of a "feature startup" is dead. What used to be a niche subscription business is now an individual Epic level of work. The smallest viable business becomes what two or three years ago was a mid tier enterprise. It is no longer "look at this tool I maintain", but "we take this specific approach using these hundreds of tools merged together to solve a problem in a specific way that nobody is going to compete with. Not because they can't compete if they wanted to, but that the competitions approach diverges in fifty different chosen ways that they are targeting a different market segment essentially."
I adhere to the idea that this is software's "Tower of Babel" moment where everyone just fundamentally ships things in completely diverging architectures, because creating a ground up architecture is no longer something that needs to be avoided for an economically viable business mode that in the past two decades would have otherwise incentivized people into industry standards. In a world where "taste" is the focus, single ingredients in the recipe aren't enough.
totetsu
5 days ago
Are you saying laya copied from jev, and released in two weeks? If so I don’t thinks it’s quite as simple a story as that. https://xtxinversexty.com/layas-prior-art-claim-is-absurd/
cgio
5 days ago
It a paradox when the article is claiming the prior art is absurd, but then goes on to analyse the one side and compare it to another for which most of the values (except scaling the concept) are unknown. And even for scaling, it uses the first, pre-laya instance to judge the limited schema, while overlooking that Laya is just doing this scaling. Important to note that prior art is not having built the exact same thing.
kensai
5 days ago
There is definitely more to the story. There is a huge financial interest for each side to discredit the other. Fact of the matter is, we still don't know who will prevail. These are cutting edge tech stacks and they were just released.
janalsncm
5 days ago
Presumably the training recipe and training dataset itself cannot be easily copied in a week or two. So if they want to shut down these competitor models they need to make it obvious how they are better than them.
rosegroove
5 days ago
[dead]
mattstir
4 days ago
> I'm not sure what this means for AI startups if their innovations can be copied by OSS so quickly
This particular "innovative" concept already has a rich, open research background. What Jev appears to have done is scale that up a bit and isolate good training data, which results in a great product but not really something impossible to imitate. The only major difference currently is that the open source decision models need to be fine-tuned as they're not trained off of the entire internet yet.
cobzilla
5 days ago
This is precisely why we need strict government regulation of open weight models. :wink:
cung
5 days ago
Wait so a startup that didn’t innovate much should have a bigger moat? If their work can be copied in 2 weeks maybe they don’t have anything?
scotty79
4 days ago
Didn't llaya come first? So the entire Jev's innovation is encapsulating idea in a cheap service?
tukHelix
5 days ago
Today we might need to evaluate “innovation” in a new standard, and have a different expectation for what innovator would be awarded. Getting public attention in such an era where innovation happens every a few days could’ve already been something precious. And that attention would allow TypeSafe to be heard easily next time. Like OpenAI, Anthropic, or any others, they launch frequently but still each time they launch something new, that would hit headlines. I think that’s the “surplus” flown to innovators today.
_menelaus
5 days ago
The moat is the RL synthetic data pipeline they set up to train jev. Open sourcing that would be the coup, not the model architecture and training scripts, which are trivial.
cjonas
5 days ago
Is it likely not just distilled from one of the flagship models? That sounds like an afternoon of work and a few thousand dollars in tokens?
locknitpicker
5 days ago
> I'm not sure what this means for AI startups if their innovations can be copied by OSS so quickly (what, like 2 weeks?).
My thoughts too. It sounds like a minor feature being framed as a whole new business.
Then again, Dropbox and Docker are too.