ricardobeat
a day ago
Everyone is doing this to emulate Jev, but...
I took a random book excerpt with 23,000 words (±30k input tokens) and used it as context. Jev still responds in 800ms, sometimes 500ms. That's in the neighbourhood of 20-50,000 tok/s prefill, which is obviously not possible with normal LLMs, not even Cerebras is this fast.
HarHarVeryFunny
21 hours ago
It seems people are just guessing at the architecture behind Jev. Obviously the functionality itself is easy to replicate, but why Jev seems to be making such a splash (beyond the doh! factor of it's huge applicability) is the ultra-low cost and speed, which may be due to architecture.
The Laya model compared in TFA shows one way Jev may be getting it's speed and low cost - by using a BERT-like bidirectional model rather than an auto-regressive one (LLM).
prometheus1992
a day ago
was the answer correct?
i have tested jev for my use cases and its horrendously wrong, but then the follow up from jev's team is "oh, you need to boil the question down further". it's a spiral of how much do you wanna dumb down the ask so that it answers it correctly. i'll pass for now.
also, 30k input tokens is a lot.
amluto
a day ago
I imagine it’s not so hard to optimize a model for this use case.
Off the top of my head, I would skip all the modern linear attention / state space stuff and use classical attention. But run prefill in a fully sliding-window mode so that “state” tokens simply don’t attend to far away tokens, or maybe also allow everything to attend to the first few tokens (and train like this). Now prefill is almost embarrassingly parallel, and you can make it fully parallel by duplicating work at block boundaries. (I’m not saying this is an awesome architecture if you want excellent results, but I’m also not convinced that Jev gives excellent results…)
The let queries attend to everything.
And architect the stack around this. Don’t try to cache the KV data — process the queries as you go so that the each input block and layer’s K and V data is computed, attended to, and discarded.
I’m curious whether Cerebras actually is a good device for this. Cerebras is kind of low on RAM, but if you don’t need to store KV data, maybe the entire computation fits on the die.
reissbaker
12 hours ago
20k tok/sec prefill on B200/B300 isn't particularly noteworthy for medium-sized models like GLM-5.3-Flash, vLLM and SGLang achieve it on a reasonable number of models, especially at NVFP4.
50k tok/sec is pretty impressive though.
But... when you were doing your measurements, were you using the same random book excerpt? If you were potentially getting even partial cache hits for your 50k tok/sec measurement, it would taint the benchmark: pretty much any inference provider running any LLM would be able to hit those numbers.
mmastrac
a day ago
That's not true. I ran Cerebras as an experimental ultrafast Jev and it was faster.
ericpauley
a day ago
This has “/dev/null as a service” vibes…
manojlds
a day ago
Do we have a reliable way to count tokens for Jev yet btw?
stingraycharles
a day ago
Also, it processes all questions you ask it in parallel, which is also not possible with normal LLMs.
Xorlev
a day ago
Sure it is.
The prefill is the only blocking part, and you can prefill the whole context up to the point where they diverge, then prefill each question and decode the one token in parallel for each question.
If you batch vLLM calls with the same prompt prefix to the same process, it'll deduplicate the prompt prefix across batched requests (+/- the block size) and decode in parallel for each.
That's with a vanilla LLM. If you modify the LLM you can pull that in-graph, but it isn't really necessary.