Introducing System One Models and Jev

804 pointsposted 7 hours ago
by albelfio

265 Comments

jacobgold

6 hours ago

First, congrats to the team on launching something genuinely interesting and new.

Seems like a more accurate title would be "Jev: Trading general purpose generation for fast typed inference" or something like that.

This is interesting, but the speed comparison seems misleading? A generative model that can output code in a Turing-complete language can do anything a computer can do.

Jev can only generate structured output, right? This is probably super useful for classification/routing/scoring, but it's nothing like the code generating models we're all using today for code and automation.

Also "can't hallucinate" seems wrong? Sure, it can't emit an invalid type, but it can still emit a completely wrong valid value. You can enforce structured output from an LLM too, with an appropriate harness, etc.

Assuming there's no funny business, the Doom demo is cool.

dbbk

5 hours ago

When they say "can't hallucinate" they mean they produce a confidence value for every result, so you could see for example it has 0.1 confidence, and you can disregard the result - that'd be different from hallucinating where it believes it's correct

8note

2 hours ago

if it puts a high confidence value on a wrong answer, thats still hallucinating, no?

llm hallucinations are high probability tokens that are incorrect vs the real world

dozerly

16 minutes ago

Yes, there is no magic sauce here that makes stochastic output binary if that’s what people are looking for.

rpunkfu

a minute ago

It’s not what people are looking for, but what they wrongly claim.

resonious

9 minutes ago

Right and so maybe we should stop saying "can't hallucinate" when it can by definition.

janalsncm

5 hours ago

Technically speaking when you send the prefix “The capital of France is “ into an LLM it will also produce probabilities across its whole vocabulary.

sothatsit

3 hours ago

The probability values don’t really represent confidence in modern LLMs though, especially after RLHF and RLVR.

System One says they use RLCD, Reinforcement Learning for Calibrated Decisions, which presumably has accurate probabilities as an explicit optimisation goal.

jiggawatts

3 hours ago

… which they could provide in their APIs but are vehemently opposed to because it makes distillation much easier, and faster.

bigglebear

2 hours ago

Yeah. Yet another reason why open-weight models are better. If I want to use the logits, I can.

CompleteSkeptic

5 hours ago

that's right, but because these models are probabilistic, it's also possible to be confidently wrong (and all future models will be smarter still and still have that possibility)

orbital-decay

4 hours ago

Yeah but what stops it from producing confidently incorrect outputs...

zenlikethat

3 hours ago

Nothing, but imagine using LLMs for a classification task

People out there are so resigned to the models being unreliable that they are really doing things like hallucinating deliberately, and then matching the hallucinations to embeddings -

https://softwaredoug.com/blog/2026/08/10/hypothetical-classi...

You could do that or you could just... use a model that will never produce unreliable outputs in the first place.

threecheese

3 hours ago

But we're going from "Apple" to "Apple: 99% - trust me". It could still be an image of an orange :)

zenlikethat

3 hours ago

It's pretty darn smart. If you did want to hack on it in earnest and find out for yourself, send me an email - nathan@typesafe.ai

janalsncm

5 hours ago

I don’t think it’s misleading if you compare on the use cases they suggested. It’s faster and cheaper (no idea if higher quality), so it’s immediately interesting for certain things.

And if you buy their RLCD claims, this might be even better than huge models that know a bunch of irrelevant things.

WhitneyLand

5 hours ago

What was misleading was the original title:

"Jev: New frontier model 40-400x cheaper and 20-200x faster"

I'm not the gatekeeper of who gets to call themselves a frontier model, but I don't think most people would count Jev in that group. It sounds false.

If their specific claims hold up, then it would make more sense to say something like:

"Advanced the speed/cost frontier for structured decisions"

sroussey

4 hours ago

I dunno, I would consider Waymo and Tesla to have frontier models.

I think AlphaFold and related are also frontier models.

Being an LLM does not seem like the qualifier for frontier.

riknos314

2 hours ago

This is likely still an LLM (in the purest definition of a language model with relatively many parameters) since the inputs are natural language, just not a generative LLM as the output is something other than more language.

nickdonnelly

an hour ago

The inputs aren't natural language. https://docs.typesafe.ai/primitives

cooljoseph

10 minutes ago

The inputs are natural language, they're just also structured into a tree. The first example on that very page shows natural language instructions:

    questions = {
        "refund_requested": Noul(
            instructions="Does the customer request a refund?",
        ),
    }

janalsncm

4 hours ago

Large language models are not the only type of model.

alfalfasprout

4 hours ago

How is this not a frontier model? It's bleeding edge in its own niche. It's not a frontier LLM; however, applicable to many of the things people use LLMs for.

bigglebear

2 hours ago

It's nothing like a traditional LLM and so should not be compared to one. It's a heavily constrained, tiny model that can only produce a probability score or a yes/no answer over pre-defined selections. It has no long-context capacity.

I mean, imagine comparing this thing to Astra, it's hilarious. They don't even tell you what the max input size is, and they only allow 10 possible answers to choose from for the Choice mode. It's probably like a 1billion param model. They say it's "not small", but there's zero reason to believe that.

I suspect someone will be able to recreate this within a week by piecing together open-weight models.

janalsncm

an hour ago

> It's nothing like a traditional LLM and so should not be compared to one.

Frontier LLMs are expensive jack of all trades. You can absolutely compare them to purpose-built tools on any domain they touch. Engineering is all about assessing tradeoffs.

riknos314

2 hours ago

Has LLM become so synonymous with Generative Transformer that other high-parameter count models that interpret language need a different name?

For all we know this might be a non-language-generative transformer e.g. a transformer where the decoder produces confidence scores rather than language. Please provide more likely architectures if you know them, I'm genuinely curious.

Flere-Imsaho

6 hours ago

> Jev can only generate structured output, right? This is probably super useful for classification/routing/scoring,

My first thought was that it would be ideal for robotics? As in control of limbs, general planning, route finding, etc.

copperx

3 hours ago

Um, Isn't SELF DRIVING the elephant in the room?

aryamccarthy

an hour ago

Only if you think that everyone cares about self-driving. Lots of niches require structured domains; self-driving is just one that has a lot of capital thrown at it.

CompleteSkeptic

5 hours ago

I'm biased but I wouldn't call it misleading - generating text is super awesome and flexible, (we describe that in the blog post - and I personally use string models all the time) but it's true you pay a high tax for autoregressive generation

> Also "can't hallucinate" seems wrong? Sure, it can't emit an invalid type, but it can still emit a completely wrong valid value.

that is likely true of all ML! perhaps we could debate semantics, but I don't think it's fair to say a random forest "hallucinates" in the way LLMs do

WhitneyLand

5 hours ago

His claim was that the title is misleading, not sure how it's relevant to that claim that you use "string models" (full LLMs).

The original title before it changed less than an hour ago was:

"Jev: New frontier model 40-400x cheaper and 20-200x faster"

I'm going to agree that was misleading.

And on the second point:

>>Also "can't hallucinate" seems wrong? Sure, it can't emit an invalid type, but it can still emit a completely wrong valid value.

>that is likely true of all ML! perhaps we could debate semantics, but I don't think it's fair to say a random forest "hallucinates" in the way LLMs do"

Also going to disagree here, and I don't think it's semantics.

Type safety is not factual correctness.

CompleteSkeptic

5 hours ago

> Type safety is not factual correctness.

I very much agree with this and want to hone in on where do actually disagree. Would you say a linear classifier hallucinates?

InsideOutSanta

an hour ago

A hallucination in the context of LLMs is generally understood as an incorrect answer presented as factual. If you claim that "x can't hallucinate" in the context of LLMs, you're saying that x always gives accurate answers. It does not matter whether the answer is type safe. If its value is incorrect, it's a hallucination.

thduabmd

2 hours ago

No. Your launch post puts “0%” on a hallucination chart, then explains that the number comes from guaranteed schema matching.

You’ve already agreed that this doesn’t establish correctness. An approve for an unauthorized action still meets the schema guarantee.

That’s why I find the messaging misleading. You’re acknowledging the limitations in these replies while defending the broader reliability pitch.

Even granting that each answer is calibrated individually, that doesn’t establish calibration of the decision that combines them.

Sure, I can threshold a composite score, but there may be many wrong answers with the same score. An unauthorized action doesn’t become acceptable because it scores highly on the other dimensions.

I still have to define the constraints and test which wrong actions get through the complete workflow on my own data. That’s a substantial part of the work being pushed back onto the developer.

bigglebear

an hour ago

User input: "Hey, have your human support agent call me, tomorrow at 5pm."

Model input: "Does the user want to speak to a human support agent?"

Output: Yes.

I imagine that your model would produce this, and I think it's fair to say this is a hallucination. A human would caveat it with: "Yes, but not right now.", your model is incapable of that. Yes is technically correct, but within the context of being in a live chat, a human would understand that the caveat is required.

elcomet

5 hours ago

Hallucinations were defined in the context of text generation models so your question does not really make sense.

IMO your system can make mistakes that are similar in spirit to hallucination (i.e. answering with a false answer instead of abstaining to answer).

bigglebear

an hour ago

And furthermore, because the model is forced to answer in a boolean (if in boolean mode), if the user input is outside of the range of a boolean, it's forced to hallucinate. It can't abstain.

8note

2 hours ago

id say yes. a linear classifier that classifies between red and yellow balls will hallucinate on blue.

linear regressions hallucinate in the simpson's paradox.

the model output can be quite confident and not representative of reality

WhitneyLand

5 hours ago

Let's say classifiers don't hallucinate. To make a fair comparison we should constrain LLMs to the same classification task. In that case, no, LLMs also don't hallucinate.

- Give Jev and LLM the same input

- Lock down both to approved/rejected/unknown (LLM restricts on decoding)

- Both can be wrong, but neither can hallucinate (invent an another option).

seizethecheese

5 hours ago

Just to be sure that I understand, you're saying that your model "can't hallucinate" because it only outputs a single thing, right? In this way, an LLM can't hallucinate either if I prompt it to do a classification task with a discrete set of possible outputs, right? (Assuming I reject non-conforming output. Actually, maybe what you're saying is that your system can't output non-conforming output?)

zenlikethat

2 hours ago

Yeah that's precisely correct.

For e.g. classification tasks, even in 2026 people are doing things like hallucinating deliberately, and then matching the hallucinations to embeddings -

https://softwaredoug.com/blog/2026/08/10/hypothetical-classi...

With TypeSafe it just picks the class (actually probabilities across classes), reliably every single time.

zozbot234

5 hours ago

From a quick look at this it looks like it could easily generate natural language text by following a structured representation like UMR (Uniform Meaning Representation) or the similar representation the Abstract-Wikipedia folks will be working on for generic encyclopedic text (which will be heavily informed by Universal Dependencies). These are basically linguistically principled and frame-based counterparts to a programming language AST, that can be then converted to natural language (in a broadly language-independent way, to the extent that semantics and pragmatics make that feasible) via some sort of NLG rendering.

(To be clear, this one raw model does not support outputing a full AST directly - it wants to output "choice" among fixed options, "score" on a sliding scale, or a true/false answer (all of these with confidence scores attached), so building the AST/structure would be a code-driven (or even perhaps outside LLM-driven in some more challenging cases) multi-step affair where the model would essentially be playing a "game" of building the structured output step by step and getting a revised partial state back. But one could expect this to lead to interesting results.)

dfee

5 hours ago

> I'm biased but I wouldn't call it misleading

- @CompleteSkeptic

Very strange.

vvzz

5 hours ago

I feel like the power of the approach presented here is that it gives a model a proper "language" to describe computations directly vs moving tape silliness.

I foresee this to be the path moving forward - giving AI models understanding of the computation directly(as well as compositional rules) This feels like a short path towards total software in many areas.

bigglebear

2 hours ago

Agreed. It's a wildly dishonest presentation of their product from many perspectives, which is a shame because it might actually have some good use cases.

The comparison between LLM speed and Jev speed is misleading, because they're using autoregression to generate all of the type names, all of the schema, etc. A closer comparison would be if the LLM was purely outputting the raw numbers. Even then, comparisons to LLMs are pointless because you could train a transformer on the same sort of task that Jev is doing and get even better performance yet again, and a smaller model. I suspect this is some form of stripped down diffusion language model.

You really have to do a lot of hand holding here, and map out your problem space manually, and very carefully, to get any sort of accuracy. For example:

> Keep each Score to one dimension. If a description says “punctual and smart and experienced”, the question is measuring three things, and an input that is high on one and low on another can’t be placed. Confidence drops and the score means less. Split it into one Score per thing and combine them in code

If you don't perfectly represent the distributions of possible answers then you'll likely get garbage results. As far as probabilistic state machines are concerned, I'd say creating the distributions of possible answers, and their hierarchy, is the actual hard part.

One of their examples is:

- "state": "I have asked three times now. Can I please just talk to a real person?"

- "Is the customer asking for a human agent?"

Imagine the users request is: "I want your human agent to call me tomorrow at 5pm."

Human conversation is fuzzy, getting useful reliable results out of this is going to be a challenge. Of course, you could add follow up checks like: "Do they want that now, or later?" -> if later -> "Do they want that tomorrow, or the day after?" and so on... But now you're building an LLM out of if statements. I am skeptical of whether this model has much utility for fluid language interpretation - I suspect it'll only be useful for scenarios where you've tightly constrained the answer space but want to use fuzzy language to describe it. Like:

- Question to human: "Would you like a support agent RIGHT NOW?"

- Their response: Yes | Yeah | Mhmm | ye sure (any possible yes signal)

Model input: "Did they ask for a support agent?"

Still... a tiny LLM could accomplish this sort of thing without problem. And that doesn't stop someone from saying: "No, not right now. But tomorrow." - and the tomorrow would get missed. I think this is why people haven't really tried this approach much already.

Also their Doom demo is on structured state, not on images. Meaning, the enemies must be being served to the model as coordinates (or the exact angle of projectiles that hit the player), otherwise it'd have to scan every pixel of the 360 degrees to know whether an enemy is in front of the crosshair or not. You can see from the map below that it's also choosing travel checkpoints/destinations through walls. So they've severely cooked this to make it look far more capable than it is in practice, and any speed advantage that is offered here is not factoring in the shortcuts it is taking, the training on the map, and the fact that it can cheat because the structured state it is using is not bound by obstructions.

Here is their docs by the way: https://docs.typesafe.ai/ - so you can understand how it works.

cfowles

38 minutes ago

Wasn't really till seeing this home assistant demo they have (https://www.loom.com/share/18c4dbcf8db546dfb2d7f2ef018e78e4) that the value really clicked for me.

Seems really cool.

ramoz

11 minutes ago

Guess I'm a bit less impressed seeing that for some of the more intelligent driven work -- splitting requests in the video -- they had to kick out to an actual anthropic model.

futurisold

4 hours ago

This, combined with contracts, could make a lot of things so much fun now!

For those who don't know (which is probably everyone but me), I ported the design-by-contract pattern in Python and combined it with LLMs. This was early 2025. I originally wrote about it here: https://leoveanu.com/2025-03-01-dbc/ . Contracts are a core feature of SymbolicAI ever since. The community seems to have loved it too (https://news.ycombinator.com/item?id=44399234).

I think I'm starting to glimpse the implications and it's gonna change agentic workloads if it holds up to scrutiny. It's too early for me to tell anything other than jot down some rough thoughts.

In short, you get blazingly fast semantic branching you can use in control flows. For contracts, I can now directly take the data model that you have to design and convert it into Jev's expected format. Or I can use Jev for semantic branching in postconditions.

If my understanding is correct, that should be doable, but I need to think more about it. It could be that with Jev I can finally “compile contracts” and better chain them into workflows, which is something I always wanted but didn't know how to do properly.

Eager to test. On the waiting list.

zenlikethat

3 hours ago

love it. send me an email and i'll try to get you moved up on the list? nathan@typesafe.ai

maltalex

2 hours ago

This is a very promising idea - a model that takes arbitrary text input (which can be a complex json), plus a set of questions (yes/no, multiple-choice, or score) and quickly (milliseconds) and cheaply ($0.042/MTok) answers those questions.

Unfortunately, none of this is explained in the announcement, but the documentation [0] is pretty good.

[0]: https://docs.typesafe.ai/concepts/how-to-build-with-system-o...

big_toast

6 hours ago

It seems like the docs[0] are a better explanation? The comparison to llm tokens is kinda confusing.

It looks like the model takes as input a state (structured text? not sure if multi-modal) and a question (as a "Choice", "Score", or "Noul") with some additional augmentations possible. Then outputs the question's answers as appropriate (e.g. a choice, accompanying probabilities, confidence).

Edit: On the AI primer page, it looks like they do the RLCD on a pre-trained base model?

[0]:https://docs.typesafe.ai/concepts/system-one

CompleteSkeptic

6 hours ago

CEO here - that is right!

I do agree that the comparison to LLM tokens is hard to understand (also because output tokens are not comparable).

But yes, text or structured state (like a JSON with multiple pieces of text in) -> decisions out (e.g. choice maps to "match" statement, "score" maps to sorting, "noul" short for bernoulli maps to if-statements)

Flere-Imsaho

5 hours ago

Hi - first congratulations, System One looks really promising.

The Doom demo really help me, at least, to understand how System One differs from LLMs. However the first demo (Side-by-side demonstration) - I'm struggling to understand what is going on here!

davideg

3 hours ago

I was confused at first too, but it makes more sense when you read about their primitives. E.g. https://docs.typesafe.ai/primitives/noul

The demo is showing System One producing its output in parallel very quickly and for little cost compared to an LLM generating its answers token-by-token. The "noul" type is used to evaluate a yes/no question and return the probability that the answer is yes.

So this demo is showing System One offering much more nuanced responses and specific probabilities compared to an LLM's more crude responses (e.g. LLM shows "true" or "false" compared to "0.9" or "0.07" probabilities that the answer to some question is true).

potatoman22

4 hours ago

I think that's to demonstrate its speed

mckngbrd

5 hours ago

here is how I attempted to explain it to my company's AI group chat, is this roughly accurate?

"instead of autoregressive string output it instead outputs structured type-safe 'decisions' with probabilities/confidence scores, each generated in parallel

so sort of more like a Large Classification Model than a Large Language Model? or, maybe better to think of it as a sort of "shift left" in the LLM's transformer architecture, allowing you to replace the predefined token vocabulary of an LLM with a prescribed set of 'decisions' that need to be made based off the input context; and exposing those probabilities directly so they can be integrated into the system logic, instead of just sampling from top-K.

all of this while still being instruction-tuned (!!!)"

It's always been possible to build classification pipelines using LLM embeddings as the input. seems like this is a much more sophisticated / useful application of that concept

CompleteSkeptic

5 hours ago

very accurate!

the one nuance I'd get into is I'd call it "zero-shot" over "instruction-tuned" (the latter often implies a particular distribution), but very safe for sharing

ttul

5 hours ago

For many day-to-day computing use cases, Jev seems far better suited than an autoregressive language model, if for no other reason than it is not wasting compute thinking about anything other than how to spit out a decision.

Do you have an architectural explainer yet for Jev or are you holding that close to your chest and letting the magic rip for now?

ianbutler

5 hours ago

I see this super interestingly as the "subconscious" to the llms "conscious" for lack of better terms. I'm super interested in this for broad and rapid decision making in the context of consumer agents so will be signing up for sure.

zenlikethat

6 hours ago

> the model takes as input a state (structured text? not sure if multi-modal)

Input, and criteria/instructions can both be defined as structured input (JSON). This ends up being pretty powerful because the model is trained to understand structure.

e.g.: https://docs.typesafe.ai/primitives/advanced#structured-inst...

> not sure if multi-modal

just JSON... for now :)

> outputs the question's answers as appropriate

correct!

big_toast

5 hours ago

I assume this isn't really for consumers/individuals currently? Kinda feels like an improved magic 8 ball.

I can't really intuit how I should think about when the model will be accurate. Is there somewhere to read more about that? I assume customers would just have some tests or talk to you.

zenlikethat

2 hours ago

> I assume this isn't really for consumers/individuals currently?

Unless they're hackers, no. It's not really a chat interface, it's meant for consumption by machines and composing into higher level systems (pairs great with LLMs).

> Is there somewhere to read more about that? I assume customers would just have some tests or talk to you.

We're going to release some more info on evaluations over time, and yeah, join the waitlist! We offer faster access in exchange for good memes

dgellow

7 hours ago

Side note: it took me more time than I would like to admit to realize that Diogo Almeida isn’t a satirical version of the name Dario Amodei

bogzz

6 hours ago

That would have to default to Wario Amodei.

clayhacks

6 hours ago

I feel like should be Cario Amodei. The D to C flip a rotation of the M to W flip

jakintosh

7 hours ago

It wasn't until the demo videos that I realized the post wasn't satirical.

skerit

6 hours ago

So in theory you could feed it incomplete text, and then ask it for the probabilities of what the next character could be?

mckngbrd

5 hours ago

I think the joke here is getting missed

CompleteSkeptic

6 hours ago

you could, but it the model is not optimized for text

this is complex, but generating text is highly complicated and requires mode dropping to make long cohesive text

vatsachak

6 hours ago

If you provide it an AST of the english language, yes.

consumer451

14 minutes ago

Super cool! Instantly joined the waitlist.

I can see exactly how I could use this right now to improve my agentic rag. In two months I am supposed to deal with a giant corpus, while still maintaining responsive chat UX. I have been working my butt off to make our first big client happy. This could really help solve the chunk ranking problem.

lubujackson

6 hours ago

After much fumbling around with prompts and evals, this is exactly how I am using LLMs in production, to narrowly make choices and return structured data. Any deterministic work gets pulled out of the prompt and my goal is to narrow the model output to be as clearly defined and as minimal as possible.

Jev's focus on structured I/O and confidence scores are game changing. If this does at all what it claims, I think this is going to quickly become the new standard approach for agentic systems.

CompleteSkeptic

5 hours ago

we hope so! the bigger hope is to not just eat LLM market share, but to allow for people to use AI much more in the inner loop of software

copperx

3 hours ago

I'm sure you've thought of self-driving. How does the model work in that space?

ramon156

6 hours ago

This sounds good but so far all claims just sound like marketing terms. I'd love to see real proof. e.g. "RLCD" and "parallel sampling" have nothing to back it up.

also "70-500ms vs 3-329 seconds" are apples-to-oranges unless the LLM baseline is doing comparable work (e.g., long chain-of-thought). If Jev is skipping generation entirely for a narrow structured task, of course it's faster.

Nonetheless i want this to be true, so I'm looking forward to Jev

Edit: I really have to say that I like their manifesto https://typesafe.ai/manifesto

why_only_15

6 hours ago

They have various benchmarks, e.g. how much time it takes them to do wikipedia page -> page games. Jev seems to take the same or fewer hops but in ~10x less time and for ~10x less money.

It's totally reasonable to compare against LLMs doing chain of thought if it gets comparable performance.

yunwal

2 hours ago

> If Jev is skipping generation entirely for a narrow structured task, of course it's faster

I think this is reasonable if people are actually using LLMs to solve this type of narrow structured task, which they are. The evidence is that every LLM provider has some method of forcing the output to conform to a json schema in their documentation.

vatsachak

6 hours ago

It's not an LLM though it's a frontier model on structured data

bigglebear

an hour ago

> I really have to say that I like their manifesto

Their manifesto: "you only build on top of it if it's trustworthy." - the irony of this while putting out the most misleading, dishonest marketing campaign I've seen in months for their first public appearance doesn't exactly scream "trustworthy" to me.

BoorishBears

6 hours ago

Did you see the video where it plays Doom, it made it click for me

simianwords

6 hours ago

BTW it was not multi model playing doom, it was passing structured input and getting structured output. Its not what I thought: frames of video passed and real time game play.

yieldcrv

6 hours ago

so what? put an LLM on Cerebras and get its responses faster, and put Jev on Cerebras and gets its responses even faster

mushufasa

6 hours ago

I would love for things like this to be accessible via hubs like open router or AWS bedrock. It's hard to justify adding new model vendors directly with all the heightened concerns about privacy and security, but if bold new capabilities are added to a centralized already-vendor like AWS, technical people can adopt them without going through a whole compliance/purchasing/vendor review process. And an extra middleman tax is well worth it when the cost savings of the model itself can be one-two orders of magnitude.

varenc

an hour ago

I think the trouble is that Typesafe APIs don't fit into the normal OpenAI-style API that every other regular LLM provider users. You're not just providing unstructured text and getting unstructured text back. It would take a different request and response format than every other model on Open Router. Though you could shoe-horn it in some way, it'd be hacky.

But agreed it'd be very useful to see it deployed on other hubs, and it seems worth it to provide the bespoke API format. Perhaps Typesafe's API will end up becoming the standard for a new type of structured model, the way OpenAI's API did.

cheeze

6 hours ago

Isn't openrouter the exact opposite of caring about security and privacy?

I guess you can choose your provider still? But isn't the point that the lowest bidder is doing inference?

ajmurmann

6 hours ago

You can set privacy requirements and define an allow list. To me the main value prop is that I get one bill for all models and can quickly try new models without signing up anywhere or changing my code. Oh! Also you can pass an array of models and if the first provider is down it automatically falls through to the next provider. More useful than it should be...

LeBit

6 hours ago

I always setup guard rails so that only zdr providers are used.

oblio

6 hours ago

The thing is, in this climate it's hard to believe such tech will remain secret for long.

So, assuming this is not vaporware, this would raise the tide for everyone because it shows what's possible.

bregmandiv

4 hours ago

I'm trying to parse it down to what we had before vs what is new here.

We already had encoder models that skipped text generation for giving us a numerical output that could be computed as a probability. we also got no hallucinations and faster inference for free there. So we already had

1. "unstructured state in, probabilistic decisions out" 2. "orders of magnitude faster and more efficient"

What was hard there was to train the model head without ML expertise, and considerable amount of data.

This seems like this is a democratization of those encoders? The addition over existing encoders seems to be coming from being able to specify the output shape (up to a cardinality of 255). It is unclear to me if this is possible using Jev without additional labels for fine-tuning.

If so, that is still very impressive, but I think the faster inference and 0 hallucinations might come for free, from it not being generative.

dozerly

13 minutes ago

Very cool. LLMs have been borderline unusable as functions for the longest time, very excited for this direction.

alphazard

4 hours ago

There's a whole lot of information on this page that doesn't tell me anything about what this actually is. Can anyone spell out what the architecture is here?

They claim it's not an LLM, which I read as "not an auto-regressive token generator". I assume they are still using a transformer, otherwise they would be talking about the thing that's not a transformer, instead of all the fluff on the linked page. But they emphasize parallel generation, so is it like a text diffusion model?

bigglebear

an hour ago

I would guess a tiny stripped down text diffusion model. It only has 32k context, and for choice mode it can only select from 10 choices.

tacoooooooo

3 hours ago

Sounds like its essentially a generalized zero-shot classifier that takes and option set at runtime and works on unstructured inputs.

you pass in your "prompt" and options (described in natural language) that it can respond with, in addition to your input. it gives back that option set with a probability assigned to each one

dthedavid

an hour ago

Looks promising. I'm building an AI video editor and multi tool calls take >30s using Gemini. This would be a a game changer if Jev can take that down to single digits at p95.

activehuman

an hour ago

I can see the value in this but looks like there's going to be trouble in communicating the difference between this and a regular LLM, and also proving the potential cost savings in using this to replace existing systems that are using LLMs with frameworks like langgraph, as this can't be a drop in replacement and would require a significant amount of re-architecting/reengineering of systems to get the type system to work

nickstinemates

2 hours ago

We've already started using it for some pretty powerful decision tree stuff. We're just scratching the surface. We shipped an extension for Swamp[1] a few minutes ago and the combination is great!

The one downside is that the context window is very small (32k.) So some initial ideas we had for initial evaluation of code reviews won't fit yet in the window.

1: https://swamp-club.com/extensions/@swamp/typesafe-ai

jawns

6 hours ago

I could see this being fantastic for classification tasks. Last year I shifted from using LLMs for bulk data classification tasks (1M transcripts) to generating embeddings and categorizing based on cosine similarity. It saved a ton of costs and time, but wasn't as accurate as LLMs. This seems like it can give me Terra-level classification ability with the cost/speed I need.

pjm331

6 hours ago

yup just joined the waiting list with a very similar use case in mind

abeppu

3 hours ago

I think this is a great direction -- for some kinds of users. And this makes me wonder if the 'vs' framing is misleading.

Yes, I think it's a mistake that many organizations are cramming LLMs inside of automated pipelines where the extreme generality/flexibility of the model is at odds with the fact that you're using it for a very specific task that gets repeated over and over, and needs a very specific structured output to be successful. But specifying your task carefully (as well as deciding what counts as your input state representation etc) seems like a form of programming. Something (a person or a model working in a relatively unrestricted way) will need to produce a configuration/specification for this system.

So rather than Jev vs Claude I imagine that using Claude/ChatGPT/whatever interactively to define / refine your Jev config which then runs in prod might be the happy combination?

dinobones

5 hours ago

This is a good product but the naming/branding is pretty unfortunate.

Typesafe.AI sounds like some typescript/structured output type of tool…

What even is “system one” ?

IMO the product/tech is really there, just needs better communication.

toddmorey

5 hours ago

I mean, it's a structured output model that (apparently) can't hallucinate. I don't mind the name.

flyinglizard

4 hours ago

It can't hallucinate, but it doesn't mean it can't make wrong decisions. Just because it adheres to a specific output format at all time, while LLMs have the output format at their mercy, then the claim of not hallucinating is made technically true.

I think that this specific part is not super interesting if your harness just recovers from invalid LLM outputs.

The latency and cost - yes, those are super interesting.

albelfio

7 hours ago

caspar

34 minutes ago

I'm not sure the authors realize this is way more than "just a cool demo": if this holds up, it's going to be huge for game QA work.

Instrument your game to output properties of entities near the player and the output is the various control inputs - moment to moment gameplay gets solved. Maybe augment with a tick-by-tick controlled stepping mode if particularly twitchy - an LLM can take care of the higher level reasoning then.

magicmicah85

6 hours ago

The doom demo is also in the article, for anyone that doesn't want to go to X.com. :)

ErneX

6 hours ago

thih9

6 hours ago

The doom video is also in the article itself (headline: "Doom").

I suppose this is the same video as the one from the parent comment, but I don't know for sure - I don't have a twitter account and the above link doesn't work for me.

ErneX

6 hours ago

I linked to the tweet that has the video because if you are not signed in you cannot see the whole thread of tweets.

I can see the individual tweets in the browser while not signed in though.

einpoklum

6 hours ago

But when their system is given the instruction "do not fire, simply dodge" - it doesn't "simply dodge", it actually gets close to the fleshy pink demon rather than keeping its distance. Or am I misunderstanding?

anthonypasq

5 hours ago

i think its just telling the model that it cant output a shoot action

vatsachak

6 hours ago

It could be used for coding if you gave it an AST.

If you work at TypeSafe please try this.

Side note: This is probably how LLMs would perform with better encoders and next-latent prediction, so eventually those will beat this architecture out. Still amazing though.

ramon156

6 hours ago

I've implemented tree-sitter in pi before, and while it works, I have no real proof it saves me tokens, or is more accurate. I think a better implementation is a model that's trained for AST's, not just "use tool, see what happens".

I'd love to do research on this when I have the time.

vatsachak

6 hours ago

Cool project!

That's what I was insinuating through "better encoder"; the model creating more efficient representations of ASTs using something like JEPA

CompleteSkeptic

5 hours ago

the hard part for coding is actually state engineering (e.g. getting your dependencies in context) - we haven't even tried it yet (because my philosophy is we should automate the easy tasks before the hard and we've been working on getting the model smart on the former)

we do think there's a lot of potential though and do want coding themed releases soon

hunterbrooks

5 hours ago

I could see Jev being great at finding key symbols in codebase before a code generation/code review task. I sent you guys an email (to hello@) about using Jev in Code Review for www.ellipsis.dev.

Escapado

6 hours ago

I saw the CEO reply elsewhere in the comments to some other question. Maybe he can shed some light on it. My gut feeling is that this is non-trivial and they did not get this to work (yet?), otherwise I can’t come up with a good reason as to why they would not demo that as I assume half of the crowd here (myself included) would line up as customers.

vatsachak

6 hours ago

Yeah it would be quite trivial to try and implement an auto regressive AST generator for STLC with Jev provided that you had bounded variable names and integers.

As you said, if it worked, they would have demoed it haha

8note

2 hours ago

im not seeing it.

youd ask it to pick a location on the ast to add something from the grammar?

i dont see how this stays confined well enough? make a new output space every time? does that end up auto-regressive?

strich

44 minutes ago

Huh this looks fantastic. The Doom demo really sold for me that this could be a great tool for accelerating QA at my gamedev studio. Signed up for early access.

passive

an hour ago

While I understand that accelerating development isn't necessarily the target for this, and it's not at all intended to generate code the way many of us are...

I think this could be pretty decent in CI? There's a lot of "flakes" I've mediated that this could have handled much more efficiently. Maybe observability as well, triggering elevated logging and other initial measures?

edot

2 hours ago

Very cool! Can you explain when I would use this vs. training a standard ML model on my data? Suppose I had a fraud dataset with features like customer ID, amount, merchant, online or in-person, etc. - I can't imagine that a general model like Jev would predict this more accurately or cheaply than even a basic XGBoost model trained on my dataset (one that I could build in a few minutes by asking Codex to build it). Where does Jev add value here?

tensegrist

5 hours ago

what is the…epistemic status, for lack of a better way to put it, of the probabilities? what do they mean? what (probabilistic) guarantees do we have about, say, the responses to

- is the capital of france paris?

- it is august. is it raining in paris?

(forgive the examples; they're probably not semantically the sort of thing jev is trained to work on. but i figure the point translates to various kinds of questions that come up in "inner loop of agentic pid controller" contexts)

a normal text-generating model if asked to produce a number will also do that just fine. i assume in jev's case it was actually rled to essentially learn to express priors over things using its implicit world model, which definitely ought to help, but can we say more?

mixtureoftakes

5 hours ago

Doom demo is beyond impressive, even scary

bigglebear

an hour ago

It's very misleading. If I'm actually playing a game I don't get the coordinates of enemies sent back to me so that I can feed into my mouse to snap my crosshair to. It's looking through walls too, because it's working off structured state in text form. You could re-create this whole demo without using AI. Have an LLM generate the state machine for you and no model is required to run it.

postalcoder

4 hours ago

This has the potential to be huge for computer use.

OpenAI has been teasing how fast computer use is with their models running on Cerebras chips but the difference here is a burning hole in your pocket.

xynelius

3 hours ago

The Doom demo looks impressive but was it a fine-tuned model? It's the difference between a cool demo and revolutionary tech.

copperx

3 hours ago

Shouldn't self-driving be a piece of cake if it works this well for Doom? Or what am I missing?

hamishwhc

an hour ago

The model doesn't have image input capabilities (yet, it seems from the post), so for the Doom demo, a harness is extracting a bunch of structured information from the game (map layout, enemy locations, player ammo, health, etc) and providing it as a massive JSON blob to the model so it can make its decisions. This model _could_ be hooked up to make the decisions for a self-driving car, but it would need to be fed a structured blob of the situation around it, so all the computer vision problems of self-driving are still there. And that's before you get into the confidence and accuracy of this model.

pantelisk

2 hours ago

I think the doom demo uses a text representation of the world and it's basically, "projectile coming your way" -> "Strafe". "Enemy ahead" -> "shoot. So it works well when spawned in a room of enemies (as we see in the video).

If self driving is red means stop, green means go, and stay in your lane - then it would work great, but having to actually think and test which maneuver is optimal for a given situation while weighting safety, road rules, random unexpected actions and getting to your destination, I think it's a much bigger problem. A bigger model specifically trained on that maybe would do great, but then the output is not the constraint anymore.

But I haven't tried the model, so I 'm just ballparking and could be very wrong.

jamilton

an hour ago

Driving is more complicated than Doom, and it doesn't look that great at Doom to me.

mmastrac

an hour ago

Is this a Markov/Diffusion model with some sort of external Engram memory? If so, this could be extremely interesting.

himata4113

6 hours ago

They never show exactly how they use it? Only a bunch of animations of it 'working'. Would like to see the actual code used for the demos!

ricardobeat

6 hours ago

The doom demo shows the program state / query.

wxw

5 hours ago

> Input tokens: $0.042 / MTok ($42 per billion tokens).

> Output tokens: FREE (too cheap to meter).

Insane. The video demos are really compelling, in particular the speed.

> Structured outputs slot into ordinary software as fuzzy decision rules: classify, route, score, extract, or branch where hand-written logic is too brittle. The surrounding code constrains their freedom, making them easier to compose into reliable systems.

I buy this vision. A lot of LLM integration I see these days is ultimately exactly this. OpenAI-style structured outputs works decently but this would be a great improvement in cost, latency.

CompleteSkeptic

5 hours ago

thanks a ton!

constrained decoding (OpenAI-style structured outputs) make models dumber unfortunately - the short+dense version is that simply masking logits is insufficient because if ever a model was assigning probability to an invalid token, the model is by definition confused. you'd be better off erroring IMO

tylermarques

5 hours ago

We had early access and found it to be pretty useful. Having a second form of verification, where you can ask multiple questions (in the form of Nouls) raised our confidence in the outputs of other models. [0] IMHO This type of model works incredibly well in concert with LLMs, not as a replacement.

[0] https://goodstartlabs.com/research/verification-is-the-bottl...

filearts

3 hours ago

If we could come up with a system to classify the probabilities across a large number of candidate words (or components thereof) then this could actually be good at producing text, one element at a time. We could call these elements 'tokens' and picking the right one could be called something like 'decoding'. Crazy idea but hear me out...

On a more serious note, it will be fascinating to see how this different spin on modelling inference will create new paradigms or slot into existing ones.

initsecret

6 hours ago

> [others] Output tokens: ~5x more expensive than input tokens.

> [them] Output tokens: FREE (too cheap to meter).

I'm very confused by this.

varenc

an hour ago

The output tokens are just responses to your inputed questions and their probability. So relatively few output tokens. No unstructured text back in the response.

quotemstr

6 hours ago

They're not doing autoregression, so all the outputs are computed in one big forward pass. Very cheap.

ambicapter

6 hours ago

I think OP is confused about "others" vs "them".

initsecret

6 hours ago

they’re talking about two totally different things, right?

CompleteSkeptic

6 hours ago

it's our output tokens that are free (under the system one / jev column)

warpspin

5 hours ago

Haven't seen any docs or so. Is this actually a general model, or does it need training on the the data set it answers? Finding it suspicious you never see some kind of prompt.

Edit: never mind, found https://docs.typesafe.ai/introduction/quickstart by now

CompleteSkeptic

5 hours ago

1. yes a general model 2. no training at all 3. but it is focused on "System 1" tasks (more human judgment, less math reasoning)

zenlikethat

5 hours ago

It's very generalized. Can't wait until everyone can see it.

Wazzymandias

an hour ago

This looks and feels a lot like productionized conformal prediction

torginus

5 hours ago

I was thinking about something similar (maybe) - generally speaking, embeddings for LLMs tend to learn real world concepts - things like 'fruit' or 'France' or 'city' as directions in embeddings.

But in things like programming, most concepts are abstract - 'if hungry eat an apple' in programming terms would look like

'if hunger > 50 {apples--; hunger-=30;}'

and compilers work with 'concept erasure' - to them, tokens (which are like llm tokens) look like

'if var1 > 50 {var2--;var1-=30}'.

They don't care about how these things map to real concepts. So all the embedding directions used to encode real-world concepts are just noise to LLMs when programming. This greatly reduces dimensionality and training costs. So does a token representation tuned for programming constructs, rather than natural language would probably have a more efficient encoding.

ta988

5 hours ago

Current models go beyond the simple embedding because you start to encode groups of concepts in the context-aware part of the model (attention heads or any other method). So it is never simply words/tokens in isolation anymore.

bjconlan

2 hours ago

You know you're too old when you see the company name and think! Oh I wonder what Martin Odeskey , Jonas Bonér and co are up to. Wait, didn't they become lightbend... Altho this comment takes away from what these guys are doing which legitimately sounds interesting.

preommr

4 hours ago

This will be insane for tool usage, and probably where the major economics for day-to-day usage will be.

The goal is going to be to use llms to distill operations down to some dsl, and pass it into something like Jev.

cooljoseph

3 hours ago

A few questions:

1. Do you provide any kind of largest common subtree caching for cheaper input?

2. Have you tried auto-generating Lisp programs structurally?

3. Have you tried augmenting a Lisp language with a `choice` function that makes choices given a prompt, the environment, and the continuation stack?

zenlikethat

2 hours ago

(1) Nope, it's always the same input token cost

(2-3) No, but that's kind of a sick cook ... Want to get access and try it? nathan@typesafe.ai

cooljoseph

14 minutes ago

Thanks for the early access! I was testing the Lisp idea out in the playground, but I don't think the model is smart enough right now to generate actual code. I tried having Jev finish generating the code for a Fibonacci number function, but it kept wanting to create a literal number instead of refer to a variable which is a number. This happened both when I gave Jev the current program as a string and when I gave Jev the program as structured data.

Maybe I'm just not doing a very good job at prompting Jev, but I think right now it's not quite capable enough to generate Lisp code.

Link: https://console.typesafe.ai/playground?share=shr_148e1248984...

Mentlo

4 hours ago

Hm, would be good to understand the architecture better. Is this answering just from a world model informed prior? How informed is it by the information in the prompt? I can't see this maintaining calibration across all domains and all types of structured output.

Is there anything published on how it maintains calibration? Or when you say "outputs calibrated probabilities" you mean "as calibrated as frontier LLM models, just cheaper" - which is a different claim; as LLM's aren't particularly well calibrated

hoppp

4 hours ago

This is amazing. I really could use this.

I like the idea of System one models but all LLMs so far work as system 1 thinking because humans generate speech subconsciously with system 1.

System 2 thinking requires consciousness which AI does not have, so even reasoning models are still system 1 thinking as system 1 in humans has reasoning with heuristics.

Its limited but most people navigate the world with it completely, so it's enough for AI.

2001zhaozhao

4 hours ago

Hasn't there been a lot talk about Astra's opaque reasoning capabilities (being able to think through complex questions without using a chain of thought)?

Given that, can't you just replicate Jev by telling Astra "here is the question, you must make a multiple choice decision / output a score between 1-10, please answer directly in a single word, no reasoning allowed"?

(Edit: Ok, Jev is much cheaper in input tokens so these two aren't directly comparable at all)

CompleteSkeptic

4 hours ago

the edit is right - jev would be cheaper, faster, and more self-consistent (in general)

we actually use astra (and fable) in this way for our evals: evals.typesafe.ai

someone on the team cooked hard on that and it shows example traces comparing our model to opus/sol

iamgopal

an hour ago

If I understand correctly, it can play chess and rubic cube better than LLM ? ( may be go too ? )

altcognito

2 hours ago

If it is so cheap, why such a limited release?

johnecheck

4 hours ago

This makes me think of Expressions of Change [1], a project that aimed to make updates to a program a first-class primitive in a programming language. A model like this can't output code directly, but perhaps it would be well suited to select from the small set of discrete operations on code envisioned by the EoC author?

[1]: www.expressionsofchange.org

whazor

5 hours ago

A question I have, with the type { output: string }, would the model not become a LLM? And if it does, shouldn’t it cost the same as a LLM for output?

speedping

5 hours ago

I don’t see this as an option in their website

You could theoretically ask “what is the next appropriate character?” and add the entire ascii charset but i doubt it’d work well and you’d be implementing autoregressive churn across network latency…

CompleteSkeptic

5 hours ago

strings (and all sequential data structures) are not allowed at all - this is how we make sure all outputs can be computed in parallel (thus no output token cost)

nightshift1

2 hours ago

The whole page reads like it was vibe-written by an AI. If I'd built something as disruptive as this claims to be, I'd have spent at least fifteen minutes writing the announcement myself. Every time I see 'we' in an announcement like this, I picture one guy alone in his basement.

CompleteSkeptic

2 hours ago

unfortunately all hand-written :( my chief-of-staff does unironically handwrite em dashes though

lwansbrough

2 hours ago

For what it’s worth, I didn’t get that impression, and even noticed a couple typos ;)

2001zhaozhao

4 hours ago

Funny how the authors are asserting that "doing the right task > data > compute > algorithms" while simultaneously releasing AI model for calibrated decision making, which if they work, would mean that "compute > doing the right task"

moffers

6 hours ago

So is it a structured data-based language model? Or is there a model and a harness? Hopefully they’ll open up and explain more.

CompleteSkeptic

6 hours ago

it is just a model, no harness yet ;)

it is a structured data model, but technically not a language model (it doesn't generate language)

jrickert

6 hours ago

Signed up for the beta! :) would love to put this through some real-world shootouts against traditional LLMs to see where this type of model really excels.

I’m guessing it might be able to replace maybe 40-70% of LLM calls for a given pipeline depending on the business task, cutting the API costs on those calls by an order of magnitude.

sim04ful

6 hours ago

This sort of stuff almost sends shivers down my spine, it's like i'm looking 5 years into the future.

zenlikethat

4 hours ago

join the discord! we love forward thinkers

darpa_hr

6 hours ago

There was no "AI Winter"

adroitboss

6 hours ago

I am positive I know exactly how this works, I made something similar a few months back. But the problem is without generation you are extremely limited in the use cases. And while the model can't hallucinate, it can still be wrong. It just can't make up data.

dennisy

6 hours ago

Are you able to share how it works in that case?

adroitboss

5 hours ago

I'll tell you this. Output isn't too cheap to meter, there is no decoder.

krackers

5 hours ago

So an encoder-only model with a classifier trained on the heads or something? DeepSeek recently switched to an encoder-decoder architecture in an attempt to get the best of both worlds (fast prefill while preserving generation capability), I wonder if that might be the future?

mokre

6 hours ago

That was the first thing that come into my head. OK I can train very simple model, that can generate json's for specific tasks, so what? How we can be sure that this "limited use cases" not just overfitting for particular outputs (or even distillation?)

Except this, this thing looks like revolution.

tidewave

6 hours ago

Congrats on the release!

Finetuning a language model for decision classification (with probabilities) is already well-understood. What specifically changes in the training objective with RLCD? Are its benefits isolated from Jev’s new architecture/parallelism?

bqsile

3 hours ago

If it work as good as they say it does, confidence score + really fast response when you want very fast response, basically.. To me it is a crime against humanity to not open source it. Just get the money from cloud inference and cloud agentic sessions or whatever but open source it. This tech, a good harness, a good model provider, and you have basically a AGI building machine.

bqsile

3 hours ago

Golem, if you read this, add me on battle.net (europe) Sansviande#2540 and let's talk. Give me 1 minute.

Invictus0

3 hours ago

“Crime against humanity” buddy please

jceg

6 hours ago

> We deliberately chose not to publish performance against public benchmarks. In fact, we plan to only have one-off evals when we make product updates.

lol, I bet they would publish them if their score on those benchmarks were good.

scottyah

7 hours ago

Wild that it doesn't generate text. I wonder how its technology compares to Tesla's FSD stack.

Havoc

6 hours ago

Will need hands on to truly tell, but the doom demo seems very promising. If it can play that with text descriptions of where stuff is by distance and degrees in a 3D context then many GUI automation tasks should be easily doable

woggy

3 hours ago

Can this be used in conjunction with a text-generating LLM for better quality code generation?

pixelmelt

6 hours ago

Interesting concept, I can't see a reason to use a generalist classifier over an api rather then just training my own? If it was open weights I would probably mess around with it.

petesergeant

6 hours ago

This is basically a zero-shot classifier that can accept raw text (or structured text) as an input, and is able to classify that text as accurately (they claim) as a frontier-level LLM. I have workflows this would be useful for, looking forward to it showing up on OpenRouter.

_boffin_

2 hours ago

Any relation / inspiration to GLiClass?

gok

6 hours ago

So... a classifier model?

entrep

6 hours ago

This puts the human even more out of the loop I'll guess?

zenlikethat

4 hours ago

That's kinda the goal. Imagine all the automation in the world being able to embed intelligence directly inside it - factories could route based on more complicated questions, hardware could anticipate your needs. Customer support could be done without humans 90% of the time.

findjashua

3 hours ago

Would it be fair to say that this is tailored for tool-selection subagents?

bthornbury

6 hours ago

Is the tradeoff of the parallel output that we don't get arbitrary string generation? like output # of tokens is fixed ahead of time?

Either way, really cool and impressive.

zenlikethat

4 hours ago

Yeah, it doesn't output strings, just decisions/answers.

copperx

3 hours ago

Non-hallucinated ones at that.

_davide_

6 hours ago

What's the difference compared to just taking an embedding and feed forward a simple net trained for the task?

elcomet

5 hours ago

The technology and the results are very handwavy. What is RLCD exactly ? What are scores on benchmarks compared to LLMs ?

This website does not inspire confidence at all, it all sounds like a marketing piece. I wish it was true, some kind of text-prompted classifier with LLM performance would be cool, but I can't trust it with what we are given.

pennomi

6 hours ago

> Extraordinary claims require extraordinary evidence so see below for the receipts.

Yes, that’s the kind of attitude I want to see in these model releases

ramon156

6 hours ago

But the evidence is not there...

pennomi

6 hours ago

Indeed, they talk as skeptics but don’t offer a ton of evidence, other than a couple videos of demos. A live demo would be far more convincing.

simianwords

6 hours ago

They gesture at not using benchmarks for some reason...

meric_

5 hours ago

https://typesafe.ai/blog/antibenchmaxxing

But also effectively this is a classification model. It excels at specific certain types of workloads, and obviously will fail at others. Not really sure how one benchmarks this tbf. I can see their argument on why this requires a novel specific eval for whatever your usecase is. A consistent "global" benchmark might be hard to do

darksaints

5 hours ago

Okay, so it doesn't output text, that much is understood. What are the inputs like? I'm assuming maybe a text input? maybe an AST definition? Really hard to tell how this works at all from the demos, especially since we can't really try it out.

CompleteSkeptic

5 hours ago

inputs are structured program state. there is an example at around second 30 of the doom demo

(though ideally everyone gets off the waitlist and can try it out for themselves )

hi_hi

5 hours ago

If I’m understanding correctly, this will work well for self driving cars?

zmmmmm

4 hours ago

The eval is baffling me

> we assume there is a correct compute graph (a “workflow” represented in code) and use the predictions of the largest, smartest, and most expensive external models as reference probabilities. ... Rephrased: every model gets the same workflow. We test how they compare to the average of the smartest models (in this case, Astra and Fable).

They assume there is a correct graph, but they don't compare to that, they compare to the average of the smarts models? So the smartest models are getting it wrong but you compare that anyway as a benchmark? So the outcome is "how much of a Fable am I getting" etc. Why not compare the actually correct thing?

But then even on this hand constructed eval, the first plot is showing Jev at less than Sonnet 5 accuracy. It is barely better than Luna. There are two Opus 5's and two Sonnet 5's without explanation. What is the plot showing?

I gave up.

bilsbie

5 hours ago

I’m not understanding what this is. It’s a faster cheaper LLM?

andai

6 hours ago

Why did they pick the name System One? It's not really explained what "System One tasks" and "System One shaped queries" are. Things that need a fast response?

Does this imply it's a very small model? I couldn't find anything about the model itself.

ernsheong

3 hours ago

This is potentially huge and can crash the Big Two's stock prices or block their IPOs completely.

pama

5 hours ago

Is there a downloadable technical report somewhere?

bfeynman

6 hours ago

Super intrigued by this - large scale automation using LLMs is quite annoying due to deprecation cycles of models from frontier labs and cost of running your own being prohibitive when you have a blend of them.

seinecle

6 hours ago

Can this be used in practice to write code?

zenlikethat

4 hours ago

Not in the traditional sense of a coding agent, but we think there's a ton of opportunity in using it for context management ("do we _really_ need to pass all these tokens to the agent?"), semantic linting ("how does this score against this AGENTS.md: <...>"), etc.

bananaflag

6 hours ago

Funny how it can do everything but not chat. Sort of how when I was a kid I thought of a medicine that could cure any disease except the common cold.

totallygeeky

6 hours ago

Woof, that page is hard to read. I don't understand what they've done to the way text is rendering but it's not great for my eyes.

phenomen

6 hours ago

If you zoom in (especially on the large title), you'll see that the text is a semi-transparent gray with a black internal outline. It seems like all the typography is SVG-rendered. Actually insane. I've never seen this before. Not even the most vibeslopped websites have that.

erichocean

6 hours ago

I could put this to use today.

I think we'll see a bunch of different architectures over the next five years.

poly2it

3 hours ago

Is there a bottleneck which would hinder putting this architecture in charge of a humanoid? Would it be able to operate continuously, for example in conjunction with an LLM for long-term reasoning? Doom seemingly works extremely well.

8note

2 hours ago

defining the workflow such that the operation is a set of relevant questions

hspeiser

5 hours ago

this might finally be smart enough and fast enough for jarvis. hard to feel like iron man when your assistant takes 8 seconds to decide to pause your music

hunterbrooks

6 hours ago

um what is going on with the outfit changes in the launch video...

https://x.com/CompleteSkeptic/status/2099925682726002904

Gecko4072

6 hours ago

Can't tell if they're just having fun or if it is ai-generated. On the verge of not being able to tell. Voice sounds a little synthetic.

CompleteSkeptic

6 hours ago

definitely not AI-generated - this is my real wardrobe

we also thought the voice at the end was AI-ish, but apparently that's a real voice actor but slightly sped up

jbonatakis

6 hours ago

The whole video seemed generated to me

kylehotchkiss

5 hours ago

"While Jev gives up string generation, it’s optimized for structured outputs and can’t hallucinate"

Ouh! Any open weights models that can do this yet?? If not, how much longer? I have a Mac Studio coming soon.

yieldcrv

6 hours ago

oooooh it can play Doom!

forget LLM benchmaxxing sidequests, I'm sold on the real benchmark

charcircuit

6 hours ago

Parallel inference where you don't want a subagent seems niche. But there is a lot of random things where businesses ultimately want some kind of score instead of generating something.

I think the interesting thing would be seeing if prompt injections still work with this kind of model.

CompleteSkeptic

6 hours ago

we have played with this! the fascinating thing we've found so far is that adversarial examples for our model are quite different from that of LLMs so that they work even better together

whalesalad

6 hours ago

What is it about the rendering of this page that is so... off? It almost looks like the entire thing is a <canvas> element.

edit: looks like a framer export where there is a text stroke being applied :|

esafak

6 hours ago

Looks like a great model for NLP.

larodi

6 hours ago

"is this the real thing or is just fantasy"

kypro

6 hours ago

> Outputs

> LLMS > Strings / generated text. Strings are flexible and can be anything: chat responses, code, hallucinations, refusals, or even type-safe structured values. To be used by software, responses need to be parsed + validated. There is also always some risk that the AI goes off the rails.

> Jev > Type-safe structured values. Possible outputs and structure are defined in advance. The model never makes type errors. All answers are accompanied with calibrated probabilities and confidence scores.

I mean, this isn't even remotely comparable to LLMs so why compare? Also, why are they bringing up AGI given there approach is so restrictive that what they're building literally cannot have the creativity required for AGI? The video is 100% marketing slop...

The bulk of the application of LLMs is that they generate reasonably reliable text which doesn't need to be defined in advanced. I'm sure there is a niche for this and congrats to the team, but please let's not hype this as if it's the next big thing in AI...

bqsile

20 minutes ago

Creativity is not required for AGI, that's maybe the only thing that is not required for AGI actually.

What a sad world would you live in if you don't keep creativity for the humans.

mkrishnan

6 hours ago

If this is true means, AI Stock bubble burst. (For good)

quotemstr

6 hours ago

It looks like a specialized encoder-only(-ish) transformer with scalar and ordinal output heads. Acausal in effect, maybe? Probably not even autoregressive?

I'd use this as a tool an LLM can use for specialized tasks. It's not AI in itself.

mkrishnan

6 hours ago

If this is true, then AI Stock Bubble burst (for Good)

kart23

3 hours ago

This makes me kind of nervous for the whole AI thing now. Are people gonna lose their jobs, etc.? so much of the economy is now built on top of LLMs.

colordrops

3 hours ago

What's different about this particular model that worries you?

kart23

3 hours ago

it doesn't require nearly as much compute as normal LLMs. anything depending on increased datacenter and compute spending would be threatened.

invalidOrTaken

2 hours ago

a world where people eat so they can feed Big Computer sucks. We need Little Computer, driving robots in the fields.

kart23

2 hours ago

I definitely agree with this, but the transition is gonna suck for some.