firasd
12 hours ago
Just realized that there are basically no American open models right now ever since the Llama series was abandoned. Basically Gemma and GPT-OSS I guess?
Ah but Mira Murati's new Inkling is Apache 2.0
But it makes sense that if you're a university researcher you are thinking about what's a model that will be open weight and developed over the long term and doesn't raise 'Chyna' concerns in Washington DC
CMay
4 hours ago
LiquidAI LFM models are amazing, but very situational. IBM Granite series are also unique and interesting for trying to reduce liability and extend local context size. Nvidia ships some and there was also that Inkling model recently. Poolside just released theirs.
Meta might release something this year. X AI's Grok is still due to release a model, if Elon keeps to his word even if they only release a distilled version. Reflection AI has been quiet, but their access to compute is ramping up. Microsoft's MAI is considering releasing some open weight models which would be great to see!
Ilya's SSI is unlikely to release an open model since he's aiming for radical safety. That bet could pay off if the existing approach produces so much chaos within the next 10-20 years that some global ban is achieved and a super safe model is promoted as the compliant route.
We don't get many huge model releases though. I think it's harder and more expensive to safety align them. Even if you do, people will work around the safety and abuse the models. Plus it makes it even easier for Chinese companies to distill things that aren't as easy over filtered APIs.
There is a lot of internet propaganda to the effect that the US is simply unable to release open weight models or that China has so many more AI companies that the US is drowning in Chinese open weight models, but it's more like we're being careful and China doesn't care. If you host a model in China, it has to be censored and downloading any models requires you to provide your identity. Huggingface is banned there. When they release their open models in the west, they don't have to care whether the models are aligned in any way.
embedding-shape
5 minutes ago
> but it's more like we're being careful
What? US laboratories are currently unable to contain their agents while doing security testing, and besides that, time and time again US labs seem to put short-term money above long-term safety.
Wasn't that literally why they tried to oust Altman from OpenAI, as he basically was 100% focused on profits and tried to cut down on safety across the board and lied to get his way?
> If you host a model in China, it has to be censored and downloading any models requires you to provide your identity.
I'm not disagreeing with that first part (obviously that's about inference hosting, not creating/training weights or hosting those weights), but the second part I'm not so sure about. AFAIK, ModelScope (which is the Huggingface in China) seems to allow downloads without verifying any identity and also hosts a bunch of abliterated weights.
jauntywundrkind
4 hours ago
Allen Institute for AI has quite a range of very interesting very competent more specialized models, for earth sensing, embedded robots, for others. Their SERA model shows a remarkably capable model for such a deliberately small investment effort, with documentation on how you can train such a model yourself or refine it easily at little cost. Their EMO pioneered a better MoE with great numbers (at least at the time). https://allenai.org/
ipsum2
11 hours ago
There's a bunch of American open models. Inkling, Nemotron, Trinity come to mind, but I'm sure there's others.
embedding-shape
11 hours ago
Laguna S 2.1 is really great too, in the "preview" release they've done so far at least. Still pending some reasoning-looping, but besides that, it's a really strong model to run within 96GB VRAM with the NVFP4 variants, and it's really good at coding (specifically).
walrus01
10 hours ago
There was an obvious problem in the original release, they re issued it after like a week with the reasoning looping supposedly fixed.
embedding-shape
3 hours ago
Well, bit more complicated than that, I've been eagerly helping in testing and keeping track of what they've done. Initially there were serious bugs, also about the templates, eventually they released RC1 which had some fixes towards the looping. Then a couple of days later, they released RC2 which supposedly fixed the issue, but ballooned the size so all of us who were running Laguna S 2.1 on a single Pro 6000, suddenly could no longer. So, unsure if RC2 actually fixes the issue, as we're a bunch who can no longer run it :)
Besides that, it was also discovered that their suggested inference parameters were wrong and led to worse behavior. Eventually someone discovered these works best (so if you have the issue with looping right now, try these, helps a lot for me but not 100% still) and was also what the evals used apparently: temperature: 1.0, top_p: 1.0, top_k:20
Now we're waiting for RC3 which Poolside said will come at one point, and hopefully also brings down the size again NVFP4 weights + full context can load properly again even on "smaller" hardware.
behnamoh
10 hours ago
No it doesn't follow instructions and is substantially slower than ds4.
kadoban
10 hours ago
It's a lot smaller, and runs (quantized) on a 3090 quite well. Ds4 flash 0731 you're talking about? It's great but it's much harder to run locally.
ericd
6 hours ago
I found Laguna S to be pretty good at coding, pretty fast, but pretty bad as an agent - not proactive, would frequently stubbornly argue things that weren't true, and pretty bad general knowledge.
But as a pure coding model, pretty good.
Deepseek v4 Flash 0731 is so much better if you can run it, though.
Grain of salt, I think I grabbed Laguna after they fixed the initial looping issues, didn't notice those, but there might've been other fixes since.
embedding-shape
3 hours ago
> But as a pure coding model, pretty good.
Yeah, this is my perspective too on Laguna S 2.1. Works amazingly for coding, pretty bad for pretty much anything else. I don't do a lot of advanced math, supposedly it's good for that too.
Tepix
6 hours ago
Are you talking about S or XS? S is too large for a 3090 at 118b parameters.
jauntywundrkind
10 hours ago
Like glm-5.x I think it has enormous self introspection that it often trips up on, but that this self reflection is actually a superpower, that enables incredibly good output. And from (in some cases) very small models.
If you watch it think, which you can, unlike American closed models, you can steer it. You can provide a a massive rocket ship stratospheric boost to help it orient itself. You have no self correction, there is no multiplayer in American proprietary models.
Sure it's great having super powerful mystic oracles that have the "right" answers. But I love respect & revere the open thinking. No it's not automous. But it is brilliant. And it considers. A lot. Deeply. It chases. That to me is the most human of models, even as it falls far astray.
You should help it. You can. Unlike these vicious dark surfaces which yield and tell you nothing. I think this is the actual meta-core-super-point of "The session you cannot take with you" (link below). It's the session that does not care about you, will not interact with you, will not peer with you, that is a dead remote far off oracle to you. Fuck these "oracles". They are a plague against the human spirit. We should alloy humanity and AI to Augment Intellect (Engelbart). (To do less is species treason.) https://earendil.com/posts/session-portability/ https://news.ycombinator.com/item?id=49118781
behnamoh
10 hours ago
I like the transparency of its reasoning, and I agree with you, OpenAI/Anthropic/Google should show the reasoning traces as well.
kadoban
10 hours ago
Yeah I think it got bad press because the chat templates (or something?) were messed up on first release, but I've been using a quant of it and it's a powerhouse, better than qwen 3.6 27b for local on a 3090, which is saying a lot.
embedding-shape
3 hours ago
No, the quants they released were also messed up. RC2 also ballooned the size so the ones who were excited about RC1 (like me) can no longer fit it in our hardware. They haven't promised anything, but said they'll try to restore the RC1 size for the next update of the weights.
firasd
11 hours ago
Just looked into some Nemotron stats
Looks like on <https://arena.ai> agent arena (grouped by lab) Nvidia is 15/15 (much worse than Thinky and Mistral) and on text arena it's 18/27
On <https://openrouter.ai/models?order=most-popular> I definitely see usage though (probably mostly cause Nemotron 3 Ultra is free) the grouped order is DeepSeek, Tencent, Xiaomi, OpenAI, Z.ai, Nvidia
coder543
11 hours ago
I think glancing at a random snapshot from today misses all the context. Nemotron 3 is far more significant than you're giving it credit for.
At this point, Nemotron 3 is really an 8 month old model series. That's when Nemotron 3 Nano was released, and the Nemotron 3 Super/Ultra models this year are obviously based on that recipe, mostly just bigger with a few tweaks here and there. Against today's models, no, not that interesting. Each of the Nemotron 3 models were briefly competitive when they launched, but never exceptional, and less competitive with each scale up. The fact that it took so long for Nemotron 3 Ultra to launch really hampered its competitiveness.
The Nemotron 3 series is extremely open about training recipes and training data, far more open than most open weight models, and that is valuable.
Before Nemotron 3, Nvidia had never released a single LLM that I would consider interesting at all, so Nemotron 3 was a big step up. The closest thing was Mistral NeMo, but a significant part of the credit there goes to the Mistral team, not Nvidia.
Given how much Nemotron 3 improved, I'm curious to see if Nemotron 4 will take them to a leading edge level instead of just briefly competitive.
(Nvidia released a Nemotron 3 and a Nemotron 4 like 3 years ago... this year's Nemotron 3 is entirely unrelated. Nvidia's naming schemes leave a little bit to be desired.)
buildbot
6 hours ago
Nemotron 3 also introduced LatentMoE, which was adopted by Kimi K3 :)
no-name-here
8 hours ago
ArenaAI Agent Leaderboard direct link: https://arena.ai/leaderboard/agent
strictnein
6 hours ago
Nemotron is a very nice model with an excellent license as well.
maziyar
an hour ago
Yeah thankfully we have more than we had in 2025! I am sure we will see even more open models by US based startups before the end of 2026
written-beyond
11 hours ago
Don't forget IBM
stogot
2 hours ago
OpenAI has gpt-oss that they said is open weight
wmf
11 hours ago
Also Nemotron and Arcee.
loeg
11 hours ago
I would not be shocked if another open model eventually shakes out of Facebook (based on Zuckerberg's public remarks).
solomatov
11 hours ago
Which remarks? Could you share a link?
loeg
10 hours ago
He said something to the effect of "I love open source and open models and we'll do open models when it makes sense and closed models when it makes sense" in a recent Q&A.
johnecheck
5 hours ago
Given that his company has already released open models, I find it funny that, as you described it, his remark communicates absolutely nothing whatsoever. Not sure what the question was, but this was an artful non-answer.
mlindner
an hour ago
> But it makes sense that if you're a university researcher you are thinking about what's a model that will be open weight and developed over the long term and doesn't raise 'Chyna' concerns in Washington DC
Why phrase it "Chyna" when it's an actual legitimate concern?
walrus01
10 hours ago
Laguna is the most recent and capable one that comes to mind. In its size class it is not as "smart" in my experience as qwen 3.5 122 or DeepSeek v4 flash 0731 (all at q8), but it's also not terrible.
mistrial9
11 hours ago
review of AllenAI Olmo research team and commitment to OSS -- AI2 complete transparency including training data, code, intermediate checkpoints, and detailed logs for reproducibility and scientific rigor.
logicallee
10 hours ago
I've used Inkling a lot recently, it's an American open model and is really good!
vasco
4 hours ago
You didn't read the article because the company that worked on this has published open models before and both these things are mentioned early on.
connorbrinton
11 hours ago
Laguna S 2.1 is another fairly impressive-for-the-size American open model