chvid
7 hours ago
“The fact that a 17GB file can do all of this stuff on my home machines is a miracle. Once again, I’m delighted and amazed at how much progress local models have made this year.”
I think that should be the blinking headline - this shows what can be done with consumer hardware.
CMay
4 hours ago
For me that moment was Gemma 4 12B QAT. You're not suddenly going to start throwing your hardest programming problems at Gemma 4 12B QAT, it is still 15B parameters less. It's more that, aside from pelican art which isn't what local models are for, I didn't see anything on Simon's post that it couldn't assist with or largely succeed at.
It can run 80-100t/s on a laptop, can understand images natively and do bounding boxes, read tiny text, understands audio natively as well and can transcribe or translate anything you say, can do accurate long context retrieval with pretty large context windows, tool calling, excellent reasoning and is very token efficient.
It's only 7GB including the mmproj or 8GB with MTP. The Qwen 3.8 27B model Simon was using is ~18GB with MTP+mmproj, rather than 17GB alone. The point is not really that you compare these models directly, but that Gemma 4 12B QAT was really a special moment in model releases deserving of a similar reaction relative to its size, but was mutilated by Google themselves, Unsloth and Llama.cpp.
The overall appreciation I think we're seeing this year in particular is that people are easily surprised when multiple things are improving simultaneously which produce seemingly exponential changes. It isn't just that models are getting smaller, or that reasoning is getting better, or that speculative decoding is becoming mainstream, or that models can understand audio and images better now, or that they can reliably call tools which expands their capabilities, or that context windows are getting larger, or that accurate retrieval is improved, or that.... and so on. It's all of them narrowing in at once that is starting to make local models incredible and truly useful for far more use cases on the existing hardware people already have.
podocarp
3 hours ago
Out of the loop here. What did Google and unsloth and llama do to mutilate Gemma? I can understand Google shenanigans but llama and gunsmith is kind of surprising.
CMay
2 hours ago
Google provided incorrect settings and an imperfect template.
Unsloth modified the template and then finetuned their own version of the model to optimize for some benchmarks as a means of validating quants.
Google and Llama.cpp then adopt template changes by default, so anyone downloading the new model or even using the original model will now automatically be using it incorrectly.
Llama.cpp also uses the same inference setting defaults regardless which version of the model you use and some settings are simply defaults it uses for all models.
Then even if you account for all of these, you have to be using Gemma 4 itself correctly, which many people do not.
All of these little changes and inconsistencies hurt some of the model's original capabilities. Even if you go directly to Google's repo and download the full float 16 weights with the template they have there now, you cannot simply assume you're getting the best results.
DanielHB
2 hours ago
I am very much a beginner to local LLM stuff and I find it incredibly hard to figure out how to run models optimally with the correct settings for my hardware. The number of different variations of the same model and how each quant work is super confusing as well.
When I tried to run llama.cpp directly I was getting max 9tk/s on qwen3.5-9B, then I tried LM Studio with the same model and got 77tk/s. I haven't figured out yet how to get MTP working properly in either.
car
22 minutes ago
If you are on Mac, have a look at the Llama-macOS app. They claim sensible settings for the linked model downloads. I'd expect the authors of Llama.cpp and the Huggingface folks to know this stuff.
agile-gift0262
2 hours ago
And what's the right way to use Gemma? Where can I find the correct template and settings if those aren't the ones provided by Google, Unsloth, and aren't built into llama.cpp? I discarded using Gemma 4 because it got into weird loops when tool calling
kzrdude
31 minutes ago
Some weeks ago a new official Gemma 4 release was posted that corrected some of the chat template problems. So the official release files on hugging face should be the way to go.
tarruda
2 hours ago
> It can run 80-100t/s on a laptop
That is a lot, what is your laptop hardware?
One issue I have with Gemma is that they seem to use old architectures that rely on full attention, requiring a lot of RAM for context and quickly degrading speeds as context is filled.
Qwen 3.5+ is much better in that regard with its super efficient context. Even on Macs, speeds take degrade much more slowly.
oblio
an hour ago
> transcribe or translate anything you say
Is it multimodal? How do you do transcription with it?
a_e_k
6 hours ago
Like the old proverb: "The marvel is not that the bear dances well, but that the bear dances at all."
bitwize
4 hours ago
Indeed. LLMs resemble human intelligence in more or less the same way that the output of the TI-99/4A speech synthesizer resembles a human voice.
coldtea
3 hours ago
Not if an LLM over chat can fool most people they're talking to a human (which it can), where the TI-99 speech synthesizer voice absolutely can not.
zahlman
2 hours ago
> Not if an LLM over chat can fool most people they're talking to a human (which it can)
I keep hearing this claim, and yet I keep seeing LLM output which is trivially distinguished from human writing. I really can't understand how this gap persists; but then, there seem to have been at least some people who couldn't sniff out ELIZA, back in the day, too.
coldtea
2 hours ago
>I keep hearing this claim, and yet I keep seeing LLM output which is trivially distinguished from human writing.
That's mostly true for longer LLM output with all the sycophancy / LinkedIn bias thrown in.
Make it casual conversation or comments, and give it instructions on appearing casual, or even better kill the censoring and fixed-prompt (with an open model), and it's orders of magnitude more difficult, unless if you suspect it and try specifically tailored prompts to sniff it.
There's no shortage of people obliviously discussing with AI bots in comment sections.
zahlman
2 hours ago
At the bottom of this very submission are a bunch of dead comments that are very obviously LLM-generated.
lukan
2 hours ago
And that is proof that all LLM comments are easily found out? Also, easily found out by average humans? (this is not a average forum here)
akie
2 hours ago
Ok, you and I can easily spot LLM text. So what? The Turing test has still been passed, as is clear by people falling in love with ChatGPT, not believing something is AI, and by continuously claiming this or that is a bot.
People, many of them at least, cannot make this distinction anymore. You can, I can, but people as a whole are having problems with that.
bluebarbet
an hour ago
>You can, I can
Even this (assuming it's even true) will likely not be true in some near-term future.
>continuously claiming this or that is a bot
I see it as a contemporary form of religious thinking. Like (say) pilgrims seeing blood on a statue of the virgin, plenty of people are now seeing the hand of AI in everything they read. If you want to see something hard enough, it tends to become magically visible.
akie
21 minutes ago
The most interesting part of your reply is that you're not challenging the claim that the Turing test has been passed. I think it's a given, by now.
bluebarbet
17 minutes ago
Sure. Of course it's been passed.
shafyy
3 hours ago
Can it? I feel like I instantly recognize if I am chatting with an LLM or a human
piva00
2 hours ago
I was quite surprised on how difficult it is to tell when chatting with an uncensored LLM a friend is running (it's too big to run on any of my computers but he got some B200s). You can input your own "system prompt" to make it behave like a normal internet user and the prose writes very similarly to internet comments with none of the LLMisms from ChatGPT, Claude, Grok, etc.
IMTDb
2 hours ago
Emphasis on "feel"
tapland
2 hours ago
Can't even tell if you're real or a bot by reading one comment.
Great times!
freehorse
4 hours ago
I also believed that, but seeing qwen 27b overengineering solutions in a bit too familiar way in the article, I started doubting that.
ZaoLahma
5 hours ago
Full agree. I until very recently thought AI tools of today were limited to prohibitively expensive high end hardware hosted in data centers.
I was surprised and amazed to get "decent" (with the expectations set right / low) coding performance out of Qwen3.5-9B on a decidedly medium end Radeon 9070 paired with a 5700x3d and 32GB of DDR4 RAM.
We can finally reason with and "talk" to our hardware.
eru
4 hours ago
Yes, and we are still pretty early: AI is still advancing at breakneck speeds, and hardware is too.
DanielHB
2 hours ago
Is the hardware really getting better? It feels performance per watt is not getting better at all which is the metric that will matter eventually when supply-demand stabilizes.
As it is, it seems the improvements are about making the hardware cheaper (as in capex, not opex).
This is just feels from me from what I hear on the news and see on the products though.
eru
an hour ago
Solar power and batteries are getting cheaper and cheaper at the moment. So Watts should become cheaper in the long run.
Especially when chips are becoming cheaper (in the capex sense), then you can afford to only run them when power is cheap.
Btw, from where do you take the notion that performance per Watt ain't increasing? We are also still using what's more or less general purpose GPU hardware; we could get a lot further if we were willing to specialise more. Which would be the natural avenue to explore, if progress in general purpose hardware slows down. Google is already looking.
DanielHB
21 minutes ago
Like I said, just feels I have from the consumer-hardware space. For several generations of GPU now most improvements come from packing more transistors into a larger die than packing more transistors closer to each other.
GPUs have been getting physically bigger with huge heatsinks and fans to support those bigger dies power consumption. Just compare the TDPs:
2020 RTX 3090: 350W
2022 RTX 4090: 450W
2025 RTX 5090: 575W
Bigger dies means lower capex of course, but the similar opex (maybe slightly lower as there is less physical hardware to maintain).
I seen some specialized hardware like google's TPUs. Not sure how they compare on performance per watt with GPUs though. Regardless the manufacturing processes are still the same (EUV) which is the thing that hasn't been improving. A fully optimized specialized hardware can at most deliver a single-time linear improvement (that could be very significant, for example 30% is still huge of course) and then little compared to normal GPUs.
I don't think renewable power generation is going to massively reduce costs for data centers, especially considering power transmission hasn't meaningfully reduced in cost. If anything the only thing that I think will have significant impact for data centers would be dedicated nuclear power plants physically located right next to the data center.
In fact I expect power generation to get more expensive as demand can increase faster than supply can be established. I imagine setting up new solar farms and transmission lines to be significantly harder (as in, takes longer time due to approvals and so on) than new data centers (which requires a single large location and I assume less approvals).
jillesvangurp
an hour ago
The other implication here is that this is all software improvements and optimization. There might be a lot more wiggle room for improving quality over time. It seems the model and reasoning quality is improving faster than the hardware currently.
The over reasoning that Simon Willison highlights here is a real issue though. I've observed it with some of the OpenAI models as well. They are prone to overthinking and overengineering things.
What I would love is models that figure out their own appropriate reasoning effort given a task. I'm spending too much brain cycles worrying on what model speed, reasoning, and quality settings to pick. It's not just a cost concern it's also a time concern. Wasting a lot of time for simple UI tweaks because the model is set to high or ultra or whatever is counter productive. The last few iterations of frontier models seem to emphasize benchmarks and reasoning effort.
But of course the day to day reality of many developers is that they are trying to solve relatively simple problems compared to e.g. proving some so far unproven theorems, solving some Nobel prize level problems, etc. I'd love my tools to start making sane choices based on what I ask rather than defaulting to "boil the oceans". These tools need some kind of Auto select. Mostly Ultra is overkill and a waste of time and resources. And of course with local models, keeping simple things local is a nice option.
It's nice to have Sol Ultra extra fast as an option in my back pocket. But it's complete overkill 99% of the time. And it's not like most users make good choices here or are even capable of making good, informed choices. The models are more intelligent than the tool UX. Arguably, a local model of very modest size might be able to do better for this specific choice.
madduci
7 hours ago
Tried yesterday on my own laptop (a UltraCore 7 255H without dedicated GPU,with 32 GB RAM), it wasn't even starting thinking, even on a small context window (65k)
tylerKorhonen
a few seconds ago
> it wasn't even starting thinking
Probably stuck in prompt processing which is compute bound especially for iGPUs.
You've mentioned 3.5 - but it's actually the same model the only differences are training and implicit MTP support (affects prompt processing - can be disabled)
aphroz
6 hours ago
I think not much can run without a dedicated GPU
madduci
6 hours ago
Till now I was using successfully Qwen 3.5 and Gemma 4 at a reasonable speed
kzrdude
29 minutes ago
I think (maybe I missed something) that identical size and quant versions of Qwen 3.5 and 3.8 should run at the same speed. It’s the exact same architecture.
madduci
23 minutes ago
Tried the 4 Bit versions. It loads, bit the <thinking> output isn't even coming out.
pyrale
5 hours ago
There is no way you would run a dense 27b model on that spec. I ran 3.6 27b on a 64gb ram, 24 gb vram, and it felt like the lower limit for this model with a decent context window.
If you want a better experience, maybe wait for either a moe model (like 3.6 35b A3) or a model with less parameters (like 9b). Qwen has been releasing those in the past, so maybe we’ll have them for 3.8 too.
DanielHB
2 hours ago
From my experience if it doesn't fit on vram it is rarely worth to bother except for a few narrow tasks.
For example make an essay about something where you don't actively engage with the LLM after the initial prompt. So mostly one-shot prompts.
mobelkh
6 hours ago
were you running the MoE models? those perform better speed wise
noduerme
5 hours ago
What's the story with Mac laptops? Worth a try?
selcuka
4 hours ago
The author tested in on an M5 laptop too:
> It feels pretty slow on both the M5 Mac and the DGX Spark.
cyberrock
3 hours ago
Dense ones like this are more bandwidth-hungry, so you want to try MoE ones like Qwen3.6-35B-A3B (35 Billion params but only 3 Billion Active) or Gemma 4. Unfortunately it seems like we might not be getting a 3.8 MoE.
rawland
5 hours ago
Yes. mtplx runs it at 25 tok/sec on a M4 Max with 48GB RAM.
mdp2021
3 hours ago
Have you tried with different amounts for the "reasoning_effort (xhigh|medium|low)" parameter?
Or the "<|think_xhigh|> | <|think_low|> | <|think_off|>" tags: apart from this template detail, it is not immediately clear if reasoning_effort is deterministic (API) or is prompt engineering.
madduci
23 minutes ago
No, good point. I will have to tried it
petu
2 hours ago
What was your prompt length? It's possible it was just processing it and it's likely not fast on your setup.
madduci
24 minutes ago
Really small (<100 tokens), I wanted to test its capabilities
pdyc
5 hours ago
i have same 255h and i was able to run it with low token speed 6-8tg/s with approx similar context window 60k
madduci
22 minutes ago
Interesting. What are you using? I was using ollama
wejick
3 hours ago
In this kind of moment, I really wished hardware manufacturing and demand situation is in much state. Imagine this can be accessible by everyday people with only 6 months hardware market gap. The societal impact would be much bigger.
mhaberl
3 hours ago
I wish we could have better hardware and I think the tech is there for a few years already.
I've gone in (too many) details last night with the calcs: https://news.ycombinator.com/item?id=49324600
AgentMasterRace
7 hours ago
his 128gb Ram laptop is quite extreme
simonw
7 hours ago
It should just about be usable in 32GB.
krzyk
5 hours ago
On a consumer hardware it would be nicer. With no GPU/iGPU or a 6-8GB VRAM.
hnfong
4 hours ago
It would be somewhat slow on a CPU only machine, but it still works.
Besides, Macbooks with 32GB RAM is consumer hardware, just maybe on the higher end.
npodbielski
6 hours ago
It is. I am running it on R9700
bakraman
4 hours ago
RAM is never the issue, it's always the compute power
CamouflagedKiwi
3 hours ago
It's absolutely not for these models. There are plenty of consumer GPUs out there with 8 or 12GB VRAM - they are comparatively very fast at inference but just aren't big enough to run lots of the models you want. Also context management is a massive pain.
DanielHB
an hour ago
I run qwen3.5-9B on an RTX 3080 with 10GB of vram. It runs at ~77tk/s with around 50k context size.
As soon as I switch to a model that doesn't fully fit into vram it tanks to <10tk/s which makes it unusable for me for most tasks.
piva00
2 hours ago
RAM bandwidth is the main issue for running LLMs on consumer hardware...
spider-mario
4 hours ago
RAM is not “never” the issue. My iPhone and MacBook Air could both run larger and more capable models if they had more RAM.
tuetuopay
4 hours ago
Quite the opposite, RAM is always the issue. More specifically, high bandwidth RAM.
mhaberl
4 hours ago
what??? not true!
for inference the compute is the last thing we need more of.
memory bandwidth is the numebr one blocker, after that the inefficiencies that where introduced with MoE models (and all new large models are made that way)
Here is a quick read: https://news.ycombinator.com/item?id=49324600
geek_at
4 hours ago
and memory bandwidth
aizk
7 hours ago
Give it 6 months, the capabilities will increase even further.
marcelo-earth
5 hours ago
I thought the same thing, and I generally do a lot of animation in my work, and the results in motion graphics with Qwen are impressive, I really fell in love with it
genxy
4 hours ago
Curious how you are using it? Making blender plugins?