fooker
5 hours ago
Prediction - we are going to figure out SOTA AI performance without requiring 1TB of memory within a year or so.
Of course CXMT, Micron, and family will still be profitable, but maybe not 'surge 470% from IPO' profitable.
xyzsparetimexyz
5 hours ago
How would we do that? There's no historic precedent for that. Its fundamentally an information theory thing: what's the max amount of intelligence you can get out of 1 KB/MB/GB? There has to be a limit and I'm not convinced that it's far off.
esterna
4 hours ago
Once I start a conversation about Postgres query plans, maybe 95% of the knowledge will be almost certainly not needed, and for an inference provider there will be many more concurrent queries with reasonably large overlap. Maybe future architectures will be able to take advantage of this so not all parameters are needed in memory for almost all instances.
If we are talking about all knowledge, then I agree that the compression ratio is very impressive already.
NaiveBayesian
4 hours ago
Mixture of Experts is already used by pretty much all modern LLMs to address exactly this phenomenon.
Hopefully, future models can be trained to be even more aware of external knowledge, accessible through web search / RAG / whatever it will be then, and might not need to internalize much knowledge at all.
bbatha
2 hours ago
> future models can be trained to be even more aware of external knowledge
Then you need longer contexts, which is proving to a much more stubborn problem than general knowledge compression.
xscott
4 hours ago
All the terms are squishy, but being sloppy about it, I think there's intelligence that needs to be in the model, and knowledge that could live in a database. Right now, models are memorizing a lot of stuff they don't need to. We know how to index lots of information on (comparatively slow) SSDs or across a network.
See Karpathy's "Cognitive Core" idea. (I don't have a good link)
t-3
32 minutes ago
I was recently skimming through some 20-30 old books about Information Retrieval (https://en.wikipedia.org/wiki/Information_retrieval), and it was immediately quite striking to me that LLMs are the culmination of the integration and development of many ideas in IR that were impractical or theoretical back then. Decoupling the database from the search interface would make the relationship between modern AI and IR even clearer.
anon373839
4 hours ago
So far, frontier capability keeps getting shrunk to fit consumer level hardware. The lag time being over a year, though, precludes this from being called SOTA by the time it arrives.
There must be a limit, I agree, but there have been no signs of approaching it yet. The most recent cohort of small models have shown the biggest leap in capability so far.
londons_explore
4 hours ago
It doesn't matter. However much intelligence you can squeeze into 1 GB, people will always want more.
Standard Def TV was plenty for 50 years. But when more was on offer, everyone went for it, and now you can't even buy a 480p TV.
CamelCaseName
4 hours ago
Your example contradicts you.
2K, 4K, and 8K+ TVs have been around forever, and although 4K has become the norm, it's widely accepted that there isn't much benefit for most people above 1080p
brianwawok
3 hours ago
Depends on distance and size, there is a lovely chart. My projector at 10 feet away is night and day difference at 4k. Mathematically 8k should look identical, so I haven’t spent the money to try it to confirm.
The 1080p thing is true in the age of 40” TVs but not so much now, think Costco has a 95” TV if not 100”.
phatfish
an hour ago
I can definitely tell the difference between 1080 and 4k on my 55inch at around the recommended viewing distance. Less so on a 42 inch. A 4k image will look good even on a projector, as you say.
Personally I think 4k (with HDR) is good enough for consumer use, so agree with the parent theoretically just not quantitatively.
It did take around 20 years from DVD to 4k Bluray though.
HelloMcFly
2 hours ago
> it's widely accepted that there isn't much benefit for most people above 1080p
Is it? By whom? For larger TVs, closer distance, or a combination of the two, there absolutely is value from 4k. https://i.rtings.com/images/optimal-viewing-distance-televis...
Though I'm unlikely to ever upgrade from 1440p on my computer, personally.
t-3
10 minutes ago
Talking about resolution is pointless without mentioning screen size and display tech. PPI is better, but doesn't always work well comparing between technologies.
StilesCrisis
3 hours ago
It depends on TV size. If you've ever experienced Netflix on a 75" TV, 1080p content looks mediocre on it. So really it's better expressed as "there isn't much benefit for most TV sizes."
broeng
3 hours ago
I wont quite argue, that 1080p is always enough on a 75", but your example is more a Netflix problem; their compression results in hideous quality, and it's not really representative of what 1080p could be.
StilesCrisis
2 hours ago
Granted, maybe Netflix wasn't the best example, but I was just trying to be succinct.
oblio
4 hours ago
Yeah, but that's based on a hard limits, the physics of the eye. The example was poorly chosen.
t-sauer
4 hours ago
Your example also counters your point: A lot of people do not care about 4k, and even less do care about 8k or above. We reached a point where most people are completely happy with the quality they get.
lelanthran
4 hours ago
> Your example also counters your point: A lot of people do not care about 4k, and even less do care about 8k or above. We reached a point where most people are completely happy with the quality they get.
And when (not if) we get to the 4k (or 8k) equivalent of LLMs, they'll just be baked into hardware and we'll have near instant responses while running locally.
There is no future where OpenAI, Anthropic, etc survive with their current business model; at some point we will hit a point where training a new model is done only every 5 years or so, and in that scenario we aren't going to be running models of pricey server GPUs with latencies measured in seconds and full responses measured in minutes.
We'll be running locally with sub-millisecond latencies and responses measured in milliseconds. There is no way for any big company to compete with the current business plan of selling inference or subscriptions.
seb1204
2 hours ago
I dare say many are watching on handheld devices and not big screens. At least a sizeable percentage I would say. My guess.
phoghed
4 hours ago
Likely constrained by the fact people are mostly watching low bitrate Netflix streams.
sureglymop
4 hours ago
Don't forget that it's only an assumption that scaling more results in better models. There may be a ceiling to that. If that is hit and the best performing possible model can run in little VRAM, your argument here doesn't hold anymore.
In a way it is actually the same thing that makes us accept AI as working in the first place. It only needs to be good enough for human perception. The same is probably true for compute.
oblio
4 hours ago
I'd argue we've already hit the ceiling. Can you truly tell the different between SOTA models from 9 months ago and those from today? There are some improvements but they're mostly marginal.
Plus there is a chance the actual scaling that matters is beyond our reach. Think instead of TB models, PB or ZB models. We don't even have that kind of information. Humanity in its entire history hasn't generated 1ZB of information.
xyzsparetimexyz
an hour ago
> Can you truly tell the different between SOTA models from 9 months ago and those from today
On a task that corresponds to the benchmarks, yes absolutely.
flohofwoe
3 hours ago
FWIW I still play PC games on 1080p (also has the nice side effect that I don't need to buy an overpriced highend GPU, nor use upscaling hacks like DLSS which trash image quality).
lelanthran
4 hours ago
> Standard Def TV was plenty for 50 years. But when more was on offer, everyone went for it, and now you can't even buy a 480p TV.
And yet when we hit 4k, that's were people just stopped buying higher res. 8K is still useful, but only when the screen is so large that it doesn't fit in the room :-/
bogdan
4 hours ago
There's barely any media in 8k. There's so many ppi you can squeeze before your eye can't tell the difference anymore.
odyssey7
3 hours ago
Historical precedent: Quickselect, or any other discovery of a surprising algorithmic speed-up.
LLMs haven’t been on the scene for very long. There’s still lots of room for efficiency discoveries. Plus, I keep hearing quantum is going to be a big deal in the next few years.
_glass
3 hours ago
Super interesting, when Quantum computing would enable really large models, or much larger contexts. But QRAM is even more behind than pure fault-tolerant gates.
EarthMephit
3 hours ago
I haven't tried it, but colibri is meant to allow running GLM 5.2 the 744B model in 32GB by using prefetches and streaming from fast SSDs.
kcb
an hour ago
Terrible performance on $20,000 worth of GPUs.
fooker
an hour ago
I'm claiming that the max amount of intelligence you get out of a 100GB model is unlikely to be that much lower than what you can get out of a 1TB model.
yieldcrv
4 hours ago
It’s not just about small models, that’s only one part of evolution
Some groups are baking models into silicone, Deepmind has an example, it gets 18,000 tokens/sec on Llama 3.1, not sure about parameter size
fooker
an hour ago
> Some groups are baking models into silicone
While some other groups are baking silicone into models :)
intrasight
4 hours ago
I think this is the future - at least it will be for on-device models. Apple, for instance, will "bake silicon" once a year for their current model, and use that chip in all their devices.
petesergeant
3 hours ago
> How would we do that? There's no historic precedent for that. Its fundamentally an information theory thing
Maybe? Human science history is absolutely littered with examples of things that were “constrained” by a fundamental law … until they weren’t, because we’d misunderstood or misapplied the law.
lelanthran
4 hours ago
> How would we do that? There's no historic precedent for that. Its fundamentally an information theory thing: what's the max amount of intelligence you can get out of 1 KB/MB/GB? There has to be a limit and I'm not convinced that it's far off.
Sure, it might be very near using the current approach, but... it might also be might be very far off because we are using the wrong approach.
I mean, look at the max amount of intelligence you can get out of a human brain powered by two bananas...
leoc
3 hours ago
Though even if there are big further breakthroughs to be made here, they won’t necessarily be made soon, and if they aren’t made in the next ~10 years then they won’t really affect the outcome of chip investments made now or in the immediate future.
bebe8393jrir
2 hours ago
Just replicate dogs brain. It is the ultimate state of art (far smarter than people). But you need a few kb of memory to replicate resoning skills.
Dogs are very smart!
raincole
5 hours ago
And by the time the models that require 1TB memory will be better.
It's like people saying mobile chips are going to be better than PC chips... until they realize PC can be made of mobile chips too if that comes true.
fooker
an hour ago
No, it's like saying a chip that consumes 500W is going to be better than a chip that consumes 200W.
This used to be true for the first several decades of microprocessors, and stopped being true in the early 2000s.
petra
42 minutes ago
And besides llm's memory requirements:
-HBM has a lot of room to grow(bandwidth/power/density). And an HBM memory takes 3x the amount of wafers of regular DRAM.
-HBM + better DRAM designs could increase random memory access speed by alot. I've seen research talking about 7x.
At what point it makes sense that half of consumer GPUs and most of the server GPU/CPUs start using HBM?
leoc
3 hours ago
I think the chipmakers are pretty securely set up here, the usual cyclical issues aside. Between the stubborn resistance of the data to being compressed without significant loss, and to a lesser extent the ability to do more with more memory (we’re surely at diminishing returns for LLMs already, but diminishing doesn’t mean zero) it seems like a fairly safe guess (I am not an expert on anything) that for the next decade at least you’ll still want at least 128-256GiB of video RAM to have a comfortably “SOTA-enough” experience. Then the other side of the rectangle comes from putting that VRAM in every consumer and office-worker device. Selling everything at a big premium to a few efficiency-obsessed SaaS providers is what you do when you’re temporarily supply-constrained. Sidelining the SaaS middlemen and selling to gloriously inefficient consumers who will leave their laptops closed for most of the day is the dream. This is probably not unrelated to the fact that the country which is currently making a big strategic push into memory manufacturing is also championing open models.
ImprobableTruth
5 hours ago
The fundamental issue is that for these systems bigger is essentially always better (if affordable). So if we can squash something like Kimi K3 down to run on a 'normal' system, that just incentivizes devs to increase the model size until once again we're at the limit of what can be run.
Dibby053
5 hours ago
What do you base this prediction on? It seems very unlikely, unless by SOTA you mean the current SOTA.
oblio
4 hours ago
> What do you base this prediction on?
Their posterior. Or worse, their interest in going to the moon (HODL style, not Artemis style).
tornikeo
4 hours ago
SoTA AI will move well past 1TB+ memory requirements in 1 year or so.
Also, you will be able to play with much more competent models locally. They will still feel like children compared to the adults living in the SoTA region.
clbrmbr
4 hours ago
Are you thinking of fundamental architectural changes (like the Transformer)? Or incremental (like MoE or GQA)? Are there specific neolabs or techs you are following that lead you to this prediction?
Indeed it seems quite possible we are one architecture breakthrough away from existing chip stock driving us all the way to ASI.
Nobody has done AFAIK the information theory to prove it’s not possible.
whazor
4 hours ago
Yeah, you can directly print the SOTA AI model on the chip. I also believe this can be done on older process nodes. The most important factor would be speed. How fast can you go from new model to new chip?
dist-epoch
4 hours ago
Even if we do, we are going to need a lot of RAM for the billions of agents running everywhere.
hahahaa
3 hours ago
"we" being China?
greenavocado
4 hours ago
NOPE. The opposite will happen. CXMT will flood the market thus making 1 TB models affordable.
yieldcrv
5 hours ago
that would be perfect timing
Hyperscalers are screwed, data center mania couldn’t even be completed during this massive spending spree, all while the people seemingly sitting on the sidelines are working on getting these things baked into the OS and chipsets in a way consumers wouldn’t notice