NitpickLawyer
8 hours ago
This will be interesting for a few reasons. First, depending on where the median pricing settles w/ 3rd party providers will tell us what it costs to serve a 3T model. Since it's going to be mxfp4 native, it'll take ~1.5TB of VRAM to host this, which is juuust at the limit of 8xb200s (but realistically you'll need 16x for context / throughput optimisation). Won't be cheap to host, but at least we should get some range of $/MTok for a 3T model. Then we'll be able to guesstimate if "labs are subsidising tokens on API pricing".
Also interesting to see what effort it will take to fine-tune this beast. The latest AISI benchmarks on cybersec place it above glm5.2, but still way way behind SotA closed models. Some fine-tuning might be needed here. Also, interesting to see if Cursor does another training round on it, to directly compare it w/ kimi2.6/2.7 fine-tunes (composer series) and grok4.5.
Also also, interesting to see if someone takes on distilling (proper distillation, w/ training the entire distribution) from this into smaller models. (dsv4-kimi should be really good, since dsv4 is very cheap to serve)
walrus01
6 hours ago
It will be very interesting to see what kind of 'slow' performance people get from running it on a no GPU, but tons of RAM server (like a dual or quad socket xeon with 1.5 to 3TB of RAM). For the purpose of giving it longer duration tasks to generate a piece of something and come back and check on what it has done in 4 or 6 hours. Even if the output is like 5-6 tok/s, that might be usable for some purposes.
Huge price difference in what you can do with buying a used 4U rackmount server and putting 3TB of RAM in it (64GB DIMMs x quantity 32 in a quad socket xeon, you can see some benchmark prices on eBay for sets of 16 or 32 matched 64GB ECC DIMMs) for <$30,000, vs the cost of trying to run it on real GPU hardware.
Now obviously, as of the time I write this, the full precision hasn't been released nor has anyone like unsloth run it through quantization yet to produce a "Q8" or "Q8-XL" variant of it. But I think it's going to need more than 1536GB of RAM, with a usable and large amount of context, more like 2TB and preferably 2.5 to 3TB.
I also predict that people who try to run it in Q4 and Q6 will get the worst of both worlds, less precision/lost knowledge but also not reliable output that comes out too slow. In my personal opinion if I'm going to deal with something that is smart but slow and running on limited budget hardware, I need it to be Q8.
fooker
6 hours ago
> Even if the output is like 5-6 tok/s, that might be usable for some purposes.
You'll spend ~100x more on electricity than the API cost to have it run on someone else's GPU at several hundred tokens per second.
I think some sort of extreme data privacy requirement is the only situation that justifies this, but the intersection of {needs absolute data privacy, needs to run SOTA model, cannot afford GPUs} is really really narrow. I wouldn't be surprised if this is an empty set.
walrus01
6 hours ago
There are a number of use cases where sending the contents of your context and prompts (and the resulting output) to a 3rd party service is off the table as an option, and people will compromise speed for data sovereignty. And not everyone's electricity is equally expensive, I pay about $0.075 USD per kWh. It would for example cost me about $48 a month of electricity (not counting cost of cooling) to run a quad socket Dell R940 for a month.
mdasen
9 minutes ago
That's an unusually low electric rate for the US - way below the lowest state average which is Idaho at 12.4 cents. It's certainly possible that you are getting 7.5 cents including delivery, but I've had friends say that they're "getting 13 cents per kWh" here in Massachusetts, but that's just the supply rate and the delivery is another ~18 cents.
There are parts of states like Grant County Washington that have cheap hydro power, but it's very rare for power to be that cheap in the US. Even if this applies to you, it won't apply to the vast majority of people on here who will have electric rates 2-4x higher.
Average electric rates by region:
New England 28.1 cents
Mid Atlantic 25.1 cents
East North Central 20.8 cents
West North Central 14.8 cents
South Atlantic 16.1 cents
East South Central 15.5 cents
Mountain 14.6 cents
Pacific Contiguous 26.1 cents
Pacific Noncontiguous 42.1 cents
https://www.eia.gov/electricity/monthly/epm_table_grapher.ph...fooker
5 hours ago
Great, so the other member of the set matters for you more than cost.
Do you actually need to run the state of art model at 5 tokens per second instead of a qwen or whatever 7b or 30b model at 100 tokens per second?
sm-silversight
an hour ago
>Do you actually need to run the state of art model at 5 tokens per second instead of a qwen or whatever 7b or 30b model at 100 tokens per second?
Some people like doing things they want to do. Do I actually need to buy expensive pigments from europe to make paintings of flowers? My camera produces a much more accurate representation.
walrus01
an hour ago
Very good description of it. It does seem like a bit of a rhetorical question to ask a forum that has a very high population of Linux and BSD users why they might desire to have the option to do something themselves rather than relying on an external packaged ready to go product.
walrus01
5 hours ago
Do I really need to? No, not really. The 27B full density, 35B MoE, 70B and 122B models I have in use get me 95% of the way there on a lot of things. Particularly when dealing with languages and systems where I have at least an intermediate level of knowledge on, to know whether something is going down a dead end, using a wrong method, metaphorically chasing its tail, or is producing valid output.
On the other hand, would it be cool to also have a really big thing as an ancillary tool that I could throw a request into opencode before going to bed, let it crank away and take a look at what it's done 7 hours later? Yeah, particularly if I (very much an unknown quantity at this time) could be confident that it builds high quality, syntax valid, appropriately commented and not absurd code.
light_hue_1
4 hours ago
As someone who has worked in two industries that are at the maximal end of data sensitivity and privacy this comes across as a tinfoil hat issue not a real business requirement. In such cases we've always found ways to trade dollars for the privacy we need without having to run our own inference at excruciating slow speeds.
walrus01
4 hours ago
Do you mean by trading dollars for the privacy you need as:
a) Contracting with a third-party independent inference provider who will run your choice of model on fast hardware that they own, with all appropriate data security/privacy/contractual/compliance protection in place
or
b) Contracting with the original creators of the model to run inference via their API and with assurances that all the same data protection is in place
or
c) Spending the money to buy your own inference hardware to run it on something you fully own/control at proper usable speeds?
Edit: Everything I've been writing in this thread is mostly within the context of being able to evaluate K3 and its usefulness to be self-hosted as a preliminary proof of concept or test of feasibility of a new thing, such as on <$20,000 of server hardware, before proceeding to spend 300-400k on GPU-related hardware, or external third party services/ongoing billing.
jmalicki
2 hours ago
A) is very doable with e.g. Amazon Bedrock.
They'll give you HIPAA compliance, they even have a data center for US government classified data, they can give you European data sovereignty. And with OpenAI and Anthropic models to boot, you don't even have to settle for open weights.
What kind of privacy needs do you really have beyond that?
walrus01
2 hours ago
It is not my use case but given recent political developments in international relations caused by the executive branch of the US government, off the top of my head, I could think of a lot of European or Canadian firms for which that would not be an option. No matter what they might promise about European sovereignty. For a good 'ol patriotic US domestic company? Sure.
Taunt4
an hour ago
Yes, its the US cloud act risk EU companies run up against on hyperscalers like MS/AWS.
Even for EU companies running open weights on EU stacks LLM inference on the GPU must process plaintext and I can't find any EU provider with NVIDIA H100/H200/Blackwell CC mode plus SEV-SNP or TDX, where you can cryptographically verify the workload ran somewhere the operator cannot inspect.
Personal compute is therefore the only option if you want personal autonomy privacy for IP &c. Maybe another option is to use cloud compute rented to fine tune a personal model that suits your own needs that would help bring the cost down, I don't know enough about this area to know if it kills the "intelligence" of those domains due to limited ?cross-verification within the LLM.
trollbridge
an hour ago
Interesting. So nobody would have had a problem with you running stuff on Chinese AI providers?
I have some inference I simply don't want to run on OAI, Anthropic, or Google because I don't want to run afoul of their "rules" and end up with a banned account, and this situation is only getting worse when it comes to doing fairly basic tasks like trying to secure your app against security problems.
solarengineer
4 hours ago
There are regulated sectors in countries where data sovereignty is important enough that the sector sticks to air-gapped on-prem hardware and does not use cloud services at all. They have the dollars to pay for more than what it would cost to run on the Cloud.
frognumber
3 hours ago
Having worked in / adjacent several such industries, a lot of the question depends on scale.
A trillion-dollar business can easily trade dollars for the privacy. A business with $1M to spend won't even get a phone call with OpenAI or Anthropic, who were the only* previous players in town for doing this.
Worst-case example: Bootstrapped startup working in military.
It's also the case that an open model enables many more intermediate-cost solutions. E.g. providers certified for specific applications, on-prem rentals, etc.
* Omitting Azure, which gives some privacy for some $$$ on their models, but not at the level of high-security.
amluto
3 hours ago
> Omitting Azure, which gives some privacy for some $$$ on their models, but not at the level of high-security.
If I were ranking third parties on their ability to safely handle my data without compromising it, I would rank Anthropic pretty low for things like Fable (where they more or less promise that they will misuse my data), but I want Azure pretty low in the sense that I fully expect them to be compromised.
I would tend to trust Amazon to avoid being compromised.
cobbzilla
3 hours ago
I’ve priced it out: max $135/month to run a dual Xeon 2U server with 3T RAM & 2x 22 core Xeon Gold. It’s the 2x 750W power supplies that ultimately determine opex. My power costs $0.124/kWh, the $135 assumes drawing maximum power continuously, and in that case, I can probably offset my heating bill a little bit in the winter, so maybe effectively a little bit lower.
I don’t know if that’s 100x more than I’d pay (opex-wise) with an nvidia setup, but I can say the one-time capex is a great deal cheaper. Avoiding VRAM and DDR5 (fast DDR4 should be OK) are the biggest cost savers. ECC RAM is worth the extra price. General datacenter-quality hardware has less price sensitivity, and plenty of bang for your buck.
walrus01
3 hours ago
Keep in mind that just because it has dual 750W power supplies that doesn't mean it's what its load will be, for a full CPU loaded wattage figure you'd need basically a pair of kill-a-watts plugged in inline on the feed for each poewr supply and then run stress-ng with artificial cpu stress on all cores for an hour.
Under heavy inference load you will find that the cpu usage is actually less as the bottleneck is the RAM bus speed. An older 2U rack server that is 600W load (typically a 1+1 power supply server when plugged into two kill-a-watt would show 300W on each, equal load balancing) when maxed out with stress-ng might be only 450W total running inference.
If you have 600kWh used in a month by running something 24x7 and your power is $0.15 a kWh, that's more like $90/mo (not counting cooling or any ancillary costs for the environment where it's in).
trollbridge
an hour ago
If you actually were running this thing at 80% or 100% load, then the first thing you'd want to is get a better PDU and then connect your servers to that (48V DC).
sebmellen
3 hours ago
(Context: Parent comment was edited after I wrote this comment)
Where in the world are you finding that much RAM in a racked server for $200/month?
walrus01
3 hours ago
I think he means electrical bill at his estimated wattage load of the server and his known kWh cost, not rented server/hosting cost.
sebmellen
3 hours ago
Aha, right. That makes a lot more sense.
cobbzilla
3 hours ago
At my house. I have 5Gbps fiber and could pay for 10 or 25 if I need it.
embedding-shape
6 hours ago
I dunno, K3 thinks a lot before it actually replies, and you might be in the ~1 tok/speed region or even "seconds / tokens", and with K3, you'd wait days if not weeks for a reply in that case.
Don't get me wrong, slow is sometimes better than "not at all", but depending on the performance, it might end up way too slow to even work for batched/async jobs like that.
walrus01
6 hours ago
I agree it's very likely to be painfully slow, I very much want to see some real world results from people who try it. Early testers will inform others on whether it's even worth trying. Results very much TBD right now. I don't have a system sitting around here with 2TB of greater of RAM that isn't already committed for other uses, regretfully.
embedding-shape
5 hours ago
Lets say an easy response takes 32k tokens in total, and to be generous, let's say it does 1 tok/s. This is already ~9 hours, and 32k reasoning tokens isn't even that much and as mentioned, K3 probably does the longest/most reasoning/thinking out of the available open weights models today, much like GLM. Just lowering that performance to 0.5 tok/s, would lead to ~18 hours for a simple prompt to receive an answer.
And then that's just for single prompts, what about agent harnesses, where before every tool call the model could reason a bunch?
I agree with you that real world results would be interesting, but I wouldn't hold my breath nor expect it to realistically be able to be useful. Still, people should try it, for science if nothing else :)
justincormack
2 hours ago
You can rent one in the cloud to try it
barbacoa
39 minutes ago
They are saying that AMD's new Epyc Venice CPU has 16 memory channels allowing up to 1.6Tb/s of bandwidth. Which is higher bandwidth than most non-HBM GPUs.
So full CPU local AI inference may become viable option in coming years.
johndough
4 hours ago
> But I think it's going to need more than 1536GB of RAM, with a usable and large amount of context, more like 2TB and preferably 2.5 to 3TB.
The model is known to be MXFP4 according to Kimi's release blog post, so the model weights will be less than 1536GB: https://www.kimi.com/blog/kimi-k3
Also, their previous models were native INT4, so it would be weird if they went larger now.
magicalhippo
5 hours ago
> running it on a no GPU, but tons of RAM server
Or from SSD using something like Colibri[1]. Not going to be quick, but at least runable.
walrus01
5 hours ago
It's a great concept but I think it would cross the line from 'very slow' to 'so slow it's unusable' at this size. Even if we say you have an NVME SSD that does 7GB/s reads, that's dramatically slower than being able to hold the whole thing in DRAM. Like the difference between 1.3 tok/s in RAM vs 0.1 tok/s with a colibri-like method.
edit: the results I have seen from people trying colibri with fast consumer grade PCI-E 4.0 NVME SSD are 0.1 tok/s on models that are <700B in size, things that are well under 800GB on disk. With something that's 3T in size it'll probably be a lot slower than hat.
zozbot234
3 hours ago
For single stream inference of a MoE model, the size of active sparse parameters will matter a lot more than total parameters. This is generally around half of the reported active parameter count - the other half being a dense subset that can be easily cached in VRAM even on fairly modest consumer setups. So the achievable performance may be quite a bit better than a naïve assessment might suggest.
LtdJorge
3 hours ago
On a server machine you can have more than 100GB/s of NVMe if you parallelize (RAID 0 and the like). But it's still gonna be noticeably slower.
walrus01
3 hours ago
1536GB of DDR4 ECC server RAM is somewhere between $4000-6000 USD used right now, by the time you put in parallel enough NVME SSD to approach good speeds, you'd be approaching that (and also likely running out of PCI-E bus lanes directly attached to the same motherboard to reasonably do so).
magicalhippo
4 hours ago
It claims to support using multiple devices RAID-0 style, which should boost performance, but yea probably not very useful for most.
But still fun you can run it at home.
Havoc
3 hours ago
> Even if the output is like 5-6 tok/s
On a 3T model I’d imagine you’d be closer to 0.05 tks
amluto
3 hours ago
Presumably it’s MoE and only needs to read a small fraction of the weights per token. Bonus points if you can get decent speculative decoding without becoming ALU-limited.
zozbot234
2 hours ago
Speculative decoding is not really worthwhile for sparsely-loaded models. You end up paying in both memory bandwith and compute (loading experts based on wrongly-predicted tokens) which leaves you worse off overall. It becomes viable (even for sparse MoE) once you're batching so widely that you end up having to load most of your total weights anyway.
amluto
38 minutes ago
> Speculative decoding is not really worthwhile for sparsely-loaded models.
If wonder if you can train a model to optimize this, by trying to make the expert selection sticky across a few tokens, without too much quality loss.
Another fun idea might be to try to build a model where the router chooses the expert 1-3 tokens in advance.
sandworm101
6 hours ago
Those old LTT videos of high core-count threadrippers running GPU benchmarks become more relevant each day.
walrus01
6 hours ago
The performance bottleneck is not really so much the number of cores or processing power in each core, but the memory bus bandwidth to/from the CPU. I have an older dual socket xeon server here which is a CPU-only LLM test machine with 256GB of RAM and the actual CPU stress is not much, I can even quantify this by how little it spins up the CPU fans to meet thermal load (the CPUs are operating at nowhere near their 180W per socket max capacity, compared to like, crunching prime numbers or running cpuburn).
But the memory bus speed is fully committed when generating tokens or thinking.
sandworm101
2 hours ago
Thats where the threadrippers really excelled. They had the lanes for memmory access. We might soon see the return of dinner plate-sized CPUs with thousands of pins.
walrus01
23 minutes ago
The epyc Venice SP7 socket is apparently 9324 pins
woctordho
7 hours ago
Speaking of finetune, currently a common practice is LoRA over bnb 4-bit base model, but I think it's time to replace bnb with GGUF as the base model format. GGUF is actively supporting new model architectures and more aggressive quantizations.
I've made some proof of concept in https://github.com/woct0rdho/transformers5-qwen3.5-recipe . We can finetune Qwen3.5-35B-A3B in 16 GiB VRAM, and DeepSeek-V4-Flash (284B-A13B) in 90 GiB VRAM, without CPU offload. This works well on unified memory machines like Strix Halo.
Even so, larger models like Kimi-K3 still require multiple GPUs and nodes, and there are a lot more to do compare to single-GPU training.
maxignol
4 hours ago
I dont quite understand why GGUF is better optimized. Are the performances better for the same amount of VRAM ?
qeternity
2 hours ago
I cannot be sure what the likes of Cursor have done, but I think it's incredibly unlikely that they have trained a QLoRA for Composer.
It's almost certainly full parameter post training of the original model weights.
vb-8448
6 hours ago
> Then we'll be able to guesstimate if "labs are subsidising tokens on API pricing".
No, you don't. Without training cost you can infer only the marginal cost of serving this kind of models.
Moreover, you don't know the actual size of closed models (what if Fable is a 10T model? What if it's 1T?)
lelanthran
3 hours ago
> No, you don't. Without training cost you can infer only the marginal cost of serving this kind of models.
Still useful; "are the labs marginally profitable just on the marginal inference costs?" is still a useful question to answer. After all, if they aren't even profitable on inference in isolation, then we can expect to see large price increases.
If they are able to turn a marginal profit on inference alone, then perhaps the price increases won't be so severe (or perhaps they expand the time between generations so that they spend less on training but take longer to complete training).
"Are the labs profitable at all?" is, of course, a much more useful question, but that doesn't mean that the first question is completely useless.
vb-8448
2 hours ago
I wasn't saying it isn't useful, I was saying that you cannot infer even the marginal cost of closed models.
The K3 maths can turn true only if the models size is roughly the same and the labs didn't find any better way to run inference at scale.
solenoid0937
2 hours ago
We know labs make money on inference, and we know they lose a lot of money on inference+training.
lelanthran
2 hours ago
> We know labs make money on inference
We don't really know that, for OpenAI and Anthropic. We suspect that, but as far as I know, even they have stopped claiming that they are profitable on inference.
anthonypasq
3 minutes ago
unless you think that Opus is 10T+ params, its pretty much impossible for inference not to be profitable when doing some basic napkin math on other open models, and if Kimi K3 is 3T params with the same performance as Opus then that means that China is actually way more technologically advanced than the American labs.
So which is it?
vb-8448
2 hours ago
Just out of curiosity, based on what we know for sure they(OAI+A) make money on pure inference and lose on inference+training?
nl
2 hours ago
OpenAI's financials leaked and showed this pretty convincingly.
Anthropic was probably profitable last quarter, without training costs: https://www.wsj.com/tech/ai/mind-blowing-growth-is-about-to-...
amelius
6 hours ago
> Without training cost you can infer only the marginal cost of serving this kind of models.
Which is by far the most interesting number of the two.
> Moreover, you don't know the actual size of closed models (what if Fable is a 10T model? What if it's 1T?)
If you get close in output quality, then does that matter?
embedding-shape
6 hours ago
> If you get close in output quality, then does that matter?
When you're trying to estimate/infer the costs of serving the tokens and even include the cost of training the weights in order to output tokens then yeah, why wouldn't that matter?
jvuygbbkuurx
5 hours ago
That only matters if you are an investor not if you are a consumer.
embedding-shape
4 hours ago
Well, or if you're participating in a discussion on HN where the sub-topic happens to be "if labs are subsidising tokens on API pricing" and literally the cost of serving the tokens is relevant to the sub-topic people are trying to discuss...
solenoid0937
2 hours ago
Consumers want better models too, of course it matters
vb-8448
4 hours ago
> Which is by far the most interesting number of the two.
Only if you don't have to continuously train new models, and you are not at a runway risk.
NitpickLawyer
4 hours ago
Training cost is directly impacted by inference cost nowadays. Most of the gains come from RL these days, and that is highly dependant on inference (~7:1 inference:training in units of compute). That's mainly because you want many roll-outs for each training scenario.
Of course inference efficiency is dictated by model architecture, size, etc. You can still guesstimate some of those and have an idea about cost/serve at several size tiers.
jeremyjh
4 hours ago
That isn't really relevant to GP's point. We still don't know what training costs because we don't know how much RL is done.
energy123
5 hours ago
Gross margins are insanely important, possibly the most important single metric if for some reason you were forced to choose one.
vb-8448
4 hours ago
Only if you assume that at some point, for any reason, there will be "the model" that doesn't need costly retraining.
I guess this is one of the reason Anthropic i so "active" for asking for a development break.
trollbridge
42 minutes ago
"I want my competitors to stop competing" is an interesting ask for someone who's currently charging the highest prices in the entire industry.
criley2
4 hours ago
>No, you don't. Without training cost you can infer only the marginal cost of serving this kind of models.
Are you talking about Kimi's training cost or the training cost of the model(s) that Kimi distilled?
Because Moonshot didn't even incur the majority of the training costs either
nl
2 hours ago
> the training cost of the model(s) that Kimi distilled?
The distillation process involves getting conversation traces from the model you are distilling from, and then training your model against them.
You still have to do the training!
tiagod
3 hours ago
I believe they are talking about the closed models' training costs.
I other words, the providers that will be offering K3 inference don't have any training costs to offset, so they are only charging for the inference itself. OAI/Anthropic would need to offset their R&D and training costs in order to not be selling API access at a loss.
swingboy
5 hours ago
Say a single Kimi K3 is deployed on 16 x B200s: how many concurrent users can that handle? I realize the question assumes a major simplification that everyone's prompts/sessions are the same.
serf
33 minutes ago
>I realize the question assumes a major simplification that everyone's prompts/sessions are the same.
well, exactly.
that's tough to answer without just average sampling because some users will ask the model "what's todays date" or "what color is the sky?" and some users will ask "Let's rewrite the linux kernel in brainfuck."
woadwarrior01
6 hours ago
Since the model is natively MXFP4, I think it'll be even more interesting on the hardware front. It'll comfortably fit on a 8x AMD MI355X node. I suspect that'll drive token prices down, further.
Palmik
2 hours ago
You're assuming inference providers are going to sell tokens at cost. You're also assuming that the inference providers have will optimized inference engine. I haven't seen that to be the case so far, to be honest.
Take a look at GLM 5 vs GLM 5.2 pricing -- GLM 5.2 cost more despite being the same model.
Take a look a look at DeepSeek, which hosts DS v4, profitably, yet others aren't able or willing to match the price.
trollbridge
41 minutes ago
Xioami's MiMo did match DS-V4's price, although we now know that DeepSeek set their pricing lower than they could have, due to the leaked memo, and simply decided to use "10 months to recover capex" as their yardstick. Interestingly "10 months to recover capex" is also the same price SpaceX is renting space to Anthropic and Google for.
NitpickLawyer
an hour ago
I'll be honest, I typed that message while having morning coffee, so it's just a quick reaction from my part, not a heavily researched article in a journal :)
But I do think that the median price where this settles will tell us something about the floor at which it is profitable to serve this model.
> DeepSeek, which hosts DS v4, profitably
I specifically mentioned 3rd party providers, because there can be an argument that model creators themselves are subsidising tokens to gather training data for the next model. In fact, ds are public about their gathering of data (at least on openrouter they're marked as such). So that 0.x price point for dsv4-pro is likely subsidised.
nl
2 hours ago
I think it's unclear the the DS hosted prices are profitable. AFAIK that haven't claimed that.
OTOH, the multiple providers who have settled around the same price point ($3.48/M output tokens for multiple providers with good reputations) does indicate where it is profitable: https://openrouter.ai/deepseek/deepseek-v4-pro#providers
blks
4 hours ago
If it takes so much resource to run, how does the sharing of a single llm works? There is some interface that basically submits context/cache plus current promt, from each user, doing basically time-sharing compute?
mnahkies
4 hours ago
I believe they batch requests so the same weights in vram are shared across many requests https://www.baseten.co/blog/continuous-vs-dynamic-batching-f...
rhdunn
5 hours ago
If it is a mixture of experts (MoE) model like the 2.x models, won't this reduce the hardware needed to run the model?
The Kimi-K2.6 model is 1.1T parameters with 32B active parameters. With light quantization (Q6_K) that's enough to run it (slowly) on a single 5090. On a single B200 you can have 5-6 experts loaded into VRAM at a time. Realistically that would be 3-4 to account for the context. [!]
[!] With this and other MoE models it looks like an interesting area for research would be to detect or predict which models would be needed ahead of time. That way you could schedule the load into VRAM step before the weights are needed. That way you shouldn't lose much/any performance from offloading the weights to RAM.
petu
4 hours ago
You need whole weights in VRAM for optimal performance. Don't be confused by "experts" in the name -- you don't get to load static subset of experts and blast next 100 tokens with them. In typical MoE model they get switched "randomly" on every token, so all experts have to be readily available.
> With light quantization (Q6_K) that's enough to run it (slowly) on a single 5090. Kimi K2.6 is released as INT4 already.
So 5090 with K2.6 is just gonna sit idle 99% of the time, waiting for next slice of weights to load.
5.6 Sol calculates that single 5090 in raw compute & memory bandwidth can run K2.6 at 35 t/s (256k context depth) -- if it somehow had enough memory to hold whole model in VRAM. Man, I hope HBF succeeds and Nvidia brings it to consumer cards in 5 years..
zozbot234
4 hours ago
> In typical MoE model they get switched "randomly" on every token, so all experts have to be readily available.
It's worse than that: a typical MoE model routes a separate set of experts at every layer, not just every token! But in practice, RAM offload (for systems with non-unified VRAM) and even SSD offload still work surprisingly well given some amount of caching.
You can likely recover compute intensity and throughput by batching requests together, which (in practice, depending on sparsity) will end up reusing some of the loaded experts with high probability; though the obvious tradeoff is that having to store KV caches for the wider batches may leave you with less room to cache experts across layers and tokens.
(Plus if you're batching so widely that you end up loading essentially entire model layers, MTP then becomes applicable even for a MoE model. But this typically only applies if you're doing inference on a very large scale, or if your memory bandwidth is so scarce that you have to recover compute intensity by any means feasible.)
embedding-shape
5 hours ago
> The Kimi-K2.6 model is 1.1T parameters with 32B active parameters. With light quantization (Q6_K) that's enough to run it (slowly) on a single 5090
Without leveraging system RAM and/or SSDs, I don't think you can, or how exactly are you running this, if this is something you are doing today? With CPU/expert offloading you could probably do it with a 5090 + 1TB of RAM or something like that, but absolutely not on a single 5090 entirely within VRAM.
robotswantdata
5 hours ago
Yes hybrid approaches are much better than people realise.
There are a lot of optimisations that are not in the public sphere, source working on start up in this space
embedding-shape
4 hours ago
> There are a lot of optimisations that are not in the public sphere
Sure, but if we're participating in public discussions, isn't it more fun if we talk about things people can actually read and understand, rather than secret stuff other's can say work, but no can actually validate or know how it works?
It sounds like "hybrid approaches are much better than the public is aware, because everything else is private and secret", but also: ok, so what? No one can run that anyways, (yet?), so why it matters?
rhdunn
5 hours ago
Yes, that's what I was saying w.r.t. expert offloading, i.e. ensuring that the GPU could fit the active parameters not all the parameters.
embedding-shape
4 hours ago
Alright, I guess I misunderstood. To be fair, this part:
> The Kimi-K2.6 model is 1.1T parameters with 32B active parameters. With light quantization (Q6_K) that's enough to run it (slowly) on a single 5090.
Is painting a very different perspective, even considering the latter parts it's hard to read that as "Of course offloading everything else that doesn't fit on the GPU itself". But anyways, it's been clarified now so no harm :)
walrus01
5 hours ago
I assume you mean putting only the 32B active parameters on the GPU, and the rest on a bunch of regular server DRAM like on a 768GB to 1024GB RAM server?
Because Kimi K2.6 in Q4 is about 584GB GGUF size on disk and will use slightly more than that in RAM, Q8 is 595GB.
NitpickLawyer
5 hours ago
You're talking about running this "at home" for 1 user, using a mix of VRAM and RAM (total should be ~1.5TB). That's certainly possible. It'll be slow, especially prompt processing, but doable for single users.
But my comment on running it was more towards serving this profitably at scale. You get much better throughput / unit of compute if you load everything in VRAM and serve many requests at the same time. That's how all inference providers are doing it.
rhdunn
5 hours ago
I was talking about running this on a server, hence my comments re 1xB200. Obviously, the more hardware/VRAM you have the better/faster you can run these large models. But if you are a small/medium sized company you could feasibly do it on just one B200. It all depends on how much hardware you can afford to run.
narwhster
2 hours ago
This is a very insightful eye opening take. I haven't even thought of it this way. This really is the first open model to be as big as the frontier has been until now.
I think this release is actually great both ways when you think about it. We gonna be able to learn knowledge that labs have been hiding from us (e.g. cost like you mentioned). And labs could learn from whatever optimization techniques people come up with when trying to host this model.
It's honestly just good for everyone in my opinion.
epolanski
21 minutes ago
> The latest AISI benchmarks on cybersec place it above glm5.2, but still way way behind SotA closed models.
Sota closed models don't even answer cybersecurity questions lol.
withinboredom
4 hours ago
We have a lossless compression codec (working on open sourcing it over the next couple of weeks) that reduces it down to its minimum entropy -- it cannot be compressed further. On all tested large models, it's a ratio of 1.34-1.23 -- and smaller models up to 3.76x. It also increases the effective bandwidth by the same rate.
xscott
3 hours ago
Exploring compression algorithms for weights is a good idea, and I hope you have a successful product. However, if you can prove this statement:
> reduces it down to its minimum entropy -- it cannot be compressed further.
I think you could make a lot more money elsewhere :-)
https://en.wikipedia.org/wiki/Kolmogorov_complexity#Formal_p...
withinboredom
37 minutes ago
We're not an AI company ... nor do we have any reason to use it. Just a fun idea that was fruitful.
Alpha3031
4 hours ago
That's very interesting. Does that mean you can reduce say, a 30B class Q8 from ~30 GB down to 10 GB or less?
withinboredom
an hour ago
704gb -> 564gb; 358 gb -> 270 gb; 28.79 gb -> 7.65 gb; 439 gb -> 93 gb
It depends on the total entropy of the model. Smaller models have less entropy.
XCSme
an hour ago
So nowadays the hardware and hosting providers must be in an optimization race, whoever can make the model just a bit smaller or more efficient (to fit on fewer/less powerful cards) will have a huge advantage and can make a lot of money.
I am curios what's the most profitable thing to "plant" (agriculture analogy) on the land (cards) that you have have: web hosting, vps, llms, image/video generation, etc
vorticalbox
5 hours ago
I think cursor will likely do a grok fine tune rather than a kimi one for the next composer.
they noted in their blog post they didn't focus purely on coding for grok 4.5.
NitpickLawyer
4 hours ago
Most likely, since they were acquired. But for us outsiders it would be a cool thing, to see if the delta is the same between kimi2.x + cursor data -> kimi3 + cursor data.
dist-epoch
7 hours ago
> if "labs are subsidising tokens on API pricing"
> SemiAnalysis estimates that Anthropic's current blended gross margin has risen to the mid-60% range, with the API business gross margin exceeding 80%
Of course, people will insist "they are lying", "why should we believe them, it's well known they subsidize API pricing", ...
https://newsletter.semianalysis.com/p/anthropic-3q26-profit-...
https://finance.biggo.com/news/02d45650-b569-4d12-b44d-8d6d8...
NitpickLawyer
7 hours ago
Agreed. My (somewhat educated) guess is that top labs have healthy margins on API pricing. But this release will add another 3rd party / clear of conflict datapoint in this estimation.
npn
5 hours ago
even deepseek, with their current (dirt cheap) price, can earn enough profit to cover the cost (hardware investment?) in 10 months.
likium
40 minutes ago
Will the model even be competitive in 10 months though? Seems like models that reach top 20 on OpenRouter see 50% of all token spend by day 80, and 80% by day 180.
knollimar
6 hours ago
That numner blends in training or no?
Der_Einzige
6 hours ago
Anyone who thinks that the labs are not profitable on per token API pricing is delusional and hilariously wrong.
flaburgan
5 hours ago
It all depends if you count the fixed cost of training or not. And the cost of the hardware.
flyinglizard
5 hours ago
Or even the basis of the cost of hardware. There are lease deals, capacity traded for equity, various programs by Nvidia, there's absolutely massive depreciation, etc.