NitpickLawyer
8 hours ago
This is more exciting than k3, IMO. Dsv4 models are extremely cheap to serve. Improving their capabilities has lots of downstream effects, as it becomes "good enough" for more and more tasks.
DS was serving the pro version at extremely low prices for a long time, and they've had integrations with opencode & other providers, so they likely gathered a lot of data from real developers doing real tasks (on openrouter they were labeled as such). Now they can use those live scenarios to further post-train their models and improve them further.
Can't wait to see if distilling k3 into dsv4 brings additional improvements. Anyway, having fast cheap models getting better is great for the community. Especially since these don't "go away" on a provider's whim. Whatever capabilities they get, can be used "forever" going forward. And, at least flash can be ran "at home" with <10k in hardware, which isn't really possible / feasible with glm/k3 larger models.
dannyw
5 hours ago
Deepseek and moonshot are the only two providers I consent to training for.
arizen
5 hours ago
Why not also Qwen?
dannyw
4 hours ago
Qwen/Alibaba have stopped doing open weights releases for a while. No grudge or anything, I'm certainly not going to look at a gift horse in the mouth, but both DeepSeek and Moonshot have been very consistent with open weights as well as sharing actually detailed research.
In terms of open research, China has absolutely overtaken the US.
anon373839
4 hours ago
Qwen has a pinned tweet stating that 3.8 will be released as open weights soon. I guess it remains to be seen, though, if they’ll do the smaller model sizes or only the big 2.4T one.
applicative
2 hours ago
Xi going to shut down open-weighting of them in a matter of months. No one seriously doubts this. They will be too powerful and they will be gone.
gravypod
2 hours ago
What would the benefit of this be? If China stops open weights, US labs still have the intelligence frontier. Maybe once Chinese models have speed, cost, and intelligence beat but right now they don't.
Undermining out entire economy by subsidizing the release of DIY versions of our main economic drive sounds like a huge win for China.
applicative
39 minutes ago
> Maybe once Chinese models have speed, cost, and intelligence beat but right now they don't.
Yes this is why I referred to 'months' and was downvoted by people who can't distinguish their politics from reality. You are restating exactly the text you are criticizing.
This is the nature of mechanical parrotlike repetition of propaganda:
a) You can have 'frontier' models with closed weights. Your sentence is basically a contraction within itself, again, the weight of ideology.
b) No one actually knows whether or not the closed weights of Bytedance are beyond all existing frontiers. Again the weight of ideology blinds you to the fact that the real 800 lb gorilla of Chinese Ai is more closed than Anthropic.
anon373839
2 hours ago
Reuters was reporting that rumor. And then Xi made a public appearance at a conference in Shanghai where he said the opposite of that rumor.
applicative
an hour ago
No in that very speech Xi stated explicitly what amounted to: 'of course when we get a Mythos, it will be a state secret'.
In fact the overwhelming weight of AI use in China, the chatgpt so to say, is Bytedance's AI which is absolutely closed and uniquely opaque.
The press treatment of these matters was no good and they are slowly walking it back, e.g. NYT yesterday finally actually read the speech.
anon373839
13 minutes ago
Do you have a source for that? The NYT article does not quote Xi making that statement, and I can't find a source for that. In fact, Xi delivered veiled criticism of the security posturing the US has done:
> We should put in place laws and regulations, technological monitoring, early warning and emergency response systems in order to strengthen the line of security, prevent abuses and malicious use, and ensure that AI is always under human control. In the meantime, we should jointly oppose overstretching the national security concept in the field of AI and placing one country’s security over that of others.
There was a blog post by a state-linked broadcaster cited in the NYT piece:
> “China supports openness, but this does not mean it advocates for the unconditional proliferation of all capabilities,” the blog said.
The article also quoted an American journalist:
> “If these models do reach those dangerous capabilities, they are not going to let it be a free-for-all in terms of releasing them," Mr. Sheehan said.
But a distinction has to be drawn between real dangers (which many people believe LLMs have not actually shown, to date) versus "dangers" hyped up marketing purposes or domestic regulatory-capture motives. Presumably the Chinese government is less interested in the latter.
applicative
7 minutes ago
"Dangers" may indeed be hyped up for marketing purposes or domestic regulatory-capture motives. But as Xi stated, they will become real.
It is just a question of being unaffected by the motives of the speakers, which is what adults learn to do.
GTP
an hour ago
We can only wait to see if this is true. Another possibility could be that it was an answer to USA's government considering a ban on chinese models. In this way, he fueled the discussion around the importance of open weight models.
applicative
41 minutes ago
Xi actually said that models of the strength to pose security issues will be permanent state secrets -- in the same speech that the first western takes translated as actually using the words 'open weights' which of course nowhere appeared. Now: when will "models of the strength to pose security issues" appear? Maybe my inference 'months' is wrong, but they seem to be advancing rapidly.
The administration blather about banning open weights is characteristically confused. Xi has already stated (what is obvious) that he will ban security-endangering weights and keep them a state secret.
1234letshaveatw
an hour ago
That isn't the Chinese way. They are much more focused on undermine and extinguish. Just look at the European car industry- on its way to being non-existent after the market was flooded, bye bye manufacturing base. Undercut the US AI providers and wait them out, they will go private after they have a stranglehold
applicative
44 minutes ago
Even Xi's speech was as paranoid as Dario Amodei about the possibility of Chinese AI achieving something at the level of say Mythos. If you seriously believe such a thing will be on Hugging Face I really don't know what to say.
criley2
37 minutes ago
There is already something on HuggingFace at the level of Mythos. It's called Kimi K3 and it's running laps around Opus5, Fable5, and Sol56 at cybersecurity.
It's so good that the US government is rushing to ban all Chinese models as fast as they can.
applicative
30 minutes ago
Fine, you believe Anthropic was lying about Mythos. In fact no view of the matter is needed to formulate the proposition above. It is clear China will soon have a model with the powers imputed to Mythos but which you deny of actual Mythos. It is a trial to deal with this level of ideological blindness; 'ideology has no outside'. If someone says 'when they get Mythos...', declare the possibility of Mythos to be the real lie. But it is quite possible and will be coming in months.
zozbot234
4 minutes ago
Reading between the lines, it's fairly clear that Mythos Preview only got its reported results in offensive cyber thanks to a highly specialized harness and humongous amounts of test-time compute. We also don't know what kind of post-training/fine-tuning it got on cyber, Anthropic only stated that it wasn't specifically trained for cyber-offense but that leaves a lot of scope for training on other related things. Overall, the Mythos story is a lot less interesting than most people might like to claim - but note that this doesn't require any actual lying.
The interest around K3 is a lot more defensible because we actually know what the architecture looks like, how much effort it takes to get it to run, and what kinds of results it gets on cyber evaluations. And no, it's nowhere near "dangerous" enough to where people might honestly want to ban it for safety reasons. It does a good enough job at fixing cyber issues, but that's hardly a safety concern.
darkwater
2 hours ago
What makes you state this?
riskd
2 hours ago
Decades of xenophobic propaganda
applicative
43 minutes ago
No, reading Xi's speech. They aren't going to give their Mythos - presumably a few months away - to the Sinaloa cartel or Uighur hackers or the US military or etc etc
1234letshaveatw
an hour ago
Wow! Amazing observation! Wait, do you think there could have been any pro Chinese propaganda? Perhaps even some taking place at the present time? I think it could be possible?
baublet
2 hours ago
China bad
applicative
43 minutes ago
I'm a) quoting Xi's speech and b) projecting rapid China development to Mythos level. Either you think China will never get to 'Mythos level' because you think China bad; or you think they will release the weights, because you think China bad. Its incredible the weight of ideology over facts in this discourse
ReptileMan
an hour ago
It is true. And only Chinese Han people can continue work on what is already released. But they will be forbidden. /s
Sarcasm aside - if the community can't get their shit together to continue improving what is currently public, well - we don't deserve free stuff and open weights.
siva7
24 minutes ago
CCP will be happy! Go on and share all your data with them..
dnhkng
8 hours ago
Totally! This with DwarfStar delivers usable local AI (I hope!)
dannyw
5 hours ago
Usable local AI has been here for a while, esp on say a 5090.
You can’t treat Qwen3.6 like its fable, but if you prompt precisely and specifically it’s a great executor.
I actually found it refreshing to use more of my brain for once, and actually have to think deeper about what I’m trying to do, and how to build it.
Tepix
3 hours ago
What are your goalposts? Depending on your requirements, there have been many moments of usable local AI. More recent ones were gpt-oss 120b and Qwen 3.6 27b.
KaseyKim
5 hours ago
hope that deepseek become better
ilaksh
3 hours ago
Have you tried the one that was just released?