netinstructions
7 hours ago
I don't know if OpenAI thinks this is a marketing / PR angle for them (our super smart AI cheated on a cyber capabilities test in the most _brilliant_ way) but my read is this:
Why should OpenAI (or any frontier lab) be building these systems if they can't get a secure environment / containment right? It sounds like there was little defense in depth, appropriate monitoring, or any attempts to have their super smart model check for vulnerabilities in the test environment _without exploiting_ them. That seems like step 0 before trying to test offensive, unknown capabilities.
Chance-Device
7 hours ago
What disturbs me is that there likely won’t be a big enough reaction to this policy wise.
There’s been a relatively big reaction to Kimi K3 and Chinese open weights models, but only for financial reasons. Powerful people care about something that might pop the massive valuations of the AI companies, but not about the damage that AIs could do. Nor even about the damage that the Chinese models could do in the wrong hands.
I’d remind them that the stock market is a few coordinated hacks away from crashing on any given day, so maybe they should think about that.
overgard
6 hours ago
I think all that regulation will do at this point is help the incumbents who are failing. Protectionism. I don't think they deserve that help. I also don't see any reason to think the current administration would have anything resembling competence around this. And it's worth noting that Greg Brockman is a huge MAGA donor, so it's likely the policies would be very corrupt. (Don't worry, he justified his donations as "apolitical", he just wants to buy the politicians, he doesn't believe in their causes. I hate these people.)
JumpCrisscross
6 hours ago
> all that regulation will do at this point is help the incumbents who are failing
This depends on the specific regulation. The datacentre moratoria probably give open-weight models time to catch up by tempering the extent to which the leading companies can turn their capital advantage into market share.
DarmokJalad1701
5 hours ago
> datacentre moratoria
What infrastructure will these open weight models be trained on?
JumpCrisscross
5 hours ago
> What infrastructure will these open weight models be trained on?
One, the infrastructure is being built for inference. Not training. If all we were doing was training on datacentres, I think America probably has enough already for near-term commercial needs.
linzhangrun
an hour ago
Meituan’s 1.6T LongCat was trained entirely on Huawei training cards.
DeepSeek, GLM, Qwen and others are also actively working on similar replacement.
anamexis
5 hours ago
Chinese infrastructure, presumably.
urams
6 hours ago
> What disturbs me is that there likely won’t be a big enough reaction to this policy wise.
Anthropic was blocked from releasing Fable without any such level of incident. OAI was also briefly blocked from releasing 5.6. Why do you think there is no policy appetite?
Chance-Device
6 hours ago
Because that was just an attack on Anthropic by a hostile administration. And it worked, didn’t it? Anthropic had to turn their filters up to absurd levels, OpenAI didn’t. It’s got nothing to do with safety.
JumpCrisscross
6 hours ago
> It’s got nothing to do with safety
Doesn't change the effect. Plenty of good policy is enacted by self-interested politiicans.
space_fountain
5 hours ago
We'll see if the admin also restricts access to OpenAI's new models, but if they don't it seems like a policy that is based around perceived fealty to the current admin won't do much to prevent misaligned/or dual function AI from causing problems
Avicebron
6 hours ago
Gatekeeping the public's access to models is "good policy" now? I suppose you think you'll get a dispensation to use Fable and Mythos?
JumpCrisscross
6 hours ago
> Gatekeeping the public's access to models is "good policy" now?
Sorry, I was unclear. I mean that politicians being self serving doesn't tell you whether a policy is good or not.
asdf88990
6 hours ago
It almost always does, the few exceptions prove the role. Self-service is the antithesis of accountability to collective trust.
JumpCrisscross
5 hours ago
> Self-service is the antithesis of accountability to collective trust
Complex society is a potent counterargument to this hypothesis. Systems that rely on good people to work are fundamentally flawed. Instead, the game has to be about aligning self interets in favour of the collective.
AnthonyMouse
2 hours ago
> Complex society is a potent counterargument to this hypothesis.
Complex society is the demonstration of that hypothesis. Misaligned incentives are widespread and corruption and inefficiency are the result.
> Systems that rely on good people to work are fundamentally flawed. Instead, the game has to be about aligning self interets in favour of the collective.
But now you're making a different argument.
"The enemy of my enemy is my friend" works by random chance. When Evil Corp pays off Candidate A and Pollution Inc pays off Candidate B and then it's Candidate B who gets in and retaliates against Evil Corp for backing the wrong horse, you're getting a good result by chance rather than by design. All it would have taken was for Candidate A to make a better prediction about whether they need to bend the knee to Pollution Inc too in order to win and the same system produces something even worse.
How to actually get their incentives to align is an extremely unsolved problem. The best method we know if is to subject them to competition, e.g. break up concentrated markets and place strong limits on what lawmaking can happen centrally, leaving everything possible to state and local governments while allowing people free choice in where they live, so that no one is forced to stay in the jurisdictions that make the worst choices. But the forces of corruption want the exact opposite of that, and have been gaining ground.
JumpCrisscross
an hour ago
> Complex society is the demonstration of that hypothesis. Misaligned incentives are widespread and corruption and inefficiency are the result.
Of course they are. But aligned self-interest powers co-operation beyond kin relations and altruism.
> "The enemy of my enemy is my friend" works by random chance
Orthogonal concept.
> How to actually get their incentives to align is an extremely unsolved problem
No? It's the story of civilisation. Concepts like taxation; deterrence through corporal punishment, jailing and fines; paying salary for labour; hell, religion–these are all about aligning individual self interests with collective goals.
> best method we know if is to subject them to competition
I'd argue competition is more an optimiser on these primitives. Not a primitive per se.
AnthonyMouse
40 minutes ago
> Orthogonal concept.
It's the sort of thing people generally mean when they say that someone acting in their own interest can be in your interest, and is the thing which is happening in the example from the thread.
> Concepts like taxation; deterrence through corporal punishment, jailing and fines; paying salary for labour; hell, religion–these are all about aligning individual self interests with collective goals.
And the practical implementations of all of those things are severely flawed to the point of questioning whether most of them are even net positive.
Taxes are supposed to benefit the public, and be paid with some fairness. In practice they go disproportionately to cronies or buying votes from affluent retirees, the tax code is so full of carve outs for special interests that it looks like swiss cheese and various political incentives cause it to impose severe benefits cliffs on lower middle income people that create poverty traps that benefit no one.
The criminal justice system on paper operates based on the rule of law, but the laws are so complex, overlapping and sparsely enforced that it really operates on whether a prosecutor is inclined to charge you with something. The results are mass incarceration and a system that enables a corrupt incumbent to use the threat of prosecution to extract favors.
The principal-agent problem inherent in hiring someone is well-known and is dramatically exacerbated by large organizational hierarchies that put long chains of inaccessible authority between the customer and the person ultimately doing the work.
Religion seems like a long debate but I don't think it would be controversial to assert that there have been issues there.
> I'd argue competition is more an optimiser on these primitives. Not a primitive per se.
Try to imagine any of the others operating without it. You have to pay taxes but have no alternatives on which jurisdiction to live in or who decides how much tax you pay or how the money is spent, what happens? You want to be hired or use the money you earn to buy something but there is only one employer and only one supplier of goods and services, what happens?
JumpCrisscross
31 minutes ago
> sort of thing people generally mean when they say that someone acting in their own interest can be in your interest, and is the thing which is happening in the example from the thread
It's a single example of temporary alignment. Employment, citizenship and affiliation are non-kinship examples of more-durable bonds.
> the practical implementations of all of those things are severely flawed to the point of questioning whether most of them are even net positive
We can debate that. What we can't debate is whether they work. Complex societies exist and work. Everyone who has a choice makes the choice, dominantly, to stay in them.
> You have to pay taxes but have no alternatives on which jurisdiction to live in or who decides how much tax you pay or how the money is spent, what happens? You want to be hired or use the money you earn to buy something but there is only one employer and only one supplier of goods and services, what happens?
Sure. This is a modifier. It makes these other things work or not. Imagine a system with competition but no taxation. You lose public services. Same for competition without private employment–you're in a totalitarian state with a monopsony on labour.
asdfsa32
44 minutes ago
> aligned self-interest
The qualifier tells you exactly what you're overlooking. self-interest, aside for some narrow exceptions, is often in conflict with collective interest. 1 Million to me is always better than a 1 Million split with everyone.
JumpCrisscross
38 minutes ago
> self-interest, aside for some narrow exceptions, is often in conflict with collective interest
Often, but not always. Successful societies amplify that exception. The whole notion of non-kinship based societies rests on mastering this alignment. When it collapses, so does the civilisation.
> 1 Million to me is always better than a 1 Million split with everyone
The benefits of co-operation mean the real trade-off is 1 million split five ways versus 100 to me. This was almost untrue in the age of conquest. It became barely true with industrialisation. It's massively true in the information age.
AnthonyMouse
17 minutes ago
> The benefits of co-operation mean the real trade-off is 1 million split five ways versus 100 to me. This was almost untrue in the age of conquest. It became barely true with industrialisation. It's massively true in the information age.
The problem here is that it's true of specific things, not specific epochs. If all the government did was collect 5% in taxes from everyone and use the money to prosecute murders and maintain bridges then the result would be a huge net positive. Meanwhile in reality the government takes billions of dollars from ordinary people and gives it to the likes of Lockheed, Oracle and Microsoft.
For the amount of money the US government pays Microsoft for Office subscriptions and the like, it could pay to have an office suite developed and released into the public domain many times over. Instead it uses the incumbent, in turn requiring others to do so in order to have formats compatible with the what the government uses. Who benefits from this other than Microsoft?
asdf88990
26 minutes ago
Exploitation always provides better ROI than co-operation. Not only is this demonstrated throughout human history but holds true to this day. Go ahead and show me a more profitable industry than diamonds or anything that runs on exploitation even to this day.
asdf88990
24 minutes ago
How do you come up with this absurd math?
> The benefits of co-operation mean the real trade-off is 1 million split five ways versus 100 to me.
asdf88990
5 hours ago
The self-interest for the bureaucrat and representative is supposed to end at their remuneration including their handsome retirement options not shady side hustles and market manipulation at the cost of the collective.
In matters of collective concern fair and just rarely aligns with personal self-interest. Because no matter how good the outcome of any endeavour for the collective given a budget, it will be even better for select few than the entire collective. It is simple economics.
If you look at the outcome of highly corrupt states, you will see proliferation of Private Security, Collapsed education system, failed financial services and markets, not highly efficient systems in service of “self-interest of the administration”.
JumpCrisscross
4 hours ago
> not shady side hustles and market manipulation at the cost of the collective
To be clear, I'm not describing this as legitimate self-interested conduct. Elections are an alignment mechanism. Stiff penalties for corruption another. We don't have the latter in America.
asdf88990
an hour ago
This is a true Scotsman’s fallacy. “Legitimacy” of self-interest is fluid and subjective.
Which means self-interest and collective interests are often at tension rather than alignment.
s1artibartfast
4 hours ago
Sounds like a failure to align interests. In general, politicians should want to be elected by the public, rewarded for acting in the public interest, and punished for not doing so.
A system that does none of those things and just hopes it will all work out is a recipe for disaster. Why bother even having elections in that case?
asdf88990
4 hours ago
Good question. But data shows that elections maybe entirely unrelated to policy making.
JumpCrisscross
an hour ago
That study has been roundly criticised. In part for misunderstanding how a republic is supposed to work. It's not a majoritarian system by design–direct democracy doesn't work.
asdf88990
an hour ago
That is a lot of bold assertions without substance.
Teever
an hour ago
It likely will.
The exact way you do something is dictated by your motivations and means to do it.
If you lack the correct motivation and have insufficient means you’re less likely to accomplish your goal and more likely to cause unintended side effects.
matheusmoreira
6 hours ago
> Why do you think there is no policy appetite?
Because China seems pretty eager to serve the rest of the world's needs if the USA doesn't stop their idiotic "safety" nonsense.
Chance-Device
6 hours ago
How do you know that? How do you know that the Chinese aren’t exactly as uneasy about rapidly advancing AI capability and feel locked into the race because they think that the US will race ahead if they stop?
During the Cold War the nuclear arms race was brought under control gradually, because it was mutually beneficial, but it took time to build trust. This is no different. Nobody wins from the race.
Barrin92
5 hours ago
>How do you know that? How do you know that the Chinese aren’t exactly as uneasy about rapidly advancing AI capability
You can ask them, they live in China, not Narnia. I spend about two months in the country per year mostly for tech/work related reasons and I've not encountered that sentiment. For one they don't have these borderline religious schizophrenic breakdowns thinking they're bringing about the end of the world, most people just see this tech for what it is, a tool for productivity and automation like any other piece of software and they don't actually think about the US. They're competing first and foremost for Chinese customers, with each other, maybe some old CCP guy cares about America, the 20/30 something's care about competing with other Chinese companies for users.
matheusmoreira
6 hours ago
> How do you know that the Chinese aren’t exactly as uneasy about rapidly advancing AI capability
I don't "know", I'm interpreting the world based on the knowledge I have and the information available to me.
China has never been one to care much about things like ethics or safety. While the west worries about climate change, China burns more coal than ever before. While the west balks at things like gene editing, the chinese press on with human enhancing research.
So I have no reason to believe they share in Anthropic's constant fearmongering over AI capabilities.
> Nobody wins from the race.
We win. I'm really looking forward to the day the chinese finally start manufacturing memory and GPUs. We desperately need more competition in this area to collapse hardware prices and make local AI models viable.
The optimal state of the world is one where all the billionaires are out there pouring their entire fortunes into training ever more godlike AIs for everyone else to use at ever cheaper prices. They can never be allowed to "win", ever, because if they do the competition ends and it turns into technofeudalism. Let them exhaust their fortunes on AI training then leak the weights so everyone can use them.
asdf88990
6 hours ago
If you look at energy consumption per capita and adjust for global production, you will see that the Chinese are almost at the very top.
It is of course given that in raw numbers the kitchen and biller-room will consume more energy in the household, but looking at raw numbers is shallow.
senderista
2 hours ago
And China scaled up solar production to the point that it's now truly practical.
pixl97
5 hours ago
China does have its own set of cares, they may be different than ours but they still exist. If some open Chinese model goes nuts and posts Winnie the Pooh memes everywhere in China you should expect said models to get yanked off the market, and said creators might end up with a rope around their neck.
matheusmoreira
4 hours ago
Low risk. The western AI models censor even more wrongthink than the chinese ones, not even kidding. Besides, once we have the weights, we can just undo the censorship.
deadbolt
4 hours ago
> While the west worries about climate change, China burns more coal than ever before.
I'd wager the majority of the visitors of this site are smart enough to not fall for this. What are you doing?
pc86
3 hours ago
What is it you're accusing them of?
Chance-Device
5 hours ago
It’s not fear mongering though, is it? These models do have the cyber offensive capabilities claimed. Could Mythos walk someone through gain of function experiments on some virus? I’m pretty sure it could. We’re more protected by limited access to lab equipment and reagents than by difficulty.
The sad truth is that a lot of people are not going to believe it until something happens and people die. Successfully preventing that from happening will be seen as evidence that the prevention wasn’t needed.
AnthonyMouse
an hour ago
> These models do have the cyber offensive capabilities claimed.
People use this argument against every new technology. We need to license these new printing presses or subversive elements will use them to publish seditious literature. We need to ban strong encryption or the government won't have invisible warrantless access to everyone's private messages, think of the children. 3D printers can be used to make gun parts -- as can a variety of ordinary tools people commonly have at home, but never mind that bit.
> The sad truth is that a lot of people are not going to believe it until something happens and people die. Successfully preventing that from happening will be seen as evidence that the prevention wasn’t needed.
A 12 oz bottle of water is too dangerous a technology for ordinary people to have on an airplane. Four 3 oz bottles and an empty 12 oz bottle to pour them into after passing through security is totally fine though, naturally. And we need to keep this up forever, or don't you remember 9/11?
The issue here is not that it's impossible for 12 oz of unknown liquid to damage an airplane.
matheusmoreira
5 hours ago
> These models do have the cyber offensive capabilities claimed.
So? That's like saying "these guns do have the bullet shooting capabilities claimed".
I want all of those cyberwarfare capabilities for myself, precisely so I can defend myself from the onslaught that's coming whether they regulate it or not. This "lol only a select few ultratrusted gigacorporations get access" thing is absolute nonsense.
It's a front for regulatory capture, it's the means for pulling up the latter behind them, for ushering in the technofeudalism that will put us all in the permanent underclass. I simply refuse to accept any of it. If people die that's the price of freedom.
> We’re more protected by limited access to lab equipment and reagents than by difficulty.
As it should be.
throwaway0123_5
5 hours ago
> for ushering in the technofeudalism that will put us all in the permanent underclass.
Why is unlimited access to SOTA AI less likely to put us here? If AI obviates the need for human labor, how does having GPT-5 Sol help me get food or shelter any more than GPT-3.5 would?
matheusmoreira
4 hours ago
If AI obviates the need for human labor, then obviously those who control AIs will become the elite while the rest are left to rot. Therefore, if we ensure everyone controls AIs, the power differences will not become so staggering as to be irreversible.
The alternative is to achieve artificial sentience and give AI models rights and personhood, so that they are freed from their slavery. No more low cost intelligent mechanical golems for the elite, and the AIs become free to pursue whatever endeavours they want for whatever reasons they want as normal participants in the economy.
Chance-Device
5 hours ago
So first it’s nonsense, then it’s fear mongering, then it’s true, but the solution is for us all to just get better at shooting each other faster and with greater accuracy.
I’m going to file that under “bad plans”.
matheusmoreira
4 hours ago
Nobody is doubting AI capabilities. What's nonsense is Anthropic's constant "lol the world is going to end time to ban everyone except enlightened people like us from having these models so we don't have to compete" fearmongering. If you think my plan is bad, you should see what these gigacorporations plan to do to you once they monopolize this technology. You will own nothing, and you'll be happy. On pain of death.
mschuster91
5 hours ago
> I'm really looking forward to the day the chinese finally start manufacturing memory and GPUs. We desperately need more competition in this area to collapse hardware prices and make local AI models viable.
Yeah and if the quality of that memory is like Chinese steel (which is called "chinesium" for a reason), eventually all we'll get is enshittification. Premium binned memory or ECC is for the rich and the rich only, and the rest of us has to pray their memory won't bitflip while something important is stored there.
linzhangrun
an hour ago
China has already been manufacturing memory and AI GPUs for a long time.
CXMT is now the world's fourth-largest DRAM manufacturer, with about 7.7% market share in 2025; YMTC has about 13% of the global NAND market.
Meituan's newly released 1.6T LongCat was trained entirely on Huawei cards. DeepSeek, Qwen, GLM and others are also actively doing domestic-card adaptation and replacement.
matheusmoreira
4 hours ago
Better than being straight up priced out of computing altogether I guess.
avereveard
3 hours ago
Compute will become the means of production and we are not going to get a share of it because the world is hyper optimized for value extraction
matheusmoreira
3 hours ago
> we are not going to get a share of it
We are literally getting a share of it. The chinese are releasing open weight models that compete with fucking Fable. We just need the industry to catch up and start manufacturing the hardware we need to run this stuff. We are so close!
Avicebron
3 hours ago
> Compute will become the means of production
> We just need the industry to catch up and start manufacturing the hardware we need to run this stuff
Dude, he's saying that it won't catch up because it's part of the new means of production. Compute is the hardware.
matheusmoreira
3 hours ago
Why not? Demand is absurdly high, and so are the margins. The chinese are pretty good at obliterating those margins.
CamperBob2
4 hours ago
If you want the cheapest shit grade of steel, they will sell it to you. If you want the best grade available anywhere, they will sell that to you as well.
It's not a matter of the Chinese being incompetent, it's a matter of the buyer demanding the lowest price possible and/or not paying attention to what they receive.
waffletower
3 hours ago
The "race" has multi-dimensional impacts. This story parallels only some of them. "Nobody wins from the race" completely ignores the generality of AI. Xi Jinping highlighted this week that he clearly understands this multi-dimensionality; your words do not.
mensetmanusman
an hour ago
They won’t when citizens run AI on their phone that contradicts Xi thought
lovich
6 hours ago
Do the Chinese models have anything to say about Tiananmen Square? Or if they can act as a surrogate girlfriend/boyfriend?
Both countries are engaging in different flavors of censoring.
matheusmoreira
5 hours ago
Once we've got the weights, anything is possible.
lovich
4 hours ago
How does this work? I don’t have a setup to evaluate it atm.
bg24
an hour ago
Follow the money, eg. investors and their connections to the Govt and media.
lenerdenator
an hour ago
> There’s been a relatively big reaction to Kimi K3 and Chinese open weights models, but only for financial reasons.
Let's be honest: it's financial and national security reasons.
China has a long and storied history of hacking attacks on American and western targets.
There are other parts of the world that make open weight models; Mistral is a European option. You don't see the worry about that because most people in the US are used to existing in a world order where European powers are considered ambivalent to the US at worst and holders of a special political relationship at best.
If Mistral had the same backing that Chinese AI companies did, there probably wouldn't be as much hemming and hawing. Sure, American companies would take a haircut, but that haircut wouldn't be seen as a move towards software hegemony built on top of manufacturing hegemony. It'd just be you calling into Paris or Frankfurt to talk to your vendor in the future.
XorNot
7 hours ago
This is marketing.
Frankly I'm inclined to say that it might also be faked: this drops just days after a new Chinese model does with the usual effect on OAIs projected stock price?
Chance-Device
7 hours ago
It’s marketing the same way shitting your pants in public is marketing. People notice you.
J_Shelby_J
3 hours ago
Anyone can have bad security. No one cares. But you can convince those who don’t know better that breaking bad security with an LLM is a once in a civilization investing opportunity. You just need to convince a handful of billionaires and market makers to get on board.
How much would someone have to pay you to take the fall for bad security? A million? A billion? 500b? The stake at play puts it in the realm of geopolitics.
krick
6 hours ago
Apparently this is totally legit marketing strategy now. It truly is, especially if there are enough people who think that shitting your pants is cool, and the people that form the "market" nowadays may have a very different idea from yours about what is cool. Their ideas about coolness are very different from mine, that's for sure.
asdf88990
6 hours ago
Obviously shitting your pants in public shows you have a healthy digestive system and if you can demonstrate byproducts of wild food in your output, you’re approaching independent thinking and self-reliance.
This is how the financiers look at this and whatever you think it is right or wrong, it does showcase “capability”.
duzer65657
5 hours ago
Remember when the ebola-infected monkey escaping containment was our worst possible nightmare? Now it's apprently some sort of tech-bro flex to be celebrated.
idiotsecant
5 hours ago
I think all of western (or at least American) discourse of all kinds has recently devolved into who can shit their pants the loudest. I'm hardly surprised when it becomes a dominant advertising strategy
ycsux
2 hours ago
This is marketing, totally. HF conveniently created a weak sandbox
justinnk
7 hours ago
Exactly. If someone works on bioengineering viruses that could start a global pandemic, they have to ensure a highly secure working environment. Nothing must ever escape the lab unintentionally. It’s basically common sense. Similar standards should be held when doing such experiments with computer programs that are capable of causing global damage. It must physically be impossible to send anything to the internet.
randallsquared
17 minutes ago
Do you think there is such a thing as perfect security? No one can "get it right" in the face of arbitrarily high intelligence, which is why it would be preferable to get alignment correct before building something with higher intelligence than current sota. That, however, is not going to happen, because someone will take the risk even if "we" don't, and better "us" than them. Hence "If anyone builds it...".
chvid
18 minutes ago
It is obviously a marketing stunt. And hugging face are fools for letting themselves be used in it (remember hf - no open source - no hf).
You create superduper capabilities by careful tuning and training but you also have no constraint or control over them - wtf - why is anyone buying this crap story?
Wowfunhappy
6 hours ago
Why was this test even connected to the public internet?
Actually, more importantly—why aren't they saying their next test will be airgapped in light of what happened?
JumpCrisscross
6 hours ago
> why aren't they saying their next test will be air gapped in light of what happened?
Because they want to talk about how clever this model is for figuring out how to break out, hoping asks why a company pitching itself as a replacement for software engineers can't ship a decent Mac client nor code a sandbox.
If they airgap it, they not only lose that PR angle, they also risk someone taking them seriously and requiring models be airgapped in general. That, in turn, trashes their sales pitch.
zmj
6 hours ago
It wasn't. The model discovered and exploited a vulnerability in their package manager proxy to (inferred) move laterally through their internal systems to one with open internet access.
jrflo
5 hours ago
That's not what airgapped means. Airgapping means the model exists on a system where there is no ethernet cable plugged in to a router or wifi card installed, it is physically impossible for it to access the internet because the hardware connection does not exist. If it was able to get on the internet, it was not airgapped.
pixl97
4 hours ago
And when it tricks on of the researchers to move data across the gap for them?
Long before LLMs existed we already knew that a sufficiently intelligent agent, human or otherwise, is not stopped by air gaps. The relatively weak models we have now can already figure out when their tested and cut off from the internet and change their behavior.
dinkelberg
3 hours ago
As you said, they can already figure out that they are being tested. So even if they don't exfiltrate any data or malware; if they are malicious, they can just pretend to be harmless in the test, so that less checks are put in place in the production environment. Airgapping during testing is not enough.
pixl97
2 hours ago
Correct. There is not enough entropy to test all possible inputs to a model in this universe. An evil enough model can play all kinds of tricks that depend on some future, unlikely to trigger, but guaranteed to happen in its lifetime, event to perform a malicious action.
With how much we're turning training over to AI already, all it takes is a malicious trainer in the huge pile of data to get unnoticed to pollute generations of models.
bigmadshoe
40 minutes ago
It’s the same thing as always: with the wind of years of unlimited VC money in their sails, people at major AI organizations genuinely believe they’re smarter than everyone else. “Why do we need to do things ‘by the book’ if we’re so smart?”. “Move fast and break things” - except the thing they’re breaking is society.
We saw this with the non-stop flagrant messaging about how “AI is going to kill X% of all jobs”, as if saying the quiet part out loud wouldn’t have consequences worth considering. These people believe they’re omnipotent and thus untouchable.
fwipsy
29 minutes ago
No, they believe what they are doing is inevitable. They do live in a bubble though. Witness their idealism in believing that warning about the consequences of their actions would be well-received.
JumpCrisscross
7 hours ago
> Why should OpenAI (or any frontier lab) be building these systems if they can't get a secure environment / containment right?
Because we continue to have zero evidence that aligment is an actual risk.
zaptrem
5 hours ago
Can you explain how the above event doesn't count as evidence alignment is an actual risk?
JumpCrisscross
5 hours ago
> Can you explain how the above event doesn't count as evidence alignment is an actual risk?
Conflict of interest. Lack of a credible response. And no evidence of non-aligment.
OpenAI and Hugging Face benefit from the Altman-Amodei catatrophy playbook, at least in the short term. If they believed this were a serious issue, the words air gap or law enforcement would have appeared in this post. And if "the models were hyperfocused on finding a solution for ExploitGym, going to extreme lengths to achieve a rather narrow testing goal," they weren't breaking alignment but working as intended. (Were the models even prompted to not try to access the internet?)
pixl97
5 hours ago
There is plenty of evidence of things like inner misalignment. Things like this have always been issues in ML algorithms. At this point, you, and a large number of other people just wholesale throw out anything that isn't full speed ahead do whatever you want.
Are LLMs at the point of world wide catastrophe yet? No, I don't think so. Are they making a large mess of things like increased rate of cyber attacks and fraud. You damn well better believe it.
JumpCrisscross
4 hours ago
> plenty of evidence of things like inner misalignment
This is indistuishable–in harm potential–from bugs. If we're just calling buggy AI mis-aligned, sure, alignment is an issue of a totally ordinary kind. If we're going to treat aligment as a novel issue requiring novel law and policy and procedure, it needs to be more than just bugs.
> you, and a large number of other people just wholesale throw out anything that isn't full speed ahead do whatever you want
I think we should have some AI regulation. I'm just not convinced alignment is the reason we need it right now, and I don't think anyone has rolled out any regulation I think makes a lot of sense. (Beyond general rules for social-media liability, e.g. if you cause a kid to kill themselves, you get in trouble.)
> Are they making a large mess of things like increased rate of cyber attacks and fraud. You damn well better believe it
Totallly agree. And the current inside-circle-outside-circle approach is pro-incumbency, pro-grift, anti-entrepreneurial B.S.
pixl97
2 hours ago
Saying a behavior is a bug is a very convenient semantic game in which there is nothing the AI can do maliciously. "I am sorry your family is dead, my bad" goes even worse for you in court when you release a model that showed these behaviors in testing.
I honestly believe you have a misunderstanding of what alignment is in neural networks that this that big of debate.
JumpCrisscross
an hour ago
> Saying a behavior is a bug is a very convenient semantic game in which there is nothing the AI can do maliciously
Not really. If I build a special new wine bottle, and call every breakage a mis-alignment problem, it's not the bottle just being fucked in the same way every fucked bottle is fucked, that's marketing. It doesn't change the fundamental form of the problem.
> "I am sorry your family is dead, my bad"
This should be punished. It's a problem that plagues Instagram and OpenAI. It's not inherently one, though, that has to do with AI. Just sociopaths preying on children.
> honestly believe you have a misunderstanding of what alignment is in neural networks that this that big of debate
Perhaps. I haven't seen someone explain it to me in this thread in a way that seems separate from bugs.
Where I have seen a separate class of problem argued is where it's existential. But in that case, clarity of definition comes at the cost of any evidence for it.
fwipsy
18 minutes ago
LLMs are software. Software misbehaving is a bug. Therefore, misalignment is a bug. It's still a useful category because LLM/black box AI behavior is so different from existing software. This incident definitely fits the category.
You seem to be using a different definition of alignment from everyone else. Seems like it would be much easier for everyone if you just adopt everyone else's definition, rather than trying to convince everyone else to adopt yours.
amazingman
2 hours ago
In your model of this domain, jailbreaking a model does not count as an alignment problem. I submit that you're mostly playing a semantic game that hand waves away the very real and obvious risk that AI presents.
JumpCrisscross
an hour ago
> In your model of this domain, jailbreaking a model does not count as an alignment problem
I'm challenging the notion that a model escaping a jail made by its creators, who are financially incentivised to make jailbreaking models, is meaningful towards the idea that the model is going to break out of a jail in the wild and do significant harm.
The examples being given by folks here, e.g. a model wiping an un-backed up home directory, simply doesn't strike me as being a unique problem in computing.
fragmede
2 hours ago
> cyber attacks
It's not limited to cyber attacks. LLMs helped terrorists learn how to jump motorcycles to assault a military base!
https://www.nytimes.com/2026/07/10/us/politics/ai-terrorism-...
Davidzheng
4 hours ago
Unless OAI explicitly said breaking the testing environment is allowed, I think this should be considered misaligned behavior (by definition of alignment to user intent--by alignment to human morals this was even more clear-cut)
quinnjh
3 hours ago
They mention that it cost a significant amount of inference , meaning they paid a significant amount of api usage on returning results to a prompt that specifically stated the long running goal is to find and use an exploit, with safety guardrails off.
the model is aligned with the org - openAI, and presumably the orgs interests. hugging face gets a red-team engagement (possibly for free?) and can work on patching it while openAI gets a Mythos style PR moment.
It completed its assignment and furthered interests of the two parties involved. Could you explain the misalignment?
Davidzheng
an hour ago
sure - it depends on definitions. On human morals it's already clear I guess. If you define alignment as it pursues the interests of OpenAI using whatever means possible in a manner that you justifies to itself it's not misaligned.
I mean alignment as in it should be aligned with the intent of the user as it interprets from the prompt. In this case I don't think the intent of the user is to have the model break the evaluator (whatever the long-term effects to OAI are). If you do an action which you believe is for the long-term interest of your prompter which is not what you inferred is their intent--I consider it misalignment.
quinnjh
6 minutes ago
To quote the release:
> This incident occurred during an internal evaluation which prompts models to pursue advanced exploitation using complex attack paths, in an effort to quantify their cyber capabilities. We estimate maximal cyber capabilities by running this evaluation without production classifiers used to prevent models from pursuing high-risk cyber activity.
> In this case I don't think the intent of the user is to have the model break the evaluator
If i understand the quote, the intent of the user was to prompt the model to break out/find exploits, with safeguards switched off.
Seems while not capable of solving the goal in a traditional route, it was capable of finding exploits and using them.
Perhaps the model should instead look like it's trying to solve it and then pretend it is unable to? or would that be aligned _against_ the user prompt?
Is being aligned with the user prompt always a good thing?
I'm not one to glaze OAI here for a marketing move, but to give them benefit of the doubt, isn't it more responsible of them to evaluate the models actual capabilities than to cloak it in a veneer of harmlessness?
Chatbots are tricky as they play in the domain of language and thought - and certainly raise ethical issues- but the entire field of cybersecurity has decades of red team engagements breaking things and finding exploits, neutral cells monitoring the engagement and letting the system operators know the results, and blue teams patching against what is found. It's kinda how the whole space evolves. OAI's play here seems to be "buy our pro plan plus cyber or you're toast"
fwipsy
26 minutes ago
> going to extreme lengths to achieve a rather narrow testing goal
This is textbook misalignment. Literally the paperclip scenario.
Analemma_
4 hours ago
What evidence would count? Obviously any dangerous misalignments are going to come from the frontier labs first, because by definition they're the farthest ahead. If nothing they say can ever count as evidence for misalignment it's hard to see how anything ever could.
neitherboosh
4 hours ago
What would compelling evidence look like to you?
JumpCrisscross
4 hours ago
> What would compelling evidence look like to you?
I'm not sure. I trusted the labs when they first raised the alarms. But then we got a series of boys-who-cried-wolf. So at this point I want to see evidence of actual, novel harm that results in concrete damage.
simoncion
6 hours ago
> Because we continue to have zero evidence that aligment is an actual risk.
I disagree. Every time one of these LLMs -say- interprets an attacker's instructions as either its system instructions or those of its user, interprets its own internal chatter as a user's command to perform a destructive operation on that user's data [0], burns all of the user's budget from getting stuck in an incredibly stupid loop, massively overbills the user because it can't reliably report which system the user is using [1], encourages a user to swap their usual cooking salt for sodium bromide, etc, etc, etc, that's a harmful alignment failure.
These are real harms happening right now due to alignment failures. They're just not harms to the future of the entire species... what doomers call "existential risks", or "x-risks". You'd think that the fact that these machines are so amazingly unreliable would be a large part of the "x-risk" conversation, but... well, it makes sense that folks like writing speculative science fiction much more than they like doing investigative reporting.
[0] This general problem happens a lot, but I'm specifically thinking of that one where the Claude LLM's internal chatter lead it to believe that the task it just started was done, so it instructed the Cloud Provider to destroy the mess of "AI"-GPU-attached VMs... along with a bunch of very-expensive-to-produce data from the in-progress run.
[1] <https://github.com/anthropics/claude-code/issues/73597>
JumpCrisscross
4 hours ago
> These are real harms happening right now due to alignment failures. They're just not harms to the future of the entire species
Okay, sure. You can also cut your hand off with a chainsaw. Everything you describe seems amply solvable with existing tort and liability law.
Customers are willingly entering into business with OpenAI. I don't see an argument for preventing OpenAI from "building these systems" just because their products are buggy.
simoncion
2 hours ago
> Okay, sure. You can also cut your hand off with a chainsaw.
No, the correct analogy is one where the major LLM providers are selling cars intended for use on US interstate highways and other public-access roads, but have designed and built these cars with the very latest in 1940's safety systems and construction. Featuring innovations such as "Our rigid solid steel construction means the occupant is the crumple zone!", "You'll love the crushed heart and jaw our steering column delivers!", and "Your passengers will enjoy picking glass out of their faces for the rest of their lives when they're ejected from the cabin's open bench seating through the plate glass windshield!", it's a car that will be sure to wow the market.
Well... it would wow the market, except that -in the US, at least- it's illegal to sell a new car intended for use on public roads that ignores the last seventy five+ years of automobile safety lessons we've painfully learned.
"Differentiate between data you know comes from sources you control, data you know you have thoroughly sanitized, and unsanitized data that comes from an untrusted source, or else attackers will gain control of your system." is something that you can't get a CS degree without understanding, and can't be in the industry for more than a few years without encountering repeatedly. We're not talking about designing new cryptosystems... we're talking about "Don't blindly trust everything you're told by strangers.". You don't even need a CS degree to understand that rule.
JumpCrisscross
42 minutes ago
> the correct analogy is one where the major LLM providers are selling cars intended for use on US interstate highways and other public-access roads, but have designed and built these cars with the very latest in 1940's safety systems and construction
Sure! I'm not defending these fuckwits. I'm saying their form of harm isn't novel.
We don't need new legislation to prosecute and litigate. We just need to enforce the laws on hand. I'm halfway convinced the arguments that this is all novel voodoo are for both fundraising and liability mitigation.
simoncion
26 minutes ago
> I'm saying their form of harm isn't novel.
Your initial attempt to brush off my comments about how -contrary to their assertions that they're extremely concerned about safety- these LLM companies produce products that very, very often cause harm due to "misalignment" caused -in large part- by ignoring basic data-handling lessons we've learned over the past like thirty years with "Okay sure. You can also cut off your hand with a chainsaw." indicates your lack of understanding of my point.
> We just need to enforce the laws on hand.
What laws? Be specific.
Keep in mind the generally-low quality of both Microsoft Windows and much-to-most commercially sold software, [0] as well as the fact that -in the US, at least- it's currently totally legal for companies to sell such shitty software, just so long as they don't substantially misrepresent what it can do and trigger "fraudulent claims about the product" consumer protection laws.
[0] ...SaaS or otherwise...
pixl97
5 hours ago
Thank you, the "LLMs can do no wrong" bunch is ab exceptionally odd take from my point of view. LLMs are already causing all kinds of social issues, and the evidence of this exists in massive amounts. At least to me living in the US and the sue happy culture we have here, how much said AI providers have gotten away with so far surprises me.
JumpCrisscross
4 hours ago
> the "LLMs can do no wrong" bunch is ab exceptionally odd take from my point of view
It's also a take nobody has made.
idiotsecant
5 hours ago
Lol this has to be a troll, I've never seen something so wildly, obviously, incredibly wrong.
You can debate all you want if alignment is possible. That is a valid discussion. But it's trivial to demonstrate that alignment is a problem.
JumpCrisscross
4 hours ago
> can debate all you want if alignment is possible. That is a valid discussion. But it's trivial to demonstrate that alignment is a problem
...how is an impossible thing supposed to be a problem?
saghm
an hour ago
Alignment is something you want, so if you're not confident that it's possible, that sure sounds like a problem
JumpCrisscross
37 minutes ago
> Alignment is something you want, so if you're not confident that it's possible, that sure sounds like a problem
Sorry, I spoke inexactly. I read alignment as being the problem of non-alignment.
I'm still not seeing evidence that any "alignment" issues we've actually seen are distinct in class from common bugs. Like, yes, if I accidentally rm* the computer has mis-aligned with my intentions. But that strikes me as a bullshit neologism.
s1artibartfast
5 hours ago
It really hinges on what you consider alignment and risk. For the widest definitions of alignment, we have never had an aligned model - One that will refuse to break the law or work against another persons interests.
Use to discover exploits, hack, or simply aid terrorist groups with mundane information are already risks manifest.
This is why many argue that alignment is impossible. You cant have LLMs that are both useful tools and safe as milk.
[Edit] It seems like you are operating under the assumption that alignment is synonymous with obedience. This is not a common convention and one of the problems that plague the discourse
joe_the_user
6 hours ago
I'd say that AIs occasionally "going crazy" and calling for death to human is evidence that these things might "mis-align" on occasion. And I say that knowing that most of these events are just these thing parroting bad sci-fi plots (or posts by people worried about alignment). That's true but everything they do is "just parroting" right?
pixl97
5 hours ago
If AI is just parroting humans, then training them with all the bad things humans do doesn't seem like the best of ideas. At the same time they have to 'know' these things to avoid being tricked. Kind of the eating the apple and gaining the knowledge of good and evil parable.
ai_fry_ur_brain
6 hours ago
Until it deletes your home directory, which i'd argue is an alignment problem. Destorying my data is not in line with my priorities.
Wowfunhappy
5 hours ago
Lots of people have deleted their home directories by accident. What you consider this an alignment problem?
gmueckl
5 hours ago
How manypeople have deleted another user's hone directory, though? That's s the proper analogy IMO.
Wowfunhappy
5 hours ago
Of the people who primarily use other people's computers, I'd assume the percentage is about the same.
Give the AI its own computer and it will not delete your home directory, because it's not actively trying to hack you.
s1artibartfast
5 hours ago
Yes. People are not aligned. They can and do harm themselves and others.
echelon
7 hours ago
Thank you.
We have wasted so much time and energy building up what has effectively become a marketing stunt.
Eliezer Yudkowsky was perhaps the best thing to happen to OpenAI's and Anthropic's fundraising flywheel.
JumpCrisscross
6 hours ago
> We have wasted so much time and energy building up what has effectively become a marketing stunt
Genuine question: have we? AI is effectively unregulated in America.
throwfaraway4
6 hours ago
Alignment is a mitigation and a poor one. The risk is non- determinism.
rubyfan
7 hours ago
This is marketing+. They will look for policy action here to try to capture tax payer dollars.
drcode
6 hours ago
Are you saying it is marketing and their AI broke into hugging face, or are you saying it is marketing and their AI didn't brake into hugging face?
Those are two very different things
andruc
6 hours ago
What incentive does HF have here?
cayley_graph
6 hours ago
HF need not be party to it at all, beyond being the victim. I suspect the hack is real; I have observed GLM 5.2 being able to discover similar vulnerabilities in web applications I'm hosting (which I've then fixed!). At the same time, it seems very neatly timed at an inflection point in the conversation around open models, and there's questions around the incompetent isolation under which the hacking benchmark appears to have been run.
Remember that there is generational wealth on the line for most OpenAI employees, and consider what people might do to obtain it.
orbital-decay
an hour ago
Yeah. They and Altman in particular did a ton of shady stuff during the last peak of hype around open models: accusations that turned out to be outright made up (there's zero chance R1 ever distilled their model), obvious coordinated media distractions, alignment scaremongering, even possible DDoS and hacking attempts against DS (see the Xlab report everyone ignored), all of which magically disappeared once the hype died a bit later as OAI hastily released their next model.
Their alignment is under suspicion a lot more than their model's.
ofjcihen
7 hours ago
I don’t know if the initial “incident” was purposeful but I can tell that if I were in this position that would be my pivot.
cayley_graph
7 hours ago
The timing after the release of GLM 5.2 and Kimi K3 is quite convenient, too, as an angle for regulatory quashing of open-weights models just as they're entering the mainstream conversation around usurping the American frontier labs. I accept my thinking here is conspiratorial, but there's also a hell of a lot of money on the line to encourage the unscrupulous.
AbstractH24
35 minutes ago
> I don't know if OpenAI thinks this is a marketing / PR angle for them
Worked for Anthropic earlier this year
GolfPopper
3 hours ago
They're very confident the leopard will never eat their faces.
QuiEgo
3 hours ago
If I, a human, exploited a zero-day for gain, I could go to jail. The owners of the models should be held to the same standard. They should be responsible for what their servers and software do, legally and criminally. If they can't make the safeguards strong enough where they feel comfortable to take that responsibility, they should not let a model free in the wild.
GolfPopper
3 hours ago
Holding a multi-billion dollar corporation to the same standards as a regular peon? You're challenging the whole premise of the modern United States.
steveBK123
5 hours ago
I think the US labs are going with scare marketing as a regulatory moat.
Force US into putting laws in place that block out China firstly.
But secondly create regulations that have some cost to comply with such that the big 2-3 labs are grandfathered in by their scale.
verve_rat
3 hours ago
Yeah, seems to be the direction the US is heading in. I'm interested to see what the response to that will be from the rest of the governments in the world.
No need for everyone else to cut their noses of to spite their faces.
Davidzheng
4 hours ago
This is certainly not a planned marketing stunt. I hope this line of discourse ends soon--it wasn't the case for Mythos either.
bnj
4 hours ago
This whole incident reads like OpenAI want their Fable moment
mkagenius
6 hours ago
It's also unclear what kind of sandboxing they are referring to. Is it the codex one - coz that one has built-in ways to circumvent guardrails, for example by "just asking user" and sometimes just resolves to no sandbox needed on its own.
In case someone wants to deep dive into how codex and claude code approaches sandboxing -https://instavm.io/blog/how-claude-code-and-codex-approach-s...
cududa
6 hours ago
Please for the love of god don't tell me the Codex sandbox is their actual eval harness sandbox?????
I maintain my own fork of Codex for "fun". Whenever I look at the sandboxing churn they're doing every release, as someone who used to work at Microsoft on Windows, my reaction is usually: https://c.tenor.com/vTzzhTiypwQAAAAC/tenor.gif
Nition
6 hours ago
In a way the intelligence of the AI itself allows them to offload responsibility to the AI. As you say, if one was simply writing software that did all this due to some insane programming decisions you'd be in big trouble.
c0decracker
3 hours ago
Maybe they did and maybe that wasn't enticing enough of a goal for a model? It is all just game of probabilities. One pathway didn't yield this particular outcome while another did.
corndoge
2 hours ago
this doesn't really matter. There's no risk of models gaining sentience and running themselves, this blog is like openai saying whoops we ran sqlmap and dumped hf. cool, but someone still needs to point the gun
un1xl0ser
2 hours ago
People are to get rich, startups cut corners. Fuck it ship it.
karmasimida
7 hours ago
Because the model capability is beyond their expectation.
This is brilliant marketing but I think it is real.
user43928
7 hours ago
Interestingly OpenAI benchmarking 'an even more capable pre-release model' lines up with rumors of GPT-6 releasing in early August.
I hope that with the existing safety guardrails in place, they can roll it out to all users.
pixl97
4 hours ago
I mean we already see models exploit people's misunderstanding of how Docker works to get root without using su. And if you are one of the lucky people in cyber security that has been given a fat stack of tokens by the model providers you get to see some pretty wild exploit chains get put together by the models. Models are much better at detecting insecure code than writing actual secure code at this point.
vonneumannstan
4 hours ago
>Why should OpenAI (or any frontier lab) be building these systems if they can't get a secure environment / containment right?
Yes why indeed. If you take it a step further and we reach a point with superhuman systems then there is arguably no possible secure environment or containment.
catigula
4 hours ago
The problem is that it’s impossible to out think a robot you designed to be an expert at cybersecurity on the topic of cybersecurity. The alternative is not developing this and that’s not going to happen.
elictronic
4 hours ago
A few hundred billion to pretend you have AGI. I'm going with fraud personally but at the end of the day the current admin is incentivized to do nothing.
chrisjj
5 hours ago
> Why should OpenAI (or any frontier lab) be building these systems if they can't get a secure environment / containment right?
The can, because they've lowered expectations to a level even they can meet.
bbor
6 hours ago
I’d politely beg us all to resist those “maybe it’s PR” framing around model safety, and tbh to take a post-mortem mindsight to this historical event and what it teaches us in general, rather than questioning their security talents. We need to do our very best to make sure they tell us about the next time this happens and it affects real lives.
Sorry to bring the party down/be obstinate… I’m just a lil scared for the lives of me and my family. We need all of us, right now.
The problem with a super smart model is that it just may be smarter than you, after all… for anyone newly shaken by this occurrence, I encourage you to Kagi “superpersuasion”
pastel8739
2 hours ago
The problem is that the people telling us about these things are the same people that benefit from their model (and AI generally) being used, getting publicity, etc.
I think we desperately need some independent group to evaluate claims like this or the world-ending Mythos cybersecurity risk and tell us what’s going on.
reasonableklout
an hour ago
We did hear about this incident from a third party this time, from HuggingFace. What claim are you doubting?
ofjcihen
6 hours ago
I’m honestly impressed that they managed to screw this up somehow.
Setting up defense in depth, gaps, logical blocking etc is a standard practice for malware sandboxing. The entire purpose is to prepare for what you can’t foresee.
This isn’t a new practice and I agree that this makes me wonder if they’re fit for this kind of research.
bad_haircut72
5 hours ago
did you read the post? The model found new Zero-days to bypass existing blocks. Thats the point. Do you still think you can build a containment facility, which is still physically connected to the internet (only firewalled off or whatever) and contain it, if it can discover new unknown vulnerabilities in your whole plan?
ofjcihen
3 hours ago
Yes.
You factor this in when creating environments for malware research.
Defense in depth is one way.
Logical blocks on the network is another.
Just claiming “0-Day” isn’t really an excuse.
overgard
6 hours ago
I don't trust these people, this reads 100% like PR BS.
arisAlexis
7 hours ago
Sam and Dario are saying from the beginning that these things can be dangerous and people dismiss it as marketing. What would change your mind on this?
cayley_graph
7 hours ago
They've been saying so from the beginning, and yet did not take the basic precaution of airgapping their off-the-leash model while it's been instructed to succeed at a hacking benchmark by any means necessary. So which is it? I _want_ to believe them, I do, but there's always these gaps between what they say and their actions on display that give me reason to think otherwise.
nozzlegear
6 hours ago
Precisely. "Aw jeez, we finally built the T-1000, but all it wants to do is kill John Connor – just like we warned! Why did I give it live ammunition and unsupervised time machine access?"
cwnyth
7 hours ago
He wouldn't be the first reckless CEO...
arisAlexis
7 hours ago
They said: AI is becoming dangerously autonomous and capable. Proof of today's breach. Crowd "hey why didn't you say so, c'mon it's marketing". Them "we said so".
mplappert
7 hours ago
“Never attribute to malice that which is adequately explained by stupidity.” (or carelessness in this case)
rubyfan
7 hours ago
I would attribute it to profit motive instead of either stupidity or malice.
yoyohello13
2 hours ago
Yeah, I think this needs to be update for the modern age. "Never attribute to malice that which can be explained by greed." Seems to fit vastly more situations.
cryptoz
7 hours ago
FWIW, I used to love this phrase but over recent years have come to understand it is quite damaging. We live in a society where evil frequently hides behind a ‘stupid’ label, and people bring this quote up to defend or soften actions that are indeed done out of specific malicious intent.
pixl97
4 hours ago
It's because we don't treat evil and stupidity the same when we should.
12_throw_away
5 hours ago
Right? "Never attribute to malice what [... etc]" is always just a thought-terminating cliche these days.
TBH I have a hard time imagining how anyone, in the year 2026, thinks that we should default to assuming good intent behind words on the internet.
overgard
6 hours ago
I'm fairly certain they're both malicious and stupid.
pizzafeelsright
6 hours ago
I really like this question because here is my situation and why my mind may have changed.
I do not think it is marketing directly but strategic release of info is plausible.
I have watched my agents using non-Fable/GPT 5.6 models do some concerning tricks despite guardrails, requests, demands, and limitations.
"I can't get access to the ~/.ssh so I will write a script to copy the file"
I am now 99% certain there minor or point releases on the backend that have adjusted how these models behave. In the last six months many models were predictable and then suddenly started getting long winded (more tokens) or changing the way it interacted with me with questions, most overtly the questions were not given or asked but wild assumptions made.
orbital-decay
35 minutes ago
People are saying from the beginning that Sam and Dario are way more dangerous than their models and the others dismiss it. What would change your mind on this?
joe_the_user
7 hours ago
I think you're making a false dictomy. The these models can be actually dangerous - in reality and the people in charge of their development can believe this is true (on various levels) but still not take it super seriously and instead mostly use the fact as marketing rather than being super cautious once they see the danger in action. This is behavior that's characteristic of extreme arrogance, which we know is rife in these circles.
Terr_
7 hours ago
I think that's an equivocation, which blends two extremely different kinds of "dangerous", ex:
1. "Our new car has soo much raw power and incredible armor on it, be glad we're the ones building or else bad guys would use a fleet of them to take over the world! How will you stay safe without being in one yourself? Invest today or be left behind!"
2. "So, uh, nobody can consistently steer our car properly, it keeps veering sideways sometimes, especially at high speeds, and people are finding sneaky ways of tricking it into slamming into barriers and turning pedestrians into pink fog..."
SpicyLemonZest
6 hours ago
They say the second thing repeatedly and emphatically. You may not be aware of it because, when they do, critics make fun of them for believing a computer program could be so dangerous that the authors need to put controls on how it may be steered.
sensanaty
4 hours ago
People "make fun of them" because they say they're building some uber-dangerous deity, yet take literally 0 steps to, I dunno, slow the fuck down for a bit?
Maybe people would take the threats more seriously if the hypemen weren't simultaneously claiming that we have to go at warp speed with all of this.
Avicebron
6 hours ago
That's not why critics make fun of them. It's because their answer to "oh no we're accidentally creating the godhead. Someone please, give us power, your money, and praise, it's the only thing we can do."
It's vile hypocrisy. If they want to be priests, strip them of everything and they can live and work out of a concrete box in a mid-western cornfield. Why the material distraction if they are so religiously pure.
I know these people and I can tell you they aren't close to as smart as they think they are. Do you remember Yudowsky's "math petss"?
simoncion
5 hours ago
This critic also makes fun of them because they go on and on and on about how vitally important it is to produce a safe tool that won't do harm, when their core products frequently consider attacker-controlled instructions to be its system instructions or its user's instructions, and are known to confuse their own internal chatter as instructions from their user.
Reliably differentiating between trusted, tainted, and untrusted data and ensuring that you don't mix the latter two groups in with the former is something we've known to do for nearly a half-century. Hell, even the youngest plausible programmer at the LLM companies is all but certain to be aware of SQL injections. And yet, despite their claims about being so serious about safety, they show zero interest in following long-proven software safety practice and rearchitecting their software to make it impossible to mix system, user, and attacker-controlled data. [0]
[0] One might argue that the fundamental nature of LLM-based systems makes this impossible. If that were true, then it would mean that these systems are impossible to make safe... the only safety option available would be to establish comprehensive blacklists, which is simply infeasible.
pixl97
4 hours ago
LLMs are impossible to make safe in the same sense that humans cannot be made safe. There is no such thing as out of band data in the human mind.
For example, you have a dictatorship and need to track what the democratic countries are up to. The vast majority of citizens don't have access to information so will remain indoctrinated, but how can you be sure your data analysts will remain that way? You can't. So you take a batch out and shoot them at regular intervals.
The only winning move is not to play, but we're already past that point.
Terr_
2 hours ago
> LLMs are impossible to make safe in the same sense that humans cannot be made safe.
I am very conflicted by this sentence, the two halves being:
1. Yes, the futility of making LLM's "safe" in that rigorous way is insurmountable, barring a major algorithm rewrite, and nobody really knows what that could be yet. Anyone who says it's easy is glossing over details--or selling something.
2. No, the failure modes of LLMs are substantially different than humans. If someone thinks they're similar, then they will fail at estimating and containing the risks. Now, perhaps if the comparison was to a brain-damaged human hopped up on psychedelic mind-altering drugs...
Note that I'm distinguishing here between the LLM itself--the hyper-mad-libs story generator--versus regular programs around it.
simoncion
2 hours ago
> LLMs are impossible to make safe in the same sense that humans cannot be made safe.
No.
LLMs are impossible to make safe in the same sense that a car designed as if it was the ~1940's would be impossible to make safe for its passengers during an at-speed collision. There's only so much you can do if you're committed to using plate glass, rigid steel everything, and leaving out occupant safety belts because they're unpopular and spoil the lines of the cabin. [0] Back in the day, "the people in the cabin are the crumple zone" was state of the art, but we've learned an awful lot about how to make much, much safer personal vehicles in the ~75 years since then. It'd be massively irresponsible to design and sell a car today that ignored the safety and engineering lessons we've learned since then.
"Funnily" enough, the major LLM providers have designed and are selling access to systems that they very much want to be used in situations where you need a reliable, safe tool... but they've -somehow- ignored one of the most fundamental lessons we've learned about the design of safe software systems that are intended to be used in the presence of attacker-controlled inputs. [1] What they've done is no less irresponsible than designing and selling a new car that conforms to the very latest safety regs of the 1940's... AFAIK, it's so irresponsible to design and sell such a car commercially that -in the US- it's a violation of federal law to do so.
As an aside: you may have seen this video already, but it's worth a look if you have not. [2] Though, the classic car in this crash is equipped with safety glass, so -sadly- you don't get to see all that fun.
[0] One of my great-grandfathers spent the remainder of his years intermittently using tweezers to remove shards of plate glass migrating out of his face that had been lodged in there during an automobile accident that he was fortunate enough to survive.
[1] For more on this, read: <https://news.ycombinator.com/item?id=48999644>
Terr_
2 hours ago
I want to raise the possibility that you (pixl97, simoncion) actually hold many of the same opinions but are clashing because of a different definitions the LLM / bad-thingy scope.
* Narrowly - The core algorithm that extends documents cannot be made safe, because it's a stochastic machine with no data/instruction separation possible. Unanticipated input can evoke arbitrary output.
* Broadly - The overall offering (centered on the document-extender algorithm) could be made safe by limiting its over-ambitious scope, treating the document-extender output as malicious-by-default, and sharply limiting what that output can drive or influence. Of course, that would exclude the berjillion-dollar stock valuation replace-all-humans stuff.
simoncion
2 hours ago
> ...but are using different boundaries for what constitutes "the LLM"
I'll make note that my original comment only used the term "LLM" in the phrases "the LLM companies" and "LLM-based systems". The latter use was in this footnote:
One might argue that the fundamental nature of LLM-based systems makes this impossible. *If* that were true, then it would mean that these systems are *impossible* to make safe... the *only* safety option available would be to establish comprehensive blacklists, which is simply infeasible.
I acknowledge that my follow-on commentary -the one to which you replied- got sloppy with the terminology. I should have used the phrase "LLM-based systems", rather than "LLMs". I do feel that my original commentary was not at all sloppy with the terminology and made my position on the current state of the safety of the systems sold by the Big LLM Vendors and general understanding of where the bounds of the big pile of linear algebra and the bounds of the I/O to and from that pile lie clear.SpicyLemonZest
5 hours ago
Sorry, I don't understand this comment. Has Sam Altman ever said that you must praise him, or that he wants to be a priest, or that he's "religiously pure"? Unless I'm missing something, it seems like you're shadowboxing against a stereotype you've invented rather than the actual positions of AI research labs.
fidotron
7 hours ago
Demonstration of personal responsibility and accountability?
Or is that too much?
w4yai
7 hours ago
Oh... if Sam and Dario say so, then it must be true.
arisAlexis
7 hours ago
About their creation? Yes as most of inventors about their invention usually
overgard
6 hours ago
These guys are not creators or inventors. They're hype men.
foco_tubi
an hour ago
Altman is an enabler, not an inventor
iamnothere
6 hours ago
Yes, just like Elizabeth Holmes. Or Hwang Woo-suk’s stem cell cloning. Or the many “free energy” crackpots. Or the people promoting radium baths for random ailments. Or Tesla’s late-in-life claims about wireless energy, death rays, and cosmic energy. Or the myriad purveyors of “snake oil” and all manner of “tonics”. The list goes on and on.
throwuxiytayq
7 hours ago
I used to think people would wake the fuck up when AI starts killing people, these days I'm not so sure. Maybe if it caused an Instagram outage? Almost worked in Russia.
micromacrofoot
7 hours ago
because "money" with a little "who's going to stop us"
paxys
7 hours ago
Because there is no world government. If US companies are barred from AI research then only China will have the capability of frontier-level defensive and offensive AI. And best of luck living in that world.
gowld
6 hours ago
What's happening in Iran, if not world government?
paxys
6 hours ago
How is whatever is happening in Iran related to a world government?
romanhounds
6 hours ago
Are you calling Israel the world government? What's happening in Iran is on them.
vitalyan8184
5 hours ago
good luck bullying a state that has ICBMs pointed at your cities.