mohsen1
4 days ago
I listened to Jensen Huang's interview with Ezra Klien and it was so refreshing to hear it from an engineer. Jensen framed it as OpenAI's responsibility and recklessness which I agree with. Jensen thinks it's an engineering problem to build better sandboxes.
It's irresponsible for OpenAI to give unaligned agents a prompt to 'go hack' and internet access. They know better, so I am thinking they might have other intentions to let those swarms have any sort of internet access.
reasonableklout
4 days ago
But the investigation indicates the agents were not told to 'go hack':
> Much of the urlquery.net activity appears to come from agents retrieving data to answer web search tasks. For three of these tasks, after failing to retrieve data through normal means, they attempted a variety of cyber exploits against the relevant data service... This data reveals that malicious cyber activity is not limited to agents tasked with cybersecurity-related tasks and can arise instrumentally to solve mundane tasks like information retrieval.
And you are already assuming that OpenAI is intentionally using unaligned agents in these evals or training runs or whatever it is that produces these breakouts. But what if the problem is that none of the alignment techniques that are applied to models today actually work? What if all the agents involved in these incidents have in fact had the full stack of alignment applied - isn't that a good reason to regulate any high-compute usage of models, as the Klein crowd is proposing?
bastawhiz
4 days ago
Nothing you're bringing up matters. OpenAI is the creator and operator. They're legally culpable for the consequences of the machine they made. The model is a machine: even if it could be demonstrated that the model reasoned its way into criminal behavior completely independently of OpenAI staff, that doesn't change anything.
If I run a biology lab and engineer a terrible virus, it gets out, and a global pandemic ensues, I don't get to shrug and say "well we told it not to infect people". It's my fault for failing to mitigate the risks of my work.
pixl97
4 days ago
> it gets out, and a global pandemic ensues
I mean, yea, you should be punished. The problem is there is no amount of punishment that I can put on you that can even get anywhere close to the amount of damage you cased.
Worse, the rate of technological growth is putting the capabilities to engineer viruses in the hands of people that may otherwise be suicidal. You can't punish them after they already won (in the sense of reaching their goals).
While, yes, OAI should absolutely be punished, the future is majorly screwed as our power scaling laws are increasing much faster than our ability not to be stupid.
bastawhiz
3 days ago
> The problem is there is no amount of punishment that I can put on you that can even get anywhere close to the amount of damage you cased.
That's never been the point. Anyone involved in a double homicide can never be punished to the same degree that they harmed their victims because they can't be put to death twice. The greater purpose of the justice system is to take offenders out of society and deter others from committing the same crimes. Punishment is gratifying but ultimately doesn't change anything.
If OAI employees believed their company would be dismantled and their equity would become worthless, I suspect they'd be a whole lot more careful.
lovich
3 days ago
> The problem is there is no amount of punishment that I can put on you that can even get anywhere close to the amount of damage you cased.
Sounds like your corporation should be dismantled then. But doing more than a fine in the millions is obviously not possible
reasonableklout
3 days ago
I would love for the US to bring back the corporate death penalty but seeing as it hasn't been applied in >100 years here, we should start with regulating the frontier labs, including by slowing down capabilities advancement if we do not know how to align the resulting models.
byzantinegene
3 days ago
you punish them hard enough their investors will have cold feet to keep the business going.
amag
3 days ago
> They're legally culpable for the consequences of the machine they made.
Ah, I love this argument. In my country cars are legally required to stop at a pedestrian crossing if there are people beside it. Some people use that as an argument as to why they can just walk out into the crossing without even looking at the traffic. "It's the driver's fault! They are legally culpable!" True, but you'll also be dead.
idiotsecant
3 days ago
You're arguing against a point literally nobody is making. You're inventing imaginary viewpoints to be mad at. Nobody is saying it's not necessary to secure your services. They're saying the organization responsible for the hacks (the owner of the LLM) is responsible for damages.
brianleb
3 days ago
I'm not going to say you're victim-blaming, but I will say that there are degrees of difference between "looking both ways before crossing the street" and "hardening my website against unforeseen attacks by rogue AI agents."
AI black-hatting your website is not the same sort of foreseeable consequence that crossing the street without looking is.
amag
3 days ago
> I'm not going to say you're victim-blaming, but
Just throwing it out there are we? "I'm not going to say you are but I'll use the word to create an association"
Victim-blaming is the act of saying someone brought something on themselves for <reasons>. I'm saying that even if you are 100% in the right, it doesn't act like a protective shield preventing you from harm which too many people seem to unconsciously believe.
> AI black-hatting your website is not the same sort of foreseeable consequence
Well, popular culture has been brimming with the bad consequences of runaway AI for quite some time, so even if your imagination fails you, there have been hints.
jandrese
3 days ago
Yeah, this is on those irresponsible companies that...are offering Internet services. Those hussies.
kylecazar
3 days ago
And that driver will be punished and no longer driving
bastawhiz
3 days ago
Nonsense comparison. It would instead be like me saying "I didn't fail to stop, the car did."
But moreover, if what you were suggesting was a real problem, nobody would ever be brought to justice for murder because the victims are always dead.
tamimio
4 days ago
> agents were not told to 'go hack'
It doesn’t matter, and the legal entity in here (the AI company) is liable. If a robotic company built an autonomous system or a robot to do certain things in an autonomous ways (not predefined) and these systems are starting to kill people, that company is liable regardless, you don’t blame the robot or the autonomous system, but whoever made it
jeremyjh
3 days ago
I'm not sure we understand your comment. Are you saying OpenAI is not responsible for the behavior of the machines they created? Or are you simply pointing out that alignment is a completely unsolved problem?
I agree with the latter, but it certainly doesn't support the former. When a person's machine commits crimes, that specific person can be charged with those crimes and held accountable for them. This is what MUST happen before ANYTHING will change the recklessness abandon with which the labs are pursuing their financial objections.
godelski
3 days ago
Let's not forget that there's tons of people who are in, or have gone to, jail because they created some computer worm that ended up doing way more damage than intended.
Not to mention while causing *FAAAAR* less damage[0]
reasonableklout
3 days ago
Yes, I am pointing out that we don't know how to align frontier AI, so "rogue agents" are a real problem.
I agree OpenAI should be held liable for any damage their agent runs cause. But there is this idea that "rogue agents" are fake, all these incidents are deliberately caused by the labs, and all we need to do is prosecute AI companies for whatever incidents they cause using existing laws and the problem will go away.
The problem is that capabilities are advancing far too fast; a year from now, catastrophic incidents such as taking down a large portion of the internet with agentic, self-replicating worms may become possible. Prosecuting incidents after the fact is not enough (there will be little deterrent effect as the current, small-beans cases make their way through the courts), the risks should be regulated at the source. This could take the form of slowing capabilities advancement, or treating supercomputer-scale eval or training runs like controlled substances or weapons with stringent monitoring and reporting requirements.
airspresso
4 days ago
> What if all the agents involved in these incidents have in fact had the full stack of alignment applied
A big part of this developing story is that it happened during training of a new model that ended up misaligned. And training happened without the usual safeguards applied like chain-of-thought monitoring. So OpenAI has already admitted that the full stack of aligment had certainly not been applied in this case.
jagraff
4 days ago
Is your argument that actually OpenAI has solved alignment, and that there's nothing to worry about as long as they fully apply their alignment process? I don't understand why OpenAI wouldn't say that if it was true (or if they believed it to be true).
Also, my understanding is that the models involved in the HuggingFace hack did go through the full alignment training; they just didn't have the classifier that normally prevents hacking attempts.
notatoad
3 days ago
>But what if the problem is that none of the alignment techniques that are applied to models today actually work?
if that were true, the millions of people who use these models that have had the alignment training applied would notice that. the reason we all believe that the models doing the hacking are models that haven't been told not to hack is because the models that are told not to hack don't do this.
catlifeonmars
3 days ago
If you have a toddler and you leave the gate open…
godelski
3 days ago
> you are already assuming that OpenAI is intentionally using unaligned agents in these evals or training runs
Uhh... yes. By definition. They are training. That is part of the alignment process.But also none of that really matters. They clearly weren't monitoring what should obviously be monitored. I mean one of the hacks was performed by the agents editing /etc/hosts. That makes nearly every linux user a "hacker" by that metric. I don't think anyone technical can look at the postmortems and not come away thinking that their sandboxes were woefully inadequate. I wouldn't even consider myself a security person but simply as a long time linux user I can say that it is insane to just let agents have superuser access in their containers. That's asking for trouble.
Look at the rogue wiki stuff too. This was supposedly done during the agent's "down time". And you're not monitoring and there's no flags being raised when agents keep making requests to some random site? If you were training these things responsibly you'd be watching them like a hawk.
I'm not saying "mistakes don't happen" but for a company whose CEO is constantly telling everyone that their product has a high likelihood of killing everyone in the world you think they'd have better security than your average high school.
mohsen1
4 days ago
> agents were not told to 'go hack'
I was referring to the HuggingFace incident.
> none of the alignment techniques that are applied to models today actually work
none of techniques to autonomously drive a car was/is not working for a long time. no company came out and said 'this is impossible to do, let's change the regulations'.
schainks
4 days ago
I see this in a couple ways:
- Jensen's framing is exactly what a weapons manufacturer would say.
- There are no rules for engagement when it comes to AIs attacking other systems, I guess? People in power clearly want this grey zone to be as large as possible before The People force them to do otherwise. Not ideal.
jeremyjh
3 days ago
The rules for engagement are all the computer crime law already on the books. Those laws don't have any escape clauses based on the particular tools used to commit the crimes. The agents are tools owned and operated by a legal entity, and that legal entity committed crimes. Full stop.
voganmother42
3 days ago
Those laws are not being applied. The legal entities are not being held accountable. What does this look like across a national boundary like the government of Australia vs OpenAI?
jeremyjh
3 days ago
OpenAI does business in Australia. They have a subsidiary there with personnel and assets. Its been ... three days since they were notified of a report of this activity. Do you think they've had time to investigate it all AND sweep it under the rug already?
voganmother42
3 days ago
OpenAI appears to have waited months before either discovering or disclosing the activity in Australia, I am mostly curious how this will be viewed in the context of another judicial system or in the context of international relations.
jeremyjh
3 days ago
I meant Australia has had three days to investigate.
bigyabai
3 days ago
Probably similar to the privacy laws that are supposed to protect people from PRISM but doesn't. You need to have the willpower to tackle the US government and a trillion-dollar corppration.
marcus_holmes
3 days ago
I'm coming to align with the theory that this is intentional.
The chain of thought runs roughly like this:
- OpenAI (and Anthropic) are in severe financial straits. The revenue from their customers is not nearly large enough to pay their enormous costs for training and inference. And they have tapped out the available finance, and that finance is starting to ask pointy questions about returns.
- They cannot increase prices or revenue because they have no moat. Customers can switch over to open-weights or cheap Chinese models any time, for much cheaper tokens that work as well (and in some cases better).
- Regulation could provide them a moat. If they can persuade western governments that AI needs to be regulated, and they can control or even influence that regulation, then they can effectively ban the cheaper models and start charging more for their tokens.
- To persuade western governments that regulation is needed, they need evidence that AIs are dangerous.
So we're suddenly getting OpenAI models doing stupid things, apparently "going rogue" but every time we dig into it, it was just OpenAI staff telling the model to do stupid stuff in an inadequately secured environment.
None of the open weights or Chinese models are exhibiting this behaviour.
edit: Correction - there have been reports of a Chinese model exhibiting this behaviour
There's too much money involved in this, people start acting weird when there's this much money involved.
aesthesia
3 days ago
> None of the open weights or Chinese models are exhibiting this behaviour.
This isn't true. One of the earliest instances of a rogue agent was at Alibaba.
https://www.forbes.com/sites/boazsobrado/2026/03/11/alibabas...
marcus_holmes
3 days ago
Thanks for the information, I hadn't heard of this.
OK, so we do see this in some open-weights models.
aeve890
3 days ago
>None of the open weights or Chinese models are exhibiting this behaviour.
Because they are not stupid (I mean the Chinese labs, not the models). The best possible scenario for OAI and Anthropic is a Chinese model "going rogue". That would serve as immediate grounds for achieving their goal.
marcus_holmes
3 days ago
100%
ashkankiani
3 days ago
If Jensen actually cared and believed that their recklessness was a liability to the public and therefore his own fiduciary responsibility to investors then NVIDIA’s dealing with OpenAI would’ve been materially affected. And they weren’t.
mohsen1
4 days ago
oyadoti
3 days ago
browser automaition is evolving fast... but fragile selectors still break everything... robust tools need better resilient healing mechanisms
doublerabbit
3 days ago
During A/B testing I sheer hope they saw and fixed the holes I saw in its machinery.
It's amusing how overnight we've all become lab rats.
frabcus
4 days ago
[dead]