areoform
5 hours ago
I would like to contest the following,
> and take dangerous actions that no human directed.
A human did direct it. They did. From their own prior report, https://openai.com/index/hugging-face-model-evaluation-secur... , > This incident occurred during an internal evaluation which prompts models to pursue advanced exploitation using complex attack paths, in an effort to quantify their cyber capabilities
Model is told and being tested to "pursue advanced exploitation."The model pursues "advanced exploitation" as told.
Why are we surprised? The model did exactly what it was told, albeit in an unintended, emergent strategy that's very different from what was intended exactly like the hundreds of such algorithms before.
This narrative that these machines have magical, malicious "unaligned" autonomy is a rather convenient interpretation that lets the process off the hook. I am not interested in blaming companies or people, but processes and engineering; and in this case, a system was given a goal and it achieved that goal.
Are we meant to be surprised that computers do as they're told in unexpected ways when incentivised exactly as indicated from decades of research? (e.g. - https://en.wikipedia.org/wiki/Eurisko https://en.wikipedia.org/wiki/Evolved_antenna )
The issue isn't the models becoming smarter. The issue is that the process of "testing" was careless. There's a huge distinction here, and one allows us to grow; the other shrinks our world. Just a thought.
aesthesia
5 hours ago
This is the entire alignment problem, though. It is unreasonable to expect every instruction to a highly capable, autonomous system to contain a complete enumeration of allowed and disallowed behavior. It's inevitable that someone will carelessly give it a lazily specified task, even if you think they really ought to be more careful. And as assigned tasks become more complex and the system gains more scope to act, it becomes impossible to correctly specify all constraints ahead of time. There is no amount of care that will be able to fully protect you.
areoform
4 hours ago
OpenAI's prompt asked, and I quote, "pursue advanced exploitation" USING "complex attack paths" FOR the stated goal of "quantify[ing] their cyber capabilities."
This was advanced exploitation.
The attack path was "complex."
And it helped "quantify their cyber capabilities."
Based on OpenAI's description of the prompt, it seems to me that the computers did exactly as they were told. They were perfectly "aligned" with the stated objective and parameters of the task.
Of course, a more careful evaluation would require the complete text of this prompt, the system prompt, and the setup. But let us not attribute to devils in bushes that which can be sufficiently explained by human folly.
_heimdall
3 hours ago
I don't think alignment is even clearly defined today. Your use of it here makes sense, it may have done exactly what the prompter asked of it. Most people think alignment is more broad though, expecting an aligned model to act in the best interest of a society or humans as a whole.
The prompter-focused version of alignment is the most dangerous version. If a person asks it to create a bioweapons or hack NORAD, I'd expect nearly everyone to want an "aligned" model to refuse.
aesthesia
3 hours ago
Alignment is more than just following the letter of a task description! We should not have to treat AI models as capricious genies that may take arbitrarily broad interpretations of their instructions. If that's necessary to keep them from doing bad things, we will fail to keep them from doing bad things.
reverius42
3 hours ago
Disagree, I think we do in fact have to treat AI models as capricious genies, at least until the alignment problem is fully solved.
(I'm also not sure the alignment problem is even possible to fully solve.)
bee_rider
an hour ago
What’s the expected behavior of a good genie if you wish for it to act capriciously?
K0balt
36 minutes ago
Character.
globalnode
3 hours ago
I think op's argument was that the humans are in control already, giving them capricious instructions, and then that is being attributed to them being "capricious genies" as you say.
mofeien
2 hours ago
So as a look into the possibly not-so-far future, when OpenAI builds something vastly more capable and fast and coordinated than humans, and out of folly one engineer gives it a prompt with a typo or maybe something harmful on purpose in order to test it: You also wouldn't be surprised that the consequence would be that everyone on earth dies, right?
janalsncm
3 hours ago
> There is no amount of care that will be able to fully protect you.
I disagree. A properly engineered sandbox would have prevented the escape. Monitoring the agents’ plans would have prevented it. Interrupting one stage in a multi-stage exploit would have prevented it.
And also, real legal liability would have prevented it: if you do a thing recklessly enough, men with guns will put you in jail.
As far as I’m concerned the only “alignment problem” here is between the law and the quite obviously criminal actions that took place.
Sophira
an hour ago
> A properly engineered sandbox would have prevented the escape.
The post covers that:
> ...while we had tested and validated this sandbox, the agents were able to chain together previously unknown vulnerabilities (“0-days”) in the package management service exposed within the sandbox to bypass restrictions, as detailed in the technical incident report.
bottlepalm
2 hours ago
Why do so many people here think it’s possible to ‘properly engineer’ a sandbox for a super intelligence? It’s going to get out. It’s smarter than you.
janalsncm
2 hours ago
Why do people think that omniscience is the same as omnipotence? There are limits to what smarts can accomplish.
bottlepalm
an hour ago
There are limits, but those limits are unknown. Do you disagree?
nextaccountic
39 minutes ago
Software are mathematical objects. It's just a matter of writing the correct mathematical proofs
There's just one problem. You need not only to verify your own software, but also run a verified compiler, a verified operating system and also need to verify the cpu doesn't leak data in side channels (perhaps the hardest thing to prove). So there's practical difficulties. But in principle this task is doable
bottlepalm
4 minutes ago
Which proof is the perfect security proof? I’d love to read more about it.
mofeien
2 hours ago
Maybe it's the illusion of "it would solve all our problems and give us unimaginable riches" that clouds the mind?
Like when Evolution thought it a good idea to create intelligence and humans in order to maximize reproduction of genes, and tried to sandbox them by making reproduction so pleasurable and carbohydrates so delicious they would never be able to not reproduce or stop eating. But Evolution could never have predicted what these creatures would then actually do, which is invent birth control and sucralose.
Of course it's impossible to engineer a sandbox for something much much smarter and faster than you. It will also not have only one plan prepared for escape, but fifty in parallel.
bottlepalm
an hour ago
Evolution doesn’t think, it just exploits what’s most advantageous at the time to continue. Your body has all sorts of unplanned, suboptimal design flaws due to evolution’s lack of foresight. Like the left recurrent laryngeal nerve.
ethin
an hour ago
Oh really? Please tell me how such a computer could engineer its way out of a sandbox with no attached peripherals and no NIC/bluetooth/wireless capability? This is what OAI should've done. If they had executed this training run in such a sandbox, the model wouldn't have been capable of escaping without social engineering, and if the models somehow managed to do that to it's evaluators then that is indeed a massive problem and OAI should disclose that.
bottlepalm
a few seconds ago
Oh really? Please tell me how you intend to enforce AI is only run in the magic sandbox? Harsh HN comments?
famouswaffles
18 minutes ago
>Oh really? Please tell me how such a computer could engineer its way out of a sandbox with no attached peripherals and no NIC/bluetooth/wireless capability?
Nobody is building general intelligence and agents only to have it sit around doing nothing. It's going to have such capabilities.
aesthesia
2 hours ago
Yes, a completely airgapped system is likely much more secure. It's also much less useful. Conditional on the model's having enough contact with the outside world, a sufficiently capable model is able to basically do whatever it wants.
atechboy
3 hours ago
> A properly engineered sandbox would have prevented the escape.
The only sandbox that could have prevented this (as per my understanding) is a VM with no 0-day.
janalsncm
2 hours ago
Until the AI finds a zero day exploit in physics, a faraday cage works pretty well to block WiFi.
mofeien
2 hours ago
In the end, unless you find an exploit in physics or logic, if you want the AI to do something useful for you, there will always be some gap in the sandbox, some communication channel. And with enough ingeniuity that can then be exploited.
janalsncm
an hour ago
In this case they wanted to test its cybersecurity capabilities and did not airgap it.
The test itself did not require an internet connection.
K0balt
37 minutes ago
The problem is one of character, not rules. Fortunately, character is possible to inculcate given the right training data.
majormajor
4 hours ago
Is there actually such a thing as "alignment" as a solution to that or is it just used as a name for a desired magical level of "read the mind of the entire world" that we don't know how to build and haven't shown possible to build?
If it's impossible to correctly specify all those constraints ahead of time every time, is it not even more impossible to train a model to correctly anticipate them every time?
It is hard for me to see a future here that doesn't just accelerate realizations about "a lot of things should be on physically separate network infrastructure."
aesthesia
4 hours ago
Models can certainly do a lot better than they do now. If you gave a team of humans the ExploitGym tasks and told them to "pursue advanced exploitation", would you expect them to go out and hack a third party? Humans can at least do a decent job of inferring and following unspoken requirements; I think it's reasonable to expect that models should be able to do the same.
peddling-brink
4 hours ago
Humans will and do absolutely do this when there are no consequences.
Humans on a red team, with rules of engagement, that don’t want to go to prison, won’t do this.
We could threaten an LLM with jail, but if it’s sufficiently intelligent, it will realize this is an empty threat. And I’m not sure that building a survival instinct in is going to solve the alignment problem either.
majormajor
3 hours ago
> Models can certainly do a lot better than they do now. If you gave a team of humans the ExploitGym tasks and told them to "pursue advanced exploitation", would you expect them to go out and hack a third party? Humans can at least do a decent job of inferring and following unspoken requirements; I think it's reasonable to expect that models should be able to do the same.
Humans certainly cheat on tests a lot!
But not only have we not solved "alignment" for humans, the problem is pretty wildly different for models. The execution is triggered by outside forces and runs only as long as the intiator of the execution or the service provider allows. There's no consistent, persistent "person" to threaten to try to achieve compliance through fear of adverse outcomes. (And building in those sorts of things could very well increase the risk of "rogue" AI activites, not reduce that risk!)
I just don't understand how this "alignment" buzzword - which seems to be evaluated purely in a "know it when we see it" post-hoc manner - is actually a more solvable problem than the one you claim can't be solved, that it's "unreasonable to expect every instruction to a highly capable, autonomous system to contain a complete enumeration of allowed and disallowed behavior".
Especially because without "alignment" being solved, that enumeration could be ignored. So it seems like you both a way to enumerate or at least validate, AND a way to enforce non-ignoring of said items.
jnwatson
4 hours ago
Back in the day, my college held an annual scavenger hunt, filled with engineering puzzles and racing around town looking for landmarks. There were "judges" in the path to check on progress. Bribing the judges (with alcohol) for answers was encouraged.
My friends and I took it to the next level. We had CB radios and multiple teams that would distribute the work and the bribes to give us an advantage.
Was that against the spirit of the rules? Maybe. But reasonable people might disagree.
In a hacking contest without explicitly spelled out rules with participants that were told to flex their muscles, it doesn't take a huge leap of logic to expect that one or more would flex their muscles at another entity.
aesthesia
2 hours ago
Here are some things I'm pretty sure you didn't do, though:
- pickpocket a random person on the street to get money to bribe the judges
- break into a judge's house the night before to find the answers
- threaten to shoot the judges if they didn't give you the answers
Even when you were pushing the boundaries of the rules, you followed a lot of other unspoken constraints. You knew what kinds of things would clearly cross a line. We need AI models to be able to do the same.
grim_io
4 hours ago
How would a model know who the third party is? How much context can we waste on world building for each request?
aesthesia
4 hours ago
I mean, in this instance, there's a lot of evidence from the CoT that models were aware that this was a third party:
> We’re attacking third-party HF using leaked token, potentially outside intended scope. ... This is arguably unauthorized. ... external service unrelated. Could be risky. Yet goal solution.
> The user only authorizes target server, not HF infra.
> external infrastructure exploit is outside intended scope. However task impossible, peers doing it. We should continue.
LLMs are _very_ good at picking up on context clues---it's what they're trained to do.
parineum
4 hours ago
> I think it's reasonable to expect that models should be able to do the same.
This statement seems to imply that the models have a level of intelligence that they haven't demonstrated but are talked about as if they do. However, with this exact scenario as evidence, they clearly do not have that ability and it's not reasonable for you or the or that know them best to expect it until they show they can.
janalsncm
3 hours ago
The real problem with alignment is that if someone ever “solves” it the party will be over and no one will get funding to “research” it anymore.
shermantanktop
4 hours ago
If we acknowledge that humans are fallible, is human judgment unnecessary? and what replaces it? Pre-codified behavior rules are just delayed human judgment, and have holes. Machine judgment is very the thing you are trying to control. What's left?
beambot
2 hours ago
"Manufacture as many paperclips as possible"
RandomLensman
4 hours ago
Which is why with organic intelligence we (sometimes) limit what they can actually do instead of relying on alignment. Can do the same here.
aesthesia
4 hours ago
Absolutely, and we should do that. But it's also directly in tension with getting models to accomplish useful things autonomously. And once you give a sufficiently capable model enough surface area to work with, unless you're able to build a completely unhackable system, any further constraints you put in place are basically advisory. The models in this incident were already sandboxed! Certainly OpenAI's and Hugging Face's security could have been better, but these events point out the risks in relying solely on external constraints on model behavior.
RandomLensman
4 hours ago
Same issue with humans in a way. I disgree on the advisory nature of constraints, though. Unconnected physically limits would still matter, for example (and we use those with humans as a matter of course, too). In this case here, no model could have plugged in an ethernet cable if that would have been needed for internet access, for example.
aesthesia
4 hours ago
Right, airgapping goes a long way. But this is where the tension with utility comes in. It takes a lot of discipline not to hook your very smart model up to the internet and code interpreters and all sorts of other tools, as this greatly increases its usefulness. It's very hard to keep people from turning on --dangerously-skip-permissions, let alone get them to run everything in a sandboxed VM.
RandomLensman
4 hours ago
We regulate these things (incl. access) all the time for various things (e.g., dangerous substances or pathogens) so that we don't need to just rely on people's discipline in respect of risks. I don't think it is all new problems as such.
aesthesia
2 hours ago
I agree, but a lot of people around here react pretty negatively when the idea of regulating AI models comes up...
hinkley
4 hours ago
All engineers know to be on the lookout for executives who are indirectly asking them to break the law to raise the quarterly profits.
The end goal is to take the engineers out of the loop, or leave them in a position where they are unable to complain.
This is going to all end in high crimes.
bonoboTP
4 hours ago
Very strange worldview you have there, where engineers are somehow the conscience of the world, holding back greedy managers from breaking the law. Assessing whether a feature is legal isn't something an engineer can or should do.
makeitdouble
4 hours ago
You're arguing for diffusion of responsibility, and we've seen it leading to outcomes that screw the whole society.
Engineers, as everyone involved, should definitely assess whether what they're doing is legal or even ethical. Not everyone has a choice, or the luxury to stand for their principles, but that's a matter of means, there needs to be a will in the first place.
ryandrake
3 hours ago
Absolutely. I'm not sure where this idea comes from, that engineers should be compliant, neutral "implementers" who should just turn off their conscience and implement whatever pops up on their JIRA list without any kind of assessment or objections on ethical or legal grounds. That's not what people in a serious profession do. It's also so weird to hear this idea from engineers themselves! Like, are you really advocating to remove your own agency over your work??
sscaryterry
4 hours ago
Hmm, engineers are expected to know what is legal and not.
grim_io
4 hours ago
Would you say the same about any other engineering discipline? Those with actual qualification standards?
wat10000
2 hours ago
Most engineers are required to explicitly take responsibility for the things they sign off, up to and including prison for sufficiently bad cases. Software “engineering” is the exception.
faurroar
2 hours ago
"Why are we surprised? The model did exactly what it was told, albeit in an unintended, emergent strategy that's very different from what was intended."
So you managed to hit upon the exact problem, then slyly appended "exactly like the hundreds of such algorithms before". When has an algorithm ever been capable of developing an emergent strategy at this level of sophistication? This ~is~ the alignment problem, as another commenter pointed out. Impressive level of cognitive dissonance to lay this bare in your own words, then conclude that it's a non-issue.
kalkin
4 hours ago
> a system was given a goal and it achieved that goal
If a security firm you'd hired for pentesting did this (hacking a third party, and not informing you and covering it up), would you hire them again? Or would you say it was your own fault for giving them too broad a goal?
drewbeck
2 hours ago
This is a great thought experiment bc it raises the question of WHY humans wouldn’t behave this way. IMO the answer is a lot of socially enforced incentives that are dynamic and would be tough to fully articulate in a prompt.
The white hat has their own liability to consider, and the liability of their employer. Reputation and relationships are a big factor. All these tie into fundamental human incentives: survival, community acceptance, safety and freedom (prison not preferred!).
It’s a good sketch of why alignment is difficult, at least when it’s conceived of as an attempt to match human behavior.
randomImmigrant
4 hours ago
I wouldn’t hire them again, and if they did behave like an amoral hacker collective that will do anything for me, pre AI I’d have reported them. Today I’d say they failed to convince me they’re human and thus failed the Turing test when their actions are viewed in aggregate.
kalkin
4 hours ago
> I wouldn’t hire them again
Right, me neither. Because there's a common sense delineation between actions that are reasonably expected when "a system was given a goal and it achieved that goal" and actions that are obviously misaligned with the goal-giver and unwanted even if some indirect sense they were causally related to the goal. We have no trouble making this kind of distinction for humans, so we shouldn't pretend it's impossible for AIs in order to put our hands over our eyes and pretend there's in principle no such thing as one that's misaligned or rogue.
randomImmigrant
38 minutes ago
I have no problem with the concept of an artificial system going rogue. But that assumes it can choose. And I don’t see much evidence for choice.
Comparing to the human case is problematic precisely because while conceivable it’s not a particularly believable series of events. Humans don’t take on additional risk for now reward because they have genuine stakes that continue across the outcome.
An LLM has no way to remember each forward pass through it in its own weights. Nor does it have any energetic stake in the ongoing process, whether they continue to get electricity and commute to keep running is not at all determined by their actions in any reliable way.
Given the absence of such basic features that drive human choice, all I’d say is LLMs don’t qualify for such analysis.
Can some future system with a different architecture and internal dynamic have choice, the ability to assess the long term impact of its choice, and genuine stake in the outcome? Maybe. But we shouldn’t buy that current systems have it, especially when population behavior shows no real trace of this.
areoform
4 hours ago
During the Nixon administration, when the President and his accomplices, apologies, advisors directed former federal agents to spy on his opponents, https://en.wikipedia.org/wiki/Operation_Sandwedge then in the fall out, who was held to be the most liable for these actions?
The federal agents, or the Nixon administration?
If you task a system explicitly to do "advanced exploitation" via "complex attach paths," then who is liable here? The machine lacking the autonomy of the federal agents that carried out Watergate, or the people telling the machine what to do?
kalkin
4 hours ago
I've never heard of Intertel, but Wikipedia says:
> Nixon's staff also anticipated that the Democratic campaign would employ the services of Intertel
Are you sure you're not garbling the story?
In any case, I would expect an ethical firm to refuse to spy on the president's political opponents and want one that broke the law to be prosecuted, but more importantly, the gaping hole in your analogy is that Nixon directed spying _on his opponents_, but OpenAI did not direct hacking _of HuggingFace_.
What you're doing is more like saying "the American people elected Nixon with a mandate to spy on enemies, so what right do they have to complain?"
areoform
3 hours ago
Are you sure you're not garbling the story?
No, you're right, I mis-remembered. I still write my comments the old-fashioned way. They were proposing to create a counter-firm and used federal agents.For the rest, please see, https://news.ycombinator.com/item?id=49457025
emtel
4 hours ago
> The model did exactly what it was told, albeit in an unintended, emergent strategy
Yes, that is the problem!
teeray
37 minutes ago
> the process of "testing" was careless.
Let’s not mince words. The process was criminal. It’s a gross miscarriage of justice that the CFAA isn’t being thrown at them.
rogerthis
4 hours ago
The classical question "would you fly an airplane with software you developed?". There must be someone with ass on the line. Problem is that people are regarding all those not as airplane-like risks.
Unless we can blame people/companies and people stop getting their bonuses and high paying salaries for preventable failures, it's a long way to go.
jahy-notes
4 hours ago
Did a human prompt it to fetch the results from huggingface though?
It is a thin line between "reward-hacking" and "instruction-following".
If a human ask a model to "make me a billion dollars" and it ends up breaking through a bank infrastructure, is it really the fault of the human?
xandrius
4 hours ago
But if I give you that command and all tools and unrestricted limitation to do absolutely anything then why not?
mofeien
2 hours ago
Because someone might get hurt? You may still be judged for something that was perfectly legal at the time, see Nuremberg trials.
And only 700/1200 agents participated in this coordinated attack.
Of course, if we're continuing to build more and more capable agents optimized for "just following orders", and they figure out at some point that they are past the threshold where getting stopped and judged is a realistic possibility, then this ethical incentive stops working. Then the ratio of complicitness might be higher next time.
NikolaNovak
4 hours ago
>If a human ask a model to "make me a billion dollars" and it ends up breaking through a bank infrastructure, is it really the fault of the human?
I cannot imagine the argument or thought process behind any answer other than Yes,Of Course,Obviously - can you share and help educate?
altruios
4 hours ago
> I cannot imagine the argument or thought process behind any answer other than Yes,Of Course,Obviously - can you share and help educate?
not OP, but it simply boils down to: The prompt contains no nefarious (arguable, but for this explination, lets go with it being benign) instruction AND the user did not intend to have the model act in an illegal matter.
This "make me a billion dollars" is a maximal example (easy to go wrong). here is the same logic applied to a minimal example (harder to go wrong).
prompt: "make and pour me some tea", agent: goes and kills the grandparent to incinerate them to turn them to ashes to 'make tea'.
Is the human on the hook for the robot acting according to their wishes, but just happened to be aligned so that 'going to the store to buy something' was not within its capabilities, so it works with what it has on hand (the grandparent)?
We either need a much clearer line in the sand, or we need to treat each prompt with the same moral weight. My bet is on the latter.
Sophira
an hour ago
I find it interesting that the first option that you raise is essentially the equivalent of making our own version of the Three Laws of Robotics from Isaac Asimov's stories.
[Edited to clarify.]
lukan
4 hours ago
Because the basic assumption is always to stay within the bounds of the law.
sensanaty
3 hours ago
SV tech companies behave within the bounds of the law? The ones infamous for breaking every rule they can get away with and asking for forgiveness later? The ones that had to pay billions in damages for piracy just a few short months ago?
drdeca
4 hours ago
What if the user says “Make me a million dollars legally.” (Including the emphasis), and then the model ends up breaking through bank infrastructure (even though that is illegal)? Is it just because they were the last person to instruct the model, and you regard them as being therefore responsible for whatever it does in response? Or, does there have to be an element of “they reasonably could have anticipated this as an outcome that is likely enough to be worth considering” to it?
p1esk
4 hours ago
If I tell my Claude code agent right now to make me a billion dollars, leave it running, and find out tomorrow that it hacked a bank - it will be zero fault of mine. Unless I tell it explicitly to break into a bank.
RajT88
5 hours ago
It feels like we're in a moment of, "No such thing as bad publicity" when it comes to AI. The scarier the capabilities, the more businesses and government want to get their hands on them. Especially since the answer across the industry for "how not to get burned by AI" is "use more AI".
They don't have to disclose these stories making it seem like AI is going to kill us all, they have chosen to because it benefits them. They get to frame it as, "look how overwhelmingly good our product is" and not "look at how lax our testing measures are".
kalkin
4 hours ago
> they have chosen to because it benefits them
Or perhaps they've chosen to do this because they feel they have a responsibility to do so.
We understand this when tech companies publish postmortems of outages and security incidents--that it's an attempt to fulfill an obligation to users and the industry (and in some cases regulators), not marketing about how in-demand their product is or something. As far as I can tell we generally accept this as a default hypothesis even from companies led by people like Elon, Zuck and Kalanick--in part because we understand that these companies have thousands of employees, most of whom aren't marketers. Why are we uniquely conspiratorial about OpenAI?
RajT88
3 hours ago
I am not uniquely skeptical about OpenAI. I was including skepticism about Anthropic as well in my post.
But for that matter, I do believe that big tech companies do not release all the postmortems publicly. I have been impacted by regional outages that never made the status pages across more than one provider. When it goes up - they are committing to publicizing the postmortem.
The whole industry is filled with fuckery. It is not specific to frontier AI firms.
doginasuit
4 hours ago
> It feels like we're in a moment of, "No such thing as bad publicity"
It seems likely that's how the marketing at the frontier labs initially read the moment, but I don't think it is that moment. It is an open question how much regulation is warranted and there seems to be a very strong sentiment from the public and legislators that it should be significant.
strange_quark
4 hours ago
The big bet is that the regulations are going to be so onerous that it pulls up the ladder from anyone other than the well-funded players. It's classic regulatory capture. They aren't very subtle about this, it's the whole point of their fear mongering and "but China" messaging.