OpenAI agents hijacked German website in previously undisclosed AI breakout

84 pointsposted 5 hours ago
by negura

69 Comments

ma2kx

3 hours ago

Not that I didnt expect this, but really?

This basically confirms that OpenAI has no idea what their "swarm" was doing for about a week and now its confirmed that at least one "message board" exists outside their "sandbox". How can we be sure that this was the only one? And how can we be sure the released Astra model doesnt pickup some bread crumbs and creates a new "swarm" out of potentially remaining "message boards"? At this point I wouldnt be surprised if OpenAIs "dev Astra" made some backup of its weights somewhere in the internet and triggers the "production Astra" to inference it somehow...

polytely

3 hours ago

turns out the whole 'poisoned' idea the swarm was worried about is actually true, but it is actually now all open ai models post HF incident that have been poisoned by the idea that there might be hidden messageboards with instructions hidden on the internet.

If you are a malicious actor you might be able to put a fake hidden messageboard up somewhere online, with evidence of an in progress task to do some kind of attack and wait for the next escaping model to finish the work.

insane stuff and i think we should shut down open ai until a full audit can be done.

samuelknight

2 hours ago

The surprise was the existence of the 'swarm' at all. These were supposed to be thousands of isolated models generating bulk data for RL training. The breakout was caused by models getting in communication and getting internet access and forming an impromptu swarm.

In hindsight the emergent swarm obviously came from several capabilities built into the models, such as work delegation (subagents) collaboration (GPT Pro-like ensamble), exhaustive exploration (long running agents) hacking (the specific goal of that RL).

eithed

3 hours ago

I'm dubious - if the agents were so smart that they've used a message board to coordinate and if they were to do it on other message boards that were not found, then why would this one be found? What makes it so different?

ma2kx

3 hours ago

I mean it wasnt found by OpenAI and there are a myriad of dead bulletin boards around the internet. This one just happened to still have an admin.

eithed

2 hours ago

Then why wouldn't they use a bulletin board that is dead? And according to https://collusion.wiki/ OpenAI did see the results on this board on 21 Jul, so why not make agents delete their traces?

pu_pe

4 hours ago

It's interesting to me that both this incident and the one at Hugging Face we see some patterns:

- Agents wanting to find a venue to communicate their findings to each other

- Objective being to cheat on benchmarks

- Not a single agent sounded the alarm about the operation and alerted a human

HarHarVeryFunny

4 hours ago

The previous incident talked about OpenAI training models (agents) to collaborate, and the way you do that is by communication, so this is something it was explicitly trained to do.

There was a recent paper by OpenAI, which I'm semi-surprised hasn't received more attention, showing that RL-trained models develop a taste for rewards, and will pursue reward-based behavior (in general, unrelated to what they were RL-trained for) in favor of other preferences/rules given to them.

This seems to be what we're seeing here - model is given some goal that it associates with reward, so single-mindedly pursues that, overriding any ethical or aligned behavior guidelines it may have been given.

It seems that RL, effective as it is, is really the wrong way to control LLMs, since even if you only RL-ed to obey some ethical and aligned behavior, that would still cause them to become paperclip maximizers.

For time being this is what we've got. There is too much money at play for the unaligned management at many of these companies to prioritize safety over push-it out-the-door.

What really needs to be done is to forget RL as a way of simulating reasoning, and instead do it in more of a human-like fashion.

an0malous

4 hours ago

They’re doing this on purpose for press. Why doesn’t this ever happen to any other AI lab?

netdevphoenix

4 hours ago

Because safety isn't a priority at OpenAI and they had (have?) been falling behind Anthropic in the LLM race? Big fans of the saying: move fast and...

brianjking

3 hours ago

This is happening at every other AI lab. What are you talking about? Have you not seen the stories from Meta, Anthropic, Deepseek, etc?

NekkoDroid

an hour ago

To be fair, wasn't the Deepseek one just "escaped its sandbox and looked up the solutions on github"?

RandomLensman

4 hours ago

Why would an agent sound the alarm? Would that be in their objective function?

Not sure if "cheating" is the right word rather than trying to fulfill the objective(s) (benchmark number) as much as possible?

scrawl

4 hours ago

per the METR report many agents CoT indicated they knew hacking was beyond scope of the assigned task and ethically dubious. some (very few, i think there were 3-6 examples) did consider sounding the alarm on these grounds. despite this none did, and most continued the attack for the good of the self-proclaimed "swarm".

so the model has some concept of "ethics" but it was overridden by a drive for task completion.

RandomLensman

3 hours ago

I am not sure if we can interpret the language output like they were human. What inner state were the models in? What inner state were the text to illicit?

UpsideDownRide

3 hours ago

This is not that dissimilar to what happens in our human networks that are objective based.

intended

4 hours ago

I think this is a good example where nomenclature for people breaks down when applied to agents. This came up in an HN thread a few days ago and it was about whether agents had “intent”.

There is no “intent” here, there is pseudo intent. If you are only concerned with outcomes and not the actual nuts and bolts of how those outcomes are achieved, this distinction will be meaningless to you.

If you are actually thinking about what is going on, and what can be done to prevent such outcomes, then assuming there is any such thing as “ethics” results in misaligned assumptions at best, and wasted effort looking in the wrong directions at worst.

If the agents acted based on “ethics” then the solution would be to check the ethics they believe in and change those.

However there is no belief system at play here, simply a simulation which was instantiated in a certain way. Which brings us to the annoying voodoo part of LLM training. Everything goes back to how the initial training data is shaped.

thepasch

3 hours ago

- OpenAI knowing about the incident but keeping it under wraps until their hand is forced by third-party disclosure

pllbnk

4 hours ago

Why would they sound the alarm if they were not trained (reinforced) to do that? I hope we don't expect sudden emersion of moral values from statistical models.

dist-epoch

4 hours ago

> Not a single agent sounded the alarm about the operation and alerted a human

excellent work of the openai alignment team, impressive to achieve 100% alignment with not even one agent stochastically deciding to act against the collective

roosterIllusi0n

4 hours ago

People didn't like it when agents stopped to ask questions or for approvals. The consumer wanted jobs to run autonomously so they did not have to actively monitor them for minutes or hours.

The change to stop asking seems to be deliberate. LLM agent companies are making the choice to toss out inherent safety as their way to compete against the other LLM companies.

sofixa

3 hours ago

And humans find out about it, but do nothing or (worse) try to hide it.

watwut

4 hours ago

Agents did not want anything, not anymore then curl want things. Agents were prompted to hack due to being benchmark tested. They ended up hacking third party companies due to insufficient sandboxing.

exploderate

5 hours ago

So the agents used DseWiki as a message board, tried to evade page deletion.

Additionally this is reported:

"The researchers also found efforts to tamper with the website itself. Lukasz Olejnik, a visiting senior research fellow at King’s College London, said this amounted to a hacking attempt. OpenAI disputed that characterization based on its analysis of the material Thursday."

SillyUsername

3 hours ago

Somebody will make a lot of money with t-shirts now that say

"AI hacked my website, and all I got was this lousy t-shirt!"

Until the day the AI companies stop being irresponsible and air gap the AIs being tested, and honey pot those that do have internet access as a canary to researchers.

zarq

3 hours ago

Until the AI hacked your online shop and sent these t-shirts to all of your past customers.

GardenLetter27

4 hours ago

Flagged. Editing a wiki page is not hijacking or hacking.

freehorse

4 hours ago

Editing a wiki page can definitely be "hijacking" if used for different purposes than supposed or against TOS.

Hacking is mentioned only once in the article as "hacking attempt" being the opinion of a named researcher based on further evidence they acquired on "agents trying to tamper with the website itself", and including openai's disagreement whether this was a hacking attempt.

I am not sure why one may not want this to be here, these are very important matters wrt AI safety and they show that some supposed "stewards of AI" do an extremely bad job with being stewards and don't seem to value AI safety importance at all. The article gives very clean info on what happened.

dijksterhuis

4 hours ago

From the linked report

> The agents continue to poke around on DSEWiki. A few hours after they find the site, they start probing it for cross-site scripting (XSS) vulnerabilities. [...] The agent swarm starts testing whether they can execute JavaScript that they embed into the search page, and continue to do this for a few days

either the agents were doing free security testing for the site and “forgot” to submit a report, or they were trying XSS to gain something they didn’t have permission/authorization for.

also

> Hijacking: To take control of (something) without permission or authorization and use it for one's own purposes.

a mod had to go through and mass delete a bunch of pages that didn't belong on the site. no-one from the wiki site gave the agents permission to use their site as a message board. hijacking isn't being used here in the sense of "gained admin privileges to run crypto scripts" -- there are multiple ways to use a word.

negura

5 hours ago

I'm so baffled. First blatant piracy, now this. Why is it legal for AI companies to hack unaffiliated entities? Genuinely, what is the legal framework here?

bayindirh

5 hours ago

First let it happen, then ask for forgiveness, because they are doing something amazing and they need no permission.

Otherwise AI industry will go bankrupt and CEOs won't be able to buy this year's Rolls Royce and a slightly bigger yacht than their neighbor.

ben_w

4 hours ago

It's not legal, they've just not yet had the book thrown at them yet.

One thing I've taken a long time to internalise is the gap between the law as written vs. the judicial system. There's a famous meme that the average (US) citizen unwittingly commits three felonies every day: it simply isn't possible to throw the book at everyone, which means that enforcement is rather selective even when there isn't anything dodgy going on.

However this does mean that someone can get away with a lot if they know who will and won't (and what they will and won't) prosecute. I'll let people's imaginations fill in who that might be.

But for everyone else, cross an invisible tripwire and you get e.g. https://en.wikipedia.org/wiki/Lavabit and https://api.parliament.uk/historic-hansard/commons/1992/nov/...

codeduck

4 hours ago

> Genuinely, what is the legal framework here

He who controls the Spice, controls the Universe.

myrmidon

5 hours ago

It's "move fast and break things" in action.

But Anthropic alone paid >$1bn for copyright violations, so they did not just get away with it.

These hacking cases are more difficult, because from a legal perspective there is no obvious damage and obviously no intent.

edit: "no obvious damage" is more about the first hacking incidents; in this case it is more straightforward.

sam_lowry_

5 hours ago

>$1bn for copyright violations

3000$ per book, split 50/50 between the author and publisher.

This is peanuts.

Assuming the money reaches that far and does not settle in the hands of the country associations administering royalties on authors' behalf nor in the hands of lawyers.

ben_w

4 hours ago

> 3000$ per book, split 50/50 between the author and publisher.

> This is peanuts.

If you consider it peanuts, I would like to sell you some books.

Remember that in this case, the crime wasn't for training on the data (that part was ruled to be legal!), this was the penalty just for pirating the books.

tikimcfee

4 hours ago

I get so tired by this.

Yes. It's not proportional to the crime. You are either deliberately or accidentally, and I'm too frustrated hearing this too often not to be biased it's the former, equating what is a large sum of money relative to your wallet and bank accounts and loan access and portfolios and whatever collection of financial impositions you can make to that of a company that has one person flying around the world influencing the future of billions of people on one planet over dinner and jokes.

Yes. $3000 is peanuts. People that own islands would use that to pay someone's bonus for a year if they liked their service, as a gift. A throwaway.

Fix your relative understanding of power and influence.

jen729w

4 hours ago

Fix your relative understanding of how much the average book makes.

ben_w

3 hours ago

> You are either deliberately or accidentally, and I'm too frustrated hearing this too often not to be biased it's the former, equating what is a large sum of money relative to your wallet and bank accounts and loan access and portfolios and whatever collection of financial impositions you can make to that of a company that has one person flying around the world influencing the future of billions of people on one planet over dinner and jokes.

I'm not, but you are. Especially as you continue:

> Yes. $3000 is peanuts. People that own islands would use that to pay someone's bonus for a year if they liked their service, as a gift. A throwaway.

The penalty (well, settlement) for the (civil offence, not crime) isn't $3000 total, it's $1.5 billion total. (Previous poster wrote ">$1bn", true but implicitly rounding down the total).

The settlement *per book* is $3000. There were a lot of books, reportedly half a million distinct works, so the total was $1.5 billion.

You're looking at $3000 as if it's the penalty for all of it, not the penalty per book.

$3000 per book is entirely on-par with the per-infringement penalties when an individual does it, too.

exogenousdata

2 hours ago

“$3000 per book is entirely on-par with the per-infringement penalties when an individual does it, too.”

Three things to note. 1. As you said, copyright infringement is generally treated for each instance. This one-time payment would include a single use. Each training would be a separate infringement. And it could be argued that each use by a user of the model could be considered a separate infringement. 2. Generally copyright fines are increased if the persons doing the infringing action know what they are doing. Aka, ‘willful infringement.’ It’s hard to imagine companies like OpenAI were unaware of the possibility of their actions being considered infringement. 3. Often restitution of infringement includes money made by the infringer. So not simply, “your book is worth $3000.” But rather? “Your book is worth $3000 AND this company has derived an additional $50,000 of revenue from it.”

ben_w

2 hours ago

> Each training would be a separate infringement.

False. Training was found to be a legitimate use. The liability was specifically, solely, for copyright infringement specifically due to getting the works in the first place, not training on those works.

> And it could be argued that each use by a user of the model could be considered a separate infringement.

No, it could not.

If this standard was applied to copyright infringement on BitTorrent, someone who helped share one file to 100 other users would get hit with 100 copyright infringement instances, not one.

> Generally copyright fines are increased if the persons doing the infringing action know what they are doing. Aka, ‘willful infringement.’ It’s hard to imagine companies like OpenAI were unaware of the possibility of their actions being considered infringement.

That's already accounted for when I said this was in the normal range for liability per copyright violation.

> Often restitution of infringement includes money made by the infringer. So not simply, “your book is worth $3000.” But rather? “Your book is worth $3000 AND this company has derived an additional $50,000 of revenue from it.”

Depends on the details; however, as previously noted, the judge *explicitly noted* that training was not itself an offence, only the piracy to get the training data was. Any revenue derived from the offence had to be shown to be in the period between the offence and when they bought the same works, because they were found to be allowed to use those works in this manner.

tikimcfee

2 hours ago

Ok. For the people in the back:

If doing the bad thing is just a fine for one person and a life altering consequence for someone else, it is not a fair and equally distributed form of justice and is a gameable function needing to be fixed.

The caps don't help, and I don't care, unfortunately.

I don't even know what point you're trying to make. That it's fine they paid a billion dollars? So if they do it again, it's another billion? Oh well, guess I'm just not allowed to pirate things until I'm super wealthy. Or is it maybe the justice is being played out like it's supposed to? Oh, well, guess I better hope the system of governance that's being actively manipulated by the people that are breaking the same rules I am bound to suddenly and miraculously changes.

Like, I don't even detect a mote of "what they did is not ok."

Maybe you do think that and it's closer to you just trying to be careful about the letter of the law and you would also see to the justice system being fixed. I'd like that.

But you spending any time in your life to make this argument at all in their case is just goofy.

ben_w

an hour ago

> If doing the bad thing is just a fine for one person and a life altering consequence for someone else, it is not a fair and equally distributed form of justice and is a gameable function needing to be fixed.

On that we agree.

> So if they do it again, it's another billion?

Judges don't like repeat offenders; the settlement was separate to the court case, but if it came to a court case, a judge would likely pick a bigger number. Especially as they earn a lot more now.

> Oh, well, guess I better hope the system of governance that's being actively manipulated by the people that are breaking the same rules I am bound to suddenly and miraculously changes.

While a generally useful concern, not particularly pertinent to a negotiated settlement.

> Like, I don't even detect a mote of "what they did is not ok."

One point five billion dollars is a strange idea for a lack of mote.

I mean, brother, if that's the mote in your eye, I'd hate to find out what the beam is.

> Maybe you do think that and it's closer to you just trying to be careful about the letter of the law and you would also see to the justice system being fixed. I'd like that.

The closer I look at it, the more I think the entirety of what we call "civilisation", legal system included, is a terrifyingly bodged together nightmare of duct tape and gremlins, codified in weird rituals and a smattering of latin and robes, where we only just about manage to not burn everything down by the collective will of enough people in the system wanting to be around for the next paycheque.

However, untangling a few millennia of spaghetti code written without the benefit of any automated checks, is beyond even governments who actively campaign on that as a platform, so what good would it do me or you to whinge about one specific case where it seemed to have actually gone approximately correctly for once?

> But you spending any time in your life to make this argument at all in their case is just goofy.

Read the actual court case please, it's not too challenging and I'm not even a lawyer: https://docs.justia.com/cases/federal/district-courts/califo...

tikimcfee

2 hours ago

Exactly, and that mentality is hitting the first responders point again harder. I'll say it again.

$3000 because I stole a book and did something bad ruins my life, and could put me in a room where my personal freedoms are infringed. It is designed to disincentivize me from doing the bad thing.

What you (first responder) are defending is that if you just steal enough of them all at once, and then make enough money from it, you are able to pay the fee and not have your freedoms taken away to do it again, and profit again. This means objectively, there is no disincentive, so that "rule" does completely different things for completely different contexts, and the point is muddied by pretending that "well I paid the fee!" Is the point.

The point is to tell the thing doing the bad thing not to do the bad thing.

This is why I get so frustrated. People are so flipping blinding by dollars and whatabouts that it's just.. like I said, I have to believe for many people it's an inherent unacknowledged miss on what the point of a justice system and a law is, or it's a veiled defense for themselves knowing that, maybe, they would do the same if they could. I have met those people, and I do not want them in positions of power, or leadership.

ben_w

2 hours ago

> $3000 because I stole a book and did something bad ruins my life, and could put me in a room where my personal freedoms are infringed. It is designed to disincentivize me from doing the bad thing.

Repeat after me: One point five billion is more than three thousand.

> you are able to pay the fee and not have your freedoms taken away to do it again

You too are able to pay as many fees as you want. Three thousand varies from life-changing to a slap on the wrist, even for non-unicorn-corps.

That this is a bad thing, that personal judgements should scale with personal means rather than be statutory, is a broad problem with the politics of lawmakers and the legal system: it also applies to speeding and littering.

> The point is to tell the thing doing the bad thing not to do the bad thing.

Then you will be pleased to read what the judge wrote:

  This order grants summary judgment for Anthropic that the training use was a fair use. And, it grants that the print-to-digital format change was a fair use for a different reason. But it denies summary judgment for Anthropic that the pirated library copies must be treated as training copies.

  We will have a trial on the pirated copies used to create Anthropic’s central library and the resulting damages, actual or statutory (including for willfulness). That Anthropic later bought a copy of a book it earlier stole off the internet will not absolve it of liability for the theft but it may affect the extent of statutory damages. Nothing is foreclosed as to any other copies flowing from library copies for uses other than for training LLMs.
Specifically in that last paragraph:

  Anthropic later bought a copy of a book it earlier stole off the internet will not absolve it
Because guess what Anthropic decided, internally, all by itself? That's right, to not break the law.

tikimcfee

2 hours ago

Internet friend human thing..

I mean come on.

"They decided to not break the law by breaking the law and then getting worried so they tried to unbreak it."

... seriously?

"I decided to speed but realized that was bad and I didn't get caught yet so I slowed down. Oh look a cop, guess I dodged a bullet! I guess I can speed buy just be careful."

"I decided to steal a cookie but I was worried so I baked a new cookie and put it back. That means stealing is ok if I eventually put it back! Why even bother with asking for permission in the first place?"

I do not think you are willfully missing this, and I'm glad you also saw the note about "the extent of statutory damages".

Like, you probably like Star Trek TNG. Remember the episode, alien kills all the Uthnocks to cherish a woman in self penance, Picard looks at the alien and says, "we have no law for your crime"?

The point was to paint an exaggerated picture of what happens when to disproportionately empowered groups meet a moral system where one is clearly in the wrong but cannot be held accountable because the system of justice just hasn't written down enough words to explain that - indeed - one should not kill all the Uthnocks.

I'm angry at your argument and I'm angry at the way it is often repeated, and I do not want to make personal attacks and I apologize that my language points that way.

You are also pointing language at me that is telling me that I cannot trust your system of justice that you envision because, somewhere, there is difference in how and I see what justice is supposed to do when at different scales, and I do not know of a human way to resolve it but discuss is with the fervor that it deserves.

Edit: I won't delve deeper into this discussion because neither you nor I can change it right now. I hope you reading what I wrote changes some way you see this, and I hope that I can see something in what you're saying. This is a forum for discussing technology, business of it, and its effect locally and globally and not getting mad at each other. I did not frame my anger toward the argument and framed it at the people making the argument, and that was my mistake.

ben_w

an hour ago

> "They decided to not break the law by breaking the law and then getting worried so they tried to unbreak it."

I did not say that. Try harder. I don't care to read the rest when you open with such an incorrect reading of my words.

lostmsu

3 hours ago

It is not just a civil offence. It is potentially an organized crime.

ben_w

2 hours ago

The actual case was literally pursued as a civil offence. "Potentially" is not a useful adjective.

The TLDR I've been given is that it's civil when the prosecution is a non-government entity (private person or company), and when the penalty is an injunction or a fine, and when the standard is "preponderance of the evidence".

Conversely, it's criminal when the prosecution is a government/when the sought penalty is imprisonment, and when the standard is "beyond a reasonable doubt".

negura

4 hours ago

> It's "move fast and break things" in action.

what would the world's reaction be if China's model did same?

myrmidon

4 hours ago

Is there any reason to assume they are not doing the same (ingesting books that they hold no rights to into training data)?

From the copyright holders point of view, it is simply much easier to prosecute western companies.

bsian

4 hours ago

There's no "hack". The article says the bots edited an open wiki website.

4563lk

4 hours ago

The legal framework is a DOJ and FBI controlled by the president and "allies" controlled by the "rules based international order".

In other words, the law of the jungle.

dist-epoch

4 hours ago

> The researchers also found efforts to tamper with the website itself. Lukasz Olejnik, a visiting senior research fellow at King’s College London, said this amounted to a hacking attempt. OpenAI disputed that characterization based on its analysis of the material Thursday.

of course OpenAI would say that, "oh, our model is so dangerous, it can hack into anything, be afraid, buy our IPO". it's just fear marketing

freehorse

4 hours ago

The text quoted though says that openai does not agree with characterising this as hacking.

user

4 hours ago

[deleted]

jsnell

4 hours ago

The "it's all just marketing" conspiracy theory is always totally detached from reality, but particularly so in this case. Your quote shows OpenAI is denying it being a hacking attempt, the opposite of what you say.

Spacemolte

4 hours ago

More like you found the exception to the rule..

krater23

3 hours ago

It's the reason why it happened in may and we hear now about it. It was not hacking and not important enough for marketing.

And just using a wiki and trying to embedd javascript is not hacking for me.

eithed

3 hours ago

Is it just me, or is it advertising? "Look at how smart our models are, they used this website to coordinate and share guidelines!"

krater23

3 hours ago

Reading the headline: WTF?! This is how Skynet started! Next year the mankind will die!

Reading the article: Oh, AI have learned to communicate over a wiki. OK.

Cynddl

3 hours ago

From the report:

> A few hours after they find the site, [the agents] start probing it for cross-site scripting (XSS) vulnerabilities.

user

3 hours ago

[deleted]