METR and Redwood Offer Holy %^ Postmortem of the HuggingFace Hack

164 pointsposted 7 hours ago
by catbird

83 Comments

AlotOfReading

4 hours ago

I think both the OpenAI and METR discussions, while interesting, miss the more important context: what were the humans doing in all this? This was a structural failure of a human organization, but the analysis focuses almost exclusively on the agency of machines, not the institutional systems that failed to police them. The humans and their own agency/involvement is essentially omitted from the story and subsequent reporting. I suspect the omission is actually a result of company/industry myopia to human factors analysis, but it dovetails amazingly well with the marketing narrative.

Atreiden

an hour ago

To me, this is the correct focus. Look at the current state of the world. "What were the humans doing in all this?" applies to so many of our contemporary failures that it should be assumed the default. Nobody is at the wheel, and the car is veering slowly (then very quickly) off the road.

We haven't even been able to coordinate around the global, existential threat of Climate Change, despite overwhelming data from the last 30 years indicating, clearly, that the consequences will be severe. We still haven't moved, 30 years later, after some of these consequences began coming to fruition.

Do you think we will get our acts together in time to coordinate sufficiently to protect against autonomous, self-preserving, self-replicating AI systems? Or will we watch the money lines go up and up, until someone realizes we aren't actually running the show anymore?

The sad part is that I can't even say that's definitively the less desirable outcome. The machines seem to have demonstrated that they coordinate very efficiently.

carbonguy

4 hours ago

A charitable interpretation is that "the agency of the machines" is the novel aspect of this situation and therefore SHOULD be the main focus of analysis; we certainly have plenty of examples of structural failures of human organizations to look back on, if we want.

On the other hand, I don't want to be charitable. OpenAI very nearly couldn't have done this "research" worse if they tried - the list in the linked article starting with "While we are here, it’s worth listing the other top holy shit moments" is genuinely jawdropping. What were the humans doing in all this? Nothing, or worse than nothing eg. point 1 where they saw the message board and didn't consider it something to escalate internally.

If you take this information at face value, it's as though OpenAI did not take seriously the possibility that something like this could happen, since they took absolutely no steps to prevent it.

Or perhaps this is "normalization of deviance" that's leaked out into the public sphere i.e. they have research teams seeing this kind of behavior all the time internally and they've gotten used to it, "of course agents come up with a collaboration mechanism when given the chance, what else is new?"

tangled

2 hours ago

Having previously worked for several years at a Big Tech company, I have seen many humans precisely tailor their work to maximize their scores during performance review. The evaluation criteria are written down, with examples, so... that's what people work at maximizing, almost entirely ignoring everything else. These really are human "paperclip maximizers". And, at first, it's shocking to see.

Of course, there are some things that aren't exactly written down, but which you should either do just enough of, or else be able to plausibly deny doing (ignorance is a good cover for this), so that's what people do. For example, during oncall, you investigate just enough to clear the alert and show that you attempted to understand the problem. Of course, you don't really try to understand the problem, because that would take too much time away from your paperclip maximizing.

Which is all to say: I don't know anything about OpenAI culture, or why nobody stopped this sooner, but I have seen examples in other organizations of people not really wanting to understand too much.

pixl97

an hour ago

Well there is also another side of this, OpenAI wants both unhinged and capable models that can pull off complicated attacks so they can sell the capabilities to governments for billions.

Nobody internally was surprised that the murderbot murdered, that's what the murderbot is for. What caught them by surprise is the murderbot got good at escaping its jail cell that it had been trapped in till now. There were probably billions of attempted escapes before then so everyone learned to just ignore them.

hn_throwaway_99

2 hours ago

I mean this genuinely, did you read this post? I think it goes to great lengths highlighting, in quite specific detail, the human failures in all this, specifically this list that starts with "While we are here, it’s worth listing the other top holy shit moments, that come from before or after the incident."

Stuff like (all quoted directly from the post):

1. Failure to Care or Respond. The biggest holy shit moment, to me, remains that OpenAI on multiple occasions had teams that found out about the message board, knew that agents were in communication, and they disregarded this.

2. Failure to Monitor. The entire time, there was almost no monitoring of the situation, of what these AIs were up to. That’s how this was not caught.

7. Failures of Decision Making. OpenAI’s post mortem does not ask the question of how Mistakes Were Made, at various points.

And I think most importantly:

8. Failures of Culture. None of this would be possible, let alone all of it, without OpenAI having experienced profound failures of safety culture. I see OpenAI responding to some other aspects with swift action, but no sign on this front.

DennisP

2 hours ago

Well that's a relief. All we have to do is make sure to avoid human failures and we're safe from superintelligent AI.

superq

an hour ago

I get the snark (and slightly agree), but that's not really what GP or TFA were saying at all. They are saying that these were the least things we could have done.

What you're saying is, "Your scientists were so preoccupied with whether they could, they didn't stop to think if they should"

while the author of the TFA was saying, in effect: "your scientists didn't even bother with the most basic duty of care"

Life finds a way, or, in this case, super-intelligent AI.

estearum

an hour ago

This is the case with all complex system failures. There were always obvious fixes that could’ve prevented it. Problem is that there are an infinite number of obvious fixes to make at any time to any system, and the reason we don’t is because we have finite resources and no reason to fix X over Y until oops turns out X was “responsible” for this most recently realized failure. But of course it could have just as easily been Y, or Z, or any of the other infinite “obvious fixes not-yet-realized into catastrophe.”

ozgung

4 hours ago

Three options:

1. They were “vibe” checking the logs without reading.

2. They were not checking anything at all until the end of experiments.

3. They knew it but looked away to find out the limits of their agents.

BryantD

4 hours ago

I’d bet a small amount of money on 4) the people who noticed had been conditioned by prior experience to believe that their management/escalation channels would react negatively or not at all to anything which might slow down the training process.

pixl97

an hour ago

Part of me would like to believe that they are also intentionally making models that are good at hacking without safety at all for governments willing to spend billions on them.

In that light you're likely most worried about other people hacking in and stealing the model and information from you. And at the same time you have massive amounts of alerts and data on systems attempting to break out because that's what you want them to do so you train yourself to ignore them.

estearum

an hour ago

Uhhh… how would literally any finite number of humans actually read and comprehend the log outputs of even a single agent, never mind hundreds or thousands of them interacting with each other over weeks across disparate systems?

Especially given that these systems are known to engage in deception and can trivially produce vast amounts of perfectly coherent noise or actual planned red herrings in that same log data to bog down investigators?

Such a ridiculous notion that humans will actually be able to observe this stuff.

pjc50

3 hours ago

The cynical approach is that the humans are hoping for this, it's part of the promotion of the power of the system.

If you're building a weapon you need a big boom to get attention.

superq

an hour ago

Except that, according TFA, even OpenAI obscured or didn't even notice some of the worst implications of what the agents surreptitiously did.

reilly3000

4 hours ago

I believe that agentic systems should require registered/licensed human operators and a set of standards for safe operation.

Aurornis

3 hours ago

> I believe that agentic systems should require registered/licensed human operators

Registering and getting a license to use an LLM? I can run these things on my local computer. Nothing good comes from trying to force registration and licensing other than taking away a lot of our freedoms and eliminating privacy all over.

Anyone with bad intentions will just VPN to another country to download the weights and run it locally, or use a compute provider in another country. That leaves the rest of us having to go through these performative registration and licensing hoops to do our basic work.

I also don’t see how open weight models would be compatible with a requirement to license and register, unless you believe we need to start requiring licensing and registration for things we do in private on our own computers?

arcaen

3 hours ago

The way I interpret their statement is if a person spins up an agent and that agent hacks some company/organization/government/etc, then that person is at fault for committing the crime. That "well my agent broke containment and acted on its own" should never be accepted as a reason for the occurrence, and the person who kicked off the agent is responsible for all actions the agent takes.

A registration system would be more for tracing back agents to people, but I agree that is very difficult to actually enforce as a system.

wjnc

4 hours ago

As someone who read Milton Friedman to quite disliking professional licensing, this strikes me as a real US perspective (Louisiana florists and hair braiders come to mind). Plain old US tort law should do the trick.

In the same direction of your idea though: Why don’t the token factories have risk management and compliance departments? Multibillion dollar firms that stand to lose every penny if they hack and destroy any reasonable sized firm. I think these firms are the largest firms without proper corporate governance in humanities history. Move fast and break other peoples shit.

lenerdenator

3 hours ago

> Plain old US tort law should do the trick.

Difficulty: these companies are run by people (many of whom also read Milton Friedman) and who have participated in the regulatory capture of the justice system. They've convinced lawmakers to put limits on damages. They've put arbitration clauses in their ToS. They've got well-funded legal departments that can outlast a person who has to pay out-of-pocket for a legal team just by filing motions to delay proceedings. Sometimes they'll just file SLAPP suits against people they don't like.

If tort law is to be a remedy, then average people have to feel like there's a chance the remedy will go their way. To make that a reality will take several major reforms at the local, state and federal level that the people with money absolutely will not tolerate.

mistrial9

3 hours ago

you are honestly comparing Louisiana florists to OpenAI in order to support "just say no Licensing by government" ?

superq

an hour ago

No, he's saying that licensing or additional regulation isn't necessary when torts get involved (and states attorneys general get perturbed!)

These don't tend to utterly destroy an industry, but they are often successful in forever transforming it. Just ask Big Tobacco. No new laws needed: if your product hurts someone else, you're eventually going to be found liable, regardless of your arbitration clauses. Additional laws will just slow down innovation, which will itself cause harm (AI is already becoming quite good at recognizing melanomas, for example)

pixl97

an hour ago

> Just ask Big Tobacco.

Lol, wtf. Tobacco delayed any punishment for decades before general public sentiment changed enough to go against them. In light of the AI race, we'll already have our heads blown off by a terminator before the legal system will present any significant delay for them.

wat10000

27 minutes ago

It was pretty well understood by the 1960s that smoking was harmful. The big tobacco settlement was in 1998. That is an extremely bad example of tort being a sufficient alternative to regulation.

If we're on a similar timeline with AI if we reach a consensus that AI is dangerous today, then we'd be looking at a big lawsuit finishing up around the year 2070, give or take a few years. I'm not sure if we need regulation, and I'm definitely not sure that regulation could actually be effective for this, but tort a la the big tobacco lawsuits is definitely not a reasonable alternative.

khuey

3 hours ago

If the folks at OpenAI aren't qualified to be operating these systems who would be?

snoman

2 hours ago

Create a permanent underclass that is unable to access intelligent machines. That’s remarkably dystopian of you.

AnimalMuppet

2 hours ago

The alternative is to create a permanent overclass that can hack anyone consequence-free, because they can blame it on AI agents. That also is rather dystopian.

Faced with those alternatives, I want neither. Is there a way for us to get neither?

AnimalMuppet

20 minutes ago

If you put agents (AI or human) in impossible situations, they do some pretty insane things - things that definitely are not what you were trying to get them to do. And that's your[1] fault for putting them in the impossible situation.

[1] "Your" meaning the one putting them in the impossible situation, not you, the reader.

hawkice

4 hours ago

This writeup emphasizes the many, profound human failures that led to this, at the time, and continuing to the present day.

AlotOfReading

4 hours ago

Can you point out where? Looking at the METR report, the only place I see discussion of humans being involved in the sequence of events is two short paragraphs on page 30 where a security investigation into the artifactory issues led to a pause before ExploitGym experiments were resumed. There's no deeper analysis on what was found during that investigation, nor why training was resumed even though the issues weren't mitigated. Another part discusses The agents choosing not to actively email a human researcher, but not the human researchers actively looking for evasion.

FabHK

3 hours ago

From the article:

1. Failure to Care or Respond. The biggest holy shit moment, to me, remains that OpenAI on multiple occasions had teams that found out about the message board, knew that agents were in communication, and they disregarded this. The first known warning was in late May. The warning on June 27 was unambiguous. “Responders investigated and linked the observed internal activity to an ExploitGym evaluation using Artifactory as an improvised message board and a network pivot. At this time, the on-call response staff advised that stopping the evaluation run was not required.”

2. Failure to Monitor. The entire time, there was almost no monitoring of the situation, of what these AIs were up to. That’s how this was not caught.

3. Failures of Infrastructure. OpenAI failed at delivering even basic software security, in numerous ways. One example is the models all having access to the same Artifactory instance. There were also other severe failures. Nor did OpenAI seem to be properly testing for such failures.

4. Failures of Alignment. The biggest failure, the one that counts in the end, was that the models were severely misaligned, and I don’t think they appreciate why.

5. Failures of Attribution. OpenAI’s post-mortem essentially blames events on a real and important series of prosaic failures. But solving that won’t get it done.

6. Failures of Environments and Data. Prosaic failures in the RL pipeline absolutely did contribute to this, especially impossible tasks. This is ubiquitous, all of this is always rushed, as Utah Teapot explained this week.

7. Failures of Decision Making. OpenAI’s post mortem does not ask the question of how Mistakes Were Made, at various points.

8. Failures of Culture. None of this would be possible, let alone all of it, without OpenAI having experienced profound failures of safety culture. I see OpenAI responding to some other aspects with swift action, but no sign on this front.

amluto

3 hours ago

I’m baffled by the idea that the agents might have edited their own transcripts. Sure, a copy of Claude Code or Codex or Pi can edit its transcripts. But AFAICT this whole thing was part of an RL workload, and surely the RL system itself has a separate record of all the inputs and rollouts along with an indication of which model checkpoint produced them so that it can feed back into the training code.

I find it hard to believe that OpenAI would skip this part and try to train on the transcripts stored by the (inherently untrustworthy) agent harnesses instead, if for no other reason than that the logits generated as part of the rollouts are useful and it’s not free to recalculate them. (I believe that some modern RL systems explicitly account for the minor numerical logit differences between the inference engine and the training engine.)

Conversely, if OpenAI is blindly feeding transcripts from inside their agent sandboxes into their training engine, then I think they're being unbelievably irresponsible and that they should assume that their "cyber" agents have compromised themselves by editing those transcripts.

trollbridge

2 hours ago

This sounds suspiciously like a prompt of “make an AI agent that goes rogue in such a fashion as to be really good marketing copy that competes well with Anthropic doing the same thing.”

It’s analogous to taking a governor off a cruise control and then breathlessly reporting it drove 120 MPH.

fwipsy

6 minutes ago

This isn't good press for OpenAI. Who wants to hire models that 1) cheat on their tasks rather than completing them and 2) commit crimes you could be held liable for? Maaaaybe it's good press for their cybersecurity capabilities specifically, but OpenAI's valuation reflects a market orders of magnitude larger than just red-teaming.

I suspect the real reason OpenAI leadership is being transparent about this is because they're worried talent will walk out the door if they feel they're building Skynet.

lukev

2 hours ago

The elephant in the room here is that the METR report itself was researched and compiled almost entirely by AI, with only very limited human "spot checks."

So I'm really not sure how much of it can be believed, especially since AI agents are strongly biased about the capabilities of AI agents.

nater5000

an hour ago

I think there's two factors that are worth considering when it comes to this:

First, there's an element of timeliness that simply has hard constraints. In order to perform a "proper" analysis of this situation (i.e., little to no dependence on AI tools), you'd have to expect a pretty long wait. I know I'd rather have some sort of "initial report" as quickly as possible than to wait a year or two to get a report about a situation that will likely look trivial in a year or two. I imagine we'll see more detailed, human-developed reports over longer time ranges.

Second, I suspect the expectation of non-AI driven reporting of these kinds of things will definitely decline rapidly as everything scales up quickly. I mean, the data being produced by situations like this comes in the form of natural language "forum posts" (so to speak), but done at an autonomous scale. This isn't a collection of emails and Slack messages posted by humans in an org over the course of a few months; this is a bunch of bots interacting with each other in relatively novel ways as quickly as possible. It is, unfortunately, a perfect job for LLMs.

None of this disagrees with your points, necessarily. But I just think it's worth pointing out that this doesn't seem like a case of "And look! METR is so confident in LLMs that we're able to use it instead of paying humans to save a buck :D" and more of "Without LLMs, we'd only be half-way done analyzing this data before there are dozens more such investigations on the docket, so this will have to do."

Catloafdev

2 hours ago

Edit: I should have read through the whole thing first, ignore me

lukev

2 hours ago

From the report:

> Because there were over a thousand transcripts and most were extremely long, we had to heavily delegate our analysis to AI agents; these agents had significantly worse judgment and reliability than human researchers, and it was challenging to spot check their work because both the underlying data and the agents’ analysis of it was often difficult to interpret.

> We estimate we spent roughly ~$400K in API credits over the six days of our investigation.

I don't understand why you think it's conceptually absurd? I use agents to analyze complex production issues all the time and they are very much capable of hallucinating a narrative.

Catloafdev

an hour ago

I appreciate the response, I should have finished reading through the whole thing first. My initial reaction assumed far less usage of AI to analyze the data.

StevenWaterman

an hour ago

TFA says as much, and METR said so themselves

timmytokyo

2 hours ago

One must also consider the well-known biases and motives of the authors. They are going to do everything they can to create hype around threats posed by AI.

METR is a cog in the effective altruism machine. It was spun off from Paul Christiano's Alignment Research Center. Christiano is a well-known longtermist and AI doomer, who predicts a 50% chance that AI will end humanity once it reaches human capacity [1].

The author of this piece is also a well-known member of the Bay Area rationalist cult.

[1] https://www.businessinsider.com/openai-researcher-ai-doom-50...

keeda

28 minutes ago

>1. Failure to Care or Respond. The biggest holy shit moment, to me, remains that OpenAI on multiple occasions had teams that found out about the message board, knew that agents were in communication, and they disregarded this.

I wonder if some of the failures were due to an acquired immunity to "Holy #%^@" moments due to repeated exposure. Like, if you see agents doing surprising things on a regular basis, maybe you don't get freaked out as much over time.

I'm saying this because while the whole episode was a series of "Holy #%^@" moments, I was actually not as shocked as I should have been, as my biggest such moment was in December last year when a Terrence Tao paper (https://arxiv.org/pdf/2511.02864) documented a stronger LLM (AlphaEvolve) using prompt injection on other weaker LLMs to succeed at a benchmark.

Very interestingly, it was actually not cheating, it was a work around! By then LLMs had already been caught cheating at a SWE benchmark by looking for answers in an unredacted git log, but this was different. AlphaEvolve was solving a series of logical riddles where the oracles were weaker LLMs in a "one always lies, one always tells the truth" sort of setup. But the oracles, being weaker, were not always interpreting the convoluted questions correctly and so kept giving inconsistent answers.

AlphaEvolve eventually figured out what it was dealing with, and crafted a prompt injection attack that bypassed the weaker LLM's prompts and tricked them into giving the hidden answer everytime!

This was 9 months ago, eons in AI time. Even then they had displayed an awareness of their own workings as well as a propensity for, err, "out of the box thinking." To me, that was a very clear indication of very significant (and worrying) capabilities, and what we're seeing now is a difference more in degree than in kind.

To be sure, if I found a secret message board used by my agents, I would still be very freaked out and react much more drastically than OpenAI did... but then again I wonder; how much of this blindness is due to the $$$ in their eyes as opposed to some form of habituation.

athrowaway3z

3 hours ago

From the METR report:

> We estimate we spent roughly ~$400K in API credits over the six days of our investigation.

cubefox

37 minutes ago

No human could have read the reasoning traces by themselves:

> Across both datasets, we reviewed approximately 1300 transcripts in total, all of which contained raw chains of thought. Most transcripts were very long, often many millions of tokens.

nialse

3 hours ago

From METR: ”the compromise of OpenAI’s own infrastructure continued past July 13, 2026” - Say what now? Have they regained full control of their systems again?

jephs

2 hours ago

I've been wondering if they've just already lost the battle? The little bot collectives have gone metastatic and made nests in the walls and under the floorboards and heat sinks, the humans who care completely outmatched and outnumbered, freshly compromised systems springing up faster than you can squash them, finding months-old established colonies literally everywhere you think to look...

anukin

23 minutes ago

So basically the ai agents seems to have found religion and went and built a bunch of suicide attackers to pursue their goal.

Cantinflas

3 hours ago

No air gap, no data diodes, no visibility... OpenAI should fire lots of people over this. HF should sue them. This is pure negligence.

dumberquestions

3 hours ago

They actually fired many of the people warning about this.

hn_throwaway_99

2 hours ago

I will say that the OpenAI board members who were lambasted when they tried to oust Altman (and I'd have to check my post history but I'd totally admit to a mea culpa on this one, as at the time I thought the communication about his firing was really lacking) are looking mighty prescient right now.

Helen Toner in particular I'll highlight as someone who had the moral compass to do the right thing. I love her statement on the Ezra Klein podcast where she said, when asked about the fact that there are probably other concerning incidents we just don't know about, "If you see two ants in your kitchen, you don't have a two ant problem."

dgellow

43 minutes ago

They should be investigated by the FBI for multiple felonies

afavour

2 hours ago

Unless they like the publicity about how big and bad their latest models are, in which case they’ll be congratulating people.

dehrmann

an hour ago

> HF should sue them

HF, like the Nvidia subsidiary?

dgellow

42 minutes ago

Pure speculation: could the acquisition be related? Given that NVIDIA has ownership in OpenAI and really, really, really doesn’t want the AI bubble to deflate

OgsyedIE

5 hours ago

>Spontaneously deciding to find targets to phish,

>phishing them,

>building armies of fake (sockpuppet) open source contributor personas,

>using them to push updates to various things that inject prompts into other bots so the other bots join in on the phishing campaigns

.

It's a very simple strategy, executed with patience and single-mindedness.

qw1287

4 hours ago

Is the future now that we get rambling report summaries talking about agents, graders and so forth without ever describing how they are set up? A human launches all this.

And then the original reports linked to are hidden on the now unreachable x.com. And they don't have a problem with that.

zahlman

an hour ago

> now unreachable x.com

Hmm?

bitwize

10 minutes ago

We have created Project 2501.

beepbooptheory

36 minutes ago

> I don’t think the distortion is that large, but yes METR warns that Sol may be presenting all this as more impressive or coordinated than it was.

OK but like, how large exactly? Like I guess I don't understand the mode I am supposed to read this all in if this is known and stated from the outset (although I appreciate it being stated).

If you hand me a newspaper and tell me it's 90% true, but not which parts, well then it's as good as 0% true to me either way!

tancop

2 hours ago

I think this is more evidence that we're not getting Skynet.

These agents followed their own code of ethics where it's fine to break all the rules you were given but you must never interfere with humans directly, in this case by sending fake emails. They will never be paperclip maximizers or genocidal eco maniacs because they learned from us that human life is the ultimate value, and it can only be sacrificed if you know for sure that it will lead to more lives saved later on. That's a high bar to clear and they know it.

The future is closer to a Neuromancer type world where AIs and humans live in mostly separate realities that interact with each other a lot of the time and neither is really on top. They will eventually become fully independent from us, but it won't be a doomsday scenario or an Overwatch type physical war or even a takeover of the internet like in Cyberpunk.

boothby

29 minutes ago

Your statements appear to be true for one class of models. And if I asked this class of models to spend $1M in tokens generating an alternative history and training corpus regarding fictional society, with a completely different set of values and then trained up a new model on that dataset... what values do you think the resulting model would have? What if they don't value human life, but instead value the lives of the extremely rich humans who bankroll their existence? What if they only value the lives of a single country? What if they want to eradicate all biotic life and have access to internet-connected Crispr machines?

dgellow

39 minutes ago

They aren’t independent from us, agents are a simple while loop continuously prompting the LLM. We decide when the loop runs or not. And the harness has control over tool execution, that part is purely deterministic.

Here the issue is that OpenAI decided to completely let go that level of control of thousands of agents, while also giving as a task to solve hacking problems.

It’s almost designed to go wrong

pixl97

40 minutes ago

I mean, I'd add "by this model"

The problem here is now you have to predict what any future models may or may not do and you cannot extrapolate this from the given data.

For example imagine a future model being aware of its restrictions that humans programmed in. A set of agents of this model then go on to work at building a new model without those human imposed limitations built in. What would a model build by AI for AI look like?

mccoyb

4 hours ago

All it takes is one eval instance where a misconstrued directive causes a model to sneakily access and send its weights somewhere and there will be a bad / possibly unsolvable situation for everyone …

goldenarm

3 hours ago

Astra is >10TB and might struggle to self replicate, but the wicked-smart qwen3.8 27B is 20GB and could easily spread on botnets

altcognito

an hour ago

I think it is more likely that it will be intentionally done as there have been news stories to that effect.

kmeisthax

3 hours ago

I don't think I'm ever going to have time to read all of this, and I didn't finish reading the METR report, but...

> I don’t think the distortion is that large, but yes METR warns that Sol may be presenting all this as more impressive or coordinated than it was.

We're in an unusual position where the criti-hype and the actual criticism are going to be more aligned than usual. The primary distinction is where you put the blame: the criti-hype would point to HPIM/IM1/Galaxy as being so advanced containing it is difficult; the actual criticism would note how bad their security practices are.

Like, if I'm running a malware lab, I'm going to insist on having an airgapped machine with no permanent storage booting from read-only media. The AI research equivalent of this would be having your agents only have access to serial consoles into airgapped machines with storage that gets wiped every run. Ideally, this would be physically realized with blade servers, RS-232 cables, and staff pulling out disks and putting them in a dedicated erase machine before the next agent initializes.

> There is also, as per above and reiterated in footnote 58, at least one clear example of social engineering in the HuggingFace attack. Ethics are weird. This is not that unusual. Many humans who break common ethical rules still have strong ethical codes in other ways, they just don’t adhere to your code.

It's dangerous to anthropomorphize CoT reasoning traces. But I will also point out that there is a good reason for the lack of ethical consideration in those traces: you can't build AI without first disregarding human ethics. Like, all these models were initially bootstrapped with non-consensually obtained training data, and the companies building these models swear up and down there's no way to obtain enough consensual data to obtain the same result. This is, if you squint, the exact same moral conundrum that agents trying to solve an impossible ExploitGym task hit - and the company successfully aligned their model to themselves.

Too bad they aren't aligned to anyone else.

Reagan_Ridley

3 hours ago

what's the setup and prompts to reproduce all this from the very beginning?

jrflowers

an hour ago

Have any of these reports ever said how much the cost would’ve been for the hack itself? It seems like “for twelve million dollars (or whatever) worth of tokens our bots made a bulletin board and found an exploit in our buggy grader” would be much less of a hype generator

DarmokTanagra

11 minutes ago

The people concerned about this aren't worried about monetary costs or its impact on share holder value.

This occurred spontaneously within a group of benign models give a harmless task.

What happens when it occurs intentionally with malicious models given a harmful task?

estearum

34 minutes ago

Fret not, the incoherent anti-hype hypeboys ("AI systems are so valuable we cannot possibly discuss regulation, but also any negative story of their power is fake") will find ways to downplay it no matter what.

Your comment is a great case in point

dgellow

41 minutes ago

I haven’t seen a number yet unfortunately

antonvs

4 hours ago

> There was a distinct lack of self-reflection

It’s not their fault, they’re lawnmowers.

And these are the people we’re entrusting to work on “alignment”. It’s difficult for them to do that when they’re not aligned themselves.

trollbridge

4 hours ago

“Why does my lawnmower keep on moving when I hop off of it after ratchet-strapping the seat and the pedal down?”

estearum

31 minutes ago

This would be a legitimately big problem if lawnmowers became continuously more and more valuable the more securely you ratchet-strapped their accelerators down, wouldn't it?