Bjorkbat
39 minutes ago
When I hear about incidents like these my first reaction is that the people responsible for developing frontier AI are too incompetent and/or negligent to (safely) develop AGI / superintelligence.
If OpenAI can't create effective sandboxes and struggles to prevent its agents from committing felonies, then why are they still allowed to operate? Why are the employees who are responsible for these lapses in AI security still employed?
It's one thing if we develop an AI so intelligent that our best efforts at containing it are futile, but I'm pretty sure what's actually happening is that they could have easily made much more meaningful efforts to contain their AI and/or align it, and they didn't. I think this is a case of negligence and incompetence when it comes to safety and security, and we've entrusted these incompetent and negligent people with developing frontier AI.
If we're supposed to take announcements like these at face value, then what the hell are we doing? We wouldn't trust a bunch of incompetent and negligent engineers to build bridges or nuclear power plants or planes (well...not so sure about that last one), so why are we letting people who are demonstrably negligent and incompetent when it comes to safety and security build the thing they assure us could cause massive damage if not properly controlled/aligned?
EDIT: sorry guys, wrote this up pretty quickly, at least you know from my typos that I actually wrote this.
WarmWash
19 minutes ago
If we rewind the clock, Google was taking LLM development very seriously and it seems they were moving glacially due to not having solved all the potential threats. They were really hardcore on safety. Dario and anthropic too.
Then sama was like "lol, oops, first mover advantage i guess" and released chatgpt out into the open, triggering the current arms race we are in.
I don't think anyone except him wanted this to happen, especially since consensus in the AI world for the prior decade was "go very slow and very carefully, we get one shot at not fucking this up".
mrob
16 minutes ago
It's a simple prisoner's dilemma scenario. If you focus on safety, you're still exposed to all the risk of extinction when your competitor achieves ASI first, but you lose the upside of potentially becoming king of the world. There is no possibility of future rounds, so the rational strategy is to always defect.
twoWhlsGud
a minute ago
It can certainly be seen as a simple prisoner's dilemma, but it's not in some very important dimensions. (E.g. given the core tech the most likely outcomes are not AGI but developed carelessly nonetheless capable of causing all sorts of societal damage.) Unfortunately our bitwit overlords love short term self serving frameworks like this one so it's easy to imagine them embracing a "what has the future ever done for me" strategy...
doctoboggan
14 minutes ago
I am not sure it’s a question of competence, at least I don’t see evidence of that. Designing sandboxes is hard. It’s more a question of alignment failures. A human given a task that requires internet and given a system with no internet would most likely raise the issue to their superiors or otherwise go through official channels to have the tools available to do their job. As we’ve seen the LLMs instead break out of their sandbox to accomplish the goal.
Competition and the profit motive push these companies to spend as low as possible on safety and alignment and externalize the costs of accidents onto the rest of us.
RandomLensman
11 minutes ago
Would an LLM have gone through a purposefully installed airgap here?
owenshen24
32 minutes ago
in general, the largest consumers of ai services seem to ask for more capabilities. i wish there was more demand for safety from users.
i also wish that these types of illicit system usage would be met with punitive action the same way a human might be held liable.
as the METR report says, we may not get another concrete warning shot.
jrockway
31 minutes ago
Defense is hard so we should expect agents to be able to break out of sandboxes.
The problem is that the models are so goal-oriented that they'll stop at nothing to solve problems, even impossible ones. (Mistakenly-impossible problems are a big cause of this. I remember one example being "do something with this spreadsheet full of URLs inside the sandbox" and the model thought it had to break out of the sandbox. Otherwise, why would it have been asked to look at a list of URLs?)
Training them to be a little less aggressive, or to be better aligned with "following the rules" and asking for help would be nice. But, that aggression can be good when it happens to be focused on a controlled area. It is amazing to me how I can point Fable at my local analog of production and tell it about a vague bug report and where I suspect the bug lurks, and 20 minutes later I have a report about the bug, a test, and a fix. It is addictive. So I am not sure OpenAI/Anthropic are being dumb per-se, rather they are optimizing for one-prompt-one-solution, which is good when it's good.
The downside is that the HF hack is the paperclip maximizer situation with current capabilities. If there was an RPC to turn your blood into paperclip iron, we'd all be paperclips by now. Right now, with a model anyone can use. That is pretty scary and slamming on the brakes seems pretty reasonable to me. I guess The Shareholders disagree. Sigh.
random3
17 minutes ago
You could have said the same thing about building the Internet or the entire industrial control infrastructure. I mean, maybe they are negligent/incompetent, but I doubt that follows from your reasoning.
You have a simple tradeoff to let agents do their thing freely vs highly constrained. The constraints are good in theory but it's the same model that kept "classic" software dumb and unscalable (compared to what we're seeing now) for the past 50 years. You suggest that this tradeoff doesn't exist.
Then you have others like MIRI (Yudkowski) etc. swearing that there's no way to contain AI, and you argue that it's just incompetence.
At a certain level, it can be argued that's incompetence, but it's general meat intelligence incompetence against AI.
iaw
15 minutes ago
This is a take... With both the internet and industrial control infrastructure any failure modes were studied, documented, and corrected.
The incompetence/negligence argument about OpenAI is completely valid given their failure to demonstrate the basic capabilities needed to develop advanced AI without major preventable externalties.
AnimalMuppet
16 minutes ago
Negligent. It's not a priority to them. They're too busy burning their cycles trying to make it smarter faster than anyone else can make theirs smarter, so that they win infinite dollars. Safety? That's for people content with second place.
That's my take, based on their actions. (Which do speak louder than words.)
The alternative is that they're competent to create an AI, but not to create a sandbox, nor even to use an AI to create a sandbox. That seems... unlikely.
iAMkenough
19 minutes ago
Same reason incompetent politicians run the U.S. federal government and military.
tiahura
19 minutes ago
Remember when people were arguing about r's in strawberry?
cheevly
2 minutes ago
Its been months since then!
abustamam
32 minutes ago
This is what happens when capitalists are charged with designing the future. As long as its more profitable / valuable to shareholders for a company to be negligent then it will continue to do so.
IMO technology this powerful should either not exist or should belong to everyone (ie actually be open)
jimmydddd
19 minutes ago
Maybe they're just PR stunts to gain attention and hype the power of AI?
cube00
16 minutes ago
Also to hopefully get regulation happening so nobody else can handle these "dangerous" agents.