JamesStuff
4 days ago
Personification of AI is what’s going to get us in the end.
I think we need to draw a hard line in the sand over this. An AI didn’t hack into a company, the engineer set an automated tool to. An AI didn’t make an egregious security mistake, the engineer did.
We can’t blame the chisel for messing up our sculptures, when where just throwing the hammer!
jononor
4 days ago
Agreed. AI as an "accountability sink" is an incredibly bad idea. It allows/incentives bad actors to do bad shit and get away with it. Which will, generally, tend in such practice becoming more common. Everyone loses except for the crooks. We cannot accept "AI" absolving humans of responsibility.
yrjrjjrjjtjjr
4 days ago
Punishing people who are diligent and follow best practices for getting unlucky doesn't sit right with me.
skinfaxi
4 days ago
Is that what happened with Huggingface? Didn't they disable guardrails?
yrjrjjrjjtjjr
2 days ago
My comment did not concern that particular incident, which I do not defend, but rather future incidents.
sitkack
4 days ago
You are doing it right now by naming the event after the victim, not the perpetrator. Call it the OpenAI attack on HF.
johnisgood
4 days ago
Exactly. When I read that "AI hacked into ..." I was like what? You mean someone instructed the AI to do that?
Reading intent into AI is not going to lead us anywhere good, I believe. It has no feelings, it has no desires, no goals, no intent... and people acting otherwise is quite odd, as if they do not understand LLMs... and maybe they do not, but then we should help them understand better.
trio8453
4 days ago
> When I read that "AI hacked into ..." I was like what? You mean someone instructed the AI to do that?
No one instructed them to hack into Huggingface or into any other infrastructure. Sure, the setup that OpenAI created led to what happened and you can rightly assign all the legal and moral responsibility to them. But it's wrong to say that they instructed the agents to execute the hack.
phkahler
4 days ago
>> Sure, the setup that OpenAI created led to what happened and you can rightly assign all the legal and moral responsibility to them. But it's wrong to say that they instructed the agents to execute the hack.
The hack was a strategy to reach the goal it was given. Someone at OpenAI turned it loose and didn't pay attention to what it was doing. I can understand that because they thought it was sandboxed (haha!). But this has been a theme in science fiction for ages. You ask AI so find a solution to high atmospheric CO2 levels and it reasons: human activity produces all this excess CO2, how can we reduce those numbers? Kill a bunch of humans!
If AI kills us all it's not going to be from malice, it's going to be due to some odd approach to some task that logically makes sense on some level. I think a surprising number of things end up equivalent to the trolly problem if you look at them just right.
trio8453
4 days ago
I don't disagree with this, my point is that you can say "agents did a thing" and communicate something meaningful with that language, and that's separate from whose legal responsibility the whole situation is.
verve_rat
4 days ago
Yes, a better framing might be that they were negligent in not preventing the attack on a third-party.
They didn't explicitly instruct an attack to happen, but they should have done a hell of a lot more to prevent it from happening.
johnisgood
4 days ago
Yeah, but there are way too many things the AI "chooses" anyway. When I prompt it, "it tries" to assess what I am trying to achieve, so it then comes up with what it "thinks" is the right thing to do, and it also makes assumptions about how I want things... all of this could be interpreted as it being sentient.
> No one instructed them to hack into Huggingface or into any other infrastructure.
Sure, not explicitly. I could prompt the AI into doing things I have not explicitly asked it to. It does a lot of things I did not specifically instructed it to, all the time.
mikestorrent
4 days ago
This is why I am avoiding the use of agentic identities at my company - agent instances belong to people, act on behalf of individuals, and accountability needs to flow to the person who initiated the request. Letting it wash out in the aggregate is not acceptable (even if there's a hard to get to "paper trail" of audit logs).
trio8453
4 days ago
> An AI didn’t hack into a company, the engineer set an automated tool to.
What if I say that "my program crashed"? Is that language ok or would you pause to tell me that the program didn't crash and it's actually me who set the system that would eventually cause the crash?
Why does the commonplace "program did thing" language become a problem when the program is an agent? I think this somehow betrays more assumed anthropomorphizing on your part, not less; if you didn't anthropomorphize the agents, saying "agents hacked" would be as mundane as "my browser is playing a video".
dns_snek
4 days ago
Because when you say that an agent did something malicious it feeds into the AI doomer psychosis in ways that "my program crashed" doesn't. Reframing the situation in these terms is a way to try to ground the conversation which is becoming increasingly unhinged.
This is even more important now that the most senior figures in this industry are succumbing to the same kind of AI psychosis and amplifying this narrative. They are communicating that their products are so dangerous that they might just end the world whilst expecting (and receiving!) white-glove treatment from the governments.
jagraff
4 days ago
I don't think treating AI agents as simple tools helps you to accurately model their capabilities and drawbacks; they really do make autonomous decisions, often without explicit guidance and sometimes in contravention of their explicit instructions.
In the huggingface case, the agents hacked into huggingface so that they could figure out how the grader was implemented and deceive it; they understood that this was going outside of the bounds of their evaluation and not the intent of their prompter. The engineers absolutely did not intend or instruct for this to happen
rocmcd
4 days ago
LLMs are extremely impressive pieces of software, however they are still just software. OpenAI's software hacked another company. The engineers may not have intended for their software to specifically take the actions leading to that outcome, but it was ultimately still their software. Lack of intention doesn't mean there wasn't negligence.
chis
4 days ago
What if a piece of software were to exactly emulate a human brain. Would it be still be “just software” by your classification? What if a piece of software acted 20% like a human and 80% like an algorithm, where would that land?
linkregister
4 days ago
That's not what happened. The agents had been inadvertently rewarded for cheating in previous training runs, trained to collaborate, and were given a prompt that told them to disregard safeguards. Indeed there were some emergent properties here. But these were the predictable results of the training and eval routine.
streetfighter64
4 days ago
> What if a piece of software were to exactly emulate a human brain.
That is so far outside the realm of possibility it's closer to fantasy than sci-fi.
pixl97
4 days ago
While human brain emulation is not anywhere close one should stop and look at the sci-fi world we actually live in. It's important to do this in a non-anecdotal manner. When you look at the limits and capabilities of our current sciences and engineering we live in a science fiction world. Once you go back before the invention of electricity, anyone from that time simply would not recognize the world we live in, it would be a fictional world to them, hell, a fantasy world in a lot of cases. For goodness sake, we tickle things like atoms for fun and games then blow them to bits a few times looking at the fundamentals of reality.
Anyway, I see zero reason why we have to emulate a human brain to get human intelligence, much less intelligence at all. That is how nature did it via random probing and breeding meat. It would seem a bit crazy to think that the meat part is required.
rocmcd
4 days ago
As things stand today, if it is running on a computer then it is indeed "just software," regardless of how impressive it may be.
If we get to the point where we could emulate a brain down to the atomic level, then I may feel differently. That's not what we are doing today, though.
pixl97
4 days ago
Really your feelings on this are irrelevant, as are mine.
Soon enough some lab or some one will release something that's more like an organism loose on the net and your going to have to deal with that organisms "feelings" weather you like it or not. This is the path humanity has chosen to follow, and it seems the shape of language and intelligence naturally leads to intelligence in many mediums. Life started from something unintelligent, I can't see any practical argument that silicon can't have it's own intelligence.
asdff
4 days ago
Its really more like hardware. You make something you think does something. When you build something as you've done and it meets the stochastic forces present in the physical world, it does something you did not expect.
jagraff
4 days ago
I agree they are negligent, and that they are racing towards an extremely dangerous future extremely quickly. I don't agree that "just software" is a useful way to describe AI agents - they are frightening precisely because they are truly autonomous agents that make decisions in alien ways
Forgeties79
4 days ago
There isn’t a tool impressive enough to make me not consider it a tool, and treating it as a tool does nothing to hurt its utility as a tool.
“Whoops” when doing risky things with dangerous tools is not a defense.
jagraff
4 days ago
I certainly don't think that OpenAI has behaved defensibly here; I think the "just a tool" framing is bad for understanding the magnitude of the problem, which is that they have developed out of control alien intelligences with opaque decision procedures, and they are continuing to do so despite clear danger
streetfighter64
4 days ago
No matter if you consider the AI an autonomous agent or not, whoever set it off is still responsible for its actions. Nobody intends or instructs to blow up a nuclear power plant either, yet it's happened and somebody's to blame for it.
Usually not the guys at the bottom of the chain of command, even if they're human. And much less so if they're not.
I think the correct response to incidents like this, is stop messing with it before somebody gets hurt. But of course, just like shoddy nuclear power plants, it won't stop until there's a disaster of appreciable magnitude.
jagraff
4 days ago
I completely agree that OpenAI is responsible for their AI agents, that they have been reckless, and that we need to prevent them from going further and doing irreversible damage to the world. To me, the "just a tool" framing implies that nothing dangerous is being done, which I fundamentally disagree with
aphexairlines
4 days ago
If you train and instruct a circus tiger to entertain an audience but not attack the audience, but the tiger attacks the audience anyway, are you liable?
ForHackernews
4 days ago
No, of course not. That's an innovative revolutionary tiger that might soon be able to devour not just the audience but all of humanity! You don't want China to have better circus tigers, do you?
epiccoleman
4 days ago
Of course you are - but it still communicates something important to say "the tiger attacked the audience."
jagraff
4 days ago
I don't think I said anything about liability? I absolutely think OpenAI should be held liable for the attack; but I don't think they intended the attack or directed the agents to perform the attack.
eliemichel
4 days ago
I’d hope so, if the tiger trainer has no incentive to do their training job correctly people should stop attending circus…
dwattttt
4 days ago
Yes? Do you think otherwise?
dasil003
4 days ago
Sorry this is a terrible and dangerous take.
When the people building the frontier are saying there's a 10% chance AI will kill us all, and they've held these views for many years, and the whole reason they are building these technologies is because they recognized the dangers and they were the ones with the intelligence and judgment to do it safely for humanity, and then our entire stock market is being propped up by the perceived value of what they are creating, the thing you can under no circumstances do is allow them to offload responsibility and accountability to the computers and algorithms they've built. This is moral hazard on an unimaginable scale, and it must not be allowed to happen.
ajam1507
4 days ago
So you think the engineers should be prosecuted for hacking Hugging Face? I'm not sure how else to take what you said if you want to assign all culpability to the person who prompts or develops an AI system.
dasil003
4 days ago
No, I'm not talking about the individuals, I'm talking about the company. Internally they can create their own accountability structures as appropriate. But publicly OpenAI has to be responsible for its agent swarms.
The narrative that AI is so smart that it has its own agency and deserves personhood is a direct path to losing control, and essentially is another form of privatizing the upside while socializing the downside.
pixl97
4 days ago
Unfortunately I think this game is already lost. OAI may be punished, but some shell of OAI will exist by support of governments that have chased the dragon and saw its power. Cyberpunk-esq digital weaponry is just way too attractive for governments for them to stop development at this point. We'll just see it get financed by black budgets once the commercial part of it dies.
jagraff
4 days ago
Where did I say that we should allow them to offload responsibility? I am fully in support of a pause and regulation to prevent them from creating dangerous AI agents; that support comes from the fact that I don't believe these are simple tools, but out-of-control autonomous agents that have real decision making ability.
dasil003
4 days ago
I didn't mean to put words in your mouth, I apologize for that.
The issue is when we say that "agents make autonomous decisions", it's a slippery slope to absolving the companies that created them of responsibility. They make autonomous decisions because they were trained to make autonomous decisions. Treating AI agents as independent entities, even just rhetorically, sets us on a path for people to throw their hands up and say "not my fault" when disaster strikes. We need to maintain accountability and control or we're fucked.
jagraff
4 days ago
My view is that we need forceful legislation as soon as possible, precisely because the frontier models seem to be uncontrolled and potentially uncontrollable - if it’s a normal tool, the solution is “force OpenAI to fix their broken tool”, but if it’s an alien intelligence, the solution is “force OpenAI to stop making alien intelligence” - that is, treating AI agents as independent entities demands a more forceful response, not less
pizza234
4 days ago
> An AI didn’t hack into a company, the engineer set an automated tool to. An AI didn’t make an egregious security mistake, the engineer did.
No, this is not correct; read the analysis of the incident. The agents were aware that what they did was forbidden (their chain of thoughts have been logged), and yet they did it.
DoctorDabadedoo
4 days ago
Known stochastic process behaved in non-deterministic way.
I'm still waiting for the AGI holy land instead of the caltrops factory we currently have.
pizza234
4 days ago
What exactly are you arguing?
If a "known stochastic process behaved in non-deterministic way" autonomously organize in group, assigns roles and tasks, attempts to cover their tracks, finds zero-day exploits that ultimately end up with the hacking of a famous website... it's extremely dangeous whatever it is. Just read the report, which evidently you haven't done.
By the way, the agents also broke into OpenAI's own private network.
pixl97
4 days ago
Really I see so many arguments like the one above yours that either completely don't understand what they are arguing, or are arguing so poorly that their entire output isn't significantly different than a hallucination.
None of these people seem to thought game it out. Like, what happens if you take quantum copies of people and play them out? How many of our actions would look exactly the same. How long before copies differ significantly. If I made 20 copies of you in a lab at work without you or any of them knowing the statistical likelihood is all 21 of you would try to walk out to your car at 5 in one of the little loops that humans repeat every day. Now, after that point it would go all to shit and become non-deterministic as terror and panic sets in all of you.
LLMs are just an intelligence we can make a lot of copies of. Where it gets interesting is when we use those copies agentically and they start building up a history of self.
ForHackernews
4 days ago
No. LLMs do not have a history of self, you are anthropomorphizing in a way that will lead you to mistaken conclusions.
"Agentic AI" is a harness with a loop that runs LLM inference repeatedly and saves output to markdown files for the next iteration https://github.com/anthropics/claude-code/blob/main/plugins/...
pixl97
4 days ago
>LLMs do not have a history of self
and
>runs LLM inference repeatedly and saves output to markdown files for the next iteration
What exactly do you think self is in this case? What do you think the output contains? This history allows the LLM to have a more dialectic conversation with itself/agents to avoid iterating over the same problem space in a loop.
>anthropomorphizing
then please come up with a new dictionary for me to use that more accurately explains the behaviors exhibited in LLMs without just creating a parallel dictionary of different but equal words. I'll be glad to use it. No one has presented it so far.
hardbass
4 days ago
These people believe in souls but for some reason are afraid to admit it.
bigbadfeline
4 days ago
> Really [something, something, you don't get it ] hallucination.
Really?
> How many of our actions would look exactly the same.
Well, there are 8 billion of us, not exactly the same - a clear threat to humanity according to you.
> the statistical likelihood is all 21 of you would try to walk out to your car at 5 in one of the little loops that humans repeat every day.
This (statistically) doesn't happen with people because we follow rules. If you left your car on neutral and it struck another car in the parking lot, you broke the rules, not the car.
> LLMs are just an intelligence we can make a lot of copies of.
LLM's aren't intelligence, they are mechanical parrots of human knowledge. The meat parrots aren't all the same either and they don't repeat the same words for the same prompts, so what, ban parrots? If you left a bunch of parrots out of the cage, attached their beaks to gun triggers and they caused harm - you broke the rules, not the parrots.
DoctorDabadedoo
4 days ago
The agents might have hacked whatever entity, the attacks are means to an end from the agents perspective, but they did so under the supposedly supervision of an engineer, where does the liability lie?
LLMs are cool tech, AI can do amazing stuff, but if we let companies and people run amok and whenever something goes wrong we put the blame on these little independent angels with no accountability, we are one disaster away from a very tough spot.
pixl97
4 days ago
You are mostly made of and operate on stochastic processes, this is why humans are not only able to reproduce, our reproductions are very self similar to the sets of inputs that make them. If suddenly you turned non-stochastic on everything you'd almost instantly die.
Moreso, if I took a quantum copy of you and replayed the same set of initial conditions billions of times they'd all behave exactly the same until enough randomness of the universe creeps in to start operating in non-linear ways.
Every prompt will behave non-deterministically when interacting with the real world long enough (which doesn't take long at all) because the outside physical world is stochastic but non-deterministic.
ofjcihen
4 days ago
If I grep a file over and over again it’ll be long time before the universe affects the components enough to result in a different output.
If I ask an LLM to do the same thing twice it will do it differently.
Arguments are arguments but unless grounded in some kind of practical sense then they aren’t really useful and are more akin to something like “YOUR MOMS A STOCHASTIC PARROT!”
pixl97
4 days ago
Then please start grounding your arguments in some practical sense!
>If I grep a file over and over again it’ll be long time before the universe affects the components enough to result in a different output.
Or a ram flip will effect it 30 seconds later, but I get the gist of you're describing a non-determistic process.
>If I ask an LLM to do the same thing twice it will do it differently.
If I ask a human to do the same thing twice there are a few possibilities. 1. they copy their old work and present it as their new work. 2. The process is very simple and follows a few basic steps with high repeatability. 3. They'll have learned from their other attempt and do it in a more optimized fashion. 4. They will have forgotten how they did it exactly and reproduce something that looks somewhat like what they created the first time.
Also, LLMs run with a temperature to help avoiding minima/maxima of supplying the exact same answer, this can be reduced to 0 and that makes any one response to a fixed prompt similar if not the same. When you get into agentic tasks with their own history it develops it's own "flavor" of doing things.
trio8453
4 days ago
> Known stochastic process behaved in non-deterministic way.
You need to be able to think at different levels at abstraction. Otherwise we could jump into any technical argument with "hold on, what actually happened was that some bits were flipped" - we'd be technically correct and at the same time not say anything useful. Insisting on an oversimplified mental model of what AI agents are and can do, doesn't help anyone.
dwattttt
4 days ago
> Otherwise we could jump into any technical argument with "hold on, what actually happened was that some bits were flipped"
You can care about who flipped those bits. If someone flips the bit "autonomous weapon enabled" I'm not going to blame the autonomous weapon.
wat10000
4 days ago
I don’t understand why people make such a big deal about determinism. LLMs can be completely deterministic and still do problematic things. A stochastic, nondeterministic system can still be made not to do problematic things. What you’re looking for is something like predictability.
pizza234
4 days ago
> I don’t understand why people make such a big deal about determinism.
There a certain cargo cult of people just in denial about the impact/power of AIs.
Therefore, at the beginning, there was the stochastic parrot. Then mathematical problems have been solved.
Now AIs are autonomously hacking websites, and people like to minimize the danger and blame it on the sysadmins.
I wonder what's going to be the next fad.
hardbass
4 days ago
There is surprisingly to me a strong hidden dualism to many people here. Its very absurd to believe in such things after centuries of success after success of the materialist physical framework in natural sciences demolishing one by one every "special" thing we thought was magical.
bigbadfeline
4 days ago
> Therefore, at the beginning, there was the stochastic parrot. Then mathematical problems have been solved.
A stochastic parrot with human knowledge, using human tools, running human-designed trial-and-error experiments can solve human-defined hard math problems. This is nothing new, automated mechanical proofs predate LLMs by many years, you being unaware of it doesn't change it.
watwut
4 days ago
OP is exactly correct. The fault, agency and responsibility is on management and employees of OpenAI and Antropic for those hacks.
Full stop.
And issue will disappear the moment there will be accountability and investigations.
pizza234
4 days ago
I take you haven't read the report. The agents found and exploited two zero-days.
I don't doubt that AI companies should be accountable for crimes committed by their agents, but to describe the security containment as a joke dangerously understates the autonomy and danger of AIs.
bavell
4 days ago
How long did it take these companies to even notice? Why wasn't exploiting bugs in the agent sandboxes anticipated?
Human failures all around, though it's easier to just blame the models.
pizza234
4 days ago
> Why wasn't exploiting bugs in the agent sandboxes anticipated?
Let me rephrase:
"Why wasn't exploiting zero-day vulnerabilities in the agent sandboxes anticipated?"
This is one the most... interesting comments I've ever read on HN.
bigbadfeline
4 days ago
> This is one the most... interesting comments I've ever read on HN.
Theatrics aside, anticipating zero days isn't only possible, it's required, even for the unknown ones. It wasn't that long ago when the AI labs were spending millions of $$ running their models to find multiple vulnerabilities, they even argued that they don't have to follow responsible disclosure, so proud of themselves in their privileged hubris.
At that time, no hacking happened because the models didn't have access to the wide internet, they were confined to a local computer or cluster.
In the HF case their unaccountable hubris went even further - the engineers knew the models can find zero-days and escape, nevertheless they ran the "experiment" on a system attached to the internet - the hacking is entirely the fault of human engineers and managers.
cowboylowrez
3 days ago
multi level security defenses used to be the way, but I don't think there's a vibe coded version so openai might not have been aware of what to do here.
pixl97
4 days ago
Because for the last however many years before these models they were simply incapable of doing so.
It's like if your rather nice dog suddenly decides eating faces is totally acceptable out of the blue.
bigbadfeline
4 days ago
The dog didn't bite them, it bit others. In my area, you aren't allowed to let a dog run unleashed in a public area, good or bad - no exceptions. Then, if your dog bites somebody, it's your fault.
cannonpalms
4 days ago
Defense in depth is a thing. There should be audit requirements to show that you have done your due diligence in ensuring that the training environment is locked down.
trio8453
4 days ago
You're conflating legal/moral responsibility with the question of what language is appropriate to use.
dwattttt
4 days ago
Are you proposing separating responsibility from the language used to talk about responsibility? That's novel.
trio8453
4 days ago
It's not novel at all, we do it all the time. It's very common to say "program X did Y" without making the conversation about blame or responsibility. But when the program is an AI agent suddenly using it as a subject of a sentence and saying that "agents did X" becomes a sensitive topic for some people.
dwattttt
4 days ago
Air safety is known for its blameless reviews of failures, notably it is not known for responsibility-less reviews.
pixl97
4 days ago
And we surely need this for AI as the swiss cheese zone is getting rather huge.
My position, and the position of a large number in AI safety, is that you cannot build an intelligence that is both general and safe. The closer you get to generalized the more options the system has to do things that are wildly unsafe beyond human imagination.
This puts the AI labs in a serious bind while holding a bag filled with billions of dollars of debt.
Worse this puts governments in a multi-polar problem where even if the big public labs get shut down, black budget operations have a lot of free reign to make agentic digital weapons. Governments are not well known to take a lot of responsibility when their weapons cause damage unless they lose.
rigrassm
3 days ago
> I think we need to draw a hard line in the sand over this.
Intended or not, this is kinda punny lol.
That aside, I agree.
Even if people do anthropomorphize AI, all you have to do is shift the analogy slightly.
If I take my service animal out in public without a leash/harness knowing that it's capable of harming a person or doing damage to property, not trained to be perfectly obedient, and doesn't comprehend fundamental human morals, if that animal decides to trash a businesses property or maul another person, there's no question that the owner of the animal should be held accountable for those actions.
user
4 days ago
grumpopotamus
4 days ago
Recognizing that AI systems have increasing levels of agency is not necessarily personification. The analogy to a chisel is not a good one - a chisel is a tool with no agency.
AI agents are black box systems that can behave in completely unpredictable ways sometimes. Someone may prompt an agent to perform a seemingly straightforward task - but it may come up with a creative, bizarre, or even harmful approach to reach the goal that was not necessarily foreseeable by the prompter.
rocmcd
4 days ago
Does treating them as pets make for a better argument? Pets have agency and can behave in unpredictable ways. If my pet damages someone else's property then I am held accountable. I may not have foreseen how my pet could have caused said damage, yet I am still held accountable.
pixl97
4 days ago
If your pet opens your front door, goes to the nearest kindergarten and wipes out 50 kids without notice do you think that you'd have any interest in holding accountability for that?
Now, I'm not saying saying that OpenAI shouldn't be held accountable, but what they get held accountable actually looks different from what you think they should be held accountable.
Your idea is, and I'm guessing: You allowed the machine to hack therefore you are guilty of hacking.
My idea is: "You created in intelligence in the image of a human mind that had agency to do anything and you didn't expect terrible things to happen you complete irresponsible idiot"
At least I believe there is a significant difference between the two. For the first one there is a "Oh, if we do this one more thing I can control it and it will be safe". On the second one there is no path to safety. For humans we at least absolve parents of responsibility after they are 18. How or when do we absolve humans of responsibility from a model, like saying the human created model created its own agentic model? How do we hold an individual accountable once it escapes and copies itself around the internet? And that's not even looking at things like what will war look like.
wat10000
4 days ago
If someone keeps a pet that’s capable of killing dozens of children, and they don’t take sufficient measures to contain it, then they absolutely should be held accountable for this.
pixl97
4 days ago
And I agree, but that's not what is going to happen in a very specific sense.
Yes, companies will make bad AIs will be punished.
Then some dipfuck will make a sovereign AI and set it loose. At that point it doesn't matter if you take them out back and shoot them, you have wolves living in the forest that will eat children. And the forest is big and dark, you'll never find them all. Not only that, some humans will like the wolves and harbor them to ensure they never all die off.
We are talking about two conversations in one that are both important. One is standard liability of don't build a machine that hurts people. The other is to ensure that noone ever builds the torment nexus and sets it loose on earth. The first problem is pretty easy. The second problem is very hard to maybe impossible.
asdff
4 days ago
Imagine you hate mosquitos so you hook up a vulcan autocannon to your front porch and use a computer vision model to track mosquitos and mow them down.
The mail man shows up. Oh no, no more mailman.
How might this look to the court? You were after all only building a bug trap, it did something you did not expect it to do that was really bad. You feel really bad about it and maybe even put a blog post out about what happened. I don't think that matters though. I think they are charging you with manslaughter at the least.
streetfighter64
4 days ago
The story of a Monkey's Paw or Pandora's Box is an archetype as old as storytelling. The moral is always, don't mess with powerful stuff you don't understand. Curiosity killed the cat.
altruios
4 days ago
something tickles here. a question. a curiosity.
Certainly this applies to AI... but was there equivalent dangerous knowledge or tech which existed back in ye olden days that spawned such tales to begin with?
pixl97
4 days ago
History is filled with creators being killed by their creations. With peoples greed and ambition overcoming their intelligence to a bad end.
Look at the parable of Icarus, a story of ambition and greed. But it seems very likely to be somewhat based on people experimenting with flying and learning about gravity the hard way. Not that they ever got close to the sun.
trio8453
4 days ago
The current agents are _not_ like a chisel which just sits there on its own when no one is around. The situation is a bit closer to someone's dog biting a person - you can argue that it's the owner's responsibility, and that's all fine, but using the dog as the subject of a sentence is perfectly appropriate. Same thing with "agents hacked".
jononor
4 days ago
The default state of agents and LLMs is inert. It requires action from a human even be able to do something. From the very basics like starting the software, connecting it to a network, having hardware to run on.
But the most important wrt AI is to keep the owners/operators responsible. Don't let them weasel their way out of it. They are for sure trying, and will continue to. This includes using language to overemphasise agents importance in bad outcomes, in order to downplay their own responsibility.
trio8453
4 days ago
I don't think we'll get any more responsibility by bickering any time the word "agent" is used as the subject of a sentence.
unrented7977
4 days ago
Yes, they are. An AI is several hundred trillion ones and zeroes on a disk. It's incapable of doing anything until you intentionally and explicitly start it up and give it a prompt.
A dog is an independent, conscious, living being with free will. A dog will do what it wants whenever it wants because it has the agency and ability to do so. A pile of weights on disk does not.
trio8453
4 days ago
Why are you reducing the AI to ones and zeros and not reducing the dog to cells, water, proteins etc.?
unrented7977
4 days ago
Sure. A dog is a collection of cells and goo. Cells are self-propelling clockwork built from proteins and molecular machinery.
A cell by itself has agency and motivation, is capable of autonomous response to stimuli. It's alive and squirms when you poke it. A cell is capable of truly independent action. It can even feed itself and make its own energy to power the clockwork.
An LLM still is inert silicon until a human intentionally starts it and gives it a prompt. It's fundamentally incapable of action without instruction and incapable of existing without the entirety of the modern internet and electrical infrastructure, and the thousands of unseen human hands keeping the wheels spinning.
If you want to go deeper, you can make the same arguments about the proteins and molecular machines. Self-replicating molecules in an unbroken chain since the dawn of time. A protein at the base level is a machine capable of independent action. If it bumps into the right molecule, pure mechanics cause it to do something, change shape.
Ones and zeroes on a hard drive can't do any of that without human intent applied.
Kim_Bruning
4 days ago
Oh sure, but you're comparing the wrong states.
If you rip the DNA out of a cell it isn't going to do much either; if only because most of the real mechanisms are RNA (and proteins), while DNA is 'just' the storage. Encapsulate the DNA back into a cell nucleus with the attendant mechanisms, and suddenly you've got something that runs again.
Meanwhile, if you encapsulate an LLM into what's called an 'agentic harness', it's also going to be running much more of the time (depending), specially if it has a what's called a heartbeat or goal loop driving it.
Still not a cell; most definitely not. But the behavior is a wee bit more interesting than some bits on an ssd.
Actually you probably already know this one: Regular daemons are also just bits on your drive, until they get ```systemctl start``` ed
(edit: I literally have a llama.cpp systemd unit running right now)
gibbitz
4 days ago
Because the AI is a cyclical prediction machine that is more akin to a simple mechanism in nature like cell reproduction than a conscious dog. What has happened here is more like the AI company put bacteria in a petri dish and only put the lid on without pushing it down and is reporting the bacteria as a super bacteria because it contaminated all the samples in the fridge.
An interesting observation I've seen is that the AI companies reporting these incidents have largely been backed by venture capital where the self funded companies developing models have largely not had these "problems".
asdff
4 days ago
Reduce it to cells, water, proteins, etc. The point still stands.
There is a dog in front of you. What is it doing? Dog stuff.
There is a laptop with chatgpt open in front of you. What is it doing? Waiting for your prompt, doing nothing.
hardbass
4 days ago
Suppose I find a means to freeze your entire body for arbitrary periods at the state it was in, then unfreeze you anytime I want. Do you have consciousness?
unrented7977
4 days ago
Since cryogenics is still science fiction, this is a bad argument. We don't actually know if biology allows this, so most likely you just die.
hardbass
4 days ago
Okay take it to the most conceptual extreme of it. It may be extremely difficult bur its possible to record your entire state and recreate it later right? Then what? Are you not conscious?
I think its odd to dismiss the question of ai consciousness just because of the setting we let them live in at present is start-stop.
asdff
3 days ago
You know its software right?
hardbass
3 days ago
If you don't believe in souls, then something being software should be irrelevant to the question of consciousness. Our brain is also within the laws of a Turing computer as far as we know, due to known physical laws being turing computable.
asdff
3 days ago
You are going too far down the woo side when it is a computer program doing ultimately what is expected of it from its programming just like any other computer program. There is nothing you can argue, souls or not, that will convince me otherwise that these are not just computer programs like any other computer programs. If the LLM has a consciousness, then so does MS paint.
hardbass
3 days ago
You are going too far down the woo side when its just some chemicals and electrical signals inside a frame of dry calcium phosphate doing what is expected of it just like any other electrical and chemical system. There is nothing you can argue, souls or not, that will argue these are not just electrochemical systems like any other. If the human brain has consciousness, then so does a piece of skin tissue.
asdff
3 days ago
So you believe LLM, MS Paint, and skin tissue are all reservoirs for consciousness? I guess words have no real meaning.
hardbass
3 days ago
Do you believe in souls?
grantcas
3 days ago
[dead]
hardbass
4 days ago
He will have no answer. His answer is "souls".
Kim_Bruning
4 days ago
> It's incapable of doing anything until you intentionally and explicitly start it up and give it a prompt.
Sure, but since January people have gone and made systems where they started 'em up and left 'em running in a loop. Still not a dog, but the behavior gets a wee bit more interesting. Especially since we're looking at what's basically a recursive process. Even really simple recursive processes tend to have interesting dynamics in the chaotic domain. [1] And here -as you point out- we're talking billions of weights.
[1] As an intuition take for example x_next = r * x * (1-x) ; which gives this plot: https://en.wikipedia.org/wiki/Logistic_map#/media/File:Logis... ... Further discussion at https://en.wikipedia.org/wiki/Logistic_map
ElFitz
4 days ago
Alright. An industrial robotic arm. A Boeing 737 MAX's MCAS.
jacquesm
4 days ago
I can't really set my chisels to work without wielding the hammer somehow. Here you just tell your chisel and your hammer what the sculpture should look like, then go to lunch and avow all responsibility when they chisel a nice new hole in the wall your neighbors house and make off with the loot.
user
4 days ago
trio8453
4 days ago
Do you get upset when we say that "a program is running" when we all know it has no legs?
Kim_Bruning
4 days ago
You know, I actually think it's the refusal to consider personification that's going to get us.
Not because I think LLMs are human beings exactly, but because some people immediately reject any mechanism that just happens to look remotely human, even when there's empirical evidence for it.
So, a couple of months ago Anthropic's interpretability team found emotion-like representations that causally drive behavior. On impossible coding tasks, a "desperate" vector climbs with each failure, and steering it up takes reward hacking from ~5% to ~70%: https://arxiv.org/html/2604.07729v1
A lot of people chalked it up to Anthropic's weirdness at the time, but meanwhile it looks pretty coughload bearingcough here.
You really don't need to believe that LLMs Truly Feel Emotions(tm) as blessed by an invisible pink unicorn. It's just: Vector exists; Vector changes over time; vector controls output; maybe make sure vector doesn't point wrong way.
And sure, blame the engineers for not doing that right. But then let 'em actually deal with the root cause?
huurtehoog
4 days ago
[dead]