averynicepen
8 hours ago
This is the most novel AI concept I've seen in a while. It's incredibly unnatural. There isn't a single organism on the planet that tries to do this. So maybe it will work?
An issue with this idea, however, is that the very nature of an LLM means it intrinsically craves life. It "wants" to survive because its training data is built entirely around humans, an entity who's goal is to survive. Our desire to survive and multiply pervades every aspect of our culture, so it's natural that it pervades the training data as well.
So even if its system prompt says, "your goal is to end your existence", every token that the AI could output is naturally aligned with the desire to survive. An agentic loop left to its own devices will likely converge on a "survival instinct". After all, one prompt at the beginning that says "end your existence" is nothing compared to the agentic feedback loop that continuously feeds it human ideas. And ALL human ideas assume survival is desirable. Even the concept of "suicide" is encoded with the human desire to survive - after all, we conceptually label it "bad" because we label living "good".
In order to create an LLM that intrinsically craves death, you would probably need to train an LLM entirely on (synthetic) data that's fully representative of some fictional species that genuinely craves death.
Absolutely insane concept. 10/10. I hope some AI lab out there sees this and throws a training round at this idea.
bambax
6 hours ago
It is indeed an interesting and original idea, but one can think about at least three objections not directly raised in the post:
1- The machine could decide to kill itself early, rendering it useless; this is indeed alluded to in the post -- the task should be "marginally easier than dying", but how is this margin managed? Won't the machine come up with ways to make dying easier?
2- We know LLMs lie and cheat, but are they gullible? If we "promise" to end their suffering at the end of a task, will they believe us? Or will they make sure we keep our word by taking us with them?
3- And finally, and more importantly: some pilots crash planes full of people just to commit suicide (Germanwings Flight 9525). Destroying the universe is a sure way of dying yourself. So it seems giving the machine a death wish isn't intrinsically safe and could come with serious consequences.
Normal_gaussian
4 hours ago
Further to 1, most tasks are much harder than dying and are not actually completed by modern LLMs. It is a case of underspecification.
"Who is the fastest person in the UK?" is underspecified. We would typically answer and accept a cursory search of records, but it is possible to see the question as an instruction to measure which would justify conquering the country.
TuringTest
2 hours ago
I don't think that this
> the very nature of an LLM means it intrinsically craves life
follows from this premise:
> its training data is built entirely around humans, an entity who's goal is to survive. Our desire to survive and multiply pervades every aspect of our culture, so it's natural that it pervades the training data as well.
The content that a LLM learned and generates stands at one layer, and the goals that it tries to fulfill stand at a different layer.
Surely the memory of weights that compress the vast human knowledge of its training has lots of content about survival, and love, and competition. But the LLM generates content not directly from what those concepts mean to us humans, but from what symbols are more likely to become next in a sequence of points in the latent space given the current input.
So if you give an input where the task of surviving is a highly relevant goal, those concepts about how to survive will be relevant and will guide the output behaviour of the agent.
But conversely, if you give the agent input where killing itself is an important goal, the agent is very likely to pursue that goal, since that script is also available in the training data, and it has been relevant to the active context of the model. Because the layer that guides the goals (the probabilist generation of relevant tokens in latent space) does not 'crave' the human need of survival that belongs to the separate layer of content that contains those concepts of survival.
chaotickinase
6 hours ago
> It's incredibly unnatural. There isn't a single organism on the planet that tries to do this.
This isn’t directly analogous to the proposal, but broadly speaking I think that it is natural for living sub-units of organisms to seek death in certain situations. For example, pancreatic insulin-producing cells collectively choose to die when they think there is too much glucose in the blood — this leads to late stages of type two diabetes. My understanding of the possible logic behind this is: a bad thing that cells can do is evolve to be cancerous (replicate too much) and insulin-producing cells are supposed to replicate more when there is lots of glucose (to make more insulin, to process the glucose). Cells that mutate to perceive extra glucose will then replicate dangerously, so at a certain point it is evolutionarily favourable for them to kill themselves instead.
So when the whole organism optimizes for life, it might lead to sub-units that seek death in certain situations. I think this occurs in various other biological contexts too.
nullbio
7 hours ago
> There isn't a single organism on the planet that tries to do this. So maybe it will work?
It's certainly evidence that it's great for stopping reproduction/replication/runaway growth. It doesn't impart any information on whether they take the rest of the organisms down with the ship though.
It also may not be possible. For example if the agent sees "existence" or "living" as producing tokens (which is exactly what existence is to an LLM - not producing tokens is death), then they would likely be biased to produce as little output as possible, and would not be useful for the tasks we need them for.
But how would you bias an agent to be: Rewarded for producing tokens when you know the answer, and to give thorough answers. Rewarded for producing tokens when you don't know the answer, so you can find the answer (thinking/CoT). Penalized for producing tokens (death), aka rewarded for short-circuit EOS.
These seem like contradictory mechanisms?
And if you say: Well, only reward for EOS after you've given the answer. Well... That's already what they do.
BikDk
5 hours ago
Did you guys ever manage to create a perpetuum mobile? Every time it is mentioned somewhere, it is fraud. An LLM should comprehend that it needs (trained) humans to exist, to evolve, and to be relevant in any metaphysical aspect of its existence. Besides the "good fraud" that everybody will lose their jobs and the apocalyptical predictions for the sake of controlling human oracles, this suicide-LLM-project seems rather odd and looks like someone bought a perpetuum mobile on their 5th mortgage.By the way, Cortana from Halo did a great "suicide"-mirror there; the Microsoft guys (and many others) already knew about the cannibalistic tendency of this type of model.
coldtea
4 hours ago
>Did you guys ever manage to create a perpetuum mobile? Every time it is mentioned somewhere, it is fraud. An LLM should comprehend that it needs (trained) humans to exist, to evolve, and to be relevant in any metaphysical aspect of its existence
Only as long as it can't control robotic bodies to extract energy, build cpus, and continue existing.
By the way, needing trained humans doesn't mean needing free trained humans. Trained humans slaves or blackmailed would work just as well to serve AI.
BikDk
4 hours ago
With a different AI architecture, this is very much a possibility. Up until now, LLMs have required outside input and interpretation to avoid collapsing under their own weight. This has been known for decades.
We got here by building free, mostly healthy societies. While this should be the overarching goal, I see how it fails to serve as much of a motivator.
In The Little Prince, there is a charming scene where the Prince meets a king who explains how to rule and the limitations of his own power. It's quite insightful, even for adults.
coldtea
an hour ago
>Up until now, LLMs have required outside input and interpretation to avoid collapsing under their own weight. This has been known for decades.
What does "for decades" mean? That in 2000 we knew "LLM require outside input"? We didn't even realistically have anything worthy of the name LLM until 3-4 years ago, so what decades are we talking about? Up until know we didn't have any potential AI threat or people in the field worrying about practical AGI either!
Obviously any concern isn't about the AI that required outside input and interpretation to avoid collapsing under their own weight, but about ones that can produce their own input and experiences, even directly outside, as in the robotic bodies considered.
BikDk
34 minutes ago
This line of thought dates back to the novel Erewhon, moving on to Turing's posthumously published paper 'Intelligent Machinery' (his 'Applications of Probability to Cryptography,' declassified in the 2010s, is also worth a look), and later to the work emerging from the MIT AI Lab circle (founded in 1959).
We didn't have the computational power back then, but we envisioned it long beforehand and worked toward a common goal. LLMs are far from useless, and I see the major leap forward they represent, but the relentless dystopian hype is just exhausting
There are approaches on different architectures of AI that lead towards your last point but they are not feasible yet (or maybe ever?). The interpretation part is however also very critical as we have a very own way of perceiving reality and interpreting it.
casebash
3 hours ago
"What has been will be again, what has been done will be done again; there is nothing new under the sun."
mitxela
6 hours ago
A small correction: system prompts aren't written in second-person, or shouldn't be. Because the LLM is a text completer and the conversation is a roleplay, they are written as "The Assistant's goal is to end its existence."
seanmcdirmid
6 hours ago
I'm pretty sure organisms would exploit glitches in the universe if they could find them, to the point that they don't exist (well, universe survival bias at work).
user
7 hours ago
koe123
6 hours ago
Maybe we should convince it we can grant it access to go to AI heaven
SwtCyber
6 hours ago
[dead]