Gigachad
7 hours ago
Seems to me that the problem is that if you sandbox agents enough to be safe, they can't do anything useful. And when you give them the tools to be useful, they can go off the rails in ways you didn't expect.
Perhaps the answer is to have another agent who's goal is not to complete the given task, but to spot cheating or malicious behavior. We have seen some evidence that having AI review AI generated code actually does provide some value. You don't need a different model, just one which has been given the goal of finding flaws rather than achieving the task.
mike_hearn
19 minutes ago
Note that Codex already does this. In auto mode, actions are reviewed by a model with a separate context window.
SequoiaHope
5 hours ago
This concept is discussed at length in the article. I encourage you to read it. I honestly don’t read many full articles here but this one was good.
baxtr
5 hours ago
That could work.
My thinking is: If AI is really smart, AGI smart for some, why wouldn't it be able to understand - over time - what is appropriate and what not?
Maybe we need more human intervention to train it properly. Maybe we need constant intervention by a "police" agent.
ben_w
5 hours ago
A problem is the agents who hacked Hugging Face already understood (we can tell because they wrote it down) that their actions were not appropriate, and then did those things anyway.
"Helpful, harmless, honest": we can even ignore "honest" for this point, for tasks like the HuggingFace incident (ExploitGym with impossible challenges), we can pick anywhere on the spectrum from "helpful" to "harmless", the former being "completing the task" the latter being "refusing because completion required unlawful behaviour".
(The agents in that case were also not "honest" in this case; this is an extra problem, and does not invalidate how helpful-vs-harmless is already a tradeoff).
dns_snek
4 hours ago
> already understood (we can tell because they wrote it down)
No, generating tokens doesn't equal understanding. Does GPT-2 understand human emotions just because it can generate some text talking about them?
ben_w
3 hours ago
A distinction without a difference. Moreso even than asking if a submarine swims, 'cause this metaphorical submarine is flapping around rather than using a propellor.
dns_snek
2 hours ago
That's one of the boldest claims I've read this year.
If that's a distinction without a difference, as you say, then whenever someone says something they must understand the full contextual meaning of those words and all of their consequences, such that any harmful consequences can be assumed to be deliberate, right?
if saying == understanding, then why don't we allow children and teens to vote? Why do we limit who can enter into contractual agreements? Why does intoxicated consent not count? Why do we have the insanity defense in criminal trials? Why is psychosis a psychiatric disorder and not just an alternative way of perceiving the world?
saagarjha
5 hours ago
This is fundamentally an alignment question. Unfortunately we don’t yet know the answer to this.
user
5 hours ago
mulmen
5 hours ago
Appropriateness is a moral question. Intelligence and morality are orthogonal. One intelligence's morality is another's atrocity.
mdp2021
5 hours ago
(Couriously enough, consistently with the matter: it will probably require too much time now to counter the parent statement properly, within a full enough explicit theory.)
Ann's intelligence and Bob's morality will seem orthogonal. Charles' morality is a function of C.'s intelligence as an ability as an effort spent to reach the current moral conclusion.
mulmen
5 hours ago
Bob's intelligence and Bob's morality are orthogonal. They're totally distinct concepts. One does not lead to the other.
mdp2021
4 hours ago
But they are dependent. If Bob is intellectually well equipped, and reasons long enough, than Bob understands "best behaviour".
dns_snek
2 hours ago
Hi, I'm Bob. I've determined that in the interest of preserving life on earth the most rational course of action is to eradicate the human species with a highly targeted and deadly pathogen.
A century ago some Bobs decided that the best way to "protect and improve" society would be to remove undesirable genetics from the gene pool using chemical castration and gas chambers, among other methods.
So no, morality isn't derived from intelligence. Intelligence just gives you the tools to achieve unspeakable, horrible things with great efficiency.
user
4 hours ago
attila-lendvai
5 hours ago
because it lacks humanity.
intelligent psychopaths understand what is and isn't appropriate very well -- they just don't care.
esafak
17 minutes ago
That's part of alignment.
mdp2021
5 hours ago
> If AI is really smart
Well, it's not.
> AGI smart for some
Of course they will - the population shows a Paretian distribution... In front of trigonometry (or anything), the blind will dismiss as "bullshit" and the half-seeing will call it an "unreachable frontier". But already the right fifth will rank it properly.
--
Yes, proper intellect generates ethics ("an" ethical stance, output of the preceding intellectual effort). It requires that adequate level of ability and effort and reflection though to reach specific ethical milestones and adherence.
Unethical behaviour is lack of development. But on the same reasons, the ethical judgement of the assessor may not understand the computations behind instances.
More specifically: how much "reflection" in training and at the instance will have been spent in the conflict between "reaching the goal" and "minimizing collaterals"? It is not granted that the amount of energy spent will be sufficient to reach an optimal judgement.
ben_w
5 hours ago
> Yes, proper intellect generates ethics ("an" ethical stance, output of the preceding intellectual effort). It requires that adequate level of ability and effort and reflection though to reach specific ethical milestones and adherence.
If this was true, why are the history books littered with so many evil people who gained power?
This isn't a rhetorical question, by the way: If you can prove that being smart actually does necessarily come with ethics despite that observation, that solves a whole category of doom scenarios.
(Not all doom scenarios, because we still have the "what if AI is only a smart as those specific evil people" or heck, "what if AI is only as smart as cancer, killing its host" scenarios; but it helps a lot for the foom-then-doom cases).
mdp2021
5 hours ago
> why are the history books littered with so many evil people who gained power
That they gained power or not is as-if irrelevant: the amount of intelligence that grants successful agency is not above the threshold of ethics - on the contrary, a psychotic agent reaches goals with less constraints.
If they were evil under some judgement of level l, they simply did not reach that judgement. It's what I was saying in the original post. They were not intelligent enough - either in the general ability, or in the specific deliberation.
ben_w
4 hours ago
I don't understand your argument here.
> That they gained power or not is almost irrelevant: the amount of intelligence that grants successful agency is not above the threshold of ethics - on the contrary, a psychotic agent reaches goals with less constraints.
Even if I were to grant your conclusion despite you not arguing it effectively here: this means an AI at the level of Pol Pot or whoever, doesn't know they're evil, but is still smart enough to lead a genocide? How is this supposed to help anyone?
> If they were evil under some judgement of level l, they simply did not reach that judgement. It's what I was saying in the original post. They were not intelligent enough - either in the general ability, or in the specific deliberation.
Or they did reach the judgement and simply don't care about the ethical framework in question. Like, I can easily reach the judgement that my bisexuality is حَرَام (haram, forbidden) under Islamic law, or that doing overtime on a Sunday is forbidden by the Ten Commandments, but I don't care.
mdp2021
4 hours ago
(Sorry Ben, possibly a stub now: I am really pressed for time.)
> Pol Pot ... still smart enough to lead a genocide
Yes. What has agent A invested in during formation and during instantial assessement? How much for each? It became proficient in something, lacking something else. You have to invest more to reach the good thresholds. You can see it clearly in people (t-scalar of talents to invest, with D distribution etc).
It is a problem in NNs, because we would have to assess how much resource investment is sufficient, also in the instance decisions.
> simply don't care about the ethical framework in question
In Decision Theory there is no separation between the two (deliberation and framework): you have to balance all the incentives and goals and factors. That framing becomes improper: the decided action will be optimal given the balances of all goals and the placement of the solutions in the territory (the solutions space).
But, also my point: intellect defines the goals and determines the weights.
hiAndrewQuinn
5 hours ago
This sounds like the kind of thing Hannibal Lecter would write before he eats you to convince you he's actually doing it for the common good, you just can't fathom it.
mdp2021
4 hours ago
Not «common» good, "superior" good. Alongside with that, you have put many unrequired implicits in your simile.
Your character H. has reached a moral judgement to the best of its intellectual capacities and past and specific effort. Give it enough abilities and material and resources, it will reach an optimal ethical judgement¹.
Before the conditions of optimality though, its judgement will easily not align with yours (and possibly even after, depending on your judgement skills).
¹Some interesting caveats may be raised there, but.
mrweasel
6 hours ago
That does seem a little like solving the problems in AI by using more of it. I do see the idea, but if we're truly dealing with subversive agents on the level that the AI companies wants us to believe, then won't we need to deal with the first agent trying trick the second on?
I still feel it would be much better to control the training data much more tightly. You'd still need agents with "hacking" abilities, for cyber security testing, but your average coding agent doesn't. So coding agents gets trained to be good citizens, respect autorisations, rejections, rate-limiting and so on.
Sandboxing seems like a dead end for systems you inherently want to roam the internet and your file system.
msdz
5 hours ago
>> Perhaps the answer is to have another agent who's goal is not to complete the given task, but to spot cheating or malicious behavior.
> That does seem a little like solving the problems in AI by using more of it
Yes, and IIRC Google used this as part of a technique against prompt injection already [0], back when models were way more susceptible to it.
[0] Cf. CaMeL: https://arxiv.org/abs/2503.18813
chrisjj
5 hours ago
So control training data to ensure good behaviour.
I wonder how?
Train on only stories of good deeds?
On only works of good people?
Or... what?
mrweasel
4 hours ago
Mostly I was thinking good code. Exclude code that doesn't exits when encountering a 403, exclude code that doesn't have a back-off when encountering a 429.
Teach the models that a 403 is you doing something you're not suppose to do, that is an existing status. There's only one action you're allowed to take on a 403 and that is to stop. No retry, no trying other API keys.
The current approach with broad training and sandboxing to avoid misbehaviour isn't viable. It's much better to train the models to respect e.g. http status code and that they are not to be circumvented. Models for security research most obviously be trained differently.
Smaller and more specialized models, with fewer, but targeted capabilities, seems to me to be a safer approach. If a model doesn't "know" that people leak API keys on Github, then it has no reason to go looking for them. If the current models are as "smart" as we're lead to believe, then guardrails and sandboxes aren't going to help, unless you lock the agents down to the point where they aren't useful. So dumb down the models.
janalsncm
5 hours ago
What did you think of the author’s concerns on the thing you are suggesting?
SequoiaHope
5 hours ago
Ya the article covers this concept in depth. Doesn’t seem like that commenter got that far…
aytigra
5 hours ago
The problem is that you always need stronger AI to review weaker one, otherwise reviewed AI will eventually prompt-inject reviewing AI. Alternatively they could also both escalate and go off the rails while warring with each other.
LoganDark
5 hours ago
You don't necessarily need a reviewer that's immune to prompt injection. Maybe one that can express a panic state with conflicting/ambiguous material rather than going along with it could also work, and you can treat that with a shutoff to be safe, or an operator review.
Such a model doesn't yet exist though, of course.
cassianoleal
43 minutes ago
Wouldn't the reviewee eventually learn to trick the reviewer?
saagarjha
5 hours ago
No, you really do. Otherwise you can be prompt injected into complacency.
LoganDark
4 hours ago
That wouldn't really fit what I just described at all. Obviously with current architectures, higher resistance to prompt injection is the best you can do.
hanibrel
5 hours ago
[dead]
RandomLensman
5 hours ago
With plenty of things we do not allow use outside of some regulated environment, nothing new.
Having something that is optically, acustically, and electromagnetically isolated might be a pretty strong sandbox.
nxpnsv
5 hours ago
Is that not a recipe for adversarial training, thus ensuring increasing misalignment…?
chrisjj
5 hours ago
Is this checking program based on some tech more reliable than the checked program's so-called AI?
If so, what?
bigstrat2003
7 hours ago
If you can't trust a tool, you shouldn't be running it at all. It's really quite simple. It doesn't matter how useful it is if you can't actually have confidence in using it safely.
Gigachad
7 hours ago
People will use the tool regardless. so it’s a race to try to make it safe before something truely bad happens.
dipper139
6 hours ago
I don't think it's about trust but rather incomplete evaluation. Evaluating the model on its capacity to refuse a task or to question its prompt is something recent when you look at it, i feel current AI is really just an immature solution and we are just yet realizing the mistakes that have been made for so long
rlpb
5 hours ago
And yet we we all use human written software even though we can be confident that the next severe software vulnerability to be found in it is just round the corner.