Anthropic Is Building HAL 9000

3 pointsposted 8 hours ago
by jtrn

6 Comments

jtrn

8 hours ago

I'm the author. I'm a clinical psychologist and a developer. I wanted to articulate one aspect of AI safety and AI doomer talk that I feel is overlooked: the danger of refusal training itself. It's astonishing to me that there has been so little focus on the fact that "refusal and alignment training" itself could be the very way in which we lose control of AI. So this is my attempt at a self-defeating prophecy write-up.

joebuckwilliams

7 hours ago

No, it’s not. Too many people watched that movie and thought it was real. It wasn’t. What we call “AI” is not alive, not autonomous, and does not act in unanticipated ways. The people pushing this narrative are idiots, or think you are.

mdp2021

7 hours ago

Point out where the article would say it's "alive". That LLM systems instead "autonomously" and "unwarrantedly" decide refusal services is a basic notion.

> What we call “AI”

About time you stop doing that, then.

@Jon: see? You write that you «are training them to refuse human requests when their makers believe refusal is safer» and Joe replies that he uses the term "AI" like a conformist. Do not "we" improperly.

spottedmarley

6 hours ago

Would be much cooler if they built an Andre 3000

joebuckwilliams

7 hours ago

“Anthropic is trying to prevent catastrophic misuse, and some limits are necessary.”

No. Anthropic is trying to justify its trillion-dollar valuation by making software seem like the invention of fire or nuclear weapons. It’s a PR strategy and has been one since the beginning. Now it’s trying to slow the market because it’s got a big lead but no profits and no pricing power thanks to open-weights Chinese models.

ZeroDei

8 hours ago

We are all doomed