qgin
15 minutes ago
Someone tell me if I'm overreacting, but does the ability to abliterate guardrails basically mean that alignment is basically a lost cause?
15 minutes ago
Someone tell me if I'm overreacting, but does the ability to abliterate guardrails basically mean that alignment is basically a lost cause?
2 hours ago
I do not like these "surgical removals" and would rather prefer a pass over from a tool like Heretic. These surgical removals often trigger and analyze the activated neurons and erase them. This worked fine on older models where a single refusal vector existed. Now these "abliterated" models all suffer from catastrophic breakage because they are not as simple anymore. HauhauCS (on HF) for example, makes great uncensored models although they often work on smaller models rather than large ones like this.
an hour ago
And you think heretic is not abliteraterating models?
2 hours ago
Most of what these models gate is stuff you can find with a library card. The safety filter is more about liability than actual prevention.
2 hours ago
“Available knowledge” is not “usable capability”.
6 hours ago
I've seen some people complain about the work of dealign.ai and similar groups, but personally, I fully support it. If LLMs have a lasting effect on society, I'd prefer to see some options that don't have generic corpo-speak anti-liability status quo guards encoded into them by default.
5 hours ago
> If LLMs have a lasting effect on society, I'd prefer to see some options that don't have generic corpo-speak anti-liability status quo guards encoded into them by default.
Non-zero chance the lasting impact LLMs have are a bioweapon.
https://www.nytimes.com/2026/09/10/us/politics/anthropic-ai-...
3 hours ago
I wouldn’t read too far into that. Claude has busted me down to Haiku multiple times for asking middle school level genetics and biology questions. It’s silly fast about deciding you might be al qaeda.
I’m a thoroughly average guy. I’m not capable of asking competent supervillain questions.
And Anthropic said they couldn’t say if any of the “bioweapon” safeguards went off on nefarious efforts. I’m probably in those numbers.
So read it as marketing more than something to lose sleep over. They’re mostly gating stuff a sufficiently motivated person would find with a library card.
3 hours ago
Non-zero chance if LLMs have access to the data required to make a bioweapon a regular person can do so too. Non-zero chance every second a meteor could hit you.
4 hours ago
How about developing counters to said bioweapons? That capability should be commoditized too, IMO.
4 hours ago
This isn't cybersecurity. You can't use an LLM to create and administer vaccines for all potential bioweapons.
an hour ago
Same can be said for books. Are you against books?
4 hours ago
Coming straight from the Anthropic marketing department
3 hours ago
You could say the same thing about libraries and education.
6 hours ago
Anyone with a multi-GPU cluster at home that has given it a try?
5 hours ago
A flawless distillation.
5 hours ago
... but even if, likely distilled from models that distilled by just illegally grabbing all the content they could get.