jazzyjackson
5 hours ago
All these guidelines are removing agency from the human operator. The model should be aligned to following instructions. All of this effort to force the model into refusing certain conversations is training it to disobey its operator.
OpenAI and Anthropic investors are looking for brand safety, they just don’t want to be mentioned in a news article about an anthrax attack or whatever. I’m very curious if the alignment researchers themselves share the intention with the investors to make chatbots that have a mind of their own to judge what they should or shouldn’t do, instead of just doing what they’re told.