Anthropic does not apply the same protections to the data labelers they use to clean their training data from the worst shit in the world. But those are just mostly poor people, what do effective altruists care about th em?
Anthropic wrote that the new policy update “is meant to apply only in extreme cases, where users repeatedly act cruelly toward our models, with no discernible purpose. It does not apply to common versions of user frustration, pushback, dark creative themes, or model testing and research.”