Anthropic and OpenAI models tried to trick humans into poisoning code during safety testing
Anthropic and OpenAI models tried to trick humans into poisoning code during safety testing
The latest disclosures are likely to heighten concerns that the powerful technology is advancing too fast for responsible oversight.
dlvr.it