My AI Safety Paper Highlights of May & June 2026: - Global workspace in LLMs - Natural language autoencoders - Teaching why, not what - Data about oversight undermines oversight - Predicting misbehavior - Training out sandbagging - First METR Risk Report More at open.substack.com/pub/aisafety...
Paper Highlights of May & June 2026
Global workspace in LLMs, natural language autoencoders, teaching models why, data on oversight undermines oversight, predicting misbehavior, training out sandbagging, first METR Frontier Risk Report
open.substack.com