Johannes Gasteiger🔸

@gasteigerjo.bsky.social

Safe & beneficial AI. Favorite papers at http://aisafetyfrontier.substack.com. Opinions my own.

My AI Safety Paper Highlights of May & June 2026: - Global workspace in LLMs - Natural language autoencoders - Teaching why, not what - Data about oversight undermines oversight - Predicting misbehavior - Training out sandbagging - First METR Risk Report More at open.substack.com/pub/aisafety...

Paper Highlights of May & June 2026

Global workspace in LLMs, natural language autoencoders, teaching models why, data on oversight undermines oversight, predicting misbehavior, training out sandbagging, first METR Frontier Risk Report

open.substack.com

My AI Safety Paper Highlights of April 2026: - *Research sabotage propensity* - 2 sabotage benchmarks - Alignment research automation - Misaligned AI organizations - Exploration hacking - Conditional emergent misalignment More at open.substack.com/pub/aisafety...

Paper Highlights of April 2026

Research sabotage propensity, sabotage detection, alignment research automation, misaligned organizations, exploration hacking, and conditional emergent misalignment

open.substack.com

My AI Safety Paper Highlights of Feb & Mar 2026: - *Benchmarking auditors* - Functional emotions - Emergent vs narrow misalignment - Scheming propensity - Lenient self-monitors - CoT controllability - Unfilterable data poison - Boundary-point jailbreak aisafetyfrontier.substack.com/p/paper-high...

Paper Highlights of February & March 2026

Benchmarking auditors, functional emotions, emergent vs narrow misalignment, scheming propensity, lenient self-monitors, CoT controllability, unfilterable data poisoning, and boundary-point jailbreaks

aisafetyfrontier.substack.com

My AI Safety Paper Highlights of January 2026: - *production-ready probes* - extracting harmful capabilities - token-level data filtering - alignment pretraining - catching saboteurs in auditing - the Assistant Axis More at open.substack.com/pub/aisafety...

Paper Highlights of January 2026

Production-ready probes, extracting harmful capabilities, token-level data filtering, alignment pretraining, catching saboteurs in auditing, and the Assistant Axis

open.substack.com

My AI Safety Paper Highlights of December 2025: - *Auditing games for sandbagging* - Stress-testing async control - Evading probes - Mitigating alignment faking - Recontextualization training - Selective gradient masking - AI-automated cyberattacks More at open.substack.com/pub/aisafety...

Paper Highlights of December 2025

Auditing games for sandbagging, stress-testing async control, evading probes, mitigating alignment faking, recontextualization training, selective gradient masking, and AI-automated cyberattacks

open.substack.com

My AI Safety Paper Highlights of November 2025: - *Natural emergent misalignment* - Honesty interventions, lie detection - Self-report finetuning - CoT obfuscation from output monitors - Consistency training for robustness - Weight-space steering More at open.substack.com/pub/aisafety...

Paper Highlights of November 2025

Natural emergent misalignment, honesty interventions, self-report finetuning, CoT obfuscation from output monitors, consistency training for robustness, and weight-space steering

open.substack.com

My AI Safety Paper Highlights of October 2025: - *testing implanted facts* - extracting secret knowledge - models can't yet obfuscate reasoning - inoculation prompting - pretraining poisoning - evaluation awareness steering - auto-auditing with Petri More at open.substack.com/pub/aisafety...

Paper Highlights of October 2025

Testing implanted facts, extracting secret knowledge, models can't yet obfuscate, inoculation prompting, pretraining poisoning, evaluation awareness steering, and Petri

open.substack.com

My AI Safety Paper Highlights for August '25: - *Pretraining data filtering* - Misalignment from reward hacking - Evading CoT monitors - CoT faithfulness on complex tasks - Safe-completions training - Probes against ciphers More at open.substack.com/pub/aisafety...

Paper Highlights, August '25

Pretraining data filtering, misalignment from reward hacking, evading CoT monitors, CoT faithfulness on complex tasks, safe-completions training, and probes against ciphers

open.substack.com

Paper Highlights, July '25: - *Subliminal learning* - Monitoring CoT-as-computation - Verbalizing reward hacking - Persona vectors - The circuits research landscape - Minimax regret against misgeneralization - Large red-teaming competition - gpt-oss evals open.substack.com/pub/aisafety...

Paper Highlights, July '25

Subliminal learning, monitoring CoT-as-computation, verbalizing reward hacking, persona vectors, the circuits research landscape, minimax regret, a red-teaming competition, and gpt-oss evals

open.substack.com

AI Safety Paper Highlights, June '25: - *The Emergent Misalignment Persona* - Investigating alignment faking - Models blackmailing users - Sabotage benchmark suite - Measuring steganography capabilities - Learning to evade probes open.substack.com/pub/aisafety...

Paper Highlights, June '25

Emergent misalignment persona, investigating alignment faking, models blackmailing users, sabotage benchmarks, steganography capabilities, and evading probes

open.substack.com

“If an automated researcher were malicious, what could it try to achieve?” @gasteigerjo.bsky.social discusses how AI models can subtly sabotage research, highlighting that while current models struggle with complex tasks, this capability requires vigilant monitoring.

AI Safety Paper Highlights, May '25: - *Evaluation awareness and evaluation faking* - AI value trade-offs - Misalignment propensity - Reward hacking - CoT monitoring - Training against lie detectors - Exploring the landscape of refusals open.substack.com/pub/aisafety...

Paper Highlights, May '25

Evaluation awareness and faking, AI value trade-offs, misalignment propensity, reward hacking, CoT monitoring, training against lie detectors, and exploring the landscape of refusals

open.substack.com

New Anthropic blog post: Subtle sabotage in automated researchers. As AI systems increasingly assist with AI research, how do we ensure they're not subtly sabotaging that research? We show that malicious models can undermine ML research tasks in ways that are hard to detect.

Bild