Kyle O’Brien

@kyletokens.bsky.social

AGI Alignment & Control Pretraining @ Geodesic Research | https://kyobrien.io

Geodesic mission is to develop the science of providing robustly aligned initializations for RL, where alignment priors persist through the remainder of training. Do considering applying if you want to help make sharing the world with superintelligence go well for humanity.

Geodesic Research@geodesicresearch.bsky.social · 3mo ago

Geodesic is hiring Members of Technical Staff. We're a Cambridge-based AI safety org. Our seminal work showed you can bake alignment priors into base models. Now, we want to make base models robust to the adversarial effects of long-horizon capabilities RL. EOI ~5 mins: tally.so/r/vG4G6A

Very excited to see pretraining safety efforts! We’re only now beginning to understand how promising pretraining safety and alignment interventions are. Much in the way that curating the base model is important for capabilities like reasoning, so too might it be important for safety.

Katherine Lee@katherinelee.bsky.social · 8mo ago

Here's our job posting for more info! jobs.ashbyhq.com/openai/d829b... Please tell me a little about yourself when you email!

I've joined Geodesic Research to build the open-science field of AI safety pretraining research. Our first paper is wild. TL;DR — LLMs pretrained on data about misaligned AIs themselves become less aligned. Luckily, pretraining LLMs with data about good AIs helps them become more aligned.

Alignment Pretraining: AI Discourse Causes Self-Fulfilling (Mis)alignment

LLMs trained on data about misaligned AIs themselves become less aligned. Luckily, pretraining LLMs with synthetic data about good AIs helps them become more aligned. These alignment priors persist th...

alignmentpretraining.ai

Applications to apply for the ERA:AI Fellowship close November 3rd! Participating in this Summer's fellowship was my gateway into pursuing AGI safety research full-time. I will be a research manager for the upcoming Winter fellowships. Feel free to DM me with questions. :) erafellowship.org

ERA Fellowship

ERA is a talent programme supporting early-career researchers and entrepreneurs to understand and mitigate risks from frontier AI, based at Cambridge, UK.

erafellowship.org

I like that OpenAI published this. They were able to fine-tune away GPT-oss's refusal, decreasing refusal rates to ~0%. These results aren't surprising. Acknowledging that existing safeguards don't generalize to open models is the first step in developing solutions. arxiv.org/abs/2508.031...

Estimating Worst-Case Frontier Risks of Open-Weight LLMs

In this paper, we study the worst-case frontier risks of releasing gpt-oss. We introduce malicious fine-tuning (MFT), where we attempt to elicit maximum capabilities by fine-tuning gpt-oss to be as ca...

arxiv.org