Ahmad Beirami

@abeirami.bsky.social

stealth // Gemini RL+inference @ Google DeepMind // Conversational AI @ Meta // RL Agents @ EA // ML+Information Theory @ MIT+Harvard+Duke // Georgia Tech PhD 📍{NYC, SFO, YYZ} 🔗 https://beirami.github.io/

While everyone is focused on the safety risks exposed by the OpenAI/HF breach, the incident also highlights an equally important but less understood issue. Even in the non-adversarial case, unlike classical ML, we have no idea how to measure/prevent overoptimization when building agentic harnesses.

The best career advice used to be simple: Be the best at what you do. Ignore the trends. Become the best designer, the best engineer, the best researcher. Still necessary. No longer sufficient.

We are hiring Members of Technical Staff (Research Engineers)! Current LLM agents lack reliability, creating a gap between demos and production. We solve this by automating the complex workflow of debugging, evaluation, and iteration required to make agents robust. 👇

- Iran is in a humanitarian crisis. - Thousands are reported dead in 72 hours. - We are past the point of solidarity. Empty words do not stop bullets. Action does. - The world must intervene now.

For years, Iran’s masked plainclothes regime thugs have abducted and murdered citizens with absolute impunity for wanting prosperity and refusing to fear them. Officials call it law enforcement and smear protesters as paid agents of the state’s enemies. This must end. Iranian people must prevail!

Found myself repeating this to several students at NeurIPS: When you’re choosing an internship or a job, what you work on and who you work with matter way more than the logo. Don’t optimize for brands. Become the brand!

Hiring researchers & engineers to work on –building reliable software on top of unreliable LLM primitives –statistical evaluation of real-world deployments of LLM-based systems I’m speaking about this on two NeurIPS workshop panels: 🗓️Saturday – Reliable ML Workshop 🗓️Sunday – LLM Evaluation Workshop

Woke up to this email this morning - Wow, I won a NeurIPS award?! - …runner-up, but I’ll take it. - Wait, I didn’t submit a paper. - Ah, I’m chairing the session and I’m supposed to give the award. Huge congratulations to the actual winners and runners-up!

Bild

Will be at NeurIPS Thu Dec 4 to Sun Dec 7, excited to reconnect with old friends and make new ones. If you are excited about AI engineering (orchestration, evals, and optimizing scaffolds), we are hiring! On Saturday I’ll be on panels at the Reliable ML & UniReps workshops.

The actual unpopular opinion is that the notion of senior and junior authors should be abolished. It has completely diluted the notion of scientific authorship and created this entire industry of free-riding, head-in-the-clouds, incompetent PIs/managers. List down exact contributions instead. [+]

Ahmad Beirami@abeirami.bsky.social · 11mo ago

Unpopular opinion: When a paper has a senior mentor and a junior mentee, the senior author must make sure the claims are correct and well supported. They must check every claim and gate the submission until it meets that bar.

I occasionally get messages asking how to follow my path and get into Meta, DeepMind, or similar places. That is the wrong question. Do not focus on the brand! Focus on what you want to work on, then find the opportunity that fits your goals best.

Unpopular opinion: When a paper has a senior mentor and a junior mentee, the senior author must make sure the claims are correct and well supported. They must check every claim and gate the submission until it meets that bar.

This is the recipe for many provable claims: Make enough assumptions and narrow down the claim, then prove a narrow result with caveats. Present it as broad, hide the caveats, and declare “XYZ is provable!”

Today, Industry research is focused on short term (3-6months) bets. Academics have an opportunity to balance their portfolio with medium term (1-2 years) and long term (5-10 years) bets. Putting all academic efforts in short-term basket is suboptimal!

When I worked in corporate, I was often first in the office because that routine worked for me. It was a personal preference, not a benchmark for anyone else. We should not judge commitment by hours, especially in research. We should look for thoughtful work and steady progress.

Common mistake in LLM prompting projects: jumping into full-scale pipelines (datasets and inference) without testing feasibility. Iterating at scale is expensive and time-consuming. Start with ONE example to validate the hypothesis, verify context, debug the design, then scale.

Thoughts that are explained clearly are more respected. Clarity is a scarce skill. Many of us (me included) leave out key context and make people work too hard to understand us. AI models should get better at this: not just reasoning, but communicating with the right amount of context.