Gonzalo Benegas

@gonzalobenegas.bsky.social

Research Scientist @ Open Athena | AI for Science https://gonzalobenegas.github.io/

With the generous support of The Jen-Hsun and Lori Huang Foundation, we have launched the largest live-streamed pretraining run in history, and Marin’s largest model yet: a fully open source 535B total parameter MoE model, with 18T tokens of data. Learn more at: openathena.ai/blog/huang-f...

Marin's 535 billion parameter model training run launched with support of The Jen-Hsun and Lori Huang Foundation GPU gift

Open Athena announces the launch of Marin's largest training run yet—a 535B parameter large language model—with a generous gift of compute from The Jen-Hsun and Lori Huang Foundation.

openathena.ai

Here's the tale of how @jder.bsky.social and I scaled Samudra, a neural ocean emulator capable of predicting 8 years of the ocean on a single GPU, to operate at a full 1/4° resolution (16x the size in bytes). It was quite a humbling process.

Open Athena@openathena.ai · 2mo ago

Simulating ocean climate takes a supercomputer 4,600+ CPU cores to produce 12 simulated years per day (SYPD). Samudra 2 produces 4,800 SYPD on 1 GPU at the same resolution. In a new blog, @al.merose.com reports on Samudra, a neural ocean emulator built in collaboration with NYU & MIT: bit.ly/oa-ss

I am thrilled to announce that in January 2026 I will be starting my own lab at NYU Biology! Soon enough I will be recruiting postdocs and students! Please reach out if you are interested with a CV and description of your research interests, or if you know of people who could be interested! 🧬🗽 🦊

Thrilled to see my digital art on the cover of Trends Genet. The two binary strings represent reverse-complementary DNA sequences (00=A, 01=C, 10=G, 11=T) and the connecting rectangles represent “embeddings” learned by DNA language models. Pls check out our article as well: doi.org/10.1016/j.ti...

Bild

Can DNA sequence models predict mutations affecting human traits? We introduce TraitGym, a curated benchmark of causal regulatory variants for 113 Mendelian & 83 complex traits, and evaluate functional genomics and DNA language models. Joint work w/ Gökcen Eraslan and @yun-s-song.bsky.social 🧵👇

Bild
bioRxiv Genetics@biorxiv-genetic.bsky.social · 2y ago

Benchmarking DNA Sequence Models for Causal Regulatory Variant Prediction in Human Genetics https://www.biorxiv.org/content/10.1101/2025.02.11.637758v1

Large protein language models can learn complex epistatic interactions, but how much does that help with predicting variant effects? In this NeurIPS article, we show that classical independent-sites phylogenetic models can outperform pLMs on this task. 1/7 openreview.net/forum?id=H7m...

Ultrafast classical phylogenetic method beats large protein...

Amino acid substitution rate matrices are fundamental to statistical phylogenetics and evolutionary biology. Estimating them typically requires reconstructed trees for massive amounts of aligned...

openreview.net