Johannes Hingerl

@johahi.bsky.social

ML for regulatory genomics. PhD student @ Gagneurlab johahi.github.io

Join us for our next Kipoi Seminar with Ruoyu Wang, Jian Zhou Lab, University of Chicago no recording! 👉 Title: Sequence-based regulatory code for heterogeneous and dynamic chromatin 🗓️ Wed Jul 1, 5:30pm CEST 🧬http://kipoi.org/seminar 🦋@kipoizoo.bsky.social

Excited to share Nona: a unifying multimodal masking framework for functional genomics. Models for DNA have evolved along separate paths: sequence-to-function (AlphaGenome), language models (Evo2), and generative models (DDSM). Can these be unified under a single paradigm? 1/15

Bild

Excited to share UKBBGym at #ASHG25, a new benchmark for variant effect predictors using WGS, proteomics and phenotypes from 500K UKBiobank participants. Stop by for insights on the impact of non-coding variants and how computational scores stack up against exp assays. Poster 5022W, Wed 2:30-4:30.

Bild

Happy to share that Flashzoi is now published! We enhanced Borzoi with RoPE & FlashAttention for >3x faster training/inference & 2.4x reduction in memory usage. This brings large-scale genomic analysis and fine-tuning within reach of academic budgets. 📄: doi.org/10.1093/bioi...

Flashzoi: an enhanced Borzoi for accelerated genomic analysis

AbstractMotivation. Accurately predicting how DNA sequence drives gene regulation and how genetic variants alter gene expression is a central challenge in

doi.org

Excited for a major milestone in our efforts to map enhancers and interpret variants in the human genome: The E2G Portal! e2g.stanford.edu This collates our predictions of enhancer-gene regulatory interactions across >1,600 cell types and tissues. Uses cases 👇 1/

In the genomics community, we have focused pretty heavily on achieving state-of-the-art predictive performance. While undoubtedly important, how we *use* these models after training is potentially even more important. tangermeme v1.0.0 is out now. Hope you find it useful!

Update of our protein outlier caller PROTRIDER. We now handle missing values, a widespread issue for mass spec where missing values are not a random -- and this improves outlier detection on non-missing data! Thumbs up to Daniela and George for the great work. doi.org/10.1101/2025...

Bild
Gagneur lab@gagneurlab.bsky.social · 2y ago

Excited to share that PROTRIDER, our method to call outliers on mass spectrometry-based proteomics data, is out now!! #proteomics #massspectrometry #raredisease doi.org/10.1101/2025...

Very proud of two new preprints from the lab: 1) CREsted: to train sequence-to-function deep learning models on scATAC-seq atlases, and use them to decipher enhancer logic and design synthetic enhancers. This has been a wonderful lab-wide collaborative effort. www.biorxiv.org/content/10.1...

CREsted: modeling genomic and synthetic cell type-specific enhancers across tissues and species

Sequence-based deep learning models have become the state of the art for the analysis of the genomic regulatory code. Particularly for transcriptional enhancers, deep learning models excel at decipher...

biorxiv.org

In today's poster session #probgen25. To the pop gen folks, interesting observation: The influence of a nucleotide on reconstructing others, rather than its own reconstructability, is a better predictor of function. This metric makes DNA LMs beat conservation in several benchmarks.

Bild
Gagneur lab@gagneurlab.bsky.social · 2y ago

and @pedrotomazdasilva.bsky.social will present tomorrow at #probgen25 poster 128 on dependency analysis of DNA language models. Come and see what functional relationships DNA LMs capture, from regulatory code to RNA structures. Preprint: doi.org/10.1101/2024...