Stein Aerts

@steinaerts.bsky.social

Computational biologist interested in deciphering the genomic regulatory code at vib.ai

Congratulations @seppedewinter.bsky.social @davidmauduit.bsky.social and Gabriele Partel for your vision and hard work to translate our sequence-to-function modeling research into CellTuned, let’s go

VIB.AI@vibai.bsky.social · last mo.

The @steinaerts.bsky.social lab received an ERC Proof of Concept Grant to develop CellTuned: an AI platform that combines sequence-to-function models with single-cell regulatory networks to uncover the mechanisms driving disease, and translate them into new therapeutic opportunities. 👏 celltuned.ai

Very proud of this and so cool that enhancer-level models can predict the effect of genetic variation. There is so much personal variation in terms of gene regulation in the human brain, it is fantastic to uncover this thanks to technology (whole-genome sequencing and single-cell multiomics) and AI

Alexandra P@alexanrna.bsky.social · 4mo ago

1/ 🧬 Happy to share our new preprint on modeling cis-regulatory variation in human brain enhancers across a large Parkinson’s disease cohort: www.biorxiv.org/content/10.6... Details in the thread below:

In addition to the bioRxiv this is also pilot for a new interactive preprint developed by @curvenote.com w/ support from @hhmi-science.bsky.social including directly embedded Jupyter notebooks for fig reproduction, data, models, prediction tracks, code, etc shendure.curve.space/articles/evo...

Evolutionary transfer learning enables organism-wide inference of mammalian enhancer landscapes

Understanding and modeling how the human genome encodes gene regulatory programs for thousands of cell types remains a central challenge in genomics and machine learning. However, most human cell types emerge during embryonic, fetal, and pediatric development which are inaccessible to comprehensive molecular profiling. To overcome this, we hypothesized that the mismatch in evolutionary rates between cis-acting enhancers (fast) and the trans-acting regulatory programs that interpret them (slow) creates an opportunity for ‘evolutionary transfer learning’. Specifically, models trained to predict cell type-specific enhancers in one species should generalize to the orthologous cell types and enhancers of related species. To test this, we generated a single-cell atlas of chromatin accessibility spanning mouse embryonic day 10 (E10) to birth (P0). Using combinatorial indexing1, we profiled 3.9 million nuclei from 36 staged embryos, resolving genome-wide accessibility in 36 cell classes and 140 cell types. With the goal of identifying distal enhancers for all cell classes, we trained a series of multi-output deep learning models (CREsted2), each addressing limitations of the preceding approach. An ‘evolution-naive’ model achieves strong performance on heldout peaks, but exhibited two failure modes during genome-wide inference: overprediction at tandem repeats and conflation of promoter and distal enhancer grammars. An ‘evolution-aware’ model resolves these by regrouping accessible regions based on functional coherence across syntenic orthologs, but fails to generalize across species — suggesting insufficient sequence diversity during training. Finally, STEAM (Synteny-aware Transfer learning for Enhancer Activity Modeling), our ‘evolution-augmented’ model, expands the training corpus to include enhancer orthologs from up to 241 mammalian genomes (Zoonomia3) in a synteny-supervised manner. This increases the effective data scale by up to 195-fold, markedly improving generalization across mammals despite greater label noise. We apply STEAM predict enhancers for all major developmental lineages throughout the human, mouse (HumMus) and 239 additional mammalian genomes3 (BabaGanoush), i.e. 32 × 241 = 7,712 genome-wide enhancer tracks. Together, our results unify advances in single-cell profiling, deep learning, and comparative genomics into a framework for the evolutionary transfer learning of noncoding regulatory grammars. More broadly, our work supports the view that model organisms and evolutionarily diverse genomes are indispensable resources for accelerating the AI-enabled exploration of human biology.

shendure.curve.space

Last summer I spent 4 months working at the @alleninstitute.org as a Visiting Scientist. Recently we released some preprints about the work we collaborated on, where from new multiome atlases of CNS regions we tried to decipher underlying enhancer logic with CREsted (among many other things). (1/n)

Bild

Big congrats to the entire Kaessmann lab for this spectacular achievement and beautiful insights. It was a great honour to contribute to this study and to host Ioannis in our lab, an absolutely brilliant scientist. Evolution of genomic enhancers controlling neuronal cell types is just too cool..

Kaessmann Lab@kaessmannlab.bsky.social · 6mo ago

We are thrilled that our study on the evolution of gene regulation in mammalian cerebellum development – led by @ioansarr.bsky.social, @marisepp.bsky.social and @tyamadat.bsky.social, in collaboration with @steinaerts.bsky.social – is now out in @ScienceMagazine! www.science.org/doi/10.1126/...

Hydrop-v2 is now published ! Allows generating cheap scATAC-seq training data for enhancer modeling with CREsted. Make sure to check out the 600K cell atlas of the last 4 hours of Drosophila embryo development. Fun to use bioML for technology benchmarking :)

Hannah Dickmänken@hannahdckmnkn.bsky.social · 6mo ago

Paper alert! 💻 How many cells do you need to train reliable deep learning models in regulatory genomics? We asked how data quality, sequencing depth, and dataset size affect training of sequence-to-function models from scATAC-seq. Out now www.nature.com/articles/s41... (details below)

To test the sufficiency of the TF-MINDI extracted enhancer code rules we turn to synthetic enhancer design in facial mesenchyme cells. A homeobox-ebox dimer motif (Coordinator) has been shown to be instrumental for this cell type. TF-MINDI identified Coordinator instances at varying affinities.

tSNE dimensionality reduction of facial mesenchyme TF-MINDI seqlets colored based on TF-family. The coordinator instances are circled and an arrow drawn to a PCA of those coordinator instances colored based on coordinator motif score. This shows that TF-MINDI captures multiple coordinator affinities. For each affinity bin a TF binding motif logo is shown.

TF-MINDI is out! A new method to learn cis-regulatory codes through rich embeddings of TF binding sites. TF-MINDI decomposes motif neighbourhoods, and works downstream of any sequence-to-function deep learning model. We deeply study the enhancer code in human neural development, check out the thread

Bild
Seppe De Winter@seppedewinter.bsky.social · 7mo ago

We are thrilled to share our new pre-print: “System-wide extraction of cis-regulatory rules from sequence-to-function models in human neural development”. S2F-deeplearning models can accurately encode enhancers, yet decoding these models into human-interpretable rules remains a major challenge.