1/11 Seq2exp models can predict how any DNA variant changes gene exp.However, in reality,sometimes they nail it,sometime sthey get the direction flat wrong (true for every model).Right now you have no way to tell which is which for any single prediction. So we built one (gRely)
Abdul Muntakim Rafi
@muntakimrafi.bsky.social
PhD candidate @SBME_UBC | Machine Learning | Gene regulation
Sharing Carl's thread on our preprint. Give it a read and tell us what you think—we're excited about the next steps. Grateful to everyone who kept us going. 😀
There is widespread homology across our genomes. How does this impact genomics models? Major update to our preprint! 🎉 We tested more models and more tasks; the overall findings were robust, but one conclusion on how future models should be trained differs⚠️ www.biorxiv.org/content/10.1...
1/14 Sequence-to-expression (S2E) models keep getting better at reading cis-regulatory logic. But they haven't solved it. On tasks like variant effect prediction they're still far from accurate. Remember, solving cis-regulation is the goal and we're not going to settle for less!
Why can't we explain enhancer action despite 2 decades of chromosome conformation technologies? 😬 Our new study spearheaded by Leonid Mirny's group points to a flaw in our assumptions, and to a solution from physical principles By @timothyfoldes.bsky.social 💻& @karissalhansen.bsky.social 🧪 🧵👇
must read if you are interested in MPRAs to test variants and/or training seq2exp models on them
MPRAs are the gold-standard tool for measuring how DNA sequences drive gene expression and prioritizing variant effects. In this preprint we asked: does it matter WHERE you place a variant in an MPRA? Spoiler: yes, and it might lead you to miss disease-causing variants. 1/6 doi.org/10.64898/202...
cool work!
Have you ever wondered 🤔... Does phenotypic variance respond to environmental perturbation? Does it have a genetic basis? Are mean and variance regulating loci exposed to different selection pressures? These and more questions are explored in our new preprint 🔥 www.biorxiv.org/content/10.6...
a lot of important benchmarks shown here @lxsasse.bsky.social @saramostafavi.bsky.social
Refining the cis-regulatory grammar learned by sequence-to-activity models by increasing model resolution
Chromatin accessibility can be measured genome-wide with ATAC-seq, enabling the discovery of regulatory regions that control gene expression and determine cell type. Deep genomic sequence-to-function ...
biorxiv.org
New (and hotly anticipated - at least by me) preprint from my group describing a better way to partition training data for genomic-trained models to solve the long-neglected problem of homology-based data leakage. Thread from first author @muntakimrafi.bsky.social 👇
0/ Essential reading for anyone training or using sequence-function models trained on genomic sequences! 🚨 In our new preprint, we explore the ways homology within genomes can cause leakage when training sequence-based models and ways to prevent it
0/ Essential reading for anyone training or using sequence-function models trained on genomic sequences! 🚨 In our new preprint, we explore the ways homology within genomes can cause leakage when training sequence-based models and ways to prevent it
Had a lot of fun at the CSHL Biological Data Science conference. Thanks to the scholarship from the "James P. Taylor Foundation for open science" for making it possible. #cshl
I am attending the Biological Data Science Meeting at CSHL. Will be giving a talk this Friday morning on the results from the Random Promoter DREAM Challenge. Will also be presenting a poster on a recent work where we address and solve the homology-based leakage in genome trained models.
Thrilled to share our research at the recent @KipoiZoo seminar! 🧬 We showed how chromosomal splitting of genome can cause train-test leakage through sequence homology and proposed a scalable solution to tackle it. Preprint coming soon! youtu.be/0_08qB0wLoM?...
Kipoi Seminar - Abdul Muntakim Rafi (University of British Columbia)
YouTube video by Kipoi Seminar
youtu.be
1/If you're training ML models on DNA sequences, u need to take a look at our new paper in @NatureBiotech! It contains analysis done by over 300 researchers, tells the story of how we built state-of-the-art for short regulatory DNA and developed a framework to keep improving