Babak Alipanahi

@babaka.bsky.social

computational biology, artificial intelligence and large-scale datasets to improve human health

Discovery of a 14-protein biomarker that predicts lung cancer 5.6 years before it is diagnosed, even in non-smokers, and an anti-inflammatory medicine that prevents its progression. And, challenging dogma, the proteins are not coming from cancerous cells!

Bild

Delighted to share our latest research from the 23andMe Research Team, just published in @nature.com ! We looked at data from >27,000 participants to uncover how human genetics influences weight loss efficacy and side effects of GLP-1 medications like semaglutide. A short thread 🧵👇

"[Some], model in hand, go around looking around for problems to solve. Some of them have realized that DNA sequencing is super-efficient and is generating reams of data, so they’ve latched onto this." "This is not how good science gets done. You don’t start with a tool and then look for problems."

Jason Moore@moorejh.bsky.social · 4mo ago

Is it AI that needs the dose of skepticism or the pervasive reductionism in genetics & genomics? stevensalzberg.substack.com/p/ai-is-star... #genetic #genomics #ai

🚨🧬🚨 We're looking for a specialized Research Assistant to spearhead a project aimed at exploring new RNA therapeutics in diabetes and other metabolic disorders. Junior or Senior, we're hoping to recruit a highly motivated and self-driven scientist interested in both basic and applied science. 🦄

Bild

Bluesky is the new science Twitter, new study by @whysharksmatter.bsky.social and Julia Wester concludes! "Results show that for every reported professional benefit that scientists once gained from Twitter, scientists can now gain that benefit more effectively on Bluesky than on Twitter."

Scientists no Longer Find Twitter Professionally Useful, and have Switched to Bluesky

Synopsis. Social media has become widely used by the scientific community for a variety of professional uses, including networking and public outreach. For

academic.oup.com

TF-MINDI is out! A new method to learn cis-regulatory codes through rich embeddings of TF binding sites. TF-MINDI decomposes motif neighbourhoods, and works downstream of any sequence-to-function deep learning model. We deeply study the enhancer code in human neural development, check out the thread

Bild
Seppe De Winter@seppedewinter.bsky.social · 7mo ago

We are thrilled to share our new pre-print: “System-wide extraction of cis-regulatory rules from sequence-to-function models in human neural development”. S2F-deeplearning models can accurately encode enhancers, yet decoding these models into human-interpretable rules remains a major challenge.

H

Now published in Algorithms for Molecular Biology: link.springer.com/article/10.1.... Key message: a tiny CNN model with 7k parameters can capture main splice signals across vertebrates+insect and halves the minimap2 & miniprot junction error rate. I always use this new feature now.

HHeng Li@lh3lh3.bsky.social · last yr.

Preprint on "Improving spliced alignment by modeling splice sites with deep learning". It describes minisplice for modeling splice signals. Minimap2 and miniprot now optionally use the predicted scores to improve spliced alignment. arxiv.org/abs/2506.12986

H

Now published in gigascience: academic.oup.com/gigascience/.... Key messages: SVs are highly enriched in low-complexity/tandem-repeat regions and are harder to call. They behave differently from transposon insertions. Always stratify if you study SVs.

Validate User

academic.oup.com

HHeng Li@lh3lh3.bsky.social · 10mo ago

Do you know ~60% of human SVs fall in ~1% of GRCh38? See our new preprint: arxiv.org/abs/2509.23057 and the companion blog post on how we started this project and longdust: lh3.github.io/2025/09/29/o.... Work with Alvin Qin

Excited to see this out www.nature.com/articles/s41...! Nonparametric kernel-based tests for spatially variable isoform usage in spatial transcriptomics. So many interesting examples in the CNS and cancer, we're only scratching the surface!

Mapping isoforms and regulatory mechanisms from spatial transcriptomics data with SPLISOSM - Nature Biotechnology

Differential isoform usage is identified with high statistical power from spatial transcriptomics data.

nature.com

SCIENCE SAVES LIVES. Overall pediatric cancer survival rate increased from 63% in mid 1970s to 87%‼️between 2015 & 2021. And this isn’t due to supplements, eating better or avoiding red food dye. It’s due to science & industry working together to develop & approve therapies!

H

579 high-quality human genomes from @humanpangenome.bsky.social, Arab Pangenome and individual papers (CHM13, CN1, KSA001, I002C, YAO and KOREF1). Sequences available in the AGC format (3.7GB) and FM-index in the ropebwt3 format (20.3GB). For details, see github.com/lh3/human-asm

GitHub - lh3/human-asm: A collection of high-quality human genomes

A collection of high-quality human genomes. Contribute to lh3/human-asm development by creating an account on GitHub.

github.com

A nice paper on distilling AI-based splicing models into much simpler additive models: "[...] the distilled models achieve this without modeling RNA structure or feature interactions, indicating that [AI]-based splicing models recognize exons primarily through simple additive sequence features."

AI x Bio Discovery@aixbiobot.bsky.social · 8mo ago

Interpretable Distillation Reveals that Deep-learning-based Splicing Models Suffer from Pervasive Confounders and Blind Spots [new] Splicing models rely on confounders, failing on non-reference sequences, revealing training limits.

Interpretable Distillation Reveals that Deep-learning-based Splicing Models Suffer from Pervasive Confounders and Blind Spots