Discovery of a 14-protein biomarker that predicts lung cancer 5.6 years before it is diagnosed, even in non-smokers, and an anti-inflammatory medicine that prevents its progression. And, challenging dogma, the proteins are not coming from cancerous cells!
Babak Alipanahi
@babaka.bsky.social
computational biology, artificial intelligence and large-scale datasets to improve human health
A new feature @science.org on the clusters of cells that enhance the spread of cancer, and what can be done to break them up www.science.org/content/arti...
Were you inspired by our paper on the genetics of GLP-1 drug response? www.nature.com/articles/s41... Want to make impactful discoveries with the world's best genetic dataset? We're hiring! StatGen: tinyurl.com/ys4mvhej Risk Prediction: tinyurl.com/psamt294 Data Products: tinyurl.com/2ff7eavb
Genetic predictors of GLP1 receptor agonist weight loss and side effects - Nature
Identification of genetic variants associated with the efficacy and side effects of GLP1 medications could underpin development of precision medicine approaches in the treatment of obesity.
nature.com
Fantastic results from 23andMe team led by amazing @adamauton.bsky.social on the efficacy of GLP-1 drug response.
Genomics of GLP-1 drug response and side effects With genomic and demographic data it's possible to predict magnitude of weight loss response @23andme23.bsky.social nature.com/articles/s41...
Delighted to share our latest research from the 23andMe Research Team, just published in @nature.com ! We looked at data from >27,000 participants to uncover how human genetics influences weight loss efficacy and side effects of GLP-1 medications like semaglutide. A short thread 🧵👇
A national screening programme increased the number of people diagnosed with early-stage lung cancer in England over half a decade go.nature.com/4veTAjw
Huge lung-cancer screening campaign boosts early diagnosis
A programme that offers scans to smokers between the ages of 55 and 74 detects a large number of early-stage lung tumours.
go.nature.com
"[Some], model in hand, go around looking around for problems to solve. Some of them have realized that DNA sequencing is super-efficient and is generating reams of data, so they’ve latched onto this." "This is not how good science gets done. You don’t start with a tool and then look for problems."
Is it AI that needs the dose of skepticism or the pervasive reductionism in genetics & genomics? stevensalzberg.substack.com/p/ai-is-star... #genetic #genomics #ai
Congratulations to 3rd-year undergrad Steven Tan (first author!!), @benlangmead.bsky.social, @mohsenzakeri.bsky.social, & @sinamajidian.bsky.social on their Best Paper Award at @acm-bcb.bsky.social 2025! 🏆
Langmead Lab team recognized with Best Paper Award at ACM-BCB 2025
The winning paper presents a software tool for quickly and accurately identifying which species’ DNA are present in a sample.
cs.jhu.edu
🚨🧬🚨 We're looking for a specialized Research Assistant to spearhead a project aimed at exploring new RNA therapeutics in diabetes and other metabolic disorders. Junior or Senior, we're hoping to recruit a highly motivated and self-driven scientist interested in both basic and applied science. 🦄
142,000 participants! The largest randomized trial of a multi-cancer early detection (MCED) test failed its primary endpoint genomeweb.com/cancer/grail... Had the enrollment been risk-based instead of age 50, it would likely have been very positive.
Bluesky is the new science Twitter, new study by @whysharksmatter.bsky.social and Julia Wester concludes! "Results show that for every reported professional benefit that scientists once gained from Twitter, scientists can now gain that benefit more effectively on Bluesky than on Twitter."
Scientists no Longer Find Twitter Professionally Useful, and have Switched to Bluesky
Synopsis. Social media has become widely used by the scientific community for a variety of professional uses, including networking and public outreach. For
academic.oup.com
There’s growing agreement in the research that non-smoking #lungcancer is increasing and that our approaches to understanding and early detection haven’t kept pace. This Trends in Cancer piece brings a lot of that thinking together in one place. #LCSM 🔗 www.cell.com/trends/cance...
Is it time for a new standard of care when it comes to mammograms in the era of AI? erictopol.substack.com/p/why-all-ma...
Why All Mammograms Should Incorporate A.I.
A very impressive body of evidence has accumulated
erictopol.substack.com
AI hallucinations in science manuscripts are a nuisance. Paranormal citations, or paracites, will be a nightmare. www.biorxiv.org/content/10.6... (w/ @sina.bio & @lauraluebbert.com).
Time for a thread on our Christmas preprint “Origin and evolution of acrocentric chromosomes in human and great apes”. I had so much fun with this project and paper. It will be hard to summarize in a thread, but I’ll try www.biorxiv.org/content/10.6... [1/21]
A good, enjoyable paper on AUROC vs AUPRC under class imbalance. In a nutshell, AUPRC's superiority is a myth. AUROC with bootstrapping all the way! arxiv.org/abs/2401.06091
A Closer Look at AUROC and AUPRC under Class Imbalance
In machine learning (ML), a widespread claim is that the area under the precision-recall curve (AUPRC) is a superior metric for model comparison to the area under the receiver operating characteristic...
arxiv.org
TF-MINDI is out! A new method to learn cis-regulatory codes through rich embeddings of TF binding sites. TF-MINDI decomposes motif neighbourhoods, and works downstream of any sequence-to-function deep learning model. We deeply study the enhancer code in human neural development, check out the thread
We are thrilled to share our new pre-print: “System-wide extraction of cis-regulatory rules from sequence-to-function models in human neural development”. S2F-deeplearning models can accurately encode enhancers, yet decoding these models into human-interpretable rules remains a major challenge.
Vaccines, the most impressive public health intervention in medical history, and where we could be headed if there was not efforts to negate truth, facts, and evidence A great, open-access, review and perspective by @scientificdiscovery.dev
Now published in Algorithms for Molecular Biology: link.springer.com/article/10.1.... Key message: a tiny CNN model with 7k parameters can capture main splice signals across vertebrates+insect and halves the minimap2 & miniprot junction error rate. I always use this new feature now.
Preprint on "Improving spliced alignment by modeling splice sites with deep learning". It describes minisplice for modeling splice signals. Minimap2 and miniprot now optionally use the predicted scores to improve spliced alignment. arxiv.org/abs/2506.12986
Now published in gigascience: academic.oup.com/gigascience/.... Key messages: SVs are highly enriched in low-complexity/tandem-repeat regions and are harder to call. They behave differently from transposon insertions. Always stratify if you study SVs.
Validate User
academic.oup.com
Do you know ~60% of human SVs fall in ~1% of GRCh38? See our new preprint: arxiv.org/abs/2509.23057 and the companion blog post on how we started this project and longdust: lh3.github.io/2025/09/29/o.... Work with Alvin Qin
Excited to see this out www.nature.com/articles/s41...! Nonparametric kernel-based tests for spatially variable isoform usage in spatial transcriptomics. So many interesting examples in the CNS and cancer, we're only scratching the surface!
Mapping isoforms and regulatory mechanisms from spatial transcriptomics data with SPLISOSM - Nature Biotechnology
Differential isoform usage is identified with high statistical power from spatial transcriptomics data.
nature.com
Nature research paper: Uncovering the role of LINE-1 in the evolution of lung adenocarcinoma go.nature.com/4oUHIPb
Uncovering the role of LINE-1 in the evolution of lung adenocarcinoma - Nature
Lung adenocarcinomas bearing the ID2 mutational signature display increased LINE-1 retrotransposon activity, which contributes to their fast evolutionary dynamics and aggressive phenotype.
go.nature.com
That’s a wrap on #SABCS25! Thank you to Dr. Lee Schwartzberg for presenting data demonstrating our platform’s ability to detect early stage breast cancer with high accuracy. #AI #RNA #earlydetection Learn more here: www.exai.bio/publications...
Important new, large (N>28,000 women) randomized clinical trial of breast cancer screening: age-based vs risk-based by polygenic risk score, genomics "opportunity to modernize screening" jamanetwork.com/journals/jam...
Risk-Based vs Annual Breast Cancer Screening
This randomized clinical trial examines whether risk-based screening is a safe and effective alternative to annual mammography for detecting breast cancer in women 40 years and older.
jamanetwork.com
Even though the highest-profile names in today’s corporate Cambridge are in biotech and software, the influx of defense startups hearkens back to an earlier era — which, in 1922, saw the birth of Raytheon, now synonymous with the old guard of defense contractors. www.thecrimson.com/article/2025...
As Cambridge Faces a Life Sciences Downturn, Startups Turn to a New Industry: Warfare | News | The Harvard Crimson
As biotech firms shed jobs and life sciences funding dries up, policymakers have started to see defense technology as a way to buttress the Massachusetts economy. Industry experts say Cambridge may be...
thecrimson.com
SCIENCE SAVES LIVES. Overall pediatric cancer survival rate increased from 63% in mid 1970s to 87%‼️between 2015 & 2021. And this isn’t due to supplements, eating better or avoiding red food dye. It’s due to science & industry working together to develop & approve therapies!
JASPAR 2026 is out 🎉 The new release massively expands the TF motif collections and adds a dedicated DeepLearning collection of motifs learned from deep learning models. Database: jaspar.elixir.no Paper (NAR): doi.org/10.1093/nar/... 🧵1/2
JASPAR: An open-access database of transcription factor binding profiles
JASPAR is the largest open-access database of curated and non-redundant transcription factor (TF) binding profiles from six different taxonomic groups.
jaspar.elixir.no
579 high-quality human genomes from @humanpangenome.bsky.social, Arab Pangenome and individual papers (CHM13, CN1, KSA001, I002C, YAO and KOREF1). Sequences available in the AGC format (3.7GB) and FM-index in the ropebwt3 format (20.3GB). For details, see github.com/lh3/human-asm
GitHub - lh3/human-asm: A collection of high-quality human genomes
A collection of high-quality human genomes. Contribute to lh3/human-asm development by creating an account on GitHub.
github.com
A nice paper on distilling AI-based splicing models into much simpler additive models: "[...] the distilled models achieve this without modeling RNA structure or feature interactions, indicating that [AI]-based splicing models recognize exons primarily through simple additive sequence features."
Interpretable Distillation Reveals that Deep-learning-based Splicing Models Suffer from Pervasive Confounders and Blind Spots [new] Splicing models rely on confounders, failing on non-reference sequences, revealing training limits.
Abbott Laboratories is nearing a potential acquisition of Exact Sciences Corp, in what would be its largest deal in nearly a decade, people familiar with the matter said. www.bloomberg.com/news/article...
Abbott Nears Deal for Cancer Test Maker Exact Sciences
Abbott Laboratories is nearing a potential acquisition of medical-testing company Exact Sciences Corp., in what would be its largest deal in nearly a decade, people familiar with the matter said.
bloomberg.com