Gennady Gorin
@goringennady.bsky.social
🦠🧬📊bioinformatics, statistics, and stochastic processes.
TL;DR: We've identified more than 100 cases of apparent manipulation in Thermo Fisher Scientific's antibody verification data. @sholtodavid.bsky.social @johanduchene.bsky.social reeserichardson.blog/2026/05/28/h...
How much of Thermo Fisher’s antibody data has been manipulated?
We’ve documented more than 100 instances of apparent data manipulation in Thermo’s catalog
reeserichardson.blog
I'm a former Community Notes super-user interviewed for this piece. Whether the notes are written by people or language-mimicking algorithms, the central flaw of Twitter's Community Notes is the same as it's always been: it's a "fact-checking" system that doesn't involve the checking of facts.
Today on @indicator.media: A handful of anonymous bots have taken over fact-checking on X. This isn’t hyperbole. In the first three weeks of May, just eight AI contributors wrote 50.3% of all visible Community Notes on the platform.
Efficient Stochastic Trace Generation for Transcription https://www.biorxiv.org/content/10.64898/2026.05.05.722871v1
they should invent software updates that make the software better over time instead of worse
Arc Institute’s “MULTI-evolve” learns a classical additive model, not epistasis. Preprint (w/ Gian Marco Visani and Aayush Verma): www.biorxiv.org/content/10.6... Blog: dewitt-lab.github.io/posts/2026-0...
Additivity is all you need?
Arc Institute’s “MULTI-evolve” learns a classical additive model, not epistasis.
dewitt-lab.github.io
Super excited to see this out! Fantastic collaboration with Luke O'Connor and trainees Amber Shen and Xinran Wang. Thread with details will come soon, but linear ARG provide a HIGHLY efficient representation of genotype data that can be treated as a linear operator www.biorxiv.org/content/10.6...
biorxiv.org
How every layer of science's "self-correcting machinery" failed when Iva Veseli and I simply wanted to reproduce the findings of a high-profile study on gut microbiome and autism: merenlab.org/2026/04/15/u...
Unfalsifiable by Design: A Year of Trying and Failing to Reproduce a Human Microbiome and Autism Study
The myth of open data, reproducibility, responsibility, and accountability in science, and your role in it
merenlab.org
If you use dim. reduction, you may be interested in two recent preprints we've posted on contrastive PCA: The Rayleigh Quotient and Contrastive Principal Component Analysis I & II w/ Maria Carilli & Kayla Jackson. They cover a lot of ground from theory to practice. 1/🧵
Latest from Shendure & Qiu labs (@cxqiu.bsky.social) )! We combined a new 4M cell mouse whole embryo scATAC-seq atlas (E10-P0), millions of 'evolutionarily coherent' orthologs from 241 mammalian genomes (Zoonomia), and the CREsted CNN framework (@steinaerts.bsky.social).
Friendly reminder that ordinal values admit an ordering, but no notion of distance. Without a notion of distance even the concept of linear models is ill-defined. Please do not use ordinary least squares to analyze ordinal data. For gratuitous discussion see betanalpha.github.io/assets/chapt....
Ordinal Modeling
betanalpha.github.io
You know what's better than inflating the variation of your observational model to heuristically accommodate "outliers"? Actually modeling the contaminating data generating process.
Eukan: a fully automated nuclear genome annotation pipeline for less studied and divergent eukaryotes academic.oup.com/nargab/artic... 🧬💻🧪 github.com/BFL-lab/eukan
”Early Modern Memes: The Reuse and Recycling of Woodcuts in 17th-Century English Popular Print“ by @katiesisneros, on the interplay of repetition, context + meaning in woodcuts and the parallels to meme culture of today: publicdomainreview.org/essay/e...
Ambient RNA & barcode swapping is a serious issue in single-cell genomics. Tools such as CellBender, scAR, DecontX & SoupX. We have developed CellSweep which is faster (in some cases by a lot) and much more accurate. Extensively tested and benchmarked. www.biorxiv.org/content/10.6... 1/
New COSIG update! 3️⃣4️⃣
Pseudoreplication is the practice of treating non-independent observations as if they were independent replicates. This can dramatically increase the rate of false positives. COSIG's newest entry covers how to spot them. Read it at osf.io/hyxvr COSIG 🎉now with 34 guides 🎉is available at cosig.net.
The empty drops you threw out in your single-cell RNA sequencing analysis might be hiding mysterious things 👀 check out our bioRxiv preprint! The empty drops do not contain cells. Yet we can still use them to learn interesting things about biology and technology. 1/
Empty drops in scRNA-seq uncover the surprising prevalence of sequestered neuropeptide mRNA and pervasive sequencing artifacts https://www.biorxiv.org/content/10.64898/2026.02.13.705850v1
Excited to share this preprint that describes my latest work on using GPUs to accelerate processing of RNA-seq data. The title says it all: "RNA-seq analysis in seconds using GPUs" now on biorxiv www.biorxiv.org/content/10.6... and github github.com/pachterlab/k... Figure 1 shows they key result
RNA-seq analysis in seconds using GPUs https://www.biorxiv.org/content/10.64898/2026.03.04.709526v1
How many samples should you sequence? Collect too few, and the experiment is inconclusive. Collect too many, and the costs add up very quickly. Check out our bioRxiv preprint and calculator at poweranalysis-fb.streamlit.app! 1/
DEPower
Working repo for DEPower, accessed by the Streamlit interface.
poweranalysis-fb.streamlit.app
DEPower: approximate power analysis with DESeq2 www.biorxiv.org/content/10.6... 🧬💻🧪
After years of work, the centerpiece of my PhD is published in @natmethods.nature.com! Read it to learn about the biophysical insights we can get from single-cell data! But first, I would like to talk a bit about RNA velocity and normalization. 1/
Monod fits biophysically motivated models to single-cell transcriptomics data, providing insights into gene expression dynamics. @goringennady.bsky.social @lpachter.bsky.social www.nature.com/articles/s41...
Hi everyone, I am looking for a new industry role in computational biology! Check out my portfolio of genomics, statistics, ML, and biophysics work at gennadygorin.github.io, and reach out if you have any suggestions or open roles!
Gennady Gorin, Ph.D.
Senior Scientist applying stochastic models for therapeutic discovery
gennadygorin.github.io
p-value of 1e-122 at a lfc of 0.67, huh
@biorxivpreprint.bsky.social Cannabis-induced immunomodulatory effects are profound, dynamic and complex among cell types and that transcriptional changes are regulated at least in part by epigenetic mechanisms www.biorxiv.org/content/10.6...
when your dataset is definitely real
Much of my work in meta-research is on finding problems in research, so I've seen a lot of bad practices. However, even I was shocked by hundreds of researchers publishing papers using data that is faked and has no data provenance. www.medrxiv.org/content/10.6.... Amazing work by my student Alex.
I drew a cover for 'Olenka' by Budi Darma, translated by Tiffany Tsao. It's out in July but you can preorder it now. www.penguinrandomhouse.com/books/760272/olenka-by-budi-darma-translated-with-an-introduction-and-notes-by-tiffany-tsao/
This is comical on account of well-known estate IP restrictions and also Salesforce's decadelong struggle to make Einstein AI a thing (paying $20M for the privilege)
Who owns Einstein? The battle for the world’s most famous face
The long read: Thanks to a savvy California lawyer, Albert Einstein has earned far more posthumously than he ever did in his lifetime. But is that what the great scientist would have wanted?
theguardian.com
Is this bad
We recently updated our paper demonstrating evidence of off-target probe binding affecting the 10x Genomics Xenium spatial transcriptomics platform with key clarifications, new quantifications, and approaches for evaluating custom gene panels: biorxiv.org/content/10.1... 🧵👇1/n
Widespread data leakage inflates performance estimates in cancer drug response prediction https://www.biorxiv.org/content/10.64898/2026.02.05.704016v1