Michael Kuhn

@biocs.bsky.social

Computational biologist. Research staff scientist and research coordinator in the @borklab.bsky.social at @embl.org Heidelberg.

New preprint from the lab! "Planetary structure and drivers of diazotroph communities reveal key reservoirs of nitrogen-fixation potential" Nitrogen fixation is scattered right across the prokaryotic tree — but it's heterotrophs, not cyanobacteria, that dominate the potential. 🧵

bioRxiv Microbiology@biorxiv-microbiol.bsky.social · last wk.

Planetary structure and drivers of diazotroph communities reveal key reservoirs of nitrogen-fixation potential https://www.biorxiv.org/content/10.64898/2026.07.27.741065v1

Looking for BGCs in large metagenomic datasets? Our new biorxiv preprint introduces metaSMASH, a scalable fork of antiSMASH designed specifically for metagenome-scale BGC detection and analysis : www.biorxiv.org/cgi/content/... Thanks @canerbagci.bsky.social and @kblin.bsky.social ❤️

metaSMASH: Scalable Biosynthetic Gene Cluster Detection for Large Metagenomic Assemblies

antiSMASH is widely used for biosynthetic gene cluster (BGC) detection and annotation, but its standard workflow is poorly suited to large metagenomic assemblies, where massive contig counts create severe runtime bottlenecks and complicate downstream result exploration. We present metaSMASH, a re-engineered fork of antiSMASH for metagenome-scale BGC analysis. metaSMASH preserves the original antiSMASH detection and annotation logic while introducing streaming, memory-bounded execution, record-level parallelisation, optional output filtering, and an interactive dashboard for large result sets. Across 25 benchmark metagenome datasets, metaSMASH reproduced identical BGC detection results while dramatically reducing computational cost. Relative to the default antiSMASH configuration, metaSMASH was a geometric-mean 38x faster. It also outperformed an ad hoc chunked antiSMASH workflow: in the default configuration it achieved a geometric-mean 2.9x speed-up and 1.7x lower peak memory, and with extended-analysis modules enabled it was 2.7x faster and used 3.1x less memory while completing all datasets, whereas the ad hoc workflow ran out of memory on the two largest assemblies. By substantially reducing the computational burden of large-scale metagenome analysis without sacrificing result equivalence, metaSMASH makes routine mining of assembled metagenomes more practical and provides a scalable foundation for natural product discovery from complex microbial communities. ### Competing Interest Statement The authors have declared no competing interest. German Center for Infection Research, TTU Novel Antibiotics 09.716 Volkswagen Foundation, 0072511-00

biorxiv.org

New lab preprint - Deep learning models are most often difficult to interpret black box models. What if we engineer them with interpretation in mind? Dennis Gankin carefully studied so-called biologically inspired neural networks and found an interesting phenomenon www.biorxiv.org/content/10.6...

Leveraging multiplicity in biologically informed neural networks to uncover disease heterogeneity

Biologically inspired neural networks (BINNs) embed pathway, ontology, or protein-interaction structure directly into neural networks, promising interpretable disease prediction where hidden nodes map...

biorxiv.org

I wish I could teach people just how grating it is to get the fourth clearly AI job application of the week that reads like "While I'm interested in X, to be honest what truly drew me to your lab was Y. You not only do A; you also do B."

Great to see this important study published. For some time I've pointed to the preprint version when students ask about co-evolution between humans and their microbiota. "Limited codiversification of the gut microbiota within humans"

BildBild

I've been thinking of how to make tools "AI-ready" so it's not just an empty buzzword I'd like to hear other opinions, but this is what I'm trying with SemiBin: 1) bundle skills with the tool SemiBin2 install-skills will (in the next release) install a skill for claude/codex/...

We've discussed this guide in the lab: zenodo.org/records/1845... and found it a good starting point to responsibile AI usage. Researchers still need to understand all of their pipelines and be answerable for their output, which is impossible if all of it is AI and agent generated

Expertise before augmentation: a practical guide to using generative AI during research training

This comprehensive implementation guide accompanies the framework described in "Build expertise first: why PhD training must sequence AI use after foundational skill development" (Krishnan, 2026, DOI:...

zenodo.org

From their examples, it unfortunately looks more like an abdication of domain expertise to AI rather than just handling friction points. Like the three sentence prompt to write a review article 🤢 and being proud of a host of AI agents having reviewed the papers rather than oneself

More evidence, from a large-scale study in China, that using AI hurts learning if it undermines mental effort. When homework time drops due to AI use, so do test scores. Across studies, there is a clear theme: AI tutoring in support of classes is good, using AI to "help" with homework is bad.

BildBildBild

When utilized in literature review, LLMs consistently 1. fail to mention female authors in female-led literatures, 2. insist that men are more influential or more heavily cited when this is contradicted by objective citation counts, and 3. attribute women’s work to hallucinated male scholars.

I'm attending a digital humanities event in Montreal. There was a keynote on AI and something about the talk made me wonder if it had been written by Claude. I said as much in the Q&A. As I posed the question, the speaker shifted, looking slightly uncomfortable. What he said next shocked the room +

I’ve officially resigned as Associate Editor for Frontiers in Systems Neuroscience. It used to be a reputable journal, but became a case study in how forced automation destroys academic integrity. 👇

My first EMBL project uncovering synthetic steroid bacterial metabolism is now available publicly! My favourite part is the discovery of novel desmolase enzyme - see the preprint how we have found it 🧬💊🦠

bioRxiv Microbiology@biorxiv-microbiol.bsky.social · 2mo ago

Bacterial metabolism of synthetic steroids across ecosystems reveals diverse biotransformation products, reactions, and enzymes https://www.biorxiv.org/content/10.64898/2026.06.05.730323v1

1. How common is LLM use in scientific publishing, and how does it vary across field, publisher, journal prestige, author demographics etc.? @kylesiler.bsky.social has new paper in PNAS that addresses this question on a massive scale: 7.3 million papers from Elsevier, PLOS, MDPI, and Frontiers.

The diffusion of large language models in published academic articles | PNAS

Large language models (LLMs) are rapidly changing academic research, raising questions of who is adopting these tools and under what conditions. Th...

pnas.org

Our work exploring host-microbial co-metabolism of bile acids in patients with alcohol-related liver disease (ALD) has been published in the JHEP Reports: www.sciencedirect.com/science/arti... 🪐 A big thank you to everyone involved, to the GALAXY and MicrobLiver consortia who made this possible! 🙌

Alcohol-related liver disease disrupts bile acid homeostasis and gut microbial bile acid metabolism

Alcohol overuse disrupts liver function and alters gut microbial communities, with alcohol-related liver disease (ALD) causing half of all liver-relat…

sciencedirect.com