No science today. 🚴🇫🇷 Just the sound of freewheels, endless mountain roads, and the incredible atmosphere of the Tour de France. Every stage is a reminder that persistence, teamwork, and a little bit of suffering can lead to something extraordinary. Enjoying the ride... #TDF2026 #TourdeFrance
Thilo Muth
@drmuth.bsky.social
Group Leader & Adjunct Prof 🧑🏫 | Building scalable, FAIR data platforms for public health epidemiology | Bioinformatics & mass spectrometry 📈 | Based in Berlin | Open science & real-world impact 🔬✨
(BioRxiv All) Evaluating the use of non-linear models in data-driven rescoring of peptide-spectrum matches: In mass spectrometry (MS)-based proteomics, computational tools match acquired tandem MS spectra to peptides from a sequence database. Machine learning increasingly… #BioRxiv #MassSpecRSS
Evaluating the use of non-linear models in data-driven rescoring of peptide-spectrum matches
In mass spectrometry (MS)-based proteomics, computational tools match acquired tandem MS spectra to peptides from a sequence database. Machine learning increasingly supports this task through peptide-spectrum match (PSM) rescoring, in which a classifier, typically a linear semi-supervised model, refines the initial matching score. However, Mokapot allows the user to choose among different machine learning algorithms of increasing complexity, from the default linear support vector machine (LSVM) to random forest and XGBoost. Here, we use an entrapment approach to assess the effect of this increasing complexity on PSM identification and the accuracy of the estimated false discovery rate (FDR). We show that, while more complex models increase the number of identified PSMs at a fixed FDR threshold, this gain reflects a bias towards random matches from the target proteome database rather than genuine identifications. Indeed, for the most complex model, the entrapment FDR reaches 6.3% instead of the estimated 1% decoy FDR. This bias thus yields overly optimistic FDR estimates, indicating that model complexity in PSM rescoring must be carefully balanced against this overfitting risk.
dlvr.it
Very nice Java GUI. Nice (inofficial) successor of our previous DeNovoGUI software. Also supports mapping of peptides to reference databases - cool! 😎
CasanovoGUI: a cross-platform desktop application for deep learning-based de novo peptide sequencing with Casanovo https://www.biorxiv.org/content/10.64898/2026.07.11.737889v1
Generation of peptide detectability datasets from single DIA experiment for prediction model fine-tuning academic.oup.com/bioinformati...
Generation of peptide detectability datasets from single DIA experiment for prediction model fine-tuning
AbstractMotivation. Accurate prediction of peptide detectability in mass spectrometry–based proteomics is critical for improving both protein identificatio
academic.oup.com
Glycomics • Glycoproteomics • Glycoenzymes • Glycoinformatics (AI/ML) Join BioF:GREAT (Oct. 12–14, 2026) for hands-on training in AlphaFold 3, JAAG, AI-driven glycoenzyme modeling, glycoproteomics, and glycomics. Register below ⬇️ ugeorgia.ca1.qualtrics.com/jfe/form/SV_...
After years of development, collaboration, and hard work, it's finally here! Now published in Bioinformatics: Foundation model enables interpretable open and error-tolerant searching for mass spectrometry-based proteomics academic.oup.com/bioinformati... #TeamMassSpec #DeepLearning #AI #Proteomics
Foundation model enables interpretable open and error-tolerant searching for mass spectrometry-based proteomics
AbstractMotivation. Mass spectrometry-based proteomics allows studying all proteins of a sample on a molecular level. However, mass spectra are noisy and c
academic.oup.com
Excited to share our new preprint! 🎉 ProteoDUDes improves taxonomic profiling in #metaproteomics by reducing false positive taxonomic assignments. On experimental mock communities, it cuts the error rate by ~50%, enabling more reliable identification of microorganisms. doi.org/10.64898/202...
ProteoDUDes: Taxonomic profiling for metaproteomics with false positive reduction
Metaproteomics is the investigation of the protein composition of multi-organism samples. While metagenomics answers the question which organisms are present in a sample, metaproteomics additionally a...
doi.org
Preprint: False discovery rate control for trustworthy AI-based de novo peptide sequencing www.biorxiv.org/content/10.6...
| bioRxiv
bioRxiv - the preprint server for biology, operated by openRxiv, a nonprofit organization dedicated to advancing scientific communication
biorxiv.org
New Preprint: Learning Fragmentation Physics or Exploiting Sequence Priors? Benchmarking Bias in Deep Learning Models for De Novo Peptide Sequencing doi.org/10.64898/202...
Learning Fragmentation Physics or Exploiting Sequence Priors? Benchmarking Bias in Deep Learning Models for De Novo Peptide Sequencing
Deep learning models have advanced de novo peptide sequencing, but their predictions may reflect both physics-based spectral evidence and learned peptide-sequence priors. Systematically measuring such...
doi.org
New study: #metaproteomics reveals viral proteins in #glioblastoma tissues across two cohorts (n=273). HHV-1/2/8 were more frequently detected in tumours and linked to host proteomic signatures involving mitochondrial metabolism, translation and immune pathways. 🧠🦠 www.nature.com/articles/s41...
Metaproteomic profiling reveals viral proteins and associated host proteomic alterations in glioblastoma - Scientific Reports
Scientific Reports - Metaproteomic profiling reveals viral proteins and associated host proteomic alterations in glioblastoma
nature.com
Nature research paper: Spatial distribution of the proteome in the human body and in cancers go.nature.com/4vSiZPk
Spatial distribution of the proteome in the human body and in cancers - Nature
A spatially resolved map of the human proteome across a variety of healthy tissues and cancers provides wide-ranging insights in developmental biology and oncology, and could aid the identification of therapeutic targets and development of treatments for cancer.
go.nature.com
A 'megacluster' of genes found in a common soil bacterium can produce a range of antibiotics that can act against multi-drug-resistant bacteria go.nature.com/4aRR6ii
Antibiotic cocktail made by soil bacteria can kill superbugs
Nature - Four antibiotic compounds produced in Streptomyces bacteria attack multiple parts of an essential metabolic pathway.
go.nature.com
A Platform for High-throughput and Ultrasensitive Immunopeptidomics - Molecular & Cellular Proteomics www.mcponline.org/article/S153...
A Platform for High-throughput and Ultrasensitive Immunopeptidomics
In BriefImmunopeptidomics enables untargeted discovery of peptides presented by MHC molecules, informing vaccine and immunotherapy development. Existing workflows focus either on throughput or sensiti...
mcponline.org
Together with the legends of computational proteomics (Alexey Nesvizhskii) and metaproteomics (Bob Hettich) at the 7th International Metaproteomics Symposium in Dessau. Alexey presents FragMeta for efficient metaproteomics searching. Great talks and poster- such a vibrant communit! #TeamMassSpec
Tomorrow, the 7th International Metaproteomics Symposium begins! Looking forward to 3.5 days of cutting-edge research, stimulating discussions, and reconnecting with the amazing metaproteomics community. #IMS2026 #Metaproteomics Program here: www.hs-anhalt.de/landingsites...
Program | Hochschule Anhalt
hs-anhalt.de
The computational framework “Peptonizer2000” improves taxonomic assignment in complex #microbiomes by reducing ambiguity in protein-to-taxon mapping and increasing reliability of community profiling. #Metaproteomics #Proteomics 📄https://doi.org/10.1021/acs.jproteome.5c00567 👤EVBC member: Thilo Muth
The Peptonizer2000: Bringing Confidence to Metaproteomics
Metaproteomics, the large-scale study of proteins from microbial communities, faces challenges in identifying species due to similarities in protein sequences across different organisms. Current methods often rely on simple counting of matches between proteins and taxa, which can lead to low accurac...
doi.org
Computational #Metaproteomics Preprint out ! MetaPilot: genome-aware adaptive search-space refinement for unified DDA and DIA metaproteomics Kai Cheng, Daniel Figeys bioRxiv 2026.06.12.728088; doi: doi.org/10.64898/202...
MetaPilot: genome-aware adaptive search-space refinement for unified DDA and DIA metaproteomics
Metaproteomic peptide identification is constrained by the structure and size of the protein search space. Pooled gene catalogues provide coverage but obscure genome-level evidence, and current workfl...
doi.org
Preprint Alert! PeptiDIA: A Machine Learning Framework for Enhanced Peptide Identification in Fast-Gradient Data-Independent Acquisition Proteomics www.biorxiv.org/content/10.6...
PeptiDIA: A Machine Learning Framework for Enhanced Peptide Identification in Fast-Gradient Data-Independent Acquisition Proteomics
Data-independent acquisition (DIA) mass spectrometry has become increasingly prevalent in proteomics as advances in instrumentation, chromatography, and computational analysis have enabled robust prot...
biorxiv.org
New preprint on the current bottlenecks of deep learning-based methods for de novo peptide sequencing that underuse physical properties of spectra and how to overcome this with MemNovo: arxiv.org/abs/2606.11868
MemNovo: Look Back at the Spectrum for Balanced De Novo Peptide Sequencing from Mass Spectrometry
De novo peptide sequencing from tandem mass spectrometry is pivotal in proteomics, enabling identification of novel peptides without reference databases. While recent Transformer-based encoder-decoder...
arxiv.org
Trends in Computational Metabolomics: A Perspective on Five Years of Software Development, Challenges, and Opportunities (2021–2025) | Analytical Chemistry pubs.acs.org/doi/10.1021/...
Trends in Computational Metabolomics: A Perspective on Five Years of Software Development, Challenges, and Opportunities (2021–2025)
Metabolomics software development has accelerated rapidly, yet no recent systematic analysis has quantified how the landscape is evolving across computational methods, geographies, and the research co...
pubs.acs.org
Reshaping of the fecal proteome and metaproteome in obese patients 2 years after bariatric surgery | mSystems journals.asm.org/doi/10.1128/...
Reshaping of the fecal proteome and metaproteome in obese patients 2 years after bariatric surgery | mSystems
Bariatric surgery is widely recognized as the most effective and durable intervention for severe obesity; however, its long-term molecular effects on gut microbiota-host interactions remain poorly understood. By applying shotgun metaproteomics to fecal samples collected before and 2 years after surgery, our study provides novel insights into the functional consequences of bariatric bypass procedures. We demonstrate sustained alterations in both microbial and host protein profiles, including metabolic enzymes, outer membrane proteins, and immune-related factors, revealing a long-lasting remodeling of gut ecosystem functions. These findings underscore the value of metaproteomics in uncovering molecular mechanisms underlying bariatric surgery outcomes and may ultimately guide the development of microbiome- or host-targeted strategies to optimize therapy and long-term patient care.
journals.asm.org
Mass spectrometry proteomics loves benchmarks. But an important one is rare: - Accuracy of proteome quantification when using short LC gradients. Fast MS instruments can quantify 7 - 9k proteins from 200ng samples using short separation times affording the analysis of 200 – 500 samples / day. 1/
Another hidden gem by Gelio Alves and co-workers: Taxonomic-Level Protein Quantification in Metaproteomics Using a Biomass-Constrained Expectation–Maximization Approach | Journal of the American Society for Mass Spectrometry pubs.acs.org/doi/10.1021/... It is worth to look at Gelio‘s publications!
pubs.acs.org
Early use of directed acyclic graphs to speed up analysis of mass spectrometry of peptides. Low polynomial time complexity, down from exponential complexity without DAGs. en.wikipedia.org/wiki/De_novo... doi.org/10.1002/bms.... #DAG #Peptide #sequencing
scpviz: A Python bioinformatics toolkit for Single-cell Proteomics and multi-omics analysis joss.theoj.org/paper... --- #proteomics #prot-paper
This banger from Matt @mattwfoster.bsky.social et al. just dropped. Love some whole blood proteomics fun, and really so much of what Matt has been doing with blood. Excellent work! pubs.acs.org/doi/10.1021/...
Development and Validation of a Streamlined Workflow for Proteomic Analysis of Proteins and Post-translational Modifications from Dried Blood
It is increasingly recognized that the ‘omic analysis of whole blood has applications for precision medicine and disease phenotyping. Despite this realization, whole blood is generally viewed as a challenging analytical matrix in comparison to plasma or serum. Moreover, proteomic analyses of whole blood have almost exclusively focused on (non)targeted analyses of protein abundances and much less on post-translational modifications (PTMs). Here, we developed a streamlined workflow for processing 20 microliters of venous blood collected by volumetric absorptive microsampling that incorporates serial trypsinization and N-glycopeptide and phosphopeptide enrichment and avoids laborious sample dry-down or cleanup steps. As many as 10,000 analytes (reported as protein groups, glycopeptidoforms, and phosphosites) can be quantified by liquid chromatography-tandem mass spectrometry in under 2 h of MS acquisition time. Using these methods, we explored the stability of “dried” and “wet” blood proteomes, as well as the effects of ex vivo inflammatory stimulus or phosphatase inhibition. Multiomics factor analysis enabled facile identification of analytes that contributed to interindividual variability of the blood proteomes, including N-glycopeptides that distinguish immunoglobulin heavy constant alpha 2 allotypes. Collectively, our results help to establish feasibility and best practices for the integrated MS-based quantification of proteins and PTMs from dried blood.
pubs.acs.org
Abstract submission and registration for the Proteomic Forum / EuPA 2026 are still open. Submit your abstract, secure your spot, and join the proteomics community in Würzburg. 📝 Abstract deadline: 30 June 2026 🎟 Early bird deadline: 15 July 2026 👉 Everything in one place: proteomic-forum.com
When it's time to pick the best peptides for targeted MS, Bromo doesn't just guess from sequence. It learns from millions of peptide comparisons, considers charge, and ranks precursors by their response. And it adapts faster than a bromance at a #ASMS coffee break! www.biorxiv.org/content/10.6...
biorxiv.org
(BioRxiv All) High-Speed Mass Spectrometers diminish the difference between Data-Dependent and Data-Independent Acquisition Proteomics: Data-dependent acquisition mass spectrometry (DDA-MS) and data-independent acquisition mass spectrometry (DIA-MS) have historically offered… #BioRxiv #MassSpecRSS
High-Speed Mass Spectrometers diminish the difference between Data-Dependent and Data-Independent Acquisition Proteomics
Data-dependent acquisition mass spectrometry (DDA-MS) and data-independent acquisition mass spectrometry (DIA-MS) have historically offered complementary strengths in bottom-up proteomics, with DDA providing high-selectivity spectra for post-translational modification (PTM) analysis and DIA enabling more systematic peptide sampling. Here, we asked if this is still the case for the Orbitrap Astral platform that offers high-speed DDA and (ultra-) narrow-window DIA (nDIA) capabilities across proteome and phosphoproteome applications. When DDA and DIA measurements were parameter-matched (to the extent possible), the differences in analytical performance diminished markedly. Across extensive replicate analyses, both methods continued to identify new peptides and proteins without reaching saturation, indicating that the molecular complexity of biological samples still overwhelms even the fastest liquid chromatography-MS (LC-MS) methods. Incomplete sampling also contributed to substantial peptide-level non-overlap between DDA and nDIA and data completeness was only modestly better for nDIA than DDA across many replicates. Quantitatively, DDA and nDIA showed broadly similar precision and accuracy, with nDIA offering slightly higher precision and DDA slightly better accuracy in controlled mixture experiments. MS1-based quantification outperformed MS2-based quantification, particularly for short gradients, supporting MS1 quantification as a robust and general strategy for high-throughput proteomics. In phosphoproteomic samples, DDA and nDIA identified similar numbers of phosphopeptides, but DDA retained a small edge for phosphorylation site localisation. Together, the results show that advances in acquisition speed and sensitivity are narrowing the historical gap between DDA and DIA, while also revealing that current LC-MS workflows remain far from providing comprehensive proteome coverage. Going forward, further gains in dynamic range, scan speed, sensitivity, and transparent software tools will be required to reach systematic, comprehensive and reliable measurements of complex proteomes in a single shot.
dlvr.it