Thilo Muth

@drmuth.bsky.social

Group Leader & Adjunct Prof 🧑‍🏫 | Building scalable, FAIR data platforms for public health epidemiology | Bioinformatics & mass spectrometry 📈 | Based in Berlin | Open science & real-world impact 🔬✨

No science today. 🚴🇫🇷 Just the sound of freewheels, endless mountain roads, and the incredible atmosphere of the Tour de France. Every stage is a reminder that persistence, teamwork, and a little bit of suffering can lead to something extraordinary. Enjoying the ride... #TDF2026 #TourdeFrance

Bild

(BioRxiv All) Evaluating the use of non-linear models in data-driven rescoring of peptide-spectrum matches: In mass spectrometry (MS)-based proteomics, computational tools match acquired tandem MS spectra to peptides from a sequence database. Machine learning increasingly… #BioRxiv #MassSpecRSS

Evaluating the use of non-linear models in data-driven rescoring of peptide-spectrum matches

In mass spectrometry (MS)-based proteomics, computational tools match acquired tandem MS spectra to peptides from a sequence database. Machine learning increasingly supports this task through peptide-spectrum match (PSM) rescoring, in which a classifier, typically a linear semi-supervised model, refines the initial matching score. However, Mokapot allows the user to choose among different machine learning algorithms of increasing complexity, from the default linear support vector machine (LSVM) to random forest and XGBoost. Here, we use an entrapment approach to assess the effect of this increasing complexity on PSM identification and the accuracy of the estimated false discovery rate (FDR). We show that, while more complex models increase the number of identified PSMs at a fixed FDR threshold, this gain reflects a bias towards random matches from the target proteome database rather than genuine identifications. Indeed, for the most complex model, the entrapment FDR reaches 6.3% instead of the estimated 1% decoy FDR. This bias thus yields overly optimistic FDR estimates, indicating that model complexity in PSM rescoring must be carefully balanced against this overfitting risk.

dlvr.it

Excited to share our new preprint! 🎉 ProteoDUDes improves taxonomic profiling in #metaproteomics by reducing false positive taxonomic assignments. On experimental mock communities, it cuts the error rate by ~50%, enabling more reliable identification of microorganisms. doi.org/10.64898/202...

ProteoDUDes: Taxonomic profiling for metaproteomics with false positive reduction

Metaproteomics is the investigation of the protein composition of multi-organism samples. While metagenomics answers the question which organisms are present in a sample, metaproteomics additionally a...

doi.org

New study: #metaproteomics reveals viral proteins in #glioblastoma tissues across two cohorts (n=273). HHV-1/2/8 were more frequently detected in tumours and linked to host proteomic signatures involving mitochondrial metabolism, translation and immune pathways. 🧠🦠 www.nature.com/articles/s41...

Metaproteomic profiling reveals viral proteins and associated host proteomic alterations in glioblastoma - Scientific Reports

Scientific Reports - Metaproteomic profiling reveals viral proteins and associated host proteomic alterations in glioblastoma

nature.com

Together with the legends of computational proteomics (Alexey Nesvizhskii) and metaproteomics (Bob Hettich) at the 7th International Metaproteomics Symposium in Dessau. Alexey presents FragMeta for efficient metaproteomics searching. Great talks and poster- such a vibrant communit! #TeamMassSpec

Bild

Mass spectrometry proteomics loves benchmarks. But an important one is rare: - Accuracy of proteome quantification when using short LC gradients. Fast MS instruments can quantify 7 - 9k proteins from 200ng samples using short separation times affording the analysis of 200 – 500 samples / day. 1/

Bild

This banger from Matt @mattwfoster.bsky.social et al. just dropped. Love some whole blood proteomics fun, and really so much of what Matt has been doing with blood. Excellent work! pubs.acs.org/doi/10.1021/...

Development and Validation of a Streamlined Workflow for Proteomic Analysis of Proteins and Post-translational Modifications from Dried Blood

It is increasingly recognized that the ‘omic analysis of whole blood has applications for precision medicine and disease phenotyping. Despite this realization, whole blood is generally viewed as a challenging analytical matrix in comparison to plasma or serum. Moreover, proteomic analyses of whole blood have almost exclusively focused on (non)targeted analyses of protein abundances and much less on post-translational modifications (PTMs). Here, we developed a streamlined workflow for processing 20 microliters of venous blood collected by volumetric absorptive microsampling that incorporates serial trypsinization and N-glycopeptide and phosphopeptide enrichment and avoids laborious sample dry-down or cleanup steps. As many as 10,000 analytes (reported as protein groups, glycopeptidoforms, and phosphosites) can be quantified by liquid chromatography-tandem mass spectrometry in under 2 h of MS acquisition time. Using these methods, we explored the stability of “dried” and “wet” blood proteomes, as well as the effects of ex vivo inflammatory stimulus or phosphatase inhibition. Multiomics factor analysis enabled facile identification of analytes that contributed to interindividual variability of the blood proteomes, including N-glycopeptides that distinguish immunoglobulin heavy constant alpha 2 allotypes. Collectively, our results help to establish feasibility and best practices for the integrated MS-based quantification of proteins and PTMs from dried blood.

pubs.acs.org

Abstract submission and registration for the Proteomic Forum / EuPA 2026 are still open. Submit your abstract, secure your spot, and join the proteomics community in Würzburg. 📝 Abstract deadline: 30 June 2026 🎟 Early bird deadline: 15 July 2026 👉 Everything in one place: proteomic-forum.com

Bild

(BioRxiv All) High-Speed Mass Spectrometers diminish the difference between Data-Dependent and Data-Independent Acquisition Proteomics: Data-dependent acquisition mass spectrometry (DDA-MS) and data-independent acquisition mass spectrometry (DIA-MS) have historically offered… #BioRxiv #MassSpecRSS

High-Speed Mass Spectrometers diminish the difference between Data-Dependent and Data-Independent Acquisition Proteomics

Data-dependent acquisition mass spectrometry (DDA-MS) and data-independent acquisition mass spectrometry (DIA-MS) have historically offered complementary strengths in bottom-up proteomics, with DDA providing high-selectivity spectra for post-translational modification (PTM) analysis and DIA enabling more systematic peptide sampling. Here, we asked if this is still the case for the Orbitrap Astral platform that offers high-speed DDA and (ultra-) narrow-window DIA (nDIA) capabilities across proteome and phosphoproteome applications. When DDA and DIA measurements were parameter-matched (to the extent possible), the differences in analytical performance diminished markedly. Across extensive replicate analyses, both methods continued to identify new peptides and proteins without reaching saturation, indicating that the molecular complexity of biological samples still overwhelms even the fastest liquid chromatography-MS (LC-MS) methods. Incomplete sampling also contributed to substantial peptide-level non-overlap between DDA and nDIA and data completeness was only modestly better for nDIA than DDA across many replicates. Quantitatively, DDA and nDIA showed broadly similar precision and accuracy, with nDIA offering slightly higher precision and DDA slightly better accuracy in controlled mixture experiments. MS1-based quantification outperformed MS2-based quantification, particularly for short gradients, supporting MS1 quantification as a robust and general strategy for high-throughput proteomics. In phosphoproteomic samples, DDA and nDIA identified similar numbers of phosphopeptides, but DDA retained a small edge for phosphorylation site localisation. Together, the results show that advances in acquisition speed and sensitivity are narrowing the historical gap between DDA and DIA, while also revealing that current LC-MS workflows remain far from providing comprehensive proteome coverage. Going forward, further gains in dynamic range, scan speed, sensitivity, and transparent software tools will be required to reach systematic, comprehensive and reliable measurements of complex proteomes in a single shot.

dlvr.it