We have a little new paper at ICLR led by @AntonBushuiev. Test time training for proteins :) arxiv.org/abs/2411.02109 For example, you know the sequence that you want to fold, so fine-tune ESM on it at test time to get a better ESMFold structure prediction!
Roman Bushuiev
@roman-bushuiev.bsky.social
🌏 https://roman-bushuiev.github.io/
ProteinTTT is now easy to run on Hugging Face Spaces and Google Colab. We’ll also be presenting the paper at ICLR 2026 🇧🇷 🤗 Hugging Face Space: huggingface.co/spaces/pimen... ⚙️ Google Colab: colab.research.google.com/drive/1l_h7c... 🧵👇
My time in @martinsteinegger.bsky.social's group is ending, but I’m staying in Korea to build a lab at Sungkyunkwan University School of Medicine. If you or someone you know is interested in molecular machine learning and open-source bioinformatics, please reach out. I am hiring! mirdita.org
Mirdita Lab - Laboratory for Computational Biology & Molecular Machine Learning
Mirdita Lab builds scalable bioinformatics methods.
mirdita.org
Very happy @roman-bushuiev.bsky.social and I joined the amazing team led by @hannes-stark.bsky.social to work on BoltzGen, a generative model for binder design based on Boltz-2. Excited what it will enable!
Excited to release BoltzGen which brings SOTA folding performance to binder design! The best part of this project is collaborating with a broad network of leading wetlabs that test BoltzGen at an unprecedented scale, showing success on many novel targets and pushing the model to its limits!
Excited to release BoltzGen which brings SOTA folding performance to binder design! The best part of this project is collaborating with a broad network of leading wetlabs that test BoltzGen at an unprecedented scale, showing success on many novel targets and pushing the model to its limits!
We train machine learning models on millions of proteins. But when it comes to making predictions, do we need them to understand all proteins at once? Often, we need an accurate model for the specific protein we are studying or designing. We address this with ProteinTTT arxiv.org/abs/2411.02109 1/🧵
Yesterday @roman-bushuiev.bsky.social visited our lab. He gave a seminar talk about his recent work on DreaMS (www.nature.com/articles/s41...), and provided a hands-on training to my lab members. Looking forward to what we will discover using this awesome tool!
Mass spectrometry is a key method to discover and identify molecules in biological and environmental samples. Yet, >90% of mass spectra remain hard to interpret. In our recent paper, we present DreaMS — a foundation model to interpret mass spectra of small molecules. www.nature.com/articles/s41...
Self-supervised learning of molecular representations from millions of tandem mass spectra using DreaMS - Nature Biotechnology
A transformer model is used to construct the DreaMS Atlas—a molecular network of 201 million MS/MS spectra.
nature.com
This paper represents a great effort by @roman-bushuiev.bsky.social and his brother @anton-bushuiev.bsky.social. The DreaMS foundation model for mass spectra of small molecules now opens lots of avenues for possible downstream applications. It might be a game changer for computational metabolomics.
Self-supervised learning of molecular representations from millions of tandem mass spectra using DreaMS - @pluskal-lab.org @iocbprague.bsky.social go.nature.com/4k1n5iC
🚀 Exciting MassSpecGym leaderboard update 🚀 Two new machine learning models achieve up to a 300% improvement in de novo molecular generation given mass spectra and corresponding chemical formulae. 🔥 1/n
MassSpecGym - the first comprehensive benchmark for the discovery and identification of molecules from MS/MS data. @roman-bushuiev.bsky.social @anton-bushuiev.bsky.social @josef-sivic.bsky.social @pluskal-lab.org NeurIPS 2024 paper: arxiv.org/abs/2410.23326 #ChemSky #MassSpec #AI4Science
🤝 In April 2024, brothers Roman and Anton Bushuiev from the teams of @pluskal-lab.org @iocbprague.bsky.social and Josef Šivic #CIIRC_CTU initiated a collaboration among 14 research institutes across the globe to benchmark #AI methods for the discovery of molecules from mass spectrometry data. 1/2
🤝 In April 2024, brothers Roman and Anton Bushuiev from the teams of @pluskal-lab.org @iocbprague.bsky.social and Josef Šivic #CIIRC_CTU initiated a collaboration among 14 research institutes across the globe to benchmark #AI methods for the discovery of molecules from mass spectrometry data. 1/2
MassSpecGym is the largest publicly available collection of mass spectra data with 231K spectra for 29K unique molecular structures. 33% of the dataset was generated from newly measured, in-house data. 🛡️The dataset is now certified on Polaris! polarishub.io/datasets/rom... youtu.be/G8ZnVRm0ogc
MassSpecGym: A benchmark for the discovery and identification of molecules
YouTube video by PolarisHQ
youtu.be
Check out our NeurIPS 2024 spotlight poster on MassSpecGym, a dataset and benchmark for discovering new molecules from nature 🌿. If you work on generative models for graphs/molecules, plug your model into MassSpecGym and see how many molecules you can discover! 🚀 1/6
What are the most interesting datasets and benchmark-related work for ML in drug discovery at NeurIPS? We’ll be at the conference doing short interviews with researchers and handing out some Polaris merch! Here’s who we have on the shortlist. 🧵
👏 IOCB Prague is celebrating two significant achievements in the field of scientific research. See more at ► www.uochb.cz/en/news/664/... ⤵️ #EMBO #ERC #ERCCoG
Check out our MassSpecGym dataset on @polarishq.bsky.social. 🤩
Wow! 🤩 This may be the most carefully documented dataset on @polarishq.bsky.social. Great work by @roman-bushuiev.bsky.social! Do I have any #MassSpec researchers in my 🦋 network yet? I would love to hear what you think! github.com/polaris-hub/... #mass #spectrometry #dataset #benchmark