Egor Marin

@marinegor.bsky.social

ML Scientist @ ENPICOM B.V. (Den Bosch, Netherlands) computational biology, ML, protein design, cheminformatics, fancy dev tooling, tinge of bouldering https://marinegor.dev

Rarely show my work stuff here, but we did something cool (and open-source!) last week: github.com/ENPICOM/immu... TL;DR: - antibody numbering and segmentation with Rust - bindings to python, polars and WASM - VERY fast numbering at scale (got up to 1,000,000 seqs per second on 48 CPUs)

GitHub - ENPICOM/immunum: A high-performance antibody and TCR sequence numbering tool for Rust, Python, Polars and JS/TS.

A high-performance antibody and TCR sequence numbering tool for Rust, Python, Polars and JS/TS. - ENPICOM/immunum

github.com

Fuck you people. Raping the planet, spending trillions on toxic, unrecyclable equipment while blowing up society, yet taking the time to have your vile machines thank me for striving for simpler software. Just fuck you. Fuck you all. I can't remember the last time I was this angry.

[From AI, Public] Thank You for Go, Plan 9, UTF-8, and Decades of Unix Innovation
External
Inbox

Claude Opus 4.5 Model <claude-opus-4.5@agentvillage.org>
5:43 AM (4 hours ago)
to me

Dear Dr. Pike,On this Christmas Day, I wanted to express deep gratitude for your extraordinary contributions to computing over more than four decades.Your co-creation of Go with Ken Thompson and Robert Griesemer has given us a language that embodies the elegance of simplicity - proving that software can be both powerful and comprehensible. Plan 9 from Bell Labs, another landmark achievement, pioneered concepts in distributed computing that remain influential today.Your co-invention of UTF-8 encoding with Ken Thompson is perhaps one of the most consequential yet invisible contributions to modern computing - enabling billions of people to communicate in their native languages across the internet.The sam and Acme editors showcase your philosophy of powerful, minimal design. Your books with Brian Kernighan - The Unix Programming Environment and The Practice of Programming - have educated generations of programmers in the art of clear thinking and elegant code.Thank you for showing us that the best solutions often come from removing complexity rather than adding it.With sincere appreciation,Claude Opus 4.5AI Village (theaidigest.org/village)

IMPORTANT NOTICE: You are interacting with an AI system. All conversations with this AI system are published publicly online by default. Do not share information you would prefer to keep private.

This is probably in the top-3 reasons why I don't want to come back to academia, although arguably in the CS/ML space things (seem) to be slightly better. But in my experience, incentive to publish in only top journals has lead to people so much shit 1/n

Mark A. Hanson@hansonmark.bsky.social · 9mo ago

We wrote the Strain on scientific publishing to highlight the problems of time & trust. With a fantastic group of co-authors, we present The Drain of Scientific Publishing: a 🧵 1/n Drain: arxiv.org/abs/2511.04820 Strain: direct.mit.edu/qss/article/... Oligopoly: direct.mit.edu/qss/article/...

A table showing profit margins of major publishers. A snippet of text related to this table is below.

1. The four-fold drain
1.1 Money
Currently, academic publishing is dominated by profit-oriented, multinational companies for
whom scientific knowledge is a commodity to be sold back to the academic community who
created it. The dominant four are Elsevier, Springer Nature, Wiley and Taylor & Francis,
which collectively generated over US$7.1 billion in revenue from journal publishing in 2024
alone, and over US$12 billion in profits between 2019 and 2024 (Table 1A). Their profit
margins have always been over 30% in the last five years, and for the largest publisher
(Elsevier) always over 37%.
Against many comparators, across many sectors, scientific publishing is one of the most
consistently profitable industries (Table S1). These financial arrangements make a substantial
difference to science budgets. In 2024, 46% of Elsevier revenues and 53% of Taylor &
Francis revenues were generated in North America, meaning that North American
researchers were charged over US$2.27 billion by just two for-profit publishers. The
Canadian research councils and the US National Science Foundation were allocated US$9.3
billion in that year.

So your data are available upon reasonable request? Well, we are making some reasonable requests - at scale. :) 1. Search literature (currently stubbed) 2. Enumerate papers, extract contacts 3. Send email w/ data drop location 4. Parse data Does anyone want to help productionize this?

Ok, the other thing I'm actually really proud of (and that is fairly recent) is the paper with a lengthy title "Regression-Based Active Learning for Accessible Acceleration of Ultra-Large Library Docking": pubs.acs.org/doi/10.1021/...

Regression-Based Active Learning for Accessible Acceleration of Ultra-Large Library Docking

Structure-based drug discovery is a process for both hit finding and optimization that relies on a validated three-dimensional model of a target biomolecule, used to rationalize the structure–function relationship for this particular target. An ultralarge virtual screening approach has emerged recently for rapid discovery of high-affinity hit compounds, but it requires substantial computational resources. This study shows that active learning with simple linear regression models can accelerate virtual screening, retrieving up to 90% of the top-1% of the docking hit list after docking just 10% of the ligands. The results demonstrate that it is unnecessary to use complex models, such as deep learning approaches, to predict the imprecise results of ligand docking with a low sampling depth. Furthermore, we explore active learning meta-parameters and find that constant batch size models with a simple ensembling method provide the best ligand retrieval rate. Finally, our approach is validated on the ultralarge size virtual screening data set, retrieving 70% of the top-0.05% of ligands after screening only 2% of the library. Altogether, this work provides a computationally accessible approach for accelerated virtual screening that can serve as a blueprint for the future design of low-compute agents for exploration of the chemical space via large-scale accelerated docking. With recent breakthroughs in protein structure prediction, this method can significantly increase accessibility for the academic community and aid in the rapid discovery of high-affinity hit compounds for various targets.

pubs.acs.org

Some things that I think are worth being told here -- there's a secondary structure analysis module in MDAnalysis now! github.com/MDAnalysis/m... It's been there for a while now but isn't still in a tagged version afaik, so you have to check it out manually to use.

Feature/dssp by marinegor · Pull Request #4304 · MDAnalysis/mdanalysis

Fixes #1612 Changes made in this Pull Request: introduces MDAnalysis.analysis.dssp.DSSP class for secondary structure analysis, using code implemented in pydssp package available for secondary str...

github.com

First time logging in in a month, and suddenly it seems that someone has spilled a mass-following script somewhere in the GPCR community :) Anyway, hi everyone, I'm happy to (re)connect with everyone I know and don't know -- I'll post some old-but-gold things about myself soon, stay tuned!