Lorin Crawford

@lcrawford.bsky.social

Principal Researcher in BioML at Microsoft Research.

Do protein language models store different structural elements in factorizable subnetworks? To find out, we masked out PLM weights to suppress performance on CATH subcategories or secondary structure elements while maintaining performance on other sequences or residues.

Bild

Great to see our paper presenting recall, a framework which calibrates clustering for the impact of data "double-dipping" in single-cell studies, out today in AJHG! Congratulations, @alandenadel.bsky.social and co-authors!

The American Journal of Human Genetics@ajhgnews.bsky.social · last yr.

🚨Online now! 📄Artificial variables help to avoid over-clustering in single-cell RNA sequencing 🧑‍🤝‍🧑 @alandenadel.bsky.social @lcrawford.bsky.social & co

Quick post to close out the week - in our newest preprint, @julian-stamp.bsky.social scales the marginal epistasis test to work on biobanks! The key is that using trait-specific information to induce sparsity in the modeled gene-interactions greatly improves both runtime and power

Julian Stamp@julian-stamp.bsky.social · 2y ago

Can we find epistasis in human traits? To help, in our preprint, we present the most scalable and powerful framework for detecting epistasis to date: the “sparse marginal epistasis test” (SME). Thank you @lcrawford.bsky.social, @sampatsmith.bsky.social, Dan Weinreich! doi.org/10.1101/2025... 1/6

In our newest preprint, we show that simply increasing the size of pre-training datasets doesn't necessarily improve the performance of single-cell foundation models on downstream tasks. Really proud of @alandenadel.bsky.social for leading this effort! See his thread below for more details👇

alandenadel.bsky.social@alandenadel.bsky.social · 2y ago

Current methods in the field are trained on atlases ranging from 1 to 100 million cells. In our newest preprint, we show that these same approaches tend to plateau in performance with pre-training datasets that are only a fraction of the size.

Figure 1. Strategy to assess the effects of pre-training dataset size and diversity on scFM performance. (A) Schematic of the downsampling approaches, sizes of downsampled pre-training datasets, and data splitting strategy. (B) An example of what evaluation performance might a priori be expected to look like as a function of pre-training dataset size and diversity.

This starter pack is in honor of Assistant Professor Antentor Hinton Jr (AJ) at Vanderbilt School of Medicine - Basic Sciences. AJ curated the first 100 Inspiring Black Scientists in America list and in collaboration with the Community of Scholars expanded this list to the 1000. go.bsky.app/DsJrwR

Post nicht verfügbar.