Eric Kernfeld

@ekernf01.bsky.social

Statistician and computational biologist; uw alum; jhu student. He/him. http://ekernf01.github.io

For people that have tried to hire someone recently, did you encounter this problem? www.linkedin.com/posts/olga-v...

Failed to weed out LLM-generated applications for bioinformatics role | Olga Sazonova, Ph.D. posted on the topic | LinkedIn

Well, that's it. Hiring is f*cked. I'm just opened a contract role for a bioinformatics project. I know it's a wild market at the moment - each open role gets flooded with resumes, making it hard to separate qualified candidates from the noise. I'm convinced LLMs are partially to blame. So I devised a process to select qualified, committed candidates and avoid generic applications. Instead of resumes and cover letters, I created an intake questionnaire requiring bespoke effort to reduce cookie-cutter applications. The questions focused on specific scenarios from past experience, and therefore shouldn't be directly outsourced to chatbots. I also accidentally made the application link hard to access, but decided not to fix it. After all, computational biology often requires hacking and clever work-arounds. Finally, I included an honor code asking people not to use LLMs. Initially, I felt good about my approach. I had 20 applications after two days, and the candidates were differentiating themselves: A few didn't answer all questions, a few didn't have the right background, and several candidates provided well-written, topical, plausible answers. It was only after I read a few of these "high signal" applications in a row that the alarm bells started ringing: - Multiple candidates highlighted the same methodological paper - Many cited required skills in the exact same order - A high percentage reported identical troubleshooting scenarios involving differential gene expression studies with lab-derived batch effects The nail in the coffin was repeated phrases like "my initial hypothesis was a bug in [popular tool]" and "the spurious result disappeared." Clearly, candidates were using LLMs. Instead of learning about fitness for the role, I was learning how LLMs answer my "clever" questions. So...I failed. My approach was naive, perhaps even hypocritical. After all, I use LLMs for technical documentation myself. Still, it's really disappointing. I'm no closer to a hiring process that identifies qualified, motivated, and trustworthy candidates for remote work. Now what, y'all? | 109 comments on LinkedIn

linkedin.com

The biggest challenge for AI in biology isn't just models, it's the data used to train them. Standard biological data isn't built for AI. To unlock generative AI for drug discovery, we must rethink how we generate and capture data. 1/

Hardware/wetware codesigned data loop VISTA makes use of generative model sampling and synthesis "on chip" on-board by leveraging oligosynthesis setup shown here.

BLOG ALERT! 🧬 If you do RNA-seq, ATAC-seq, ChIP-seq, or modeling thereof, you may be overlooking LD score regression methods. These should be standard tools to study genes, variants, and regions en masse, but they are hard to understand. I wrote intros:

(flowery font "Corsiva" with bright colors)

The exotic FLAVORS of LD Score Regression

Vanilla (Linkage Disequilibrium Score Regression) distinguishes whether bias is from polygenicity or population structure using GWAS summary stats
Strawberry (stratified LD score regression) quantifies which genomic regions are enriched for disease risk using GWAS summary stats + functional annotations
Cranberry (cross-trait LD score regression) estimates genetic correlations between traits using summary stats from multiple GWAS
Lime (Signed LD Score Regression) tests whether predicted allele effects directionally align with per-allele disease risk using GWAS + deep learning
Melon (Mediated Expression Score Regression) quantifies how much disease risk is mediated through gene expression using GWAS + eQTL effects

Thrilled to announce that I am joining DTU in Copenhagen in the fall, as an assistant professor of chemistry. My research group will focus on fundamental methodology in machine learning for molecules.

1/ DNA sequence models like Borzoi predict gene expression and variant effects across tissues — but how can someone adapt the model to a custom experiment? @drkbio.bsky.social, Johannes Linder and I propose a solution via parameter-efficient fine-tuning (PEFT). www.biorxiv.org/content/10.1...

Parameter-Efficient Fine-Tuning of a Supervised Regulatory Sequence Model

DNA sequence deep learning models accurately predict epigenetic and transcriptional profiles, enabling analysis of gene regulation and genetic variant effects. While large-scale training models like E...

biorxiv.org

This is a really nice piece of work. biorxiv.org/content/10.1... - The sustainability and software quality is far superior to most one-off academic software projects. It is informed by deep experience with similar problems. - It is an especially clever data split. In causal ...

geneRNIB: a living benchmark for gene regulatory network inference

Gene regulatory networks (GRNs) underpin cellular identity and function, playing a key role in health and disease. Despite various benchmarking efforts, existing studies remain limited in the number o...

biorxiv.org