How good are LLM agents at protein design? We developed BioDesignBench: 76 expert-curated design tasks for evaluating scientific agents. Benchmark, reference agents, and leaderboard: www.biorxiv.org/content/10.6... huggingface.co/spaces/Romer...
Benchmarking and behavioral characterization of LLM agents for protein design
Large language models (LLMs) are increasingly deployed as agents for scientific discovery, but standardized frameworks for evaluating their performance and behavior in scientific workflows are lacking...
biorxiv.org