Jess Smith

@nanojess.io

I do bioinformatics stuff with @fulcrumgenomics.com. I do garden stuff at home in Somerville, MA.

We definitely need to be able to measure LLM performance, but that will inevitably cause some "teaching to the test." I'm not sure how big of a problem this is though? Clearly the technology continues to improve, at least in my hands. dylancastillo.co/posts/pelica...

Are AI labs pelicanmaxxing? – Dylan Castillo

I generated 1,000+ SVGs across 7 frontier models to test whether AI labs are training on Simon Willison’s pelican-riding-a-bicycle benchmark.

dylancastillo.co