phelimb

@phe-lim.bsky.social

CEO & Founder @ Prolific.com

Researchers from @harvard.edu find that LLMs claiming "human-like" performance actually reflect a very specific subset of humanity. They cluster closest to WEIRD populations (Western, Educated, Industrialized, Rich, Democratic), diverging as psychological distance increases (r ≈ -0.70) 👇🏻

Fresh HUMAINE results are here. Gemini 3 is still first, but Mistral Large 3 and Deepseek v3.2 are making things interesting. Opus 4.5 didn't dominate, but Antropic is likely prioritizing complex reasoning/coding over the conversational fluency that this benchmark favors. prolific.com/humaine

Bild

Without minimising the seriousness of the threat raised in this paper, I'm more optimistic. This is just the latest challenge in online integrity of online research. We've been proactively adding to our suite of authenticity tools - more every week - including many of Sean's recommendations:

Prolific sets standards for authentic human data collection | Prolific

Discover how Prolific's data quality system, Protocol, sets industry standards for authentic human data collection

prolific.com

Post nicht verfügbar.