When we started @joinprolific.bsky.social, we were motivated by the question "what if MTurk was actually good"? It's sad to see Mturk closing it's doors, but it's not a surprise. It’s the result of years of degrading data quality and an unhealthy marketplace.
Just launched labs.prolific.com. Our AI research team (PAIR) have been busy - working on science of evaluation: how to measure AI systems well, grounded in real human judgement at population scale. Check out the work that’s live, but there’s much more to come!
Prolific AI Research
Prolific AI Research (PAIR) — papers, notes, and field logs on how AI systems are measured through human experience.
labs.prolific.com
This one's been a long time coming! Longitudinal projects are live now on @joinprolific.bsky.social, a dedicated way to run multi-wave research from start to finish.
Myra Cheng, a @stanford.edu PhD student, has contributed to two important studies on AI sycophancy within months of each other. I think it's worth recognising. Quick thread below:
The data quality problem isn't going away anytime soon. About a month ago we announced Prolific's 100% Human Guarantee, but the real work is in the systems that stop threats from getting through in the first place. I've shared more thoughts here: www.prolific.com/resources/ai...
Researchers from @harvard.edu find that LLMs claiming "human-like" performance actually reflect a very specific subset of humanity. They cluster closest to WEIRD populations (Western, Educated, Industrialized, Rich, Democratic), diverging as psychological distance increases (r ≈ -0.70) 👇🏻
AI pollution in human data samples is a hot topic. Some great work from @andrewgordon.bsky.social et al. showing that concerns here are (generally) overblown, with the majority of platforms empirically showing low levels of AI pollution. osf.io/preprints/ps...
OSF
osf.io
New preprint out today (osf.io/preprints/ps...). We tested whether AI agents are actually infiltrating online surveys. Spoiler alert: they aren't Thread 🧵 [1/9]
OSF
osf.io
As of today, if an AI agent is detected in your Prolific study, you'll get twice the cost of that participant back. We’re calling this our 100% Human Guarantee. Years of investing into @joinprolific.bsky.social's system has made us confident in data integrity. www.prolific.com/100-human-gu...
New working paper on online research data quality, led by @univie.ac.at, reveals that pass rates on quality checks vary wildly by source. Pretty interesting. Prolific: 90% | Lab: 80% | Bilendi: 73% | Moblab: 55% | MTurk: 9% | AI agents: 0% github.com/survey-data-... CC @jyusof.bsky.social
Lots of hard work from the Prolific team to achieve the lowest rate of AI misuse detected in this study. More to do to get this to 0, though!
The sky is not falling; high-quality platforms (Prolific, Verasight, CR Connect) have low rates of apparent bots. osf.io/preprints/ps... But also not zero; vigilance is very much needed!
The sky is not falling; high-quality platforms (Prolific, Verasight, CR Connect) have low rates of apparent bots. osf.io/preprints/ps... But also not zero; vigilance is very much needed!
We ran a controlled study of 125 verified humans vs 5 AI agents. Can agents reliably be detected? Here's what we found: www.prolific.com/resources/au...
Authenticity checks detect AI agents best | Prolific
How we tested the most accurate method for identifying agentic AI
prolific.com
Frontiers episode 1: Jerome Wynne from @Prolific in conversation with Crystal Qian, from Google DeepMind, talking about Deliberate Lab: a platform for running online research experiments on human + LLM group dynamics. www.youtube.com/watch?v=5vyi...
AI agents are becoming a serious threat to research data quality. Today we’re rolling out Bot authenticity checks on @joinprolific.bsky.social, detecting agentic AI with 100% accuracy in testing. Comes with a native Qualtrics integration! More info: www.prolific.com/resources/in...
Fresh HUMAINE results are here. Gemini 3 is still first, but Mistral Large 3 and Deepseek v3.2 are making things interesting. Opus 4.5 didn't dominate, but Antropic is likely prioritizing complex reasoning/coding over the conversational fluency that this benchmark favors. prolific.com/humaine
Lots of chatter about this paper currently. Its a stark warning, but at present I see this as a stark warning of what might come, not what is happening now. As a research community we need to see it as a call-to-arms to develop new strategies, NOT a call to abandon online sampling. Reasoning below
From my colleague; paper here dropbox.com/scl/fi/55uok...
Without minimising the seriousness of the threat raised in this paper, I'm more optimistic. This is just the latest challenge in online integrity of online research. We've been proactively adding to our suite of authenticity tools - more every week - including many of Sean's recommendations:
Prolific sets standards for authentic human data collection | Prolific
Discover how Prolific's data quality system, Protocol, sets industry standards for authentic human data collection
prolific.com