Shayne Longpre

@shaynelongpre.bsky.social

MTS @ Anthropic. 🇨🇦 Prev: MIT, Google, Apple, Stanford. Interests: AI/ML/NLP, Data-centric AI, transparency & societal impact

I’ll be hanging out at our poster on membership inference, but in the same slot Brian Lester will present our work on “The Common Pile v0.1: An 8TB Dataset of Public Domain and Openly Licensed Text” (poster 102)! [https://arxiv.org/abs/2506.05209]

A. Feder Cooper@afedercooper.bsky.social · 8mo ago

[NeurIPS '25] Really excited to present “Exploring the limits of strong membership inference attacks on large language models” (poster 1300) this morning (Friday December 5, 11am-2pm in Exhibit Hall C-E)! [https://arxiv.org/abs/2505.18773]

Who is winning the open AI race? Our new study Economies of Open Intelligence maps @hf.co 851k models' downloads 2020→2025. 1) Power rebalance: US tech ↓; China + community ↑ 2) Models size & efficient ↑ (MoE, quant, multimodal) 3) Intermediary layers ↑ (adapters/quantizers) 4) Transparency ↓ /🧵

BildBild

📢Thrilled to introduce ATLAS 🗺️: the largest multilingual scaling study to-date—we ran 774 exps (10M-8B params, 400+ languages) to answer: 🌍 Is scaling diff by lang? 🧙‍♂️ Can we model the curse of multilinguality? ⚖️ Pretrain vs finetune from checkpoint? 🔀 X-lingual transfer scores across langs? 1/🧵

BildBild

Delighted to see BigGen Bench paper receive the 🏆best paper award 🏆at NAACL 2025! BigGen Bench introduces fine-grained, scalable, & human-aligned evaluations: 📈 77 hard, diverse tasks 🛠️ 765 exs w/ ex-specific rubrics 📋 More human-aligned than previous rubrics 🌍 10 languages, by native speakers 1/

Bild

Thrilled our global data ecosystem audit was accepted to #ICLR2025! Empirically, it shows: 1️⃣ Soaring synthetic text data: ~10M tokens (pre-2018) to 100B+ (2024). 2️⃣ YouTube is now 70%+ of speech/video data but could block third-party collection. 3️⃣ <0.2% of data from Africa/South America. 1/

Bild

Very excited to release Kaleidoscope—a multilingual, multimodal evaluation set for VLMs, built as part of our open-science initiative! 🌍 18 languages (high-, mid-, low-) 📚 21k questions (55% require image understanding) 🧪 STEM, social science, reasoning, and practical skills

Bild

#AI is evolving fast, and so are its flaws. A fresh approach to finding and reporting AI bugs is long overdue. Great initiative by @shaynelongpre.bsky.social and team, transparency and accountability in AI development are essential! #AISafety #ResponsibleAI #AIEthics #MIT

Researchers Propose a Better Way to Report Dangerous AI Flaws

After identifying major flaws in popular AI models, researchers are pushing for a new system to identify and report bugs.

wired.com

What are 3 concrete steps that can improve AI safety in 2025? 🤖⚠️ Our new paper, “In House Evaluation is Not Enough” has 3 calls-to-actions to empower evaluators: 1️⃣ Standardized AI flaw reports 2️⃣ AI flaw disclosure programs + safe harbors. 3️⃣ A coordination center for transferable AI flaws. 1/🧵

Bild

Thrilled to be at #AAAI2025 for our tutorial, “AI Data Transparency: The Past, Present, and Beyond.” We’re presenting the state of transparency, tooling, and policy, from the Foundation Model Transparency Index, Factsheets, the the EU AI Act to new frameworks like @MLCommons’ Croissant. 1/

Bild

Really excellent explainer by @shaynelongpre.bsky.social‬ that clearly lays out what's at stake in the "AI crawler wars"

How we stand to lose out

As this cat-and-mouse game accelerates, big players tend to outlast little ones.  Large websites and publishers will defend their content in court or negotiate contracts. And massive tech companies can afford to license large data sets or create powerful crawlers to circumvent restrictions. But small creators, such as visual artists, YouTube educators, or bloggers, may feel they have only two options: hide their content behind logins and paywalls, or take it offline entirely. For real users, this is making it harder to access news articles, see content from their favorite creators, and navigate the web without hitting logins, subscription demands, and captchas each step of the way.

Perhaps more concerning is the way large, exclusive contracts with AI companies are subdividing the web. Each deal raises the website’s incentive to remain exclusive and block anyone else from accessing the data—competitor or not. This will likely lead to further concentration of power in the hands of fewer AI developers and data publishers. A future where only large companies can license or crawl critical web data would suppress competition and fail to serve real users or many of the copyright holders.
Shayne Longpre@shaynelongpre.bsky.social · last yr.

I wrote a spicy piece on "AI crawler wars"🐞 in @technologyreview.com (my first op-ed)! While we’re busy watching copyright lawsuits & the EU AI Act, there’s a quieter battle over data access that affects websites, everyday users, and the open web. 🔗 www.technologyreview.com/2025/02/11/1... 1/

Re: the FTC and the platforms We *should* be concerned about platform power over speech, but it isn’t censorship. As the Supreme Court said last year, the companies’ editorial decisions to moderate content are protected by the First Amendment. 1/

I compiled a list of resources for understanding AI copyright challenges (US-centric). 📚 ➡️ why is copyright an issue for AI? ➡️ what is fair use? ➡️ why are memorization and generation important? ➡️ how does it impact the AI data supply / web crawling? 🧵

Bild

Great point by @shaynelongpre.bsky.social on the AI crawler wars: "Unless we can nurture an ecosystem with different rules for different data uses, we may end up with strict borders across the web, exacting a price on openness and transparency." www.technologyreview.com/2025/02/11/1...

AI crawler wars threaten to make the web more closed for everyone

There’s an accelerating cat-and-mouse game between web publishers and AI crawlers, and we all stand to lose.

technologyreview.com