In light of everything happening this week, I wanted to reiterate that AI Flaws and Incidents should be reported in a responsible, coordinated manner, and crucially, this should be administered by a competent body as opposed to voluntary self governance, akin to analogous systems in cyber (CVE).
Avijit Ghosh
@evijit.io
Lead Technical AI Policy Researcher at Hugging Face @hf.co 🤗. Current focus: Responsible AI, AI for Science, and @eval-eval.bsky.social!
Weekend mini project! Since commentary on AI is inherently interdisciplinary, we connected the observations in the The Pope's encyclical, with decades of scholarship in Responsible AI and Ethics research, and created an interactive space with these annotations!
I miss pre 2022 AI when there wasn’t so much noise in the field. Almost no separation between my job and personal life (a bunch of which is attributable to FOMO). Can’t even go on vacation without hearing about AI a few tables across at dinner. Even the Pope is involved
Huh so I’m not the only one who noticed this! To add to this list, the Devil Wears Prada 2 also has a plot point about AI x Creatives. www.thewrap.com/creative-con...
TV Writers Worry AI Will Replace Them. Now They’re Putting Those Anxieties on Screen
The creatives behind “Hacks,” “The Comeback,” “The Pitt” and “Matlock” tell TheWrap about tackling the emerging technology in their shows.
thewrap.com
Should all our resources go towards building chatbots? What if we built systems that actually give people meaningful agency? Finally out as an accepted @facct.bsky.social paper! Joint work with my past collaborators Sourojit, Pranav and Sanjana. huggingface.co/papers/2605....
Paper page - What if AI systems weren't chatbots?
Join the discussion on this paper page
huggingface.co
This has been a massive community project, and we need you all to participate! See more: evalevalai.com/projects/eve...
🚀 Launching Every Eval Ever: Toward a Common Language for AI Eval Reporting 🚀 A shared schema + crowdsourced repository so we can finally compare evals across frameworks and stop rerunning everything from scratch 🔧 A tale of broken AI evals 🧵👇 evalevalai.com/projects/eve...
This has to be rage bait. Did we not see the South Park episode where ChatGPT suggested a business idea to convert fries to salad? (And I tried to prompt myself too)
America, you have spoken loud and clear: You do not like AI. But what if AI is the way to restart the world’s idea machine?
Who is winning the open AI race? Our new study Economies of Open Intelligence maps @hf.co 851k models' downloads 2020→2025. 1) Power rebalance: US tech ↓; China + community ↑ 2) Models size & efficient ↑ (MoE, quant, multimodal) 3) Intermediary layers ↑ (adapters/quantizers) 4) Transparency ↓ /🧵
I used to love the word “key” until AI models decided to love it and now I cringe at “key takeaways” in text material :(
It’s that time of the year again! I’ll be at @neuripsconf.bsky.social this year too :) If you’re interested in Responsible AI, AI Evals ( @eval-eval.bsky.social ) or AI4Science (Hugging Science), say hi!
🚨 AI keeps scaling, but social impact evaluations aren’t–and the data proves it 🚨 Our new paper, 📎“Who Evaluates AI’s Social Impacts? Mapping Coverage and Gaps in First and Third Party Evaluations,” analyzes hundreds of evaluation reports and reveals major blind spots ‼️🧵 (1/7)
Extremely thrilled to talk about our new paper: "Who Evaluates AI’s Social Impacts? Mapping Coverage And Gaps In First And Third Party Evaluations". This is the first big project output from the @eval-eval.bsky.social coalition! Thread below:
We have a call for posters out! Please submit your extended abstracts, it should be quick and easy. And just like last year, provocative work is especially encouraged as it makes for such interesting conversation 😈
📮 We are inviting students and early-stage researchers to submit an Abstract (Max 500 words) to be presented as posters during interactive session. Submit here: tinyurl.com/AbsEval We have a rock-star lineup of AI researchers and an amazing program. Please RSVP at the earliest! Stay tuned!
This. Copyright is a tool for protection but it’s not everything. In fact, there’s research showing that it is possible to create competitive language models using public domain data only. The proliferation of copyright respecting models would not solve the labor impact policy problem.
Imagine you time travel back to 1847 and you find the left response to industrialization is a) machines will never be as good as human weavers or b) we need to copyright loom patterns or c) it’s a speculative bubble. You’d say “Y’all. Not helping. What you need is obviously a labor movement.”
Going to San Diego for Neurips? We at @eval-eval.bsky.social , along with the UK AISI, are hosting a closed door state of evals workshop at @ucsandiego.bsky.social on Dec 8th. Request to join below! :) evaleval.github.io/events/works...
2025 Workshop on Evaluating AI in Practice
EvalEval, UK AI Security Institute (AISI), and UC San Diego (UCSD) are excited to announce the upcoming Evaluating AI in Practice workshop, happening on December 8, 2025, in San Diego, California.
evaleval.github.io
Datasets are the backbone of AI for Science, and we want to support scientific data natively on Hugging Face. The amazing @lhoestq.hf.co started a discussion on GH for this! Please engage (better still, submit a PR) so we can start supporting your 🫵 dataset: github.com/huggingface/...
Support scientific data formats · Issue #7804 · huggingface/datasets
List of formats and libraries we can use to load the data in datasets: DICOMs: pydicom NIfTIs: nibabel WFDB: wfdb cc @zaRizk7 for viz Feel free to comment / suggest other formats and libs you'd lik...
github.com
Random off the cuff observation about American AI: LLM folks seem to be concentrated in SF, but AI4Science folks seem to be concentrated in Boston. Meaning as the former gets oversaturated and the latter is only getting started, I expect Boston to be the next big AI epicenter! 💪
🌟 Weekly AI Evaluation Spotlight 🌟 🤖 Did you know malicious actors can exploit trust in AI leaderboards to promote poisoned models in the community? This week's paper 📜"Exploiting Leaderboards for Large-Scale Distribution of Malicious Models" by @iamgroot42.bsky.social explores this!
+1000. I miss life pre-AI hype when the discourse around AI was more scientific and people used to attribute papers and opinions to scientists instead of to their companies. Not all orgs block research papers and sanity check their papers via legal teams, and HF, especially so, is very distributed.
Hugging Face (thankfully) doesn't do a groupthink too much -- e.g., "Hugging Face thinks this". We're generally able to have different opinions and thoughts, which is part of the open/collaborative ethos. I don't feel strongly about this conference myself, I see pros and cons. 🧵
Hey, so can someone tell me why ChatGPT generating erotica is bad, any more so than it generating anything else? Obviously anything non-consensual or age-inappropriate is bad, but I don't see why some researchers in my timeline are up in arms about it, while Grok already does this.
We're starting a weekly paper spotlight series! Come engage with the posts and let's improve evals together! :) First up: Do Large Language Model Benchmarks Test Reliability?
✨Weekly AI Evaluation Paper Spotlight✨ 🕵️ Is benchmark noise and label errors masking the true fragility of LLMs? 🖇️"Do Large Language Model Benchmarks Test Reliability?" - This paper by @joshvendrow.bsky.social provides insights!
I had never seen people fighting with co-authors via colored latex comments on overleaf in a developing paper draft until today. One of them even pasted a google calendar link saying "can we hop on a call and hash this out". Paper writing is actually exciting sometimes!
Trying to start a new hobby and the internet is useless. Maybe AI will finally kill unstructured information retrieval for good and then we will be forced to call or visit friends for help again
More of such research please! Chatbots are not the future of science, science is
Introducing CellTransformer, a new AI tool developed with UCSF that makes it easier to explore massive neuroscience datasets and identify important subregions of the brain. 🧵
Some of these new gpt/claude wrapper startups make me wonder how much the founders are paying themselves in salary because there is no way they expect their horrible idea to actually be sustainably profitable
AI for scientific discovery is a social problem: In our new position paper, @cgeorgiaw.bsky.social and I show that culture, incentives, and coordination are the main obstacles to progress, and we are launching the Hugging Science Initiative to address this!
At a panel a couple of weeks ago I exclaimed in despair: “who decided that the only mode of interacting with AI is via chatbots?” And this trend is relentless still. www.theverge.com/news/787076/...
Microsoft launches ‘vibe working’ in Excel and Word
Vibe working is all about Office’s new Agent Mode.
theverge.com
Wife brought her 100 year old film camera to that pirates game from a few weeks ago and she got this shot that's just perfect
So fascinating (not really) to me that company execs and tier 1 AI conferences have gone in completely opposite directions as it relates to AI usage. Surely the best minds actually developing AI models know something about overreliance, productivity, and quality? Surely?