Pasquale Minervini
@neuralnoise.com
Researcher in ML/NLP at the University of Edinburgh (faculty at Informatics and EdinburghNLP), Co-Founder/CTO at www.miniml.ai, ELLIS (@ELLIS.eu) Scholar, Generative AI Lab (GAIL, https://gail.ed.ac.uk/) Fellow -- www.neuralnoise.com, he/they
I still believe that everyone is too fixated on the state of play in AI right now (which labs are ahead, how to manage costs, etc.) and not focused enough on the continued steepness of the capability curve for AI At higher capabilities (like the ones expected in the near term), a lot changes fast.
Organisations -- sponsorship opportunities are available for #EACL2027, the flagship European conference in computational linguistics, taking place in Athens in March 2027! Support the NLP community and connect with researchers and practitioners: 2027.eacl.org
The 20th Conference of the European Chapter of the Association for Computational LinguisticsAthens, GreeceMarch 9-14, 2027
2027.eacl.org
Thrilled to have been awarded the Association for Computational Linguistics 2016 test of time award for my first ever paper, written with/under the guidance of @dirkhovy.bsky.social A couple of cute things about the paper/its genesis/its outcomes
New blog post on our {ICML, ACL} 2026 papers, plus some new interesting results on open-ended learning and multi-modal retrieval! neuralnoise.com/2026/icml-ac...
Call for papers for special issue on Ethics in NLP in the journal Computational Linguistics. Discussions around ethics in #NLP and #ComputationalLinguistics is often limited by the expectations for what should be published in *CL conferences. Link to call: www.aclweb.org/portal/conte...
NLP reviews suck! But why is not clear– @aclrollingreview.bsky.social provides *a lot* of guidance for how to review, but in that, first principles get lost. So @adamlopez.bsky.social and I have written down some of our thoughts on first principles.
The Missing First Principles of Reviewing for ACL
Zeerak Talat & Adam Lopez, University of Edinburgh
medium.com
Attention #NLProc researchers, the EACL 2027 website is officially LIVE: 2027.eacl.org! 🎉 🇬🇷 Join us in Athens, Greece (Mar 9-13, 2027) at #EACL2027 📅 ARR submission deadline: Aug 6, 2026. Open to all areas of CL/NLP + related fields. Stay tuned for the detailed CfP soon!
If you are interested in privacy-preserving clinical NLP, we are recruiting a postdoc at the University of Edinburgh! The work is on LLMs/VLMs, AI privacy, and real-world health data in secure research environments. Apply by April 6th, 2026! More details: elxw.fa.em3.oraclecloud.com/hcmUI/Candid...
My amazing colleagues Sid and Michael are looking for a postdoc! 👇
We are advertising a postdoc position to work on #generative #models, #structure #induction, and MI #estimation with Michael Gutmann as part of @genaihub.bsky.social ! elxw.fa.em3.oraclecloud.com/hcmUI/Candid... Get in touch! (#ML #AI) 👉 homepages.inf.ed.ac.uk/snaraya3/ 👉 michaelgutmann.github.io
We are advertising a postdoc position to work on #generative #models, #structure #induction, and MI #estimation with Michael Gutmann as part of @genaihub.bsky.social ! elxw.fa.em3.oraclecloud.com/hcmUI/Candid... Get in touch! (#ML #AI) 👉 homepages.inf.ed.ac.uk/snaraya3/ 👉 michaelgutmann.github.io
Siddharth - Home
Sid's page
homepages.inf.ed.ac.uk
Chatted with the amazing @elissawelle.bsky.social from @theverge.com about @rohit-saxena.bsky.social’s “Lost in Time” work (arxiv.org/abs/2502.05092) and much more! You can find the full article here 👇
Why can’t ChatGPT tell time?
Check out Yu Zhao's (@yuzhaouoe.bsky.social) latest work, “Learning GUI Grounding with Spatial Reasoning from Visual Feedback” (www.arxiv.org/abs/2509.21552), done during his internship at MSR (@msftresearch.bsky.social)! New SOTA 🏆 results on ScreenSpot-v2 (+5.7%) and ScreenSpot-Pro (+110.8%)!
trend: non-NVIDIA training DeepSeek V3.1 was trained on Huawei Ascend NPUs this one is a South Korean lab training on AMD
Motif 2.6B — compact model with long context unique: trained on AMD GPUs focus is on long context & low hallucination rate — imo this is a growing genre of LLM that enables new search patterns huggingface.co/Motif-Techno...
I really needed a Deep Research MCP server to use with Claude Code and other tools — here it is: github.com/pminervini/d...
@togelius.bsky.social has thoughts on Genie 3 and games togelius.blogspot.com/2025/08/geni... Fairly close to my own, though I didn't get the preview the tech. Walking around a generated image-to-image world is not the same as playing a game. There are no game objectives.
Genie 3 and the future of neural game engines
Google DeepMind just announced Genie 3 , their new promptable world model, which is another term for neural game engine. This is a big neura...
togelius.blogspot.com
Anthropic research identifies “inverse scaling in test-time compute,” where longer reasoning degrades AI performance. On certain tasks, models become more distracted by irrelevant data or overfit to spurious correlations. #MLSky
Anthropic researchers discover the weird AI problem: Why thinking longer makes models dumber
Anthropic research reveals AI models perform worse with extended reasoning time, challenging industry assumptions about test-time compute scaling in enterprise deployments.
venturebeat.com
The amazing folks at EdinburghNLP will be presenting a few papers at ACL 2025 (@aclmeeting.bsky.social); if you're in Vienna, touch base with them!
Hm, hard disagree here. I really fail to see how this is misconduct akin to bribery, it's just a defense mechanism against bad reviewing practices. @neuralnoise.com
🚨 New Paper 🚨 How effectively do reasoning models reevaluate their thought? We find that: - Models excel at identifying unhelpful thoughts but struggle to recover from them - Smaller models can be more robust - Self-reevaluation ability is far from true meta-cognitive awareness 1/N 🧵
Inverse scaling of reasoning models a research collab demonstrated that there are certain types of tasks where all top reasoning models do WORSE the longer they think things like getting distracted by irrelevant info, spurious correlations, etc. www.arxiv.org/abs/2507.14417
Reasoning is about variable binding. It’s not about information retrieval. If a model cannot do variable binding, it is not good at grounded reasoning, and there’s evidence accruing that large scale can make LLMs worse at in-context grounded reasoning. 🧵
Sometimes, too much reasoning can hurt model performance! New research by Anthropic (@anthropic.com), by Aryo Pradipta Gema (@aryopg.bsky.social) et al.: huggingface.co/papers/2507....
Paper page - Inverse Scaling in Test-Time Compute
Join the discussion on this paper page
huggingface.co
My "Math, Revealed" series is freely available to anyone -- no paywall! -- in the thread below.
Spotlight poster coming soon at #ICML2025 @icmlconf.bsky.social! 📌East Exhibition Hall A-B E-1806 🗓️Wed 16 Jul 4:30 p.m. PDT — 7 p.m. PDT 📜 arxiv.org/pdf/2410.12537 Let’s chat! I’m always up for conversations about knowledge graphs, reasoning, neuro-symbolic AI, and benchmarking.