It has been *wild* to me to see the way that people in my department have fully gone back to business-as-usual posting on Twitter for their papers. There were a couple of brief blips where people tried bluesky and LinkedIn, but that's nearly all gone now 😔
Ivan Kartáč
@ivankartac.bsky.social
PhD student @ Charles University, Prague. NLP & computational linguistics. Working on evaluation, explainability, and reasoning. ivankartac.github.io
Neoclassical econ assumes rationality. The corollary of, "If you're so smart, why aren't you rich?" is "you're rich, so you must be very smart!" Thus, people assume that if powerful, well-compensated CEOs insist that "AI is changing everything," well then, *AI must be changing everything*. 1/
Funny how they now call >100B parameter models “small”. And apparently 30B today is “nano”.
I recommend this article about AI reasoning, where the author lets us in on his struggles w/ AI cognitive dissonance. Plus some priceless quotes from @rao2z.bsky.social. (My recommendation has *nothing* to do with the fact that I'm quoted in it too 😇) www.quantamagazine.org/is-ai-reason...
Is AI Reasoning Right for the Wrong Reasons? | Quanta Magazine
The idea that artificial intelligence can “reason” is more intuitive than ever. But intuitions can be wrong, and the science is far from settled.
quantamagazine.org
Fantastic paper demonstrating how LLM editing is gradually distorting our writing and the way we think to write: arxiv.org/abs/2603.18161
How LLMs Distort Our Written Language
Large language models (LLMs) are used by over a billion people globally, most often to assist with writing. In this work, we demonstrate that LLMs not only alter the voice and tone of human writing, b...
arxiv.org
We need a benchmark where systems generate recipes and people cook based on it and rate the results. Item Response Theory will handle annotators with poor cooking skills and too easy or infeasible meals.
We should make the topic of climate change more sci-fi, so that Silicon Valley EAs and rationalists can finally pivot to it.
This reverse chronological #NLP feed is still running, and I've made a companion feed that is ranked by engagement and recency: bsky.app/profile/did:... Feedback and ideas are welcome!
Not attending ACL in person? Follow along via my new NLP feed! It should match all mentions of #ACL2026, #EMNLP2026, #COLM2026, #NLProc, #NLP, and more.
The AI community failed to exit twitter. What implications can we draw about what people believe from this?
What it means that the AI community can't quit twitter
Some quickly jotted thoughts working through the implications of the AI community remaining on twitter
rl-blogging.leaflet.pub
I resigned from Google DeepMind bc it broke its founding promise by selling AI to the military without restrictions against killer robots or mass spying. For months, I worked to stop this but watched powerful ethicists and institutions choose silence. Here's what happened. 🧵
📢 Call for Papers: YNLG 2026 The Young Researchers in Natural Language Generation workshop is a 2d in-person event part of INLG 2026 @inlg.bsky.social in Utrecht 🇳🇱, with poster sessions, keynote talks, roundtable discussions, and a one-day hackathon. Due: August 10, 2026 ynlg-workshop.github.io
My assumption has been that in academia, humanities are much less inclined to use GenAI for their work. To what extent is this true? Or are people in humanities just less willing to acknowledge the use?
I often see arxiv pre-prints in reference lists of many papers even for sources that already have published versions. This is a really nice tool (by @zdenekkasner.cz) which will automatically replace arxiv versions or fix incomplete references through DBLP API: github.com/kasnerz/reffix
This is a great intro to experimental methods: experimentology.io I think NLP has a lot to learn from psychology in this respect, especially as evaluation becomes more and more important these days.
If you don't know how reliable your measure is, you're wasting your participants' time. Ch 8 of Experimentology argues that measurement reliability and validity deserve more attention than most experimentalists give them. 🧵 experimentology.io
I think I will be posting this after each #ARR cycle: Please 🥺🙏 let's prohibit AI review writing. Otherwise, we will get lazy reviewers claiming they wrote bullet points and used AI only to put that in prose.
#ACL2026 continues with the main conference (main + findings + demos + industry + SRW), and @ufal.mff.cuni.cz folks will present 8️⃣ papers. Stop by and check with our colleagues. All times in PDT.
Now out (for realz) in Cognition: "People Make Graded Judgments About The Inconceivable" (by Hu, Sosa, & me) Free preprint: www.tomerullman.org/papers/grade... Journal link: bit.ly/gradedInconCog @jennhu.bsky.social @cognitionjournal.bsky.social
Not attending ACL in person? Follow along via my new NLP feed! It should match all mentions of #ACL2026, #EMNLP2026, #COLM2026, #NLProc, #NLP, and more.
Heading to San Diego for #ACL2026, where I’ll be presenting two papers (see 🧵). Stop by to chat about evaluating reasoning embedded in task-oriented dialogue, or how to use small LLMs in modular neuro-symbolic approaches to syllogistic reasoning!
I'd never have guessed models commit to their final answer this early, often within the first 20% of reasoning, across math/logic tasks and model families. The rest is mostly hedging that doesn't change their mind. And turns out they encode this internally, we can decode it! 🧵👇
Are all the CoT steps necessary? In our latest paper, we find evidence for the existence of a commitment boundary, marking a sharp transition from no/mid guesses to the model final answer across various reasoning tasks and model families. Thread 🧵👇
It’s only a question of time until Anthropic makes headlines telling us they found Claude doing Zen meditation with Extended Not-thinking.
“Dimicillin” isn’t real. We made it up. Yet many LLMs still call it an antibiotic. Across 9 models and 653 drugs, we find that drug-name affixes alone can drive pharmacological reasoning. Models often rely on morphology over facts. We trace this shortcut from behavior to mechanism. 🧵
Do you sometimes have to explain to engineers that the main role of science is not to produce software? Once in a while I see people comment on some paper along the lines of “but it’s not efficient” or “I can’t use this in production” as if this was what research is about.
New blog: I am worried by NLP research culture NLG and NLP are mostly much better in 2026 than when I got my PhD in 1990. Unfortunately research culture has gotten *worse” in this period, which really worries me as I retire. ehudreiter.com/2026/06/08/n...
I am worried by NLP research culture
In most ways NLG and NLP are much better in 2026 than when I got my PhD in 1990. Unfortunately research culture has gotten *worse” in this period, which really worries me as I retire. We have…
ehudreiter.com
We've updated the preprint of our Naturalistic Computational Cognitive Science paper (arxiv.org/abs/2502.20349) — we've tried to clarify and streamline the arguments, and added some new examples: 1/5
Naturalistic Computational Cognitive Science: Towards generalizable models and theories that capture the full range of natural behavior
How can cognitive science build generalizable theories that span the full scope of natural situations and behaviors? We argue that progress in Artificial Intelligence (AI) offers timely opportunities ...
arxiv.org
this is how massive illusion of 'creativity' gets crushed... please read it to understand why models may appear to produce coherent text but are in fact Frankenstein factory ...
Ran some 🧪 to 🔬 why the Granta short story was certainly 🤖 generated. A lot of bad writing happens because AI hasn’t learned aesthetics. It has simply memorized the whole internet and called it a day. So sure, maybe you don't trust AI detectors. But you can trust your own eyes. #AISlop
Maybe one of the biggest obstacles for progress in science comes from entrenched stereotypes? In linguistics, we have, for example, (1) the word stereotype, (2) the grammar/dictionary stereotype, (3) the building-block stereotype, and (4) the speaker directionality stereotype dlc.hypotheses.org/4343
Four stereotypes that have guided morphosyntactic thinking
Thinking about language structures is made difficult not only by their incredible complexity, but also by entrenched ways of thinking about grammatical and lexical patterns. Linguists do not investiga...
dlc.hypotheses.org
In a new blog post, I argue that the anti-ai movement ought to distinguish between claims about the technology and the "project of AI," as defined by Vetsi et al. in their new paper. 🔗: doomscrollingbabel.manoel.xyz/p/the-anti-a...