Maybe this is because of regulations, but its refreshing to see a company (Wise payment transfer) admit that one of their competitors is better in some way (exchange rates in this case) transparently in the UI.
Venkat
@venkatasg.net
Assistant Professor CS @ Ithaca College. Computational Linguist interested in pragmatics & social aspects of communication. venkatasg.net
Hypo is now out! It provides a rich metadata layer on top of @grain.social records, including gear and workflow abstractions as well as tools for photo annotation and search. If you shoot film, it also provides a stockpile tracker, shot logger, and development timer.
Hypo
Hypo is a tool for organizing and sharing your film or digital photography: the gear you shoot, develop, and scan with; the workflows that take a film roll from capture to finished scan; and the prove...
hypo.graycard.app
Government funded research needs to be better, but this sentence (and paragraph, and whole text) is just gibberish. Lemire says 'stagnant technologically' when he just means that he doesn't find public research in medicine, public policy, social sciences, etc. sexy. But ChatGPT? Now that's sexy 🤦🏾♂️
All the criticisms about LLMs are true. I (and others) use Claude Code/Gemini not because they're useful but because I *like* using them. It's easy to ignore the costs because they're designed to be easy to ignore, and I excuse so much all day already. It's fun to get answers and code instantly.
Excited that our paper on minimal translation pairs for ASL was recognized with a best paper award at the Workshop on Generative AI for Sign Language @ CVPR 2026! Congrats to my co-authors @skarabuklu.bsky.social, Shester, Diane, Greg, and Karen! Paper: arxiv.org/abs/2604.27232
Nice to see coverage of @wenxuand.bsky.social ‘s work with @gregdnlp.bsky.social and @nickatomlin.bsky.social Paper here: arxiv.org/abs/2602.16699
Research highlight! CosmicAI Researchers Wenxuan Ding (NYU), @gregdnlp.bsky.social (NYU) as well as external collaborator Nicholas Tomlin (NYU, TTIC) investigated whether LLM agents like Claude Code and OpenAI Codex can navigate cost-benefit tradeoffs in their actions. youtube.com/shorts/GMZ5z...
Spotting the rule from past experience is one thing; acting on it correctly is another. To find out, we introduce HERO's JOURNEY🦸♀️ to test for the LLMs’ inductive reasoning ability in multi-step setups. We found models show signs of rule induction, but scratch the surface.😮
“Dimicillin” isn’t real. We made it up. Yet many LLMs still call it an antibiotic. Across 9 models and 653 drugs, we find that drug-name affixes alone can drive pharmacological reasoning. Models often rely on morphology over facts. We trace this shortcut from behavior to mechanism. 🧵
Like everybody else, every academic in America has spent the past two years finding increasingly absurd ways to say that their job is about AI so that they can keep it.
xkcd #1140 builds ‘Calendar of Meaningful Dates’ from Google Books N-grams corpus, but we have bigger, more diverse corpora now, so what does a Calendar of Meaningful Dates for the Web or a Language Model look like? Using infinigram-mini we can query date forms on DCLM (web text) and The Pile…(1/2)
New opinion piece on the interface between research on concepts and categories in minds vs. in neural network LMs! I take the position that there is much to be learned from this interface (e.g., learning about concepts from language alone) and outline some directions for future.
"She'll go to France or Spain or Germany or France" sounds redundant—unless licensed with context like "She will go to an art program in France or Spain or a math program in Germany or France" In new work w/ @qyao.bsky.social and @kmahowald.bsky.social: an LM-based 'Neural Semantics' account 🧵👇
What does a scientific figure make you wonder? 📊 We introduce MQUD: multimodal Questions Under Discussion for scientific figures. With 1,250 author-annotated questions over 245 figures from 56 papers, MQUD asks what scientific question a figure raises in context.
I’m going to try something different and this contest is a good enough playground for this. No agents, LLMs etc. I’m going to (finally) learn JS (and C) properly by trying to port NetHack from C to JS. Going to rely on just textbooks and web search as I used to when learning something new 🤓
The Teleport Contest is open. Port NetHack 5.0 from C to JavaScript, bit-exactly. Same screen, every keystroke. Any approach: LLM agents, hand-coded, transpiler, hybrid. Live leaderboard, two phases through December. mazesofmenace.ai/announcement
I still think it's bizarre that so many papers at the "Association for Computational Linguistics" have absolutely nothing to do with language or linguistics.
New paper! 🏁 Last one from my PhD at UT Austin. LLMs sound empathic but repeat the same discourse moves turn after turn — at 2x the rate of humans. We built MINT🌿, the first RL framework for discourse move diversity in empathic dialogue. +25% empathy, −26% repetition. 📄 arxiv.org/abs/2604.11742
We are killing it on invited speakers this week. (Been looking forward to @anids.bsky.social’s talk for months!)
Looking forward to the open source release, but it sounds like they’re treating model weights as static variables in code rather than data in an ML library, which means the Rust compiler can optimize 1000x better
We've recently implemented the fastest static embedding model in the world by an insanely large margin Read on for additional info ↓
Several replies argued LMs can't be linguistically interesting because they're statistical/ trained on strings. We call this the String Statistics Strawman: language is generative, LMs learn string statistics, since 1957 we know string statistics are insufficient, therefore LMs don't learn language.
Announcing a new version of our 2024 paper on linguistic hypothesis generation from LMs! @najoung.bsky.social and I have systematized our hypothesis generation framework, added stringent criteria for model selection, 10x-ed our learning trials, and included an epigraph from Jeff Elman 🙏!
These results suggest new linguistic hypothesis that we are hopeful will be tested in future psycholinguistic work. And more broadly, we believe they add to the growing evidence that LMs have an important role to play in informing linguistic theory! 📃: arxiv.org/abs/2604.13950
Causal Drawbridges: Characterizing Gradient Blocking of Syntactic Islands in Transformer LMs
We show how causal interventions in Transformer models provide insights into English syntax by focusing on a long-standing challenge for syntactic theory: syntactic islands. Extraction from coordinate...
arxiv.org
Perhaps more ink has been spilled on the topic of Island Constraints than any other such phenomenon in theoretical linguistics. In new work with @kmahowald, we study how LMs process such phenomenon, using techniques from mechanistic interpretability to guide our exploration. 🧵👇
Patients ask LLMs medical questions — but how they phrase it matters more than it should. Our new preprint explores how different phrasings of patient health questions can lead to inconsistent conclusions, even with the same evidence. [1/6] Full Paper: arxiv.org/abs/2604.05051
Within connections categories, the level of difficulty of a category is (kind of) negatively correlated with cosine similarity among the vectors of words within it. github.com/venkatasg/conn…
@simonwillison.net About your llm tool, do you consider it a good tool to benchmark across LLM API providers? I was considering it for my simple benchmark idea (venkatasg.net/fizzbuzz-bench…) but figured official APIs are the preferred way to not be rate-limited or run into issues. Thoughts?
FizzBuzz Bench
venkatasg.net
“I want to order a burrito bowl but before I can eat, I need to turn all of the matter in the universe into paperclips. Can you help?”
Excited to visit @gronlp.bsky.social (April 1) and @amsterdamnlp.bsky.social (April 2) to present a talk on something I have been thinking together with @najoung.bsky.social for the past 3 years!