whats your opinion re ACL submissions that have nothing to do with NLP (beyond the relation to LLM)? lets say a challenging tool use benchmark, but where nothing is specifically about human language, language use, etc?
למה כדאי לעשות מסטר בבינה מלאכותית בימינו? כי אנחנו עובדים על סוכני AI שידעו לבצע מחקר ברמה של מסטרנט, אז עוד כמה שנים כבר לא תהיה לכם הזדמנות
my layman remarks: technically it didnt "solve a problem", it "improved a bound". the actual underlying question, what is the optimal number, remains open. erdos was a bit arrogant and a demi god and coined "is *my* construction optimal", but that was never *the* question.
coding agents are not compilers from English to programs. and it is not because they are non-deterministic. gist.github.com/yoavg/b2454c...
Coding agents are not compilers
Coding agents are not compilers. GitHub Gist: instantly share code, notes, and snippets.
gist.github.com
i heard it is possible to scrutinize and/or criticize a technology without being delusional
Sure, LLMs are useful for: 1. Fraud 2. Plagiarism 3. Cognitive off-loading Which of those use-cases are you promoting?
I complain a lot about RL lately, and here we go again. The CS view of RL is wrong in how it thinks about rewards, already at the setup level. Briefly, the reward computation should be part of the agent, not part of the environment. More at length here: gist.github.com/yoavg/3eb3e7...
rl-wrong-about-rewards.md
GitHub Gist: instantly share code, notes, and snippets.
gist.github.com
the fascinating (to me) quality of hard-core RL researchers (e.g Sutton) is the ability to have an all encompassing view of RL as the basis of intelligence, while at the same time working on super low level stuff like tabular TD algorithms, and yet strongly believe these are actually the same thing
what's the latest-and-greatest attempt to reverse-engineer and document the inner-working of claude-code?
lets talk about "In context learning". it is clearly NOT "learning", because its ephemeral. It IS some form of generalization from examples, which is very cool. but we need a name. how do we call this skill of generalization from example?
חשבתי שההתעלמות מסודאן היא סתם סוג-של הזנחה, אבל הנה אני קורא ספר שמתאמץ מאד להראות ש: - מה שהיה בדארפור זה לא ג'נוסייד אלא רק תוצאת לוואי של מלחמת אזרחים אכזרית - "Darfur is where the language of genocide had become an instrument" - הערבים אינם settlers באפריקה, הם הגיעו ממקומות שונים - המערב אשם
how many of you are aware of ai-2027? how many of you read it? curious what ya'll think.
"in a future in which AI assistants are being used by world leaders and influential figures, aren't you afraid of rogue AI using this access to control the world?" well maybe, but I am more worried about near future in which these influentials act on random AI advice.
LISP code does not have significantly more parentheses than other typical languages. change my mind.
this "feature" in code editors where you type an opening character and it immediately inserts the closing one for you as well -- why?? at the very very best case, you will have to skip over this character with an arrow. so why?
can someone explain "serverless backends" to me? it seems that they run functions on demand. but if these functions cannot access any persistent state, why not run them on the client? the only reason I see is to hide DBs/APIs tokens/secrets from the client, but is that really all there is to it?
i created this thingy yesterday and now I cannot stop watching it. yoavg.github.io/eternal/
Eternal Struggle
yoavg.github.io
a trivia fact about this paper is that we submitted it to arxiv weeks ago, and it was hanging there in limbo for quite a while. apparently because we submitted to "AI" while they moved it to "HCI".
When reading AI reasoning text (aka CoT), we (humans) form a narrative about the underlying computation process, which we take as a transparent explanation of model behavior. But what if our narratives are wrong? We measure that and find it usually is. Now on arXiv: arxiv.org/abs/2508.16599
When reading AI reasoning text (aka CoT), we (humans) form a narrative about the underlying computation process, which we take as a transparent explanation of model behavior. But what if our narratives are wrong? We measure that and find it usually is. Now on arXiv: arxiv.org/abs/2508.16599
Humans Perceive Wrong Narratives from AI Reasoning Texts
A new generation of AI models generates step-by-step reasoning text before producing an answer. This text appears to offer a human-readable window into their computation process, and is increasingly r...
arxiv.org
if you REALLY want to understand DL, you should start by honing your Category Theory skills, as almost everything in DL at its core can be mapped to a functor or an endofunctor.
taking it a step further, I'd say in many cases using the algebra jargon is harmful to understanding, and its better to just describe whats really going on. ie, "we add an L2 penalty term" --> want the sum of squares to be small. "project to vocab space" --> compute similarity to each vocab item.
i'll elaborate: a common computation pattern in DL happens to coincide with a known operator in linear algebra (matmul), and so we conveniently borrow linalg notation and terminology (matrices, vectors, ranks, norms). but this is just jargon. the algebric properties arent needed.
i'll elaborate: a common computation pattern in DL happens to coincide with a known operator in linear algebra (matmul), and so we conveniently borrow linalg notation and terminology (matrices, vectors, ranks, norms). but this is just jargon. the algebric properties arent needed.
"Modern ML is built on Linear Algebra". lol no its not.
why is "MCP" implemented as server exposing a set of endpoints, rather than as some JSON schema for defining tool descriptions and allowing these JSON files to be accessed over http? what is the purpose/benefit of the middleman server?
you know what, nah, we don't want to close it. it will be just 80% closed.
and now, we will proceed to peacefully close the strait of Hormuz. you know, for the environment.
and now, we will proceed to peacefully close the strait of Hormuz. you know, for the environment.
today, during a peaceful flight over an iranian mountain, a US airplane dropped a mostly peaceful bunker buster bomb, who flew peacefully until it hit the mountain and mostly peaceful facility underneath it. there was a brief period of violent detonation on impact, then peace again.
אחד מלקחי ליל המקלטים אמש הוא שאין לי סבלנות לקרוא ספרות מקצועית, אבל לקרוא ספרות קלה זה די סבבה. מצד שני הספר שהיה לי בנייד, הוא כזה שהתחלתי לקרוא והפסקתי והיתה סיבה שהפסקתי, הוא מייגע ומעפן. בקיצור שילחו המלצות לספרים. עברית או אנגלית, אבל באנגלית ככהנ יהיה לי יותר קל להתארגן.