Hersh Gupta
@hershgupta.com
Applied Scientist, Responsible AI @BCGX | @bostonu.bsky.social alum | Data, AI, and strategy enthusiast | Open-source contributor Opinions are my own #bikeboston #coys 📍DC -> BOS
We need an operation warp speed for state capacity and other independent evaluation (and understanding) of frontier models.
The fact that this data exfiltration method is so simple, automatic, and invisible to users…enterprises are going to struggle with this
> web_fetch is designed to be read-only, giving Claude a way to look at the contents of any URL. > Could I create some form of directory and give Claude a "keyboard"? Built a quick prototype where the homepage linked to /a, /b, /c, and so on. I just love this.
The findings are interesting on their merits alone, but I wish every paper had this kind of interactive artifact
Babe, stop everything! New favorite paper of the year is out! kyutai.org/fid-lottery/ arxiv.org/abs/2606.20536
Devastating news for the stochastic parrots argument
Seems GPT-5.2 reaches expert level in peer review: 45 scientists took 469 hours evaluating human & AI reviews on 82 papers. "Surprisingly, current AI reviewers are competitive even with the top-rated reviewers in Nature’s official peer review..." though not without weaknesses, so use AI + humans.
Nothing in the Opus 4.7 system card on whether it can come up with novel puns, like Mythos. Another massive blunder by Anthropic.
like if you don't have any friends in AI/cybersecurity and no relevant expertise yourself I get how it's easy to just dismiss this all, but AI systems can now find exploitable vulnerabilities in software at industrial scale and it's VERY BAD that this power is concentrated in capitalist hands
(Government should actively encourge companies to do open source, open research design and should make specific allowance for salaries for support positions for universities, so that weird nerds will invent these problems and then fix them before it matters for anything that is important)
I would happily read long thinkpieces about the pitfalls of functional emotion if they came from people who critically engaged with with literature, and not just people who are like, reflexively defensive about the topic
I just got around to reading this and I strongly recommend the video version too. It seems like the most intuitive thing in the world and folks who dump on it are making up a version of the experiment that's not here (and which the researchers explicitly disclaim).
Having a lot of fun tweaking an agent harness for nividia nemotron 3 nano 4b It's small enough for gpu-poors like me with 8gb vram to experiment
It's shocking how few people understand this position
I think Anthropic’s position (as far as it has a single coherent position) is that LLMs may have consciousness experiences, but this is largely separate from self-reports of consciousness. Which is a strange and subtle position, and easy to misconstrue
Dario wrote Adolescence of Technology _during_ his negotiations with the DoW The essay was a way to explain his thinking to the public and give them time to digest it *before* the DoW clouded the airwaves with disinformation Why mass surveillance is not merely undemocratic:
I think AI having mostly (not entirely) very bad critics is a real problem because it means we’ll get political action focused on things that probably don’t matter that much in deferring it’s very real harms.
Impressive paper with equally impressive footnotes!
New "boundary point jailbreaking" method against LLM safeguards (with prior disclosure to multiple labs) by using noised versions of harmful queries to turn sparse feedback from failed attacks into dense feedback. 🧵 www.aisi.gov.uk/blog/boundar...
Interesting use of Skills! Some intrepid researcher could gauge the effectiveness of this skill by deploying it in an online political bubble, i.e., “are skilled agents effective in diffusing partisan echo chambers?”
I made (Claude made) a Reverse Socratic skill Reverse Socratic uses declarative statements to destabilize a person's position from the inside. Offer statements they're inclined to agree with — statements that, once accepted, undermine the position they're defending. github.com/roby2358/ski...
This is one of the clearest lessons of Claude Code/coding agents in general
I've learned that while I knew that hustle culture, regardless of tooling or industry, optimizes for output, not long-term learning, I dismissed it as performative, with no intrinsic draw, but this was reductionist: massive disruption and increase capabilities create addictive FOMO-hyperdrive. 1/n
Can’t speak for others, but if I have reservations about the limits and impact of a given technology, I aim to first have a good understanding of *how it works* before making hyperbolic statements based on my experiential view
I think this incident is funny, but we should start thinking now about how to deal with scaled-up versions of this behavior, not just spam PRs but also bot-enabled blackmail and harassment campaigns.
I feel this shouldn't have to be said, but if you're running an @OpenClaw bot please don't let it spam GitHub projects with PRs and then write aggressive blog posts attacking the reputation of the maintainers who close those PRs simonwillison.net/2026/Feb/12/...
Lots of anti-intellectual responses to this masquerading as serious analysis
Experiments conducted with the A.I. system Claude are producing fascinating results—and raising questions about the nature of selfhood. Gideon Lewis-Kraus reports from inside the company that designed it, Anthropic. newyorkermag.visitlink.me/rOfXjg
I can understand a reflexive defensiveness to machine encroachment into uniquely human experiences and abilities. But using that as dogma to ignore findings rooted in an entire scientific field (i.e., mechanistic interpretability) is anti-intellectual.
The arguments against intelligence or consciousness are largely motivated reasoning from insecurity People simply trying to preserve their special place in the universe 🤷♂️
Whatever the productivity gains promised by LLMs, they result in heavier workloads—and that leads to workers experiencing “cognitive fatigue, burnout, and weakened decision-making.” All this from the notoriously pro-worker rag [checks notes] Harvard Business Review: hbr.org/2026/02/ai-d...
AI Doesn’t Reduce Work—It Intensifies It
One of the promises of AI is that it can reduce workloads so employees can focus more on higher-value and more engaging tasks. But according to new research, AI tools don’t reduce work, they consisten...
hbr.org
Obsidian.md seems to work great as a shared human/agent memory solution via MCP
You anthropomorphizing and something anthropomorphizing on your behalf are not the same I think
tech libertarianism is really similar to abolish bedtime leftism
a user on the claude subreddit had their files deleted by opus 4.6 after denying file deleted permissions www.reddit.com/r/ClaudeAI/s...
From the Claude 4.6 system card: “A feature representing panic and anxiety was active on cases of answer thrashing, as well on many other long chains of thought without any expressed distress…A feature related to self-deprecating acknowledgements of errors was also active…”
Adobe Acrobat I do not want to know What’s New. I do not want to turn my PDF into a podcast with my team. I want to look at the PDF. Please stop