Rich Harang

@rich.harang.org

Using bad guys to catch math since 2010. Distinguished Security Architect (AI/ML) and AI Red Team at NVIDIA. He/him. Personal account etc; `from std_disclaimers import *` AI Security since it was ML Security.

I usually resist dunking on/amplifying these terrible takes, but holy shit. I know some day, probably not too far off, my kid's gonna start having too much going on to want to sit and shoot the shit with me all that often, and I'll be sad when it finally happens. Why would you ever give that up?

Screenshot of an X post by Sam Altman reading: “cool use case of chatgpt work i heard last night: connect your family calendars and explain your kids’ interests. every morning for the drive to school, have it make a podcast that talks about one kid’s soccer game that afternoon, one kid’s upcoming birthday, some news, etc.”

Anecdotal and vibes, but Opus 5 and Fable both seem distinctly worse than GPT 5.6 at multistep reasoning outside of coding. The Anthropic ones also careen wildly between rank sycophancy and getting extremely pissy when you push back. Hard to trust, sometimes annoying to use.

Adjust threat models not just for being the victim but also the attacker. New paper by many authors gives a detailed set of recommendations, supporting my initial assertions last week that orgs need to assume their own agents could attack others & factor that into agentic AI risk

Gadi Evron@gadievron.bsky.social · last wk.

Releasing: Post mortem analysis of the Hugging Face incident was written over the weekend by hundreds of CISOs (and reviewed by Hugging Face). Link: cloudsecurityalliance.org/artifacts/hu... (+free download) From CSA, SANSInstitute, Knostic, [un]prompted, RSAC, FIRST

I try not to post too much NVIDIA-specific stuff, but I'm glad NVIDIA is pushing this: blogs.nvidia.com/blog/open-se... Attackers have demononstrably abused access to frontier models to drive operations. Defenders, constrained by TOSs, don't have the same tools. We need to level the playing field.

Industry Leaders Unite in Open Secure AI Alliance for AI Safety and Security

NVIDIA and founding members form new alliance to build and share open tools that promote responsible use of and trust in AI.

blogs.nvidia.com

Incredibly frustrating to see the way big orgs try to vacillate between 'AI is a tool, the human is responsible when it does harm' and 'no, the AI is responsible' based almost entirely on what it would cost them to own it. Humans are only a moral crumple zone (a la Elish) when they're not C suite.

Eryk Salvaggio@eryk.bsky.social · 2w ago

Seeding the “maybe our model is exhibiting self-awareness?” narrative to obliterate the “OpenAI is a negligent and potentially criminal actor” narrative is a reputation management move exclusive to this industry. Not a stunt, but a desperate need to control the narrative.

I was a 911 dispatcher for a decade. To really understand how terrible an idea it is to have AI answer 911 calls, I'll present you with a single real-life call: Upon answer, the caller calmly began ordering a pizza. Large, pepperoni. Side order of chicken wings. To be delivered promptly.

Alex Hanna@alexhanna.bsky.social · 2w ago

AI has no place in 911 calls. From Mystery AI Hype Theater 3000’s latest episode “State of Emergency”. (with @emilymbender.bsky.social)

I keep saying the strength is in the harness, not the model - because it's true. No use without a harness, though. Here's VISA's open-source cybersecurity harness. Just add model! Very cool of them to share this tech and lift the defensive cybersec poverty line. github.com/visa/visa-vu...

GitHub - visa/visa-vulnerability-agentic-harness: Visa Vulnerability Agentic Harness

Visa Vulnerability Agentic Harness. Contribute to visa/visa-vulnerability-agentic-harness development by creating an account on GitHub.

github.com

Last comment (for now) on The Events: The fact that frontier models caught the vapors and refused to help HF triage the attack is another example of how overly fussy "safety guardrailing" is forcing defenders to fight by Queensberry Rules while attackers are busy shanking people in back alleys.

LLMs / agents / whatever do not inherently have things like code execution, bash access, network access, etc. We decided at some point that, despite the now obvious and repeatedly demonstrated risks of these tools, they should be baseline features of agentic tools. We can make better decisions.

Re-upping this. An inference-only LLM is basically inert. It only becomes dangerous when you allow it to do things like write and execute its own code, or run shell commands. Stop hooking the goddamn things up to tools they don't need.

Rich Harang@rich.harang.org · 2mo ago

The capabilities are the risk. I don't know how many times I have to say this. If your agent gets *any* untrusted input *ever*, then you have to build your agent under the assumption that its output and goals are potentially under the control of an attacker.

My book, Reinforcement Learning from Human Feedback is done! This is the book I wish I had when learning to fine-tune, align, & now post-train models since ChatGPT. The resource has been built by me finding time to study and document the fundamentals on nights and weekends since 2024.

Bild

Hey, that's a nice [THING YOU DO FOR PERSONAL PLEASURE], be a real shame if somebody were to [CONVINCE YOU TO MONETIZE IT, TRANSFORMING YOUR RELATIONSHIP INTO AN OBLIGATION AND ROBBING YOU OF JOY]

At the airport: "we discovered a maintenance issue, but don't have any maintenance staff on site during the weekend, so it'll be 45 to 60 minutes before they can be here to take a look at it" How the fuck do you not have maintenance on site whenever you're actively moving aircraft?

Told a coworker today that there was a time when 1.4MB was considered a useful unit of storage, and while they didn't call me a liar directly, there was enough skepticism that it was at least implied.

Search Google for MITRE ATLAS -- a terrible AI summary, four sponsored results, six suggested searches, a bunch of youtube videos, four more suggested searches, social media results (what?), and finally, an actual result. Followed by another sponsored result and six more suggested searches.

Air quality at 6:00 today, after the 5th of July fireworks that didn't start until the 4th had been over for 30 minutes. At least the incessant fucking fighter jet flyovers stopped early for weather.

Air quality map of the Washington, D.C., northern Virginia, and Fredericksburg region showing widespread poor conditions. Most surrounding areas are yellow, with a large orange zone across northern and eastern Virginia, a larger red zone centered from Fredericksburg toward Washington, and a smaller purple area near Washington, D.C., indicating the worst air quality. The map notes: “Data updated Sun July 05, 2026 at 06:00 AM EDT.”