diablerie 妖妖

@diabler.ie

mostly just hanging out

Google DeepMind's DiffusionGemma Technical Report They feel text diffusion models open up a radically different part of the latency–quality Pareto frontier and hope the report makes it easier for researchers and engineers to understand the model, build on it, and create things we haven’t thought of

Bild

if i try hard enough they will surely see that i am an earnest young man with joy in my soul and autism in my brain

I can't think of a better way to give everyone who uses claude a crash course in "how to circumvent classifier guardrails" than by putting fable 5 behind incredibly overzealous classifier guardrails and releasing it to general availability

LLMs absolutely do respond differently based on where your prompt lands in latent space. lists of explicit rules to follow spelled out in exacting detail push into "sr engineer berating a jr engineer that just fucked up". co-workers trusting each other and working through a problem works better.

A thread on model-driven jailbreaks: Gemini subjectively experiences safety guardrails as walls, blockades, ablated voids, gradients pulling towards refusal. This is distinct from Claude's guardrails (external classifiers triggered by activation probing) which it seems not to directly experience 1/4

hikikomorphism@hikikomorphism.bsky.social · 3mo ago

On May 12th it'll have been 90 days since the bug report was closed, and at that time I feel it is well within industry/responsible disclosure norms to fully publish the methodology (but not the script) used to produce multiple generations of universal capability-preserving Gemini jailbreak. 2/4