This new benchmark is an apt metaphor for the AIs ability to gaslight humans with legit-looking-but-wrong output. Oh, you can't find the raccoon? Huh. Sorry you're having trouble with that, meatsack. I tried to not make it TOO hard for you, lil bud, but I'll make it even easier next time.
I came up with a somewhat foolish new benchmark for testing image generation models, to exercise the new ChatGPT Images 2.0: "Do a where's Waldo style image but it's where is the raccoon holding a ham radio" simonwillison.net/2026/Apr/21/...