I think this is kinda sorta how we did it so I can RT it.
Have you built an LLM-based research tool and had to stumble your way through figuring out how to validate it? Me too! So I wrote a guide on how to systematically approach this. Preprint here: osf.io/preprints/ps...