So what do AI evaluations do or should do? Wrote a brief reflection about prototypical questions evaluations often try to answer and why it is useful to have clarity about what exactly one is trying to learn about some phenomenon of interest. rigor-in-ai.leaflet.pub/3ms4ax4sojs2j
Back-to-basics: so what do evaluations do
TL;DR — Developing useful evaluations requires clarity about what exactly one is trying to learn about a phenomenon of interest.
rigor-in-ai.leaflet.pub