🔥囧Robert Osazuwa Ness囧🔥

@osazuwa.bsky.social

Probabilistic machine Learning, causal inference, language models. Teach at http://Altdeep.ai & @Northeastern, work at @MSFTResearch.

Anyone know of any work that evaluates the relationship between LLM prompting strategies and generalizability? Eg, if one applies a bunch of prompting hacks to ramp up accuracy on a benchmark, you might be sacrificing the ability for that prompt to generalize to new settings?