André Cruz

@andcrz.bsky.social

🎓 PhD student at the Max Planck Institute for Intelligent Systems 🔬 Safe and robust AI, algorithms and society 🔗 https://andrefcruz.github.io 📍 researcher in 🇩🇪, from 🇵🇹

LLM leaderboards aren't passive measurements — they're mechanisms that create incentives! Why does the current evaluation protocol reward "benchmaxxing": post-training for the benchmark without improving model capability? Our #ICML2026 paper gives a theoretical answer! 🧵

Bild

At ICLR and interested in theory for LLMs? Join us at our poster to learn more about the (im)possibility of scaling laws for test-time scaling methods like Best-of-N when verification is imperfect!

There is now a whole sub-industry around LLM routing, gateways, even compute arbitrageurs (e.g. inference dot net). This is a basic but nice study on arbitrage in such settings and implications for the ecosystem (e.g. price drops, market entry, revenue cannibalization etc.) arxiv.org/abs/2603.22404

Computational Arbitrage in AI Model Markets

Consider a market of competing model providers selling query access to models with varying costs and capabilities. Customers submit problem instances and are willing to pay up to a budget for a verifi...

arxiv.org