Tiancheng Hu

@tiancheng.bsky.social

PhD student @CambridgeLTL; Previously @DLAB @EPFL; Interested in NLP and CSS. Apple Scholar, Gates Scholar.

[#ACL2026 Paper Alert] Real-world requests are often underspecified. So when should an AI agent ask a clarification question? Not always. Not never. We introduce Value of Information (VoI): a decision-theoretic framework for deciding when to ask, when to act, and when to stop.

Bild

SimBench now at #ICLR2026! Often in social simulations, the goal is not to predict what one specific person will do. It is to estimate how a group will respond, whether in pre-testing a real polling question, or in stress-testing a policy or intervention before running it in the real world.

1/7 🧵 The GPT-4 technical report featured detailed calibration curves. Since then, not a single major model release has reported calibration. The field quietly stopped measuring whether models know what they don't know. Our new position paper argues this is a mistake. Here's why.

Bild

Proud to contribute to the new International AI Safety Report chaired by @YoshuaBengio, with a fantastic international team! Every word was weighed to ensure a rigorous, evidence-based view of current AI capabilities and the risks they pose. A short summary of my section below.

Yoshua Bengio@yoshuabengio.bsky.social · 6mo ago

Today we’re releasing the International AI Safety Report 2026: the most comprehensive evidence-based assessment of AI capabilities, emerging risks, and safety measures to date. 🧵 (1/19)

I’m pleased to share the Second Key Update to the International AI Safety Report, which outlines how AI developers, researchers, and policymakers are approaching technical risk management for general-purpose AI systems. (1/6)

Bild

Instruction tuning unlocks incredible skills in LLMs, but at a cost: they become dangerously overconfident. You face a choice: a well-calibrated base model or a capable but unreliable instruct model. What if you didn't have to choose? What if you could navigate the trade-off? (1/8)

Can AI simulate human behavior? 🧠 The promise is revolutionary for science & policy. But there’s a huge "IF": Do these simulations actually reflect reality? To find out, we introduce SimBench: The first large-scale benchmark for group-level social simulation. (1/9)

Excited to share a "Key Update" from the International AI Safety Report, which I was proud to contribute to. We took a rigorous, evidence-based look at the latest AI developments. If you want a clear view of where things stand, this is a must-read. 👇

Yoshua Bengio@yoshuabengio.bsky.social · 10mo ago

AI is evolving too quickly for an annual report to suffice. To help policymakers keep pace, we're introducing the first Key Update to the International AI Safety Report. 🧵⬇️ (1/10)

We created Approximate Likelihood Matching, a principled (and very effective) method for *cross-tokenizer distillation*! With ALM, you can create ensembles of models from different families, convert existing subword-level models to byte-level and a bunch more🧵

Image illustrating that ALM can enable Ensembling, Transfer to Bytes, and general Cross-Tokenizer Distillation.

Ever notice how something that makes your blood boil barely registers with your friend? Our emotional reactions aren't universal at all—they're deeply personal. And AI needs to understand that. Excited to share our new paper: "iNews" 🧵 (1/8) arxiv.org/abs/2503.03335

iNews: A Multimodal Dataset for Modeling Personalized Affective Responses to News

Current approaches to emotion detection often overlook the inherent subjectivity of affective experiences, instead relying on aggregated labels that mask individual variations in emotional responses. ...

arxiv.org