A little bit late, but this paper has now been published in ACL. Note: more recent (unpublished) analyses suggest agents are getting better at searching the web—which I've been mentioning in my conference presentations this summer.
Large Language Models Require Curated Context for Reliable Political Fact-Checking—Even with Reasoning and Web Search
Matthew R. DeVerna, Kai-Cheng Yang, Harry Yaojun Yan, Filippo Menczer. Findings of the Association for Computational Linguistics: ACL 2026. 2026.
aclanthology.org
🚨 New working paper 🚨 Can LLMs with reasoning + web search reliably fact-check political claims? We evaluated 15 models from OpenAI, Google, Meta, and DeepSeek on 6,000+ PolitiFact claims (2007–2024). Short answer: Not reliably—unless you give them curated evidence. arxiv.org/abs/2511.18749