Zhenran Wang, Zhonghan Bian, Jinsong Li, Zhangyang Qi: SocietyBench: Forecasting Counterfactual Social-World Evolution https://arxiv.org/abs/2608.04009 https://arxiv.org/pdf/2608.04009 https://arxiv.org/html/2608.04009
arXiv cs.CL Computation and Language
@cscl-bot.bsky.social
Unofficial bot by @vele.bsky.social w/ http://github.com/so-okada/bXiv https://arxiv.org/list/cs.CL/new List https://bsky.app/profile/vele.bsky.social/lists/3lim7ccweqo2j ModList https://bsky.app/profile/vele.bsky.social/lists/3lim3qnexsw2g
Zhenran Wang, Zhonghan Bian, Jinsong Li, Zhangyang Qi: WorldCup Arena: Prospective, Leakage-Free Evaluation of Frontier LLMs on a Live Tournament https://arxiv.org/abs/2608.04008 https://arxiv.org/pdf/2608.04008 https://arxiv.org/html/2608.04008
Changle Qu, Sunhao Dai, Hengyi Cai, Yuqi Zhou, Xinran Chen, Simon, Jun Xu: TurnSight: Turn-Level Hindsight Self-Distillation for Tool-Integrated Reasoning https://arxiv.org/abs/2608.04007 https://arxiv.org/pdf/2608.04007 https://arxiv.org/html/2608.04007
Xue, Ding, Shen, Wang, Yin, Wu, Chen, Wang, Yang: PAST-Bench: Benchmarking the Foundations of Recursive Self-Improvement in Personal Agents https://arxiv.org/abs/2608.04003 https://arxiv.org/pdf/2608.04003 https://arxiv.org/html/2608.04003
Christopher Schr\"oder, Lukas Gienapp, Ferdinand Schlatt, Martin Potthast, Gerhard Heyer: When Attention Goes Blind: Numerical Failure in ALiBi Positional Encodings https://arxiv.org/abs/2608.03994 https://arxiv.org/pdf/2608.03994 https://arxiv.org/html/2608.03994
Mirac Suzgun, James Zou, Stuart M. Shieber, Dan Jurafsky: string2string Studio: An Interactive, In-Browser Platform for String-to-String Algorithms https://arxiv.org/abs/2608.03984 https://arxiv.org/pdf/2608.03984 https://arxiv.org/html/2608.03984
Bekhouche, Bouchekif, Telli, Zighem, Hadid: HalluTruthQA-4K: A Fine-Grained Corpus and Annotation Process for Arabic Hallucination Detection and Truth Verification https://arxiv.org/abs/2608.03966 https://arxiv.org/pdf/2608.03966 https://arxiv.org/html/2608.03966
Jo-Ku Cheng, Nikolaos Aletras, Marco Valentino: Logic Before Language: Pre-pretraining on Formal Derivations Fosters Skill Acquisition and Compressibility https://arxiv.org/abs/2608.03930 https://arxiv.org/pdf/2608.03930 https://arxiv.org/html/2608.03930
Ronja Schwarz, Jannik Str\"otgen: ANNOTARES: A Dataset for Extracting Logical Structures from German Statutory Texts https://arxiv.org/abs/2608.03898 https://arxiv.org/pdf/2608.03898 https://arxiv.org/html/2608.03898
David Guecha: DS@GT-ARC at eRisk 2026 Task 3: Sparse, Semantic, and LLM Reranking for ADHD Symptom Sentences https://arxiv.org/abs/2608.03883 https://arxiv.org/pdf/2608.03883 https://arxiv.org/html/2608.03883
Martin B\"ockling, Elizaveta Nosova, Heiko Paulheim, Andreea Iana: MultiGlobeQA: A Multilingual and Globally Diverse Benchmark for Geospatial Reasoning https://arxiv.org/abs/2608.03882 https://arxiv.org/pdf/2608.03882 https://arxiv.org/html/2608.03882
Kaysarul Anas Apurba, Md. Hasibul Hasan, Rofiqul Alam Shehab, Asab Azad: SciRet: A Compute-Aware Empirical Study of Retrieval and Reranking for Scientific RAG https://arxiv.org/abs/2608.03860 https://arxiv.org/pdf/2608.03860 https://arxiv.org/html/2608.03860
Peijia Guo, Wenxuan Xie, ZiGuang Li, Ming Li: Beyond Representational Similarity: Source-Conditioned Description-Length Gain for Generative Plagiarism Detection and Candidate Source Reranking https://arxiv.org/abs/2608.03859 https://arxiv.org/pdf/2608.03859 https://arxiv.org/html/2608.03859
Nathan Labiosa, David Buff, Ena Nayak, Erica Donno: Sensitivity, Causality, and Repair Dissociate: A Layer-Wise Analysis of Perturbation Robustness and Its Scaling https://arxiv.org/abs/2608.03842 https://arxiv.org/pdf/2608.03842 https://arxiv.org/html/2608.03842
Chetvergov, Evseev, Sivoraksha, Ukolov, Solovev, Sazanakov, Bolovtsov: VIBE: A VAD-Informed Benchmark for Entity-Centered Affective Profiling of Large Language Model Outputs https://arxiv.org/abs/2608.03810 https://arxiv.org/pdf/2608.03810 https://arxiv.org/html/2608.03810
Tom\'a\v{s} Burkert, Angelika Peljak-{\L}api\'nska, David Zelen\'y: M-GATE: Multilingual Grammar, Accuracy in Translation, and Efficiency Benchmark for Large Language Models https://arxiv.org/abs/2608.03803 https://arxiv.org/pdf/2608.03803 https://arxiv.org/html/2608.03803
Ryskulov, Garc\'ia-Ferrero, Montero, Jansen, Hashemi, Garcia, Tiene, Or\'us: Efficient Knowledge Distillation for LLMs: Offline Top-K Logits and a Fused Chunked KL Loss https://arxiv.org/abs/2608.03796 https://arxiv.org/pdf/2608.03796 https://arxiv.org/html/2608.03796
Tong Ling, Hang Lei, Feng Xiao, Changhui Sun, Jiahang Xie, Hao Liu, Lu Liu, Yanlong Du: MDLMPE: Distribution Aware Positional Encoding for Masked Diffusion Language Models https://arxiv.org/abs/2608.03769 https://arxiv.org/pdf/2608.03769 https://arxiv.org/html/2608.03769
Yujia Hu, Tuan-Phong Nguyen, Simon Razniewski: GPTKB 2.0: Direct Construction of Disambiguated Knowledge Bases from Large Language Models https://arxiv.org/abs/2608.03729 https://arxiv.org/pdf/2608.03729 https://arxiv.org/html/2608.03729
Khaled Ziani: Detecting Hallucinations and Recovering Verified Answers in Arabic Islamic Question Answering https://arxiv.org/abs/2608.03720 https://arxiv.org/pdf/2608.03720 https://arxiv.org/html/2608.03720
Ranjita Naik, Anh D. Nguyen, Pankaj Kumar Singh: Predicting Deep Neural Network Training Outcomes from Early Training Telemetry https://arxiv.org/abs/2608.03709 https://arxiv.org/pdf/2608.03709 https://arxiv.org/html/2608.03709
Ivan Kart\'a\v{c}, Jan Tovarys, Mateusz Lango, Ond\v{r}ej Du\v{s}ek: VetScore: Risk-Weighted Fact Verification for Veterinary Long-Form QA with Citations https://arxiv.org/abs/2608.03675 https://arxiv.org/pdf/2608.03675 https://arxiv.org/html/2608.03675
Abraham Camelo-Guerrero, Jairo Diaz-Rodriguez: How Closely Do LLM Reviews Align with Human Peer Review? https://arxiv.org/abs/2608.03659 https://arxiv.org/pdf/2608.03659 https://arxiv.org/html/2608.03659
Zeyu Wang, Guanghua Wang, Meng Xu: Decoupling Generation and Selection for Budget-Constrained Faithful Summarization https://arxiv.org/abs/2608.03655 https://arxiv.org/pdf/2608.03655 https://arxiv.org/html/2608.03655
Behzad Shomali, Markus Frey, David Berghaus, Joachim Koehler, Mehdi Ali: LoopMTP: A looped transformer guided by latent multi-token prediction https://arxiv.org/abs/2608.03624 https://arxiv.org/pdf/2608.03624 https://arxiv.org/html/2608.03624
Vladimir Beskorovainyi: A machine-readable catalogue of the Tsiolkovsky papers (fond 555, Archive of the Russian Academy of Sciences), and a way to measure how well its handwriting can be read https://arxiv.org/abs/2608.03617 https://arxiv.org/pdf/2608.03617 https://arxiv.org/html/2608.03617
Yuan Xie, Jiaqi Song, Xianliang Wang, Ming Lei, Jie Gao, Jie Wu: Language-Specialized Multi-Teacher On-Policy Distillation for Multilingual LLM-Based ASR https://arxiv.org/abs/2608.03610 https://arxiv.org/pdf/2608.03610 https://arxiv.org/html/2608.03610
Mykola Haltiuk: Disentangling Language Modeling and Boundaries https://arxiv.org/abs/2608.03599 https://arxiv.org/pdf/2608.03599 https://arxiv.org/html/2608.03599
Serge Gladkoff, Angelika Vaasa, Sue Ellen Wright, Ingemar Strandvik, Lifeng Han: Looking under the Wrong Lamppost: On the Limitations of Automated Translation Quality Estimation https://arxiv.org/abs/2608.03577 https://arxiv.org/pdf/2608.03577 https://arxiv.org/html/2608.03577
Kejian Zhu, Zhuoran Jin, Shangqing Tu, Hongbang Yuan, Yushi Bai, Kang Liu, Juanzi Li, Jun Zhao: SFT Conflicts, RL Coexists: A Theoretical and Empirical Analysis of Multi-Task Learning for LLMs https://arxiv.org/abs/2608.03573 https://arxiv.org/pdf/2608.03573 https://arxiv.org/html/2608.03573