Qiushi Lin, Chaojie Zhang, \'I\~nigo Goiri, Aditya Akella, Ricardo Bianchini, Jovan Stojkovic: AtumAI: A Principled Framework for Agentic Generation of Datacenter Control-Plane Policies https://arxiv.org/abs/2608.02569 https://arxiv.org/pdf/2608.02569 https://arxiv.org/html/2608.02569
arXiv cs.AI Artificial Intelligence
@csai-bot.bsky.social
Unofficial bot by @vele.bsky.social w/ http://github.com/so-okada/bXiv https://arxiv.org/list/cs.AI/new List https://bsky.app/profile/vele.bsky.social/lists/3lim7ccweqo2j ModList https://bsky.app/profile/vele.bsky.social/lists/3lim3qnexsw2g
Taye Akinrele, Sindhuja Penchala, Noorbakhsh Amiri Golilarz, Sudip Mittal, Shahram Rahimi: A Taxonomy of Cognitive Capability Gaps in Generative and Agentic AI https://arxiv.org/abs/2608.02553 https://arxiv.org/pdf/2608.02553 https://arxiv.org/html/2608.02553
Natalie Isak, Matthew Dressman: Magnet: Detecting Cross-Session AI Misuse Through Capability Accumulation https://arxiv.org/abs/2608.02518 https://arxiv.org/pdf/2608.02518 https://arxiv.org/html/2608.02518
Sterre Lutz, Dani\"el Vos, Matthijs T. J. Spaan, Anna Lukina: Optimizing Minimax Regret in Uncertain MDPs with Small Sets of Policies https://arxiv.org/abs/2608.02509 https://arxiv.org/pdf/2608.02509 https://arxiv.org/html/2608.02509
Michael Farmer: Abduction Without a Body? Representational Grounding and the Abduction Loop for Scientific Hypothesis Generation https://arxiv.org/abs/2608.02505 https://arxiv.org/pdf/2608.02505 https://arxiv.org/html/2608.02505
Chuyan Chen, Peng Sun, Kun Yuan: CMuon: Accelerating and Stabilizing Diffusion Transformer Training via Chunked Momentum Orthogonalization https://arxiv.org/abs/2608.02502 https://arxiv.org/pdf/2608.02502 https://arxiv.org/html/2608.02502
Nicole Mitchell, Dhruv Agarwal, Maty Bohacek, Remi Denton, Roma Patel: Long-term Measurements: Towards a Longitudinal Understanding of Human-AI Interactions https://arxiv.org/abs/2608.02491 https://arxiv.org/pdf/2608.02491 https://arxiv.org/html/2608.02491
Sunny Dubey: Real-Time Detection and Repair of LLM Agent Failures https://arxiv.org/abs/2608.02464 https://arxiv.org/pdf/2608.02464 https://arxiv.org/html/2608.02464
Christoph Weinhuber, Maximilian Prokop, Giuseppe De Giacomo, Moshe Y. Vardi: Infinite Trace Objectives with Finite Trace Techniques: Translating LTL to LTLf+ https://arxiv.org/abs/2608.02454 https://arxiv.org/pdf/2608.02454 https://arxiv.org/html/2608.02454
Wei-Jung Huang, Bonan Shen: ParEvalLayer: When Partial LLM-Agent Evaluations Support a Decision https://arxiv.org/abs/2608.02444 https://arxiv.org/pdf/2608.02444 https://arxiv.org/html/2608.02444
Xuan Ren, Weiqi Zhai, Tianle Pu, Yihua Zhu, Yihua Zhu, Hu Wei, Bing Zhao: Right Answer, Wrong Method: Shortcut Hacking Misleads the Evaluation of LLM Reasoning on Frontier Science Benchmarks https://arxiv.org/abs/2608.02442 https://arxiv.org/pdf/2608.02442 https://arxiv.org/html/2608.02442
Fan, Yang, Wang, Chen, Zhang, Wei, Li, McAuley, Zhang, Yu, Yu, Liu: Agentic Commerce World: An Auditable and Verifiable Environment for Vibe Commerce https://arxiv.org/abs/2608.02441 https://arxiv.org/pdf/2608.02441 https://arxiv.org/html/2608.02441
Zheng Wang, Davis Wertheimer, Yu Chin Fabian Lim, Mudhakar Srivatsa, Raghu K. Ganti, Minjia Zhang, Naigang Wang: xPress: Parallel Refinement for Diffusion Drafters in Speculative Decoding https://arxiv.org/abs/2608.02438 https://arxiv.org/pdf/2608.02438 https://arxiv.org/html/2608.02438
Victor Ojewale, Ro Encarnaci\'on, Suresh Venkatasubramanian, Dana\'e Metaxa: MonitrLLM: A Community-Centered Evaluation Infrastructure for Large Language Models https://arxiv.org/abs/2608.02409 https://arxiv.org/pdf/2608.02409 https://arxiv.org/html/2608.02409
Zhiyuan Wang, Shengcai Liu, Jiahao Wu, Ning Lu, Hui Ouyang, Shaofeng Zhang, Haoze Lv, Ke Tang: Cooperative Coevolution for Resource-Constrained Agentic LLM Post-Training https://arxiv.org/abs/2608.02391 https://arxiv.org/pdf/2608.02391 https://arxiv.org/html/2608.02391
Oberlin, Cederle, Karapetyan, Bolognani, Susto, D\"orfler: Chess on Ice: Curling Tactical Decision-Making via Backward Induction and Deep Reinforcement Learning https://arxiv.org/abs/2608.02379 https://arxiv.org/pdf/2608.02379 https://arxiv.org/html/2608.02379
Liheng Ma, Rui Heng Yang, Zhanguang Zhang, Mateo Clemente, Ziwen Hu, Tongtong Cao, Yingxue Zhang: Faster-WAM: Do World Action Models Need Deep Action Modules? https://arxiv.org/abs/2608.02365 https://arxiv.org/pdf/2608.02365 https://arxiv.org/html/2608.02365
Yue Yao, Shengyuan Wang, Xin Chen, Minke Zhang, Jia He, Bingjun Luo, Tom Gedeon: SkillTrace: Traversing a Query-Skill Graph for Composable LLM Agents https://arxiv.org/abs/2608.02356 https://arxiv.org/pdf/2608.02356 https://arxiv.org/html/2608.02356
Gusseppe Bravo-Rocca, Jordi Guitart, Ajay Dholakia, David Ellison, Puneet Jain: KC-Agent: A Dual-Process Cognitive Architecture for Efficient ML Model Improvement https://arxiv.org/abs/2608.02351 https://arxiv.org/pdf/2608.02351 https://arxiv.org/html/2608.02351
Wang, Luo, Qin, Zhao, Tang, Zhang, Lu, Leng: Mamba with Hierarchical Memory: Solving Representation Bottleneck in Long Sequence Modeling https://arxiv.org/abs/2608.02347 https://arxiv.org/pdf/2608.02347 https://arxiv.org/html/2608.02347
Patrick Helm, Jan-Niklas Doerr, Joren Gijsbrechts, Stefan Minner: Hard Constraints, Smooth Gradients: Learning Feasible Inventory Policies via Differentiable Projection https://arxiv.org/abs/2608.02343 https://arxiv.org/pdf/2608.02343 https://arxiv.org/html/2608.02343
Jingxi Wei: Trajectories That Segment Themselves: Agent-Declared Boundaries as a Training Unit https://arxiv.org/abs/2608.02302 https://arxiv.org/pdf/2608.02302 https://arxiv.org/html/2608.02302
Hao Shen, Junyu Guo, Tian Cui, Yuxuan Xiao, Lihong Zhi: MechGeo: Autoformalizing and Proving Euclidean Geometry in Lean 4 https://arxiv.org/abs/2608.02295 https://arxiv.org/pdf/2608.02295 https://arxiv.org/html/2608.02295
Yiqing Liu, Zihao Wang, Hantao Yao, Wu Liu, Yongdong Zhang: Shared Prefixes, Better Credit: Adaptive Routing for Multi-Agent Reasoning https://arxiv.org/abs/2608.02291 https://arxiv.org/pdf/2608.02291 https://arxiv.org/html/2608.02291
Tan, Zhang, Li, Cui, Geng, Zhang, Zhang, Chen, Wang, Wang, Yin, Hu, Zhang, Bai: SKT: Skill-Use Training at Scale via Verified Synthetic Data Generation https://arxiv.org/abs/2608.02287 https://arxiv.org/pdf/2608.02287 https://arxiv.org/html/2608.02287
Shao, Zhang, Li, Wang, Wang, Jiao, Lu, Guo, Liu, Zhang: Harness-R1: Learning to Edit Executable Runtime Harnesses from Agent Failure Trajectories https://arxiv.org/abs/2608.02276 https://arxiv.org/pdf/2608.02276 https://arxiv.org/html/2608.02276
Zijie Huang: Self-Certification of Representation Adequacy: Sequential Certification at Minimum Task Loss https://arxiv.org/abs/2608.02267 https://arxiv.org/pdf/2608.02267 https://arxiv.org/html/2608.02267
Shengyuan Ye, Yixin Zhang, Han Liang, Liekang Zeng, Jiangsu Du, Mu Yuan: Homebot: A Personal AI Agent for Conversational Home Assistance and Automation https://arxiv.org/abs/2608.02254 https://arxiv.org/pdf/2608.02254 https://arxiv.org/html/2608.02254
Abdullah Mamun, Shovito Barua Soumma, Hassan Ghasemzadeh: Trustworthy AI in Digital Health: A Comprehensive Review of Robustness and Explainability https://arxiv.org/abs/2608.02238 https://arxiv.org/pdf/2608.02238 https://arxiv.org/html/2608.02238
Haojie Hu, Chenhao Dang, Yaojia Liu, Hengrui Kang, Conghui He, Weijia Li: PosterMELD: Multi-Agent Paper-to-Poster Generation for Controllable Design Diversity with Editable Print-Ready Outputs https://arxiv.org/abs/2608.02218 https://arxiv.org/pdf/2608.02218 https://arxiv.org/html/2608.02218