Chao Peng, Ruida Hu, Ajitha Rajan, Tegawend\'e F Bissyand\'e, Jacques Klein, Cuiyun Gao: Can LLMs Test Terminal User Interfaces? https://arxiv.org/abs/2608.03743 https://arxiv.org/pdf/2608.03743 https://arxiv.org/html/2608.03743
arXiv cs.SE Software Engineering
@csse-bot.bsky.social
Unofficial bot by @vele.bsky.social w/ http://github.com/so-okada/bXiv https://arxiv.org/list/cs.SE/new List https://bsky.app/profile/vele.bsky.social/lists/3lim7ccweqo2j ModList https://bsky.app/profile/vele.bsky.social/lists/3lim3qnexsw2g
Sun, Lyu, Ma, Kuang, Widyasari, Zhang, Ma, Lawall, Lo: We Must Have Missed This Comment: Detecting and Repairing Stale Function References in Linux Kernel Comments https://arxiv.org/abs/2608.03734 https://arxiv.org/pdf/2608.03734 https://arxiv.org/html/2608.03734
Khai-Nguyen Nguyen, Oscar Chaparro, Antonio Mastropaolo: Pattern over Pixels: Measuring Pattern Completion Bias in Multimodal Code Generation https://arxiv.org/abs/2608.03691 https://arxiv.org/pdf/2608.03691 https://arxiv.org/html/2608.03691
Cunming Zhang, Yu Pei, Michail Papadakis: From Bug Reports to Browser-Executable Procedures: An LLM-Driven Agent for Web GUI Bug Reproduction https://arxiv.org/abs/2608.03598 https://arxiv.org/pdf/2608.03598 https://arxiv.org/html/2608.03598
Haowen Yang, Yun Peng, Zishuo Ding: EffiHolmes: Differential Profiling-Guided Repository Level Time Inefficiency Fix Localization https://arxiv.org/abs/2608.03558 https://arxiv.org/pdf/2608.03558 https://arxiv.org/html/2608.03558
Shahbaz Siddeeq, Muhammad Waseem, Umar Subhan Malhi, Pekka Abrahamsson: CodeAssay: A Multi-Metric Benchmark with Audited Ground Truth for LLM Code Generation https://arxiv.org/abs/2608.03535 https://arxiv.org/pdf/2608.03535 https://arxiv.org/html/2608.03535
Simos Gerasimou, Xingyu Zhao: Leveraging System-Level Observations to Inform Bayesian Learning of Model Parameters for Quantitative Verification https://arxiv.org/abs/2608.03489 https://arxiv.org/pdf/2608.03489 https://arxiv.org/html/2608.03489
Giusy Annunziata, Rudrajit Choudhuri, Anita Sarma, Gemma Catolino, Filomena Ferrucci: When AI Joins the Team! A Model of How AI Adoption Relates To Social Patterns in Software Engineering Teams https://arxiv.org/abs/2608.03462 https://arxiv.org/pdf/2608.03462 https://arxiv.org/html/2608.03462
Hao Zhou, Haichuan Hu, Ye Shang, Quanjun Zhang: Self-Evolving Coding Agents https://arxiv.org/abs/2608.03392 https://arxiv.org/pdf/2608.03392 https://arxiv.org/html/2608.03392
Erxue Zhou, Jingxiang Meng, Aofan Liu: Route-Align-Verify for Functional Correctness in Code Generation https://arxiv.org/abs/2608.03341 https://arxiv.org/pdf/2608.03341 https://arxiv.org/html/2608.03341
Yu Pei, Cunming Zhang, Jeongju Sohn, Mike Papadakis: Assessing Behavioral Validation in UI Component Test Suites Using Inferred Metamorphic Relations https://arxiv.org/abs/2608.03337 https://arxiv.org/pdf/2608.03337 https://arxiv.org/html/2608.03337
Yunqi Chen, Thomas Zimmermann, Bianca Trinkenreich: Making AI Visible, Not Vanished: How AI Policies Reshape Developer Experience on GitHub https://arxiv.org/abs/2608.03329 https://arxiv.org/pdf/2608.03329 https://arxiv.org/html/2608.03329
Wrenn, Raykov, Angelov, Kitahara, Watanabe, Sailer: Evaluating LLM Trade-offs for Enterprise Automation: Lessons from Workflow Generation in a Production Enterprise Platform https://arxiv.org/abs/2608.03311 https://arxiv.org/pdf/2608.03311 https://arxiv.org/html/2608.03311
Jikai Wang, Ningyu He, Tianming Liu, Junhai Wang, Haoyu Wang: NotDec: WebAssembly Decompilation With Inter-Procedural Type Recovery https://arxiv.org/abs/2608.03286 https://arxiv.org/pdf/2608.03286 https://arxiv.org/html/2608.03286
Chenyu Wang, Yunbo Lyu, Junda He, Zhou Yang, Chenxing Zhong, Yaniv Harel, David Lo: Fail-Fast, Restart-Smart: Early Failure Prediction and Restart for SWE Agentic Tasks https://arxiv.org/abs/2608.03222 https://arxiv.org/pdf/2608.03222 https://arxiv.org/html/2608.03222
Mir Mohammad Yousuf, Shabir Ahmad Sofi, Bisma Majid: Cross-Ecosystem Bug Classification in Quantum Software https://arxiv.org/abs/2608.03173 https://arxiv.org/pdf/2608.03173 https://arxiv.org/html/2608.03173
Yongmin Li, Yihong Dong, Jia Li, Ge Li: Efficient Grammar-Constrained Decoding via Parser Stack Classification https://arxiv.org/abs/2608.03065 https://arxiv.org/pdf/2608.03065 https://arxiv.org/html/2608.03065
Breno Cerqueira Reis Nakamura, Arlindo Flavio da Concei\c{c}\~ao: Integration Barriers in Open-Source SSI Frameworks: An Exploratory Developer Experience Probe https://arxiv.org/abs/2608.03039 https://arxiv.org/pdf/2608.03039 https://arxiv.org/html/2608.03039
Forough Majidi, Mohammad Mehdi Morovati, Foutse Khomh, Heng Li: LLM Serving in the Wild: An Empirical Study of Frameworks, Methods, and System Designs https://arxiv.org/abs/2608.03036 https://arxiv.org/pdf/2608.03036 https://arxiv.org/html/2608.03036
Thomas Bock, Audris Mockus, Bogdan Vasilescu: The Ground Is Shifting: A Reflection on the Foundations of Software Measurement https://arxiv.org/abs/2608.03007 https://arxiv.org/pdf/2608.03007 https://arxiv.org/html/2608.03007
Shuai Shao, Dingbang Wang, Yiming Zeng, Tingting Yu: ConFL: Explainable Concurrent Fault Localization via Hierarchy-Guided LLM Reasoning https://arxiv.org/abs/2608.02974 https://arxiv.org/pdf/2608.02974 https://arxiv.org/html/2608.02974
Shuai Shao, Yiming Zeng, Yu Zhao, Tingting Yu: HyperFL: Query-Adaptive Representation Learning for Software Fault Localization https://arxiv.org/abs/2608.02967 https://arxiv.org/pdf/2608.02967 https://arxiv.org/html/2608.02967
Sun, Wu, Chen, Hou, Qiu, Cao, Shao, Lu, Zhang: Don't Regenerate, Debug: A Domain-Specific Agent for Repairing Near-Miss Hardware Operators https://arxiv.org/abs/2608.02712 https://arxiv.org/pdf/2608.02712 https://arxiv.org/html/2608.02712
Yuekun Wang, Mingfei Cheng, Xiaofei Xie: PRWeaver: Evaluating LLM-Based Code Auditors against Long-Horizon Malicious Pull Requests https://arxiv.org/abs/2608.02693 https://arxiv.org/pdf/2608.02693 https://arxiv.org/html/2608.02693
Zetong Xiong, et al.: BulkPR-Bench: Benchmarking Queue-Level Governance of Interacting Pull Requests https://arxiv.org/abs/2608.02685 https://arxiv.org/pdf/2608.02685 https://arxiv.org/html/2608.02685
Osamah H. Alaini, Taher A. Ghaleb: Studying Developer Perceptions on the Potential of CI Recommendation Systems https://arxiv.org/abs/2608.02682 https://arxiv.org/pdf/2608.02682 https://arxiv.org/html/2608.02682
Salma El Yadouni (EPFL), Guanyi Li (Binome Technologies): TraceCompiler: Skill-Guided Mining and Compilation of LLM Agent Traces into Mostly Deterministic Workflows https://arxiv.org/abs/2608.02680 https://arxiv.org/pdf/2608.02680 https://arxiv.org/html/2608.02680
Waseem, Islam, Shuvo, Hasan, Kemell, Rasku, Saari, Saari, Pajasmaa, Oivo, Abrahamsson: AI Sandbox: Technical Report https://arxiv.org/abs/2608.02679 https://arxiv.org/pdf/2608.02679 https://arxiv.org/html/2608.02679
Rasvik Kudum, Max Corbett, Hitansh Paliwal, Romaisa Fatima, Thomas Jiralerspong, Sneheel Sarangi: When Policies Change Probabilities: Modular Decision-Making for LLM Code Review https://arxiv.org/abs/2608.02677 https://arxiv.org/pdf/2608.02677 https://arxiv.org/html/2608.02677
Ilesh Vora, Srikanth Thudumu, John Carlson, Harshil Kamdar, Rajesh Vasa: Intent-Level Quantum Programming with Assertion-Guided Execution and Inspectable Intermediate Representation https://arxiv.org/abs/2608.02648 https://arxiv.org/pdf/2608.02648 https://arxiv.org/html/2608.02648