Yangyang Qu, Massimiliano Todisco, Nicholas Evans: Phoneme-Aware Pronunciation Representations for L2-English L1-Background Accent Identification https://arxiv.org/abs/2609.10466 https://arxiv.org/pdf/2609.10466 https://arxiv.org/html/2609.10466
arXiv eess.AS Audio and Speech Processing
@eessas-bot.bsky.social
Unofficial bot by @vele.bsky.social w/ http://github.com/so-okada/bXiv https://arxiv.org/list/eess.AS/new List https://bsky.app/profile/vele.bsky.social/lists/3lim7ccweqo2j ModList https://bsky.app/profile/vele.bsky.social/lists/3lim3qnexsw2g
Rishabh Jain, Aristeidis Papadopoulos, Zhaofeng Lin, Naomi Harte: Candor-LR: A Dyadic Conversational Dataset for Audio-Visual Speech Recognition https://arxiv.org/abs/2609.10394 https://arxiv.org/pdf/2609.10394 https://arxiv.org/html/2609.10394
Shuubham Ojha, Carol Espy-Wilson: Teacher-Free Self-Distilled Consistency Trajectory Learning for Fast Speech Enhancement https://arxiv.org/abs/2609.10392 https://arxiv.org/pdf/2609.10392 https://arxiv.org/html/2609.10392
Rishabh Jain, Naomi Harte: AVSRBench: A Multi-Condition AVSR Benchmark https://arxiv.org/abs/2609.10366 https://arxiv.org/pdf/2609.10366 https://arxiv.org/html/2609.10366
Taejin Park, Ivan Medennikov, Kunal Dhawan, Weiqing Wang, Jagadeesh Balam, Boris Ginsburg: Pushing the Boundaries of Streaming Multi-Speaker ASR: A Systematic Study of Architectural Trade-offs https://arxiv.org/abs/2609.10265 https://arxiv.org/pdf/2609.10265 https://arxiv.org/html/2609.10265
Nan Xu, Mingxue Yang: SCNet: Enhancing GAN-based Speech Generation with Subband Condition Network and Magnitude-aware Phase Loss https://arxiv.org/abs/2609.10025 https://arxiv.org/pdf/2609.10025 https://arxiv.org/html/2609.10025
Shota Horiguchi, Takanori Ashihara, Marc Delcroix, Naohiro Tawara, Alexis Plaquet: Over-Tightening-Aware Pseudo-Labeling for Tight-Boundary Speaker Diarization https://arxiv.org/abs/2609.09965 https://arxiv.org/pdf/2609.09965 https://arxiv.org/html/2609.09965
Zhan, Wang, Hu, Zhang, Ren, Liang, Yu, Wu, Chen, Liu, Feng, Xue, Xie: SpeechAnnotator: A Context-Aware Multi-Agent Framework and Benchmark for Multidimensional Speech Annotation https://arxiv.org/abs/2609.09947 https://arxiv.org/pdf/2609.09947 https://arxiv.org/html/2609.09947
Cao, Mu, Lin, Li, Zhan, Liu, Xie, Zhang, Xue, Xie: NVV-Locator: From Transcript Tags to Acoustic Boundaries for Fine-Grained Nonverbal Vocalization Grounding https://arxiv.org/abs/2609.09940 https://arxiv.org/pdf/2609.09940 https://arxiv.org/html/2609.09940
Yuang Cao, Qirui Zhan, Jingbin Hu, Ziyu Zhang, Yunxiang Chen, Houdun Liu, Shuo Feng, Bengu Wu, Lei Xie, Liumeng Xue: Source-Adaptive Data Curation for Bilingual NVV-Aware ASR https://arxiv.org/abs/2609.09929 https://arxiv.org/pdf/2609.09929 https://arxiv.org/html/2609.09929
Zhang, Hu, Xie, Zhan, Li, Zhang, Ren, Li, Zhu, Chen, Xie: SphereVAE: Hyperspherical Latent Autoencoders for Robust Autoregressive Speech Representation Modeling https://arxiv.org/abs/2609.09903 https://arxiv.org/pdf/2609.09903 https://arxiv.org/html/2609.09903
Mingyu Zhao, Zhiyong Wu: UniStream: Multi-Expert Residual Vector Quantization for 48 kHz Causal Streaming Audio Coding https://arxiv.org/abs/2609.09866 https://arxiv.org/pdf/2609.09866 https://arxiv.org/html/2609.09866
Minu Kim, Eunjung Yeo, Kwanghee Choi, June-Woo Kim: Language Orthogonalization of Self-Supervised Speech Representations for Cross-lingual Parkinson's Detection https://arxiv.org/abs/2609.09499 https://arxiv.org/pdf/2609.09499 https://arxiv.org/html/2609.09499
[2026-09-10 Thu (UTC), 13 new articles found for eessAS Audio and Speech Processing]
Orantqing, Ji, Tong, Zuo, Fu, Cao, Li, Wu, Franz, Evan, Veyra, Pan, Lu, Yang, Xie, Tan, Shen, Yang, Wang, Teddysun, Steveyves, Zhao, Bryanytian: Omni Interaction Agent Technical Report https://arxiv.org/abs/2609.08977 https://arxiv.org/pdf/2609.08977 https://arxiv.org/html/2609.08977
Daniel Kohlsdorf, Denise Herzing, Thad Starner: Interpreting Dolphin Vocal Sequences via Multiple Sequence Alignment https://arxiv.org/abs/2609.08795 https://arxiv.org/pdf/2609.08795 https://arxiv.org/html/2609.08795
Nour Bouayed, Adrien Llave, J\'er\^ome Daniel, Pascal Scalart: Spatial Audio Coding Through Relative Room Impulse Response Estimation https://arxiv.org/abs/2609.08542 https://arxiv.org/pdf/2609.08542 https://arxiv.org/html/2609.08542
Min, Dai, Zheng, Fan, Xiang, Zhang, Song, Shi, Zhao, Li: Semantic Refinement of Universal Audio Representations through Audio-Description Alignment https://arxiv.org/abs/2609.08429 https://arxiv.org/pdf/2609.08429 https://arxiv.org/html/2609.08429
Fulvio Missoni, Katarina C. Poole, Tim Murray-Browne, Andrea Canessa, Lorenzo Picinali: Beyond Localisation Accuracy: Sensorimotor Effects of HRTF Individualisation https://arxiv.org/abs/2609.08422 https://arxiv.org/pdf/2609.08422 https://arxiv.org/html/2609.08422
Sunil Tyagi: Open-Set Vessel Re-Identification from Underwater Ship-Radiated Noise with a Raw-Waveform Selective-Kernel Acoustic Neural Network (SKANN) and a Cross-Passage Evaluation Protocol https://arxiv.org/abs/2609.07399 https://arxiv.org/pdf/2609.07399 https://arxiv.org/html/2609.07399
Ziyi Yang, Zhengding Luo, Boxiang Wang, Libin Zhang, Woon-Seng Gan: Direction-Preserving Active Noise Control with a Conditional Control-Filter Estimation Network https://arxiv.org/abs/2609.07173 https://arxiv.org/pdf/2609.07173 https://arxiv.org/html/2609.07173
Duy Vo, Kiet Anh Hoang, Hao Do: SETEAB: Multiscale approach with Squeeze-and-Excitation Temporal Enhanced Aware Block for Speech Emotion Recognition https://arxiv.org/abs/2609.06101 https://arxiv.org/pdf/2609.06101 https://arxiv.org/html/2609.06101
[2026-09-09 Wed (UTC), 8 new articles found for eessAS Audio and Speech Processing]
[2026-09-08 Tue (UTC), no new articles found for eessAS Audio and Speech Processing]
Yao Guo, Yang Ai, Hui-Peng Du, Xiao-Hang Jiang, Chen-Yuan Ning, Zhen-Hua Ling: Enhancing Neural Speech Coding with Semantic and Visual Cues https://arxiv.org/abs/2609.05076 https://arxiv.org/pdf/2609.05076 https://arxiv.org/html/2609.05076
Shrishti Saha Shetu, Emanu\"el A. P. Habets, Andreas Brendel: Discriminative Flow Matching: Beyond Time-Conditioning in Generative Restoration via Flow-State Representations https://arxiv.org/abs/2609.04525 https://arxiv.org/pdf/2609.04525 https://arxiv.org/html/2609.04525
Shreeram Suresh Chandra, Zexin Cai, Yu Tsao, Simon King, Berrak Sisman: Brain2Speech-Net: Intelligible, Real-Time Brain-to-Speech Synthesis Without Text Decoding https://arxiv.org/abs/2609.04455 https://arxiv.org/pdf/2609.04455 https://arxiv.org/html/2609.04455
Nafez, Poulaei, Feriz, Mousavi, Mahdavi, Mosayebi, Rohban: GhostWord: A Fine-Grained Backdoor Attack on Automatic Speech Recognition https://arxiv.org/abs/2609.04260 https://arxiv.org/pdf/2609.04260 https://arxiv.org/html/2609.04260
Eugenia Kim, Bolor-Erdene Jagdagdorj, Dina Pekelis, Leah Zulas, Amanda Minnich: Probing Warmth-Mediated Harm in Speech-Enabled LLMs for Mental-Health Conversations https://arxiv.org/abs/2609.04256 https://arxiv.org/pdf/2609.04256 https://arxiv.org/html/2609.04256
Yuchen Deng, Chang Sun, Hai-Tao Zheng, Feidiao Yang, Yuxing Han: CAD: Conflict-Aware Decoding to Mitigate Cross-Modal Hallucinations in Omnimodal Large Language Models https://arxiv.org/abs/2609.04247 https://arxiv.org/pdf/2609.04247 https://arxiv.org/html/2609.04247