Adithi Shankar, Gopika Krishnan, Gloria Haro, Xavier Serra, Mart\'in Rocamora: MambaVoice: Lightweight Audiovisual Singing Voice Separation Via A Hybrid Mamba-Transformer Model https://arxiv.org/abs/2609.26635 https://arxiv.org/pdf/2609.26635 https://arxiv.org/html/2609.26635
arXiv cs.SD Sound
@cssd-bot.bsky.social
Unofficial bot by @vele.bsky.social w/ http://github.com/so-okada/bXiv https://arxiv.org/list/cs.SD/new List https://bsky.app/profile/vele.bsky.social/lists/3lim7ccweqo2j ModList https://bsky.app/profile/vele.bsky.social/lists/3lim3qnexsw2g
Yizhong Geng, Yanliang Li, Jinghan Yang, Tianhan Jiang, Yingming Gao, Ya Li: SceneTTS-Bench: A Benchmark for Scene-Level TTS in Drama Dubbing https://arxiv.org/abs/2609.26255 https://arxiv.org/pdf/2609.26255 https://arxiv.org/html/2609.26255
Yeung, He, Chen, Peng, Tao, MC Cheung: Benchmarking Open-Source Speech Emotion Recognition in Naturalistic Mandarin Spine Clinic Consultations: A Pilot Validation Study https://arxiv.org/abs/2609.26054 https://arxiv.org/pdf/2609.26054 https://arxiv.org/html/2609.26054
Hongjin Song, Jiasheng Kuang, Xinyu Yang, Qiuyu Fang, Ziyu Wu, Guowu Tan, Xiang Xie: REVE: Efficient Hallucination Correction for Large Audio-Language Models via Reused Encoder States https://arxiv.org/abs/2609.26028 https://arxiv.org/pdf/2609.26028 https://arxiv.org/html/2609.26028
Jiayi Lu, Yizhong Geng, Jinghan Yang, Tianhan Jiang, Boxun An, Yingming Gao, Ya Li: From Reliable Text to Real Voices: Trust-Aware Progressive Adaptation for Low-Resource TTS https://arxiv.org/abs/2609.25951 https://arxiv.org/pdf/2609.25951 https://arxiv.org/html/2609.25951
Robert Sutherland, Stefan Goetze, Jon Barker: Challenges of Multi-Speaker Extraction for Real Conversational Speech Enhancement https://arxiv.org/abs/2609.25948 https://arxiv.org/pdf/2609.25948 https://arxiv.org/html/2609.25948
Lovro Brulec, Sahil Karawade, Leonard Kinzinger: Latent Audio Watermarking for Robustness to Neural Codec Resynthesis https://arxiv.org/abs/2609.25830 https://arxiv.org/pdf/2609.25830 https://arxiv.org/html/2609.25830
Zhe Ye, Xiangui Kang, Minhua Huang, Kai Wu, Kong Aik Lee, Chng Eng Siong: Boundary and Intra-Segment Learning for Partial Audio Deepfake Localization https://arxiv.org/abs/2609.25822 https://arxiv.org/pdf/2609.25822 https://arxiv.org/html/2609.25822
Annan Wu, Wen-Chin Huang, Tomoki Toda: NeuMark: Neural Codec Resynthesis-Robust Audio Watermarking in the Codec Latent Space https://arxiv.org/abs/2609.25719 https://arxiv.org/pdf/2609.25719 https://arxiv.org/html/2609.25719
Sungkyun Chang, Keshav Bhandari, Simon Dixon, Emmanouil Benetos: Synthesis and editing of multi-instrument audio mixtures using scalar-quantised latents with MIDI Span conditioning https://arxiv.org/abs/2609.25546 https://arxiv.org/pdf/2609.25546 https://arxiv.org/html/2609.25546
Dahong Luo, Anannya Trehan, Aritrik Ghosh, Nirupam Roy: Narrowband Voice Communication Using Streaming Neural Compression https://arxiv.org/abs/2609.25379 https://arxiv.org/pdf/2609.25379 https://arxiv.org/html/2609.25379
[2026-09-23 Wed (UTC), 11 new articles found for csSD Sound]
Jo\~ao Lima, Lucas Ueda, Paula Costa: Automated Assessment of L2 Speech Rhythm Using Low-Frequency Amplitude Modulations https://arxiv.org/abs/2609.24818 https://arxiv.org/pdf/2609.24818 https://arxiv.org/html/2609.24818
Yuancheng Luo: Fast Time-Varying Exponentiated Convolution Methods for Generative Direction Dependent Reverberation https://arxiv.org/abs/2609.24809 https://arxiv.org/pdf/2609.24809 https://arxiv.org/html/2609.24809
Huan Liao, Haonan Han, Xingwen Han, Dekun Chen, Yuancheng Wang, Zhizheng Wu: CycleSpeech: Reciprocal Alignment for Instruction-Controlled Speech Synthesis and Paralinguistic Understanding https://arxiv.org/abs/2609.24771 https://arxiv.org/pdf/2609.24771 https://arxiv.org/html/2609.24771
Mathilde Abrassart, Nicolas Obin, Axel Roebel: Understanding Hyperspherical Geometry of ECAPA-TDNN Embedding and Its Impact on Zero-Shot Voice Conversion https://arxiv.org/abs/2609.24688 https://arxiv.org/pdf/2609.24688 https://arxiv.org/html/2609.24688
Yu Zheng, Jinghan Peng, ChangHao Zhang, Jian Liu, Weiqiang Wang: MECT: Mixture of Experts with CNN-Transformer Network for Speaker verification https://arxiv.org/abs/2609.24061 https://arxiv.org/pdf/2609.24061 https://arxiv.org/html/2609.24061
Bence Mark Halpern, Thomas Tienkamp, Defne Abur, Tomoki Toda: ART-NAD: An Articulatory Inversion-based Neural Acoustic Distance for Pathological Speech Intelligibility Assessment https://arxiv.org/abs/2609.24046 https://arxiv.org/pdf/2609.24046 https://arxiv.org/html/2609.24046
Yue, Xu, Wang, Li, Zhou, Xing, Kong, Han: TTS-Guard: Black-Box Ownership Verification of Text-to-Speech Models via Adaptive Adversarial Speaker-Pair Fingerprints https://arxiv.org/abs/2609.23729 https://arxiv.org/pdf/2609.23729 https://arxiv.org/html/2609.23729
Pierre-Michel Bousquet, Mickael Rouvier: Entropy-aware logistic regression for fusion of large-scale speaker recognition systems https://arxiv.org/abs/2609.23727 https://arxiv.org/pdf/2609.23727 https://arxiv.org/html/2609.23727
Haoyue Liu, Ye Chen, Zhichao Wang, Xiaoyu Ma, Haoran Shou, Xiaoying Tang: Which Constraints Are Missing? Ask the Verifier: Graded Rewards for Constraint-Following Music Generation https://arxiv.org/abs/2609.23665 https://arxiv.org/pdf/2609.23665 https://arxiv.org/html/2609.23665
Jiaheng Dong, Xiaofeng Yu, Jean Honorio, Abhirup Ghosh, Hong Jia, Ting Dang: Listen Then Reason: Perception-Grounded Test-Time Reinforcement Learning for Large Audio-Language Models https://arxiv.org/abs/2609.23589 https://arxiv.org/pdf/2609.23589 https://arxiv.org/html/2609.23589
Heewon Oh: ArtifactBench: Lineage-Aware Evaluation of AI-Generated Music Detectors under Distribution Shift https://arxiv.org/abs/2609.23550 https://arxiv.org/pdf/2609.23550 https://arxiv.org/html/2609.23550
Paul Mo\"ise Gangbadja, Mickael Rouvier, Fabrice Lef\`evre: Beyond Encoder Fusion: Multi-View Discrete Token Augmentation for LLM-Based ASR https://arxiv.org/abs/2609.23525 https://arxiv.org/pdf/2609.23525 https://arxiv.org/html/2609.23525
Yuanxin Guo, Qiang Ji, Mengmei Liu, Yuhan Lv, Ningning Pan, Gongping Huang: LiteCASS: A Lightweight End-to-End Network for Real-Time Stereo Cinematic Audio Source Separation https://arxiv.org/abs/2609.23453 https://arxiv.org/pdf/2609.23453 https://arxiv.org/html/2609.23453
Zeyu Yang, Xinyu Zhang, Zibo Bi, Pei Zhang, Xize Cheng, Jin Xu, Baosong Yang, Satoshi Nakamura: MuLA-Bench: A Multilingual Long-Form Audio Understanding Benchmark via Multi-Tier Auditing https://arxiv.org/abs/2609.23416 https://arxiv.org/pdf/2609.23416 https://arxiv.org/html/2609.23416
Naoya Tomida, Yuki Okamoto, Keisuke Imoto: Misrecognition or Abstraction? Rethinking Outputs of Sound Event Recognition https://arxiv.org/abs/2609.23411 https://arxiv.org/pdf/2609.23411 https://arxiv.org/html/2609.23411
Liu, Wang, Xie, You, Cai, Lin, Chen, Guo, Chu, Yang, Cheng, Xu, Zhong: OmniEcho: Spatial Audio Understanding for Embodied Agents https://arxiv.org/abs/2609.23407 https://arxiv.org/pdf/2609.23407 https://arxiv.org/html/2609.23407
Goksenin Yuksel, Marcel van Gerven, Kiki van der Heijden: EquiSELD: Efficient training of equivariant sound event localization and detection networks https://arxiv.org/abs/2609.23156 https://arxiv.org/pdf/2609.23156 https://arxiv.org/html/2609.23156
Goksenin Yuksel, Marcel van Gerven, Kiki van der Heijden: Bearings: Self-Supervised Soundfield Embeddings from First-Order Ambisonics https://arxiv.org/abs/2609.23152 https://arxiv.org/pdf/2609.23152 https://arxiv.org/html/2609.23152