每日 arXiv

音频 · 语音 · 音乐 · 声学

每日新爬取的 arXiv 论文,按评分排序。中文摘要帮你快速判断是否值得深读。

2026-09-22
日期2026-09-22
已评分—
均分—
最高—

Daily Papers — 2026-09-22

38 papers on audio, speech, music, and acoustics.

1. Spoken Language Models that Think Aloud

Authors: Junyi Ao, Kainan Peng, Mingbo Ma, Shun Zhang, Zhenyu Tang et al.

Categories: cs.CL, cs.SD, eess.AS

While Chain-of-Thought (CoT) reasoning has improved the capability of language models, directly applying it to Spoken Language Models (SLMs) may introduce long silent intervals under the serial “think-then-speak” paradigm, disrupting real-time spoken interaction. To address this issue, we propose an asynchronous think-aloud framework for reasoning-based SLMs within the Thinker-Talker architecture. The framework maintains a primary reasoning stream for logical deduction and a lightweight think-aloud stream that generates short, task-grounded progress utterances conditioned on the user input and the evolving reasoning state. A dynamic balance strategy coordinates the two streams at runtime, triggering additional think-aloud speech to avoid silent gaps and canceling pending utterances when the final response becomes ready. Experiments on spoken reasoning and question-answering benchmarks show that our approach substantially reduces user-audible silence during reasoning while maintaining answer accuracy comparable to that of a serial “think-then-speak” baseline, demonstrating the potential of asynchronous think-aloud for responsive interaction in SLMs.


2. Enriching Speech Emotion Representations with Conversational Context

Authors: Arthur Peuvot, Romaric Besançon, Gaël de Chalendar, Bianca Vieru, Ioana Vasilescu

Categories: cs.CL, eess.AS

Detecting emotions is necessary for building systems that can accurately and adaptively interact with humans. Speech Emotion Recognition (SER) has become an important research focus to develop intelligent spoken interfaces. However, most studies predict emotions at the utterance level, ignoring the conversational context, along with the emotional flow and speaker interactions it carries. In this paper, we introduce ACERT (Averaged Contextual Emotion Representation through Time), a module that integrates a flexible-length window of conversational context to better capture emotional evolution in spoken interactions. To evaluate the robustness of this method, we conducted experiments on datasets spanning diverse emotionally expressive styles and contexts. ACERT outperforms current state-of-the-art (SOTA) approaches on IEMOCAP, establishes the first context-aware benchmark on SAFE, and obtains strong results on MELD for unweighted, class-balanced metrics. Ablation studies show that ACERT’s gains come from emotional and conversational continuity, rather than from speaker identity or acoustic conditions.


3. Transcribe, Translate, and Optimize: Joint Reward Learning for Speech Translation

Authors: Yanghe Dong, Wanting Huang, Weiran Wang

Categories: cs.CL, eess.AS

In LLM-based speech translation, transcription-based chain-of-thought (CoT) suffers from a mismatch between reference transcripts used in supervised fine-tuning (SFT) and model-generated transcripts at inference. To address this, we propose joint recognition and translation fine-tuning via group relative policy optimization (GRPO). We score both transcripts and translations, with translation conditioned on model-generated transcripts, and compare three token advantage strategies. Using Qwen2.5-Omni-3B across four languages, we evaluate CoT against direct speech translation (Direct ST) under SFT and GRPO, training on CoVoST 2 and testing on CoVoST 2 and FLEURS. CoT GRPO outperforms Direct ST GRPO by 1.77 and 0.83 average BLEU points on CoVoST 2 and FLEURS. Compared to CoT SFT, GRPO boosts BLEU by 0.82 and 0.67 points and reduces word error rate (WER) by 8.8% and 7.2% relatively. These results highlight reinforcement fine-tuning as an effective method to mitigate the training-inference mismatch, jointly improving recognition and translation.


4. Sonicmesh: Enhancing 3D Human Mesh Reconstruction in Vision-Impaired Environments With Acoustic Signals

Authors: Xiaoxuan Liang, Hong Zhou, Zhaolong Wei, Yansong Li, Shujian Yu et al.

Categories: cs.CV, cs.SD, eess.AS

3D human mesh reconstruction (HMR) from RGB images often degrades under poor illumination, occlusion, and non-line-of-sight conditions. Acoustic sensing provides complementary spatial cues but suffers from low spatial resolution. We propose SonicMesh, which, to the best of our knowledge, is the first acoustic–visual framework for robust 3D human mesh reconstruction. SonicMesh first converts ultrasonic echoes into range–azimuth acoustic images through an Inverse Synthetic Aperture Radar (ISAR)-based imaging process. It then introduces a cross-dimensional anatomical registration module that maps modality-specific 2D joint features into a common canonical 3D human space. The registered anatomical representations are further integrated with acoustic and visual features through a two-stage fusion network for final mesh reconstruction. Experiments demonstrate that SonicMesh achieves accurate and robust 3D human reconstruction across normal, poor-light, occluded, and non-line-of-sight environments, consistently outperforming existing RGB-, radio-frequency (RF)-, and mmWave-based approaches under challenging sensing conditions.


5. TV-AudioRemover: Joint Text-Visual Guided Sound Removal with Multi-Task Hard-Mixture Curriculum

Authors: Xinyue Guo, Jianxuan Yang, Daiguo Zhou, Jiagao Hu, Yuxuan Chen et al.

Categories: cs.MM, cs.CV, cs.SD

Visual object removal can eliminate a target from video frames, yet its acoustic trace persists in the soundtrack, causing obvious audio-visual inconsistency. Existing video inpainting models operate solely on pixels, while audio editing models, especially for the sound removal task, are typically driven by text and therefore rely on limited single-modal control, which is less effective than multimodal guidance that provides stronger semantic grounding and temporal synchronization cues. In this paper, we present Text-Visual Guided Sound Removal (TV-AudioRemover), a target sound removal framework that leverages the visually edited video together with a natural-language instruction to suppress the sound associated with the removed visual object from the original audio mixture. To acquire high-quality training data, we devise a pipeline to construct a million-scale dataset of single-object audio-visual aligned samples, from which we synthesize mixture-target pairs customized for model training. To effectively leverage visual context and follow instruction intent, we augment the model architecture with task tokens, generalizable instruction modeling, and modality-specific global guidance. We further adopt multi-task training to strengthen task-role comprehension, and employ a hard-mixture curriculum that leverages semantically similar acoustic mixtures during fine-tuning to enhance fine-grained source discrimination. To support evaluation, we present AV-Remove-Bench, a comprehensive audio-visual object removal benchmark, along with dedicated objective metrics and an MLLM-based evaluation protocol. Experiments demonstrate that our method achieves state-of-the-art performance on both subjective and objective metrics. Project page: https://yjx-research.github.io/TV-AudioRemover/.


6. From Reliable Text to Real Voices: Trust-Aware Progressive Adaptation for Low-Resource TTS

Authors: Jiayi Lu, Yizhong Geng, Jinghan Yang, Tianhan Jiang, Boxun An et al.

Categories: cs.SD

Low-resource text-to-speech (TTS) adaptation is constrained by scarce paired data and costly manual transcription. Existing fixed-voice TTS systems can provide relatively accurate pronunciation, but their synthetic speech offers limited speaker diversity and may exhibit flat prosody. Real recordings provide natural prosody and diverse voices, yet their automatic speech recognition (ASR) pseudo-labels may contain transcription errors. We find that supervision order affects content accuracy and speaker similarity. We propose trust-aware progressive adaptation: synthetic-to-real adaptation first establishes text-speech correspondences, then restores reference-speaker control using real speech. Transcript-agreement weighting uses agreement between two fixed ASR systems as a proxy for pseudo-label reliability to limit noisy supervision. Experiments with FireRedTTS3 on Burmese and Lao and OmniVoice on Burmese show improved content accuracy with high naturalness and competitive speaker similarity. Jointly considering supervision order and pseudo-label reliability when combining synthetic and real speech offers a practical path to zero-shot voice cloning in low-resource languages with less manual transcription. Audio demos are available at https://insiderx-pro.github.io/S2R-Adaptation-TTS/


7. Latent Audio Watermarking for Robustness to Neural Codec Resynthesis

Authors: Lovro Brulec, Sahil Karawade, Leonard Kinzinger

Categories: cs.SD

Existing waveform-domain audio watermarks are robust to many conventional distortions but can degrade substantially under neural codec resynthesis. We investigate whether continuous neural codec latents provide a more suitable embedding space using a restricted formulation built around frozen pretrained EnCodec. To test this, a feedforward embedder maps a multi-bit payload to an additive latent perturbation decoded through the unchanged codec decoder. Compared with AudioSeal and WavMark, our latent watermark formulation degrades more gradually under repeated and low-bitrate EnCodec resynthesis, transfers to unseen DAC, and retains high detection under most waveform distortions. Substantial EnCodec robustness emerges even without codec-resynthesis supervision, indicating that this behavior is inherent to our latent formulation and is further strengthened by codec-aware training. Learned perturbations are also preserved more strongly through codec cycling than equal-norm random controls, with preservation depending more on channel-specific allocation than temporal structure. End-to-end perceptual quality remains close to that of the frozen EnCodec reconstruction, indicating that much of the observed degradation originates from the codec carrier itself. Overall, these results show that continuous neural codec latents provide a promising embedding space for watermarks that remain robust to neural codec resynthesis.


8. MambaVoice: Lightweight Audiovisual Singing Voice Separation Via A Hybrid Mamba-Transformer Model

Authors: Adithi Shankar, Gopika Krishnan, Gloria Haro, Xavier Serra, Martín Rocamora

Categories: cs.SD

Isolating a target singing voice from a music video remains challenging, particularly in the presence of multiple vocalists and dense instrumental accompaniment. We propose MambaVoice, a lightweight audiovisual framework that leverages a hybrid Mamba–Transformer architecture for targeted singing voice separation. The model jointly encodes audio and visual streams using an attention-based band-split audio encoder and a spatio-temporal graph convolutional network (ST-GCN) for facial motion features. These modalities are fused through a multiplicative gating mechanism, enabling visual cues to selectively modulate audio representations. The fused features are processed by a hybrid backbone that combines Transformer self-attention with Selective State Space Models (SSMs), achieving efficient long-range temporal modeling with linear complexity. We evaluated MambaVoice on the Acappella and URSing datasets under challenging conditions, including mixtures with interfering singers. At 16.2 million parameters, the model demonstrates comparable performance, achieving 14.18 dB SDR on Acappella and strong cross-dataset performance on URSing, comparable to larger models at a fraction of the parameter count. These findings highlight the effectiveness of hybrid SSM–attention architectures for scalable, efficient audiovisual source separation, suggesting they are well-suited as lightweight components within larger pipelines. We conduct a perceptual study that further supports our improvements in objective metrics. We provide our implementation online.


9. NeuMark: Neural Codec Resynthesis-Robust Audio Watermarking in the Codec Latent Space

Authors: Annan Wu, Wen-Chin Huang, Tomoki Toda

Categories: cs.SD

Audio watermarking is increasingly important for tracing generated speech. Several audio watermarking methods have been proposed to embed the watermark in various domains, such as waveform, timbre feature, or latent representations, for making the embedded watermark robust against traditional digital signal processing (DSP) attacks. On the other hand, modern neural codecs introduce a different threat from DSP attacks: they resynthesize speech through quantized acoustic representations and can remove the embedded watermark evi- dence that is not aligned with codec-preserved structure. In this paper, we propose NeuMark, a codec-latent audio watermarking framework that embeds watermark evidence into SpeechTok- enizer acoustic tokens to address this resynthesis threat. NeuMark uses cross-attention to inject a 16-bit message across residual vector quantization (RVQ) layers, distributing the watermark over codec-aligned latent structure. Experimental results show that NeuMark substantially improves robustness under neural- codec resynthesis while supporting both watermark detection and message recovery. We also analyze the trade-off between reconstruction-referenced transparency and original-referenced robustness.


10. Discrete vs. Continuous: A Comprehensive Study of Unified Audio Understanding in LALMs

Authors: Jing Peng, Zichao Nie, Zhisheng Zhang, Jingran Xie, Zhiyong Wu

Categories: cs.SD, cs.AI

Large Audio Language Models (LALMs) utilize either continuous features or discrete tokens, yet the optimal representation paradigm for general audio understanding remains debated. Existing benchmarks often focus on narrow domains or evaluate encoders outside LALM contexts. To address these gaps, we systematically evaluate continuous and discrete representations across speech, sound and music. Utilizing our UniARC framework with dual evaluation strategies across model scales from SmolLM2-135M to Llama-3-8B, we analyze the dynamic relationships of data volume, model capacity, and computational efficiency. Our results reveal the pivotal role of semantic constraints in tokenization for audio understanding and demonstrate that scaling backbones fail to compensate for information loss in audio representation, especially in data-limited tasks. These findings offer practical guidance for balancing semantic density, fidelity, and efficiency in future LALMs.


11. REVE: Efficient Hallucination Correction for Large Audio-Language Models via Reused Encoder States

Authors: Hongjin Song, Jiasheng Kuang, Xinyu Yang, Qiuyu Fang, Ziyu Wu et al.

Categories: cs.SD, cs.AI

Large audio-language models may mention acoustic events that are absent from the input. A separate audio event detector can verify these mentions, but doing so requires a second audio encoder and a separate forward pass. We propose Reused Encoder States for Verifying Events (REVE), a lightweight method that uses states already computed by the target model. One readout summarizes class scores across audio frames, while another uses pooled states from four consecutive frame intervals. Class-aware score fusion combines their outputs to verify generated event mentions without encoding the audio again. On AudioSet, REVE removes 92.9% of label-unsupported mentions under a faithful-mention recall constraint. With fewer added parameters and no second audio-encoding pass, REVE achieves a reduction comparable to those of CED-Tiny and CED-Base. Its complete verification latency is about 1/18 of the CED-Base path. Results on controlled DESED mixtures and different target-model architectures further confirm the effectiveness of encoder-state reuse.


12. Challenges of Multi-Speaker Extraction for Real Conversational Speech Enhancement

Authors: Robert Sutherland, Stefan Goetze, Jon Barker

Categories: cs.SD, cs.CL

Target-speaker and multi-speaker extraction are techniques for extracting speech from a desired speaker or desired speakers in the presence of other speakers and/or noise. Neural network approaches for this task are often trained and evaluated using simulated datasets, with balanced amounts of target speech and speaker enrolment samples which closely match the target speech. However, in real multi-party conversations, participants are often silent for more time than they are speaking, and their enrolment speech samples can differ substantially from the target speech in the conversation. These factors can impact the training and evaluation of these techniques on recordings of real conversations. This work proposes a new loss function, which helps mitigate the effect of excess silence in training, improving STOI from 0.55 to 0.60, and frequency-weighted segmental SNR from 4.35 to 5.12. Additionally, the impact of the mismatch between the enrolment speech and target speech is explored.


13. Text-only adaptation in LLM-based ASR through text denoising

Authors: Andrés Carofilis, Sergio Burdisso, Esaú Villatoro-Tello, Shashi Kumar, Kadri Hacioglu et al.

Categories: cs.SD, cs.CL, cs.LG, eess.AS

Adapting large language model (LLM)-based automatic speech recognition (ASR) systems to new domains using text-only data is a significant yet underexplored challenge. Standard fine-tuning of the LLM on the target domain text often disrupts the critical alignment between the speech and text modality learned by the projector, degrading performance. We introduce a novel text-only adaptation method that frames this process as a text denoising task. Our approach trains the LLM to recover clean transcripts from noisy inputs. This process effectively adapts the model to a target domain while preserving cross-modal alignment. Our solution is lightweight, requiring no architectural changes or additional parameters. Extensive evaluation on two datasets demonstrates up to 22.1% relative improvement, outperforming recent state-of-the-art text-only adaptation methods.


14. Synthesis and editing of multi-instrument audio mixtures using scalar-quantised latents with MIDI Span conditioning

Authors: Sungkyun Chang, Keshav Bhandari, Simon Dixon, Emmanouil Benetos

Categories: cs.SD, cs.LG, eess.AS

Music creation often involves iterative refinement, changing selected musical details while retaining the rest. To support such refinement, we introduce SpanSynth-Edit, a flow-matching model for MIDI-guided synthesis and editing of multi-instrument audio mixtures using low-frame-rate scalar-quantised latents. MIDI Span encodes instrument-labelled note lifecycles as unordered event sets with continuous-valued attributes and pools each set into one conditioning vector per audio-latent frame. The model uses contextual audio for instrument-specific timbre guidance and supports editing by resynthesising the target region from revised MIDI. Experiments on single- and multi-instrument benchmarks show competitive performance and demonstrate within-frame onset control. We also discuss limitations of transcription-based note-adherence evaluation.


15. A Stem-Agnostic Approach to Hybrid AI Music Detection

Authors: Richa Namballa, François Rigaud, Romain Hennequin

Categories: cs.SD, eess.AS

The inclusion of generative audio in the music production process has led to an increase in hybrid music tracks that blend authentic human performances with AI-generated stems, challenging traditional AI music detectors which operate in a binary setting. In this work, we propose a stem-agnostic framework for identifying synthetic audio sources within hybrid musical mixtures. We introduce the inspectrogram, a novel time-frequency representation that maps localized probabilities of synthetic content across the audio spectrum. By combining the inspectrogram with a Wiener filter estimating target stem energy dominance, a single CNN model evaluates whether the specific stem is generated. Trained on rendered hybrid mixtures and evaluated across various stem classes, our model achieves strong performance on high-frequency sources such as vocals, drums, and guitar, but struggles on the low-frequency, narrow-band bass. We conclude that the quality of separation impacts the detection accuracy and identify source separation as a primary bottleneck and a crucial direction for future research.


16. Boundary and Intra-Segment Learning for Partial Audio Deepfake Localization

Authors: Zhe Ye, Xiangui Kang, Minhua Huang, Kai Wu, Kong Aik Lee et al.

Categories: cs.SD, eess.AS

Partial audio deepfakes manipulate only selected speech regions, making them difficult to be localized. Existing methods exploit boundary cues for partial deepfake localization, but primarily focus on identifying boundary positions rather than modeling the feature changes that characterize authenticity transitions. Meanwhile, the internal characteristics of continuous bona fide and spoofed segments remain underexplored. In this paper, we propose Boundary and Intra-Segment Learning (BISL), which introduces boundary learning to model feature differences between adjacent frames and distinguish authenticity transitions from general acoustic variations. In addition, intra-segment learning captures the overall characteristics of continuous bona fide and spoofed segments while enhancing feature consistency within each segment. By jointly learning frame, boundary, and segment information, BISL enables more effective fine-grained partial audio deepfake localization. Experiments on multiple localization benchmarks show that BISL achieves an EER of 2.52\% and an F1-score of 97.40\% on PartialSpoof, outperforming the compared methods, while maintaining competitive performance on HAD and improved cross-dataset performance on LPS. The code will be made publicly available upon acceptance.


17. A Temporal-Envelope Frontend with Learnable Per-Channel Energy Normalization for Whisper-Based Children’s ASR

Authors: Edem Ahadzi, Ruchi Pandey, Tomi H. Kinnunen

Categories: eess.AS

Temporal envelopes carry cues critical to speech intelligibility, yet ASR frontends based on log-mel spectrograms do not explicitly model continuous sub-band envelope structure. This limitation is particularly acute for children’s speech, where high acoustic variability demands robust feature representations. We propose a modular time-domain frontend that decomposes speech into sub-band envelopes using mel-spaced windowed-sinc filters and the Hilbert transform, with learnable per-channel energy normalization (PCEN) jointly optimized with the Whisper model. On the MyST children’s speech corpus, systematic ablations identify full-band windowed-sinc filters, Hilbert envelopes, a 25 Hz smoothing cutoff, and learnable PCEN as the best configuration. Under the same Whisper-small fine-tuning setup, the frontend reduces WER from 13.16% to 11.08%, a 15.8% relative reduction over the log-mel baseline, and outperforms the evaluated Kid-Whisper checkpoint on the same cleaned test split. These results show that temporal-envelope representations and learnable frontend normalization are effective complements to backend adaptation for children’s ASR.


18. Dual-Microphone Steerable High-Order Neural Differential Beamformer

Authors: Weilong Huang, Emanuël A. P. Habets

Categories: eess.AS

Linear arrays of omnidirectional microphones produce beampatterns that are symmetric about the array axis. For these arrays, a beamformer is considered steerable if its beampattern maintains the same shape in the semicircular plane across all look directions from 0° to 180°. For dual-microphone arrays, conventional differential beamformers are generally non-steerable and restricted to first-order, which significantly limits spatial selectivity. To address these limitations, this study presents a neural differential beamformer (NDBF) with a dual-microphone array. The contributions are as follows: (i) NDBF is steerable; (ii) NDBF achieves high-order frequency-invariant beampatterns; and (iii) NDBF enables stereo recording using only two closely spaced omnidirectional microphones. Experimental results demonstrate that NDBF outperforms existing methods while overcoming the limitations of classical differential beamforming.


19. HRTF Upsampling Across Varying Measurement Configurations with Geometry-Aware Query-Conditioned Aggregation

Authors: Xingyu Chen, Hanwen Bi, Sipei Zhao, Fei Ma, Eva Cheng et al.

Categories: eess.AS

Personalized head-related transfer functions (HRTFs) are essential for spatial audio rendering, but densely measuring an individual’s HRTFs is costly and time-consuming. HRTF upsampling reduces this burden by estimating dense HRTFs from sparse measurements. Recent learning-based methods have achieved promising performance, but many remain tied to predefined measurement configurations. In this work, we propose GeoAtt, a variable-context HRTF upsampling framework that uses a single trained model across varying measurement configurations. GeoAtt performs geometry-aware, query-conditioned spatial aggregation over the available measurements independently at each frequency bin, followed by frequency-domain modeling using Conformer blocks. The relative geometry between the target and measured directions is incorporated as an additive bias in the cross-attention. Experiments on the SONICOM dataset show that a single trained model achieves the lowest log-spectral distortion across all four canonical Listener Acoustic Personalization (LAP) challenge measurement configurations and further generalizes to configurations that are not explicitly included during training.


20. Persistent Delivery Optimization for Streaming Speech-to-Text Translation with Revisions

Authors: Zixiang Wan, Delin Chen, Wei Shi, Haihua Xu, Youxi Xie et al.

Categories: eess.AS

Revision-capable streaming speech-to-text translation (S2TT) can correct earlier drafts, but process rewards based on visible text may credit content later withdrawn. Persistent Delivery Optimization (PDO) assigns intermediate reward only to content that survives revisions while scoring final quality separately. With 7.49 h of task-specific FLEURS adaptation, PDO achieves the best BLEU on four of five directions and higher COMET than every external streaming baseline in all five directions. Relative to its History-SFT initialization, PDO reduces mean/P90 finalization-aware latency by 10.8\%/11.3\% and normalized erasure by 15.8\%, while emitting at the first permitted 2-s update and improving macro BLEU. Zero-shot evaluation on Europarl-ST and CoVoST 2 confirms that these gains are not confined to the FLEURS training domain.


21. Discrete optimal transport is a strong audio adversarial attack

Authors: Anton Selitskiy, Shaikh Akib Shahriyar, Jishnuraj Prakasan

Categories: eess.AS, cs.AI

In this paper, we investigate discrete optimal transport (DOT) as a black-box attack against modern automatic speaker verification (ASV) and anti-spoofing countermeasure (CM) systems. Our attack operates as a post-processing distribution-alignment step. Frame-level WavLM embeddings of generated speech (or another person speech) are aligned to an unpaired bona fide speech pool using entropic optimal transport and a top-k barycentric projection, followed by neural vocoding. Unlike gradient-based attacks, the proposed method requires no access to model parameters, gradients, or training data. Experiments on ASVspoof2019 and ASVspoof5 demonstrate that DOT attack substantially increases CM EER and substantially degrades ASV performance across multiple spoofing attacks. The attack transfers across datasets and remains effective after CM fine-tuning. Analysis using speaker similarity, Fréchet Audio Distance, and visualization of embedding distributions suggests that DOT succeeds by shifting source speech toward bona fide regions of the representation space rather than by maximizing speaker similarity. These results indicate that optimal-transport-based distribution alignment represents a previously underexplored attack vector for contemporary ASV and anti-spoofing systems.


22. Not Quite My Tempo: Voice Activity-aware Speech Synthesis for Lip-Synchronous Dubbing

Authors: Alejandro Pérez-González-de-Martos, Florian Lux, Angelina Elizarova, Milana Shkhanukova, Andreas Kellner et al.

Categories: eess.AS, cs.AI

Automatic lip-synchronous dubbing requires a speech synthesis model to generate alternating voice and silence patterns in the target language that match the timing of the source clip precisely to ensure an optimal viewing experience. Prior works address this problem by conditioning the speech synthesis process on lip movements extracted from the video signal. In this work, we condition the speech generation on a binary voice-activity signal, which has a lightweight representation and can be produced in multiple ways. We show that the model follows the voice-activity signal with high accuracy while maintaining natural prosody and semantically appropriate pause placement within sentences, as demonstrated through extensive objective and subjective evaluations. By randomly masking this condition during training, we make the feature entirely optional during inference, allowing editors to enforce or relax lip-sync constraints when desired.


23. SE-MSB: End-to-End Unpaired Speech Enhancement using Mamba Schrödinger Bridges

Authors: Andreas Bagge, Andreas Nymand, Michael Riis Andersen, Bjørn Sand Jensen

Categories: eess.AS, cs.AI, cs.SD

Speech enhancement (SE) models typically rely on supervised learning with paired data examples where clean speech is synthetically degraded. This paradigm limits performance in real-world scenarios where the target environment’s specific acoustic characteristics are unknown. We propose a fully unpaired SE framework that uses principled Diffusion Schrödinger Bridges (DSB) to learn a stochastic transport process between a clean and a degraded speech distribution. Algorithms for learning transport maps are computationally heavy since they require simulating differential equations during training, usually at each training step. Therefore, we propose using a high-efficiency Mamba Diffusion Model designed for end-to-end waveform processing. We compare against state-of-the-art methods for speech enhancement, both paired and unpaired, as well as a classical signal processing algorithm. Experimental results show that we are on par or better than the baselines while being orders of magnitude faster during inference. Furthermore, we show that the flexibility of the DSB formulation allows our model to generalize across SE tasks, offering a robust and efficient solution for real-world speech restoration.


24. PHONOS: PHOnetic Neutralization for Online Streaming Applications

Authors: Waris Quamer, Mu-Ruei Tseng, Ghady Nasrallah, Ricardo Gutierrez-Osuna

Categories: eess.AS, cs.CL, cs.LG

Speaker anonymization (SA) systems modify timbre while leaving regional or non-native accent cues intact, which is problematic because such cues can reveal a speaker’s first-language or geographic background and narrow the anonymity set. To address this issue, we present PHONOS, a streaming module for real-time SA that performs accent neutralization in a privacy sense: reducing accent-origin cues by converting non-native segmental realizations toward a chosen target accent domain. Our approach pre-generates golden speaker utterances that preserve source timbre and rhythm but replace foreign segmentals with native ones using silence-aware DTW alignment and zero-shot voice conversion. These utterances supervise a causal accent translator that maps non-native content tokens to native equivalents with at most 40ms look-ahead, trained using joint cross-entropy and CTC losses. Our evaluations show an 81% reduction in non-native accent confidence, with listening-test accentedness ratings consistent with this shift. PHONOS also moves outputs away from the original speaker in embedding space, suggesting lower linkability under an embedding-based proxy, while running with $\leq241\,\mathrm{ms}$ end-to-end latency on a single GPU.


25. Attack-Dependent Robustness of Neural Audio Codecs for Adversarial ASR

Authors: Jordan Prescott, Thanathai Lertpetchpun, Shrikanth Narayanan

Categories: eess.AS, cs.SD

Neural audio codecs impose a discrete bottleneck through residual vector quantization (RVQ), making them a useful class of inference-time transformations for reducing adversarial perturbations before ASR inference. We study how codec quantization depth affects defended ASR under non-adaptive, standard adaptive, and quantization-aware adaptive untargeted $\ell_\infty$ attacks. Under non-adaptive attacks, intermediate RVQ depths yield the lowest word error rates and outperform traditional compression at comparable bitrates. However, this apparent optimum is not stable under adaptive evaluation. The standard identity-gradient adaptive baseline (BPDA+EOT) can overestimate robustness, while an implementation of an RVQ-relaxed adaptive attack (SoftVQ-PGD) substantially changes the observed depth trend and largely removes the intermediate-depth advantage. Overall, neural codecs can improve defended ASR under specific threat models. However, the relationship between robustness and RVQ depth depends on the attack used for evaluation, rather than on the codec architecture alone.


26. Brain2Speech-Net: Fast and Intelligible Brain-to-Speech Synthesis Without Text Decoding

Authors: Shreeram Suresh Chandra, Zexin Cai, Yu Tsao, Simon King, Berrak Sisman

Categories: eess.AS, cs.SD

The loss of speech limits communication for individuals with paralysis. Direct neural-to-speech synthesis is challenging due to the limited availability of neural data for training speech brain-computer interfaces. Most existing systems rely on cascaded neural-to-text-to-speech pipelines, which increase inference latency and propagate errors across stages. We present Brain2Speech-Net, a single-stage neural-to-speech generation framework without intermediate text decoding. We use a differentiable phoneme bottleneck and a deep-HMM alignment mechanism to map long neural recordings into the latent space of a text-to-speech (TTS) model, enabling high-quality speech synthesis. Brain2Speech-Net is the only system in our comparison that produces intelligible speech while generating faster than real time.


27. Preserving Speech-to-Text LLM Capabilities in Speech-to-Speech Generation

Authors: Yuxuan Hu, Heng Lu, Ruchao Fan, Yao Qian, Xiaofei Wang et al.

Categories: eess.AS, cs.SD

Strong speech-to-text (S2T) LLMs already provide robust speech perception and text reasoning, but adding speech-to-speech (S2S) output is challenging: fine-tuning the backbone can degrade the original S2T performance, while attaching a downstream talker reintroduces a serial text-to-speech bottleneck. We present PRIME-Speech, a frozen-backbone S2S conversion framework that trains only speech-generation modules. PRIME-Speech synchronizes a causal audio post-decoder with intermediate hidden states of the frozen backbone, so codec tokens are generated from the model’s evolving reasoning trajectory rather than from completed text chunks. The post-decoder uses mixed hidden-state, text, and audio-history conditioning, and a training-time packing strategy with turn-level audio KV-cache and position reset stabilizes multi-turn spoken interaction without additional multi-turn S2S training data. Multi-token prediction further reduces the effective codec prediction rate and improves first-audio latency without modifying the reasoning path. Across speech translation, spoken QA, speech understanding, and multi-turn dialogue, PRIME-Speech preserves the S2T behavior of the frozen backbone while producing accurate, low-WER spoken responses.


28. SRF-SVB: Style-Consistent Singing Voice Beautifying via Rectified Flow

Authors: Wenhui Li, Biao Dong, Liwei Hu, Jiqing Han, Yongjun He

Categories: eess.AS, cs.SD

Singing voice beautifying (SVB) aims to correct pitch and rhythm of amateur singing while enhancing vocal quality, preserving lyrics and the singer’s timbre. Existing methods, however, suffer from limited generation quality and efficiency, and tend to neglect the preservation of the singer’s style. We propose SRF-SVB, a style-consistent model for SVB via rectified flow, which achieves high-fidelity and efficient beautification covering pitch and rhythm correction. Furthermore, we design a context-guided masked mel-spectrogram inpainting mechanism that effectively preserves the amateur singer’s style, including unique timbre and expressive patterns. Experiments on both English and Chinese test sets show that SRF-SVB outperforms baseline models in most objective and subjective metrics.


29. Interactive TTS: Dynamic Speaking Style Adaptation for Expressive Speech Synthesis

Authors: Wenjie Tian, Kangxiang Xia, Jingbin Hu, Xinfa Zhu, HangRui Hu et al.

Categories: eess.AS, cs.SD, eess.IV

Dynamic speaking style adaptation in multi-turn multimodal interaction remains a major challenge for text-to-speech (TTS) systems. Existing context-aware TTS (CTTS) methods typically map dialogue context to speech in an end-to-end manner. Such implicit modeling makes contextual style decisions difficult to supervise, while the entanglement of style, timbre, and content often leads to weak instruction-following and severe timbre drift across turns. To overcome these limitations, we propose Interactive TTS, a dynamic, style-adaptive framework for contextually appropriate and speaker-consistent speech generation. Interactive TTS decouples the process by explicitly modeling contextual style decisions as executable instructions. To bridge the gap between style decisions and speech generation, we introduce Iterative Rejection Sampling Fine-Tuning (Iterative RSFT) and Context-Aware Direct Preference Optimization (CADPO), which significantly enhance instruction-following and align the generated speech with conversational contexts. Extensive experiments demonstrate that Interactive TTS outperforms state-of-the-art models on VStyle and SpeechParaling-Bench. Demo is available at https://wjtian-wonderful.github.io/InteractiveTTS/


30. Same evidence, different judgments: Evidence noncommutative in vision/speech-text conflicts

Authors: Zhuoyun Li, Boxuan Wang, Xiaowei Huang, Yi Dong

Categories: cs.AI

For multimodal large language models, when images or speech conflict with accompanying text, measured text reliance can entangle modality preference with evidence position. Earlier studies of text bias often used a fixed evidence order or moved task instructions with the evidence, leaving the contribution of order unclear. In this paper, we use a paired comparison that keeps the instructions and evidence content fixed and swaps only the positions of the two sources to quantify this potential influence. Across vision and speech models, placing an image or recording after conflicting text consistently shifts answers toward its content. We also revisit previous studies and analyze why their experimental settings can lead to misleading conclusions. These findings reveal cross-modal evidence noncommutativity: the same evidence can lead to different judgments when its order changes, and placing perceptual evidence later can increase the model’s reliance on its content.


31. Training Intelligent Voice Assistant Wakeup with Controllable Synthetic Conversations

Authors: Marcin Sowański, Kacper Leszczyński, Kacper Krzywicki, Krzysztof Wodnicki

Categories: cs.AI

Wake word detection is a critical component of virtual assistants, serving as the gateway to seamless user interactions. This paper introduces a novel wake-up system that extends traditional direct keyword detection with contextual trigger detection. After an initial wake word activation, the system uses reasoning to distinguish between user commands and unrelated speech, ensuring efficient and context-aware engagement. We present a data generation architecture that produces a 62.3-hour corpus of controllable multi-speaker conversations containing direct invocations, contextual follow-ups, and non-addressed speech. Experimental results demonstrate the effectiveness of the proposed approach across diverse synthetic conversational scenarios. We release the code, dataset and trained models to promote reproducibility and further advancements in intelligent assistant technologies.


32. Lost in Speech: Trilingual Spoken Hallucination Detection Across Audio and Transcripts

Authors: Meruyert Aristombayeva, Jason S. Lucas, Chaewan Chun, Dongwon Lee

Categories: cs.CL

While text-based hallucination detection is well studied, reference-free detection of factual alterations in speech remains underexplored, especially for low-resource languages. Our spoken benchmark comprises 12,013 English, Russian, and Kazakh news samples with three synthetic alteration types and three severity levels, pairing source articles with rewrites as text, synthesized audio, and ASR transcripts. We add 290 fact-checked misinformation items collected in Russian (225) and Kazakh (65), translated into the other language and rendered through the same TTS-ASR pipeline. We evaluate fine-tuned multilingual encoders and zero-shot multimodal decoders on text, transcripts, and audio. Detectors receive only target inputs without source articles or external evidence; the task evaluates reference-free classification rather than evidence-grounded verification. Encoder degradation from source text to transcripts generally tracks per-language ASR error on the binary task. Among decoders, only Gemma-3n exceeds the binary majority-class baseline in macro-F1, on transcripts only; the other four fall below their respective baselines. Comparisons between audio and transcripts are confounded by differences in evaluation coverage and class balance. Synthetic-trained detectors achieve 0.82–0.88 macro-F1 on real-world misinformation source text; Russian provenance analysis reveals veracity-related and model-dependent machine-style signals, a key confound in synthetic hallucination benchmarks.


33. NADI 2026: The Second Multidialectal Arabic Speech Processing Shared Task

Authors: Peter Sullivan, Bashar Talafha, Ahmed Ashraf, Fethi Bougares, Haroun Elleuch et al.

Categories: cs.CL

NADI 2026 is the seventh edition of the Nuanced Arabic Dialect Identification (NADI) shared task series and the second dedicated to multidialectal Arabic speech processing. This edition comprises five tasks and eight subtasks spanning Automatic Speech Recognition (ASR), Spoken Dialect Identification (SDID), Text-to-Speech (TTS), Spoken Language Translation (SLT), and Spoken Language Understanding (SLU). NADI 2026 emphasizes realistic evaluation through low-bandwidth, mixed-dialect, code-switched, out-of-domain, and zero-shot settings, while introducing TTS, SLT, and SLU to the series for the first time. The shared task attracted 21 participating teams from at least 13 countries, with 48 test-phase submissions and 14 submitted system-description papers. Results show that out-of-domain generalization remains a major bottleneck and highlight the effectiveness of recent Arabic-specialized speech models, multimodal dialect identification approaches, and ensemble methods. Overall, NADI 2026 provides a broader and more challenging benchmark for robust Arabic dialect speech processing.


34. Rethinking Length-Based Training: Batch Composition and Loss Normalization in Speech Token Language Models

Authors: Hongjin Song, Runwu Shi, Weiqiao Shan, Jiale Luo, Yujin Wang et al.

Categories: cs.CL

Short-to-long training is a simple curriculum for speech models, but its gains can be difficult to interpret. In speech token language models, length-based training can change the shuffle policy, batch composition, token retention, and token weights under batch-mean loss. We disentangle these factors through matched comparisons. In the tested settings, short-to-long ordering shows no independent benefit when batch composition and token exposure are fixed. First-epoch grouping lowers perplexity for Mimi under batch-mean loss, but this gain is not observed under token-balanced loss. The cross-tokenizer results are consistent with a link between chunk-length variation and token weighting. This work provides a systematic analysis protocol for studying length-based training in variable-length speech models.


35. When Helpful Context Leaks: Privacy Risks in Domain-Adapted ASR

Authors: Maike Züfle, Jan Niehues

Categories: cs.CL

SpeechLLMs are increasingly deployed in professional settings where domain customisation is standard practice: users supply context in prompts with sensitive information, fine-tune on proprietary recordings, or both. We identify and systematically investigate an overlooked privacy risk of such customisation: a model adapted to recognise domain-specific terminology can be nudged into transcribing a phonetically similar word from its context or training data, even when a different word is spoken, thereby leaking private information. To evaluate this risk, we propose a technique to automatically construct benchmarks of such attacks and apply it to measure leakage rates across two customisation mechanisms, prompting and fine-tuning. Both mechanisms cause measurable leakage, compounding when combined. We evaluate a prompt-level mitigation strategy and analyse the accuracy-leakage trade-off across customisation approaches, finding that fine-tuning without context prompts offers the best balance.


36. Compositional Failure in Audio-Visual LLMs: Late-Layer Prior Dominance Under Cross-modal Conflict

Authors: Adarsh Sudheer, David Li, Omar Elbanna, Ishaan Kodarapu, Arjun Bahuguna et al.

Categories: cs.CL, cs.AI

We study audio-visual conflict as a compositional generalization test for AV-LLMs: the model must combine synchronized but semantically incompatible audio and video evidence and decide whether the pair matches. On VideoLLaMA 2-7B-AV, three alignment configurations remain nearchance on the scored exact-string Yes/No subset of AVHBench, even though their output priors shift substantially. Similarly, off-the-shelf InternVideo2 experienced a 32.3% accuracy decrease specifically under cross-modal conflict, accompanied by a 17.3% instruction-following failure. We call this failure mode prior dominance: late-layer commitment to an internally preferred answer pattern that is weakly grounded in the conflicting inputs. To explain this behavior, we conduct a mechanistic interpretability analysis and find that commitment remains concentrated at 25.5 $\pm$ 1 layers. We show that stronger temporal alignment changes answer bias, but do not improve compositional conflict resolution. Code and data to reproduce our mechanistic audit and behavioral evaluations are available at https://github.com/AdarshSudheer09/AVHBench-dmai.


37. Context-Aware Multimodal Claim Verification in Spoken Dialogues

Authors: Chaewan Chun, Delvin Ce Zhang, Dongwon Lee

Categories: cs.CL, cs.SI

Spoken factual claims often occur within multi-turn conversations, where surrounding dialogue can provide context unavailable from the claim alone. Yet most fact-checking research evaluates isolated text, leaving conversational audio under-studied. We introduce MAD2, a synthetic Multi-turn Audio Dialogues benchmark for spoken claim verification with 1,000 two-speaker dialogues, 1,230 sentence-level check-worthy candidate annotations, and approximately 10 hours of audio. We also propose calibrated multimodal fusion of a context-aware audio encoder and a dialogue-aware text model. Adding dialogue context improves verification across settings, although the gains vary by scenario. Past-only context often approaches local offline performance, suggesting its usefulness when future context is unavailable. Fusion achieves its highest mean advantage over text with full-dialogue context, but does not consistently or significantly outperform text across settings. Under full-dialogue context, exploratory subgroup results show greater performance variation across dialogue scenarios for text and fusion, but across spread styles for audio.


38. SsgCaps: A controlled dataset for the evaluation of sound scene generation algorithms

Authors: Modan Tailleur, Junwon Lee, Laurie M Heller, Mathieu Lagrange, Keunwoo Choi et al.

Categories: cs.GR, cs.AI

Sound Scene Generation is about the automatic synthesis of artificial sound scenes. We introduce SsgCaps, a publicly available dataset of human-engineered sound scenes wherein each scene matches a precisely structured prompt that guides the sampling process. The corresponding prompts are sampled from a predefined action-based typology that allows extensive sampling while retaining plausibility. SsgCaps is a sound scene dataset derived from the unpublished reference dataset for Task 7 of the 2024 DCASE Challenge edition, which contained private-and public-domain audio samples. In contrast, SsgCaps contains only public-domain audio samples, allowing us to open this dataset to the community. To make this dataset useful to the community, we first elaborate on the rationale for the prompt and dataset structure. We then perform a comparative quantitative analysis of the 2 versions of the dataset. To do so, we compare both versions to the audio synthesized by the SSG algorithms submitted to the challenge using Fr{é}chet Audio Distance (FAD) and Kernel Audio Distance (KAD) as well as perceptual ratings. This analysis shows only small differences, which enables us to recommend the open version for further benchmarking of SSG algorithms.