每日 arXiv

音频 · 语音 · 音乐 · 声学

每日新爬取的 arXiv 论文,按评分排序。中文摘要帮你快速判断是否值得深读。

2026-09-20
日期2026-09-20
已评分—
均分—
最高—

Daily Papers — 2026-09-20

33 papers on audio, speech, music, and acoustics.

1. HaikuS2S: A Cascaded System For Responding In Verse

Authors: Devangi Sharma, Sophia Judicke, Glenda Tan, Conrad Schaumburg, Shinji Watanabe

Categories: cs.CL, cs.AI, cs.SD

Expressive speech synthesis has advanced through prosody modeling, yet generating structured poetic speech, such as haiku, remains challenging. Prior work on prosody transfer improves expressiveness, and fine-tuned poetry TTS (text-to-speech) systems capture verse intonation. However, these models do not model haiku’s 5-7-5 syllable structure or line-ending pauses. We present a cascaded system, HaikuS2S, combining ASR (automatic speech recognition), LLM (large language model)-generated haiku, and TTS fine-tuning on both prose and custom haiku datasets. Our evaluation focuses on emotion similarity, speech quality, and prosody alignment. In our experiments, we see that our prosody and tonal alignment improve significantly with our fine-tuned systems, particularly the one trained on both general poetry and haiku. We also see that we maintain similar emotion similarity scores across all systems.


2. Rethinking Speech-LLM Integration for ASR: Effective Joint Speech-Text Training by Interleaving

Authors: Ruchao Fan, Yiming Wang, Rui Zhao, Liliang Ren, Keqi Deng et al.

Categories: cs.CL, eess.AS

Speech-LLM integration has shown promising results by leveraging extensive textual pretraining, yet its specific benefits for automatic speech recognition (ASR) remain unclear. We observe that as supervised ASR training data increases, the contribution of LLM priors becomes less evident, and simple speech-text joint training under-utilizes textual knowledge. We therefore propose Joint Speech-Text Interleaved Pretraining (JSTIP), an ASR-oriented pretraining strategy that constructs word-level and segment-level interleaved speech-text sequences within aligned pairs for speech-LLM architectures that accept continuous inputs. Experiments on 38k hours of ASR data show consistent entity accuracy improvement compared to ASR-only and joint speech-text training baselines. JSTIP achieves on-par entity recognition performance using domain transcription text compared to synthetic speech-text pairs, simplifying domain adaptation. Benefiting from textual pretraining and domain text data, JSTIP is competitive with open-source ASR and Speech-LLM systems in medical entity recognition. The zero-shot speech question answering behaviors further suggest that interleaving reduces the speech-text modality gap and preserves the LLM generative prior, which is likely the reason for the entity improvements on the ASR task.


3. DFD-Lab: A Modular Audio-Visual Deepfake Detection Pipeline

Authors: Jan Rybarczyk, Mateusz Roszkowski, Jacek Komorowski

Categories: cs.CV, cs.SD

Comparing audio-visual deepfake detectors requires coordinating dataset adaptation, temporal input representation, model interfaces and experimental conditions. We present DFD-Lab, a modular pipeline that separates these responsibilities while supporting shared training and evaluation workflows. We integrate three implementations: Xception-based maximum-logit fusion, ResNet with temporal LSTM fusion, and our AVFF reimplementation. Experiments cover external testing, degradation-based training augmentation and evaluation-time corruption. On a filtered subset of Deepfake-Eval-2024, models trained on FakeAVCeleb attain baseline AUROC values of 0.504, 0.538 and 0.458. JPEG50 training augmentation raises these to 0.691, 0.605 and 0.570, respectively, while all three accuracies decrease. These results illustrate why training interventions, evaluation corruptions and metric-dependent outcomes should remain distinct within a common pipeline. The contribution is the integration of audio-visual processing, interchangeable detectors and configurable experimental workflows, supported by empirical case studies. The findings highlight the challenge of cross-dataset detection and the complementary information provided by ranking and classification metrics.


4. MoSAT: Human Motion Generation from Spatial Audio and Textual Description

Authors: Shuyang Xu, Zhiyang Dou, Yiduo Hao, Zekun Li, Liang Pan et al.

Categories: cs.GR, cs.CV, cs.RO, cs.SD

Human motion is shaped by both external acoustic events and behavioral intent: spatial audio conveys environmental cues that elicit or guide a response, while text specifies the desired action and how it should be performed. In this paper, we study the novel task of human motion synthesis jointly conditioned on spatial audio and natural language, a problem that has been largely overlooked in previous research. To support this task, We introduce STAM, a dataset of motion sequences paired with spatial audio and detailed textual annotations whose rich vocabulary affords precise and nuanced specification of human motions. We further introduce MoSAT, a latent flow-matching framework for full-body motion generation jointly conditioned on natural-language intent and directional spatial-audio cues through hierarchical cross-attention before generating motion. Such a hierarchical design enhances temporally coherent and semantically aligned motion sequences. We also develop tri-modal evaluators for comprehensive evaluation on this novel task. Extensive experiments show that MoSAT achieves the SOTA performance by leveraging spatial audio’s intrinsic motion-shaping properties alongside textual semantics, enabling precise and diverse motion in various scenarios.


5. If You Hear It, Help Find It: Orthogonal Knowledge Distillation for Open-Vocabulary Audio-Visual Event Localization

Authors: Yi Xu, Cheng Chen, Wenzhuo Lei

Categories: cs.MM, cs.CV, cs.LG, cs.SD

Open-vocabulary audio-visual event localization (OV-AVEL) grounds a text-queried event in time from video, audio, and language. The supervision sources available to this task can differ in temporal-boundary reliability: on OV-AVEBench, our configured visual teacher gives more reliable boundary cues than the configured audio teacher, although the latter is a strong pretrained audio model and remains semantically informative. This is a setting-specific diagnostic rather than a universal ranking of vision and audio. We formulate the resulting challenge as supervision placement: which teacher signals may shape the localization decision, and which should remain auxiliary. Based on this view, we propose OV-OrthKD, a reliability-aware asymmetric distillation framework. Visual feature transfer shapes a decision-aligned representation, audio feature transfer enriches a complementary auxiliary subspace, a text prototype anchors seen/unseen category semantics, and an orthogonality loss limits directional overlap between the two teacher-specific projections. The student continues to use both modalities through query-aware fusion at inference, while the default training recipe keeps audio-teacher supervision off the segment-logit path. On OV-AVEBench, OV-OrthKD achieves 0.816 segment AP and improves F1@0.5 over the official fine-tuning baseline by 2.7 points overall and 3.4 points on unseen categories. Path-assignment, role-swap, corruption, and transfer analyses consistently support supervision placement as a task-specific design axis for OV-AVEL.


6. Amanous: Distribution-Switching for Superhuman Piano Density on Disklavier

Authors: Joonhyung Bae

Categories: cs.MM, cs.SD, eess.AS

A player piano can strike more keys, across more of the register, and faster than any pianist can reach. Three traditions dominate composition in that region, namely Nancarrow’s tempo canons, Xenakis’ stochastic distributions, and L-system grammars. They have developed in isolation, and none of them accounts for the instrument itself. A Disklavier does not answer instantly, because a loud note reaches the string sooner than a soft one, so music written as if the mechanism were transparent arrives distorted. We present Amanous, a hardware-aware composition system for the Yamaha Disklavier that unifies the three traditions through distribution-switching, in which each grammar symbol selects an entire distributional regime rather than adjusting parameters within a fixed one. A four-layer pipeline carries a symbol from grammar to actuation-ready MIDI. An L-system fixes the macro-form, each symbol is mapped to distributions and a tempo-canon ratio, events are sampled and time-scaled, and a hardware layer pre-compensates velocity-dependent latency and enforces the key-reset time. Convergence points close a feedback loop, letting the music’s own temporal structure trigger the next switch. Sections from different symbols remain statistically separable at the output, and ablating the L-system, the tempo canon, or the hardware compensation each degrades a distinct property. Under the modelled latency curve, pre-compensation takes the mean onset error from roughly 18 ms to below the instrument’s 1 ms scanning resolution. Structured and random textures are most separable near 25 notes/s, and melodic metrics lose their discriminative power over the 40-100 notes/s band, above which the difference is consistent with distributional content rather than melodic order. All results are computational, nothing was played or measured on an instrument, and a psychoacoustic protocol is proposed for future work.


7. Beyond Encoder Fusion: Multi-View Discrete Token Augmentation for LLM-Based ASR

Authors: Paul Moïse Gangbadja, Mickael Rouvier, Fabrice Lefèvre

Categories: cs.SD

Discrete speech tokens provide a compact interface between speech encoders and large language models for automatic speech recognition, but single-tokenization systems remain sensitive to the chosen encoder. We propose multi-view discrete token augmentation, a simple strategy that augments each training utterance by generating alternative token sequences from fixed SSL encoders, such as HuBERT, WavLM, and MMS-300M. These tokenizations are treated as complementary training views for a shared LLM decoder, exposing it to more diverse discrete speech representations without requiring multi-encoder inference. At test time, the model can operate with a single encoder. On LibriSpeech, the approach consistently improves all encoders over independently trained baselines, with WavLM reaching 3.30% WER on test-clean and 8.13% on test-other. Budget-matched controls show that the gains come from encoder diversity rather than data volume. ROVER over multi-view hypotheses further improves WER to 3.03% and 7.38%.


8. Entropy-aware logistic regression for fusion of large-scale speaker recognition systems

Authors: Pierre-Michel Bousquet, Mickael Rouvier

Categories: cs.SD

Score-level fusion based on logistic regression is widely used in speaker recognition to combine complementary systems. However, conventional approaches assign fixed system-dependent coefficients and do not explicitly account for variations in the reliability of individual enrollment and test utterances. Drawing on recent research on the entropy of deep learning-based speaker recognition models, this study incorporates an uncertainty component into the fusion process. By exploiting both system-level complementarity and utterance-dependent uncertainty, the method achieves robust performance in large-scale speaker recognition tasks that involve highly variable characteristics of the speech signal. These results demonstrate that model-entropy information provides a valuable complementary cue in large-scale scenarios.


9. LiteCASS: A Lightweight End-to-End Network for Real-Time Stereo Cinematic Audio Source Separation

Authors: Yuanxin Guo, Qiang Ji, Mengmei Liu, Yuhan Lv, Ningning Pan et al.

Categories: cs.SD

Cinematic audio source separation (CASS) decomposes a soundtrack into dialogue, music, and sound-effects (SFX) stems. Existing CASS methods, however, suffer from two critical limitations: they rely on heavily parameterized network architectures and GPU-class hardware, limiting their use in real-time and resource-constrained scenarios, and they are overwhelmingly designed for monaural signals, leaving the stereo scenario largely unexplored. We present LiteCASS, to our knowledge the first lightweight end-to-end network for real-time stereo CASS. LiteCASS combines deterministic STFT subband rearrangement with two jointly trained compact U-Nets: the first extracts dialogue, and the second separates music and SFX from the predicted non-speech component. A multi-task waveform-domain L1 loss supervises all stems. On a spatialized stereo extension of DnR v3, LiteCASS-K8 uses only 1.06M parameters and 0.72G MACs per second, while achieving the highest averaged SI-SDR among the compared CASS baselines.


10. TTS-Guard: Black-Box Ownership Verification of Text-to-Speech Models via Adaptive Adversarial Speaker-Pair Fingerprints

Authors: Xubin Yue, Zhenhua Xu, Zhebo Wang, Mengting Li, Zijie Zhou et al.

Categories: cs.SD

The rapid maturation of zero-shot Text-to-Speech (TTS) models has turned high-quality voice cloning into a widely available capability, raising acute concerns over unauthorised replication, fine-tuning and resale of proprietary speech models. Yet ownership verification for TTS remains largely open: speech is a continuous waveform whose perturbations are easily destroyed by routine signal processing, and the human auditory system imposes a much tighter perceptual budget than vision. We present \textbf{TTS-Guard}, a black-box ownership verification framework for TTS models built on \emph{adversarial speaker-pair fingerprints}. TTS-Guard(i) selects key speaker pairs in a \emph{dual} embedding space for architecture-agnostic stealth;(ii) optimises a perturbation through an \emph{adaptive curriculum} of shadow models covering fine-tuning, pruning, quantisation and distillation; and (iii) aggregates black-box queries into a calibrated \emph{Verification Confidence Score}. On five mainstream TTS systems, TTS-Guard reaches an average Fingerprint Success Rate of $96.4\%$ at a False Positive Rate of $5.8\%$, while preserving intelligibility and naturalness. The fingerprint remains effective against ten audio attacks, six model modifications, and two state-of-the-art adversarial purifiers.


11. Which Constraints Are Missing? Ask the Verifier: Graded Rewards for Constraint-Following Music Generation

Authors: Haoyue Liu, Ye Chen, Zhichao Wang, Xiaoyu Ma, Haoran Shou et al.

Categories: cs.SD, cs.AI

Constraint-following music generation asks a score to satisfy several user-specified properties at once, each checkable programmatically (key, meter, length, range, final note, rhythm, motion and form), yet no existing benchmark isolates this capability. We construct MusicConstraintBench, 2,180 items over eight constraint families, on which current models fail once a few constraints are combined. The natural remedy is reinforcement learning with these verifiers as reward, yet we observe that a reward paid only when every property holds leaves most training groups without a learning signal: over the first 50 updates, 0.550 of rollout groups score identically and receive no gradient, even though a failing score typically misses only one requested property. Under the joint criterion, rollouts for a prompt tend to fail together, so a binary reward cannot separate a nearly correct score from a malformed one. We therefore introduce MusicRLVR, which pays graded per-property credit behind a hard validation gate that rejects malformed outputs, plus a joint-satisfaction bonus, requiring no human annotation, learned reward model, or music-domain fine-tuning. On MusicConstraintBench, MusicRLVR lifts Qwen3-4B-Instruct from 0.160 to 0.807 on mixed constraints and leads every zero-shot baseline including Llama-3.1-70B at 0.380. It also generalises to property combinations unseen in training and to out-of-range parameter values, showing that verifiable rewards need not presuppose a target output.


12. Listen Then Reason: Perception-Grounded Test-Time Reinforcement Learning for Large Audio-Language Models

Authors: Jiaheng Dong, Xiaofeng Yu, Jean Honorio, Abhirup Ghosh, Hong Jia et al.

Categories: cs.SD, cs.AI, eess.AS

Large audio-language models (LALMs) are increasingly used for a broader range of audio reasoning tasks. These models typically incorporate audio representations into a large language model (LLM) backbone to enable multimodal reasoning. Recent test-time reinforcement learning (TTRL) methods further improve LLM reasoning capability by leveraging unlabelled test data after pre-training. However, the importance of the perceptual capability of LALMs remains underexplored, particularly how much acoustic evidence is integrated and relied upon during reasoning, and how this contributes to final task performance. This gap limits the development of effective post-training methods like TTRL for audio reasoning. In this work, we first analyse how audio information is integrated and utilised during reasoning process. We quantify layer-wise perceptual reliance and show that stronger acoustic reliance is associated with higher accuracy and a larger performance gain attributable to the audio input. Building on this, we propose Perception-Grounded TTRL (PG-TTRL), which aligns label-free test-time optimisation with perceptually grounded reasoning, encouraging the model to structure its reasoning more strongly on the audio input. Experiments across LALMs and benchmarks show that PG-TTRL consistently improves reasoning performance over both the base models and standard TTRL, showing the value of perceptual-grounding optimisation for test-time audio reasoning.


13. MuLA-Bench: A Multilingual Long-Form Audio Understanding Benchmark via Multi-Tier Auditing

Authors: Zeyu Yang, Xinyu Zhang, Zibo Bi, Pei Zhang, Xize Cheng et al.

Categories: cs.SD, cs.CL

Long-form audio performance is often summarized by context length and aggregate accuracy, obscuring how language, evidence, and task jointly shape difficulty. We introduce MuLA-Bench: 5,038 open-ended questions over 1,769 in-the-wild recordings totaling 1,377.9 hours, covering 16 languages and eight domains. A balanced Language x Domain semantic track supports controlled comparisons, while a complementary acoustic track preserves naturally occurring non-speech evidence. Evidence-grounded generation, shortcut checks, and language-expert review provide auditable questions without translating a shared source set or injecting target sounds. We evaluate ten audio-language models and conduct pooled diagnostics on a fixed eight-model cohort. Language rankings change across domains and tasks; acoustic-semantic performance gaps vary with the requested operation; and temporal errors can persist after the correct event is identified. Long-range retrieval is comparatively strong, while precise clock alignment and factual grounding of natural acoustic events remain fragile. MuLA-Bench thus exposes conditional failure patterns that a single long-context score does not capture.


14. VoxReason: Auditing Source-Grounded Speech Plans Before Synthesis

Authors: Mengzhe Geng

Categories: cs.SD, cs.CL, cs.LG, eess.AS

Speech systems increasingly infer how an utterance should be delivered from context, but a plausible delivery plan may not be supported by the input. VoxReason is a small public benchmark and verifier for testing this failure before waveform synthesis. Each of its 100 cases fixes the utterance, provides derived records that name the source emotion and intensity, and changes one licensed cue. A system must cite the record for its delivery decision and update only the plan fields associated with the edit. We use deterministic verifier references, not a model leaderboard. This holdout excludes every emotion and intensity combination observed in training. A prior-only predictor achieves 0.958 accuracy across plan fields but never changes its plan consistently after a cue edit. This contrast shows that plan agreement does not demonstrate source grounding. The released suite provides an auditable measurement layer for screening structured speech plans before synthesis. It is limited to derived records, not audio inputs.


15. From Scores to Evidence: Auditable Decisions Can Improve Speech Deepfake Detection

Authors: Mengzhe Geng, Yujia Lu, Manuela Kunz, Patrick Littell

Categories: cs.SD, cs.CL, eess.AS

Speech deepfake detectors usually emit one score per utterance, but a borderline score does not reveal why two examples differ during retrospective error analysis. We ask whether a final score can be calibrated from component fields while keeping those fields visible for inspection. We build a decision record with a passive detector score and a score from a probe applied to a marked copy. It also includes retrieval support held out of the evaluated family, a margin from a support-set profile, and raw neighbor closeness. A cross-fit calibrator combines these fields and two differences between raw scores into one final score. On matched ASVspoof development data, the calibrated record reduces equal error rate (EER) by 3.48 percentage points relative to the fixed retrieval-augmented rule. It reaches 8.43% EER, whereas a passive WavLM baseline reaches 6.71% on the same subset. The record is therefore not the strongest detector in this comparison. Its value is to retain inspectable component fields while producing one scalar score for retrospective diagnosis.


16. Text Scores Can Miss Waveform Use: A Qwen2-Audio Quantization Case Study

Authors: Mengzhe Geng, Jinxi Jin, Junhao Xu

Categories: cs.SD, cs.CL, eess.AS

Post-training quantization of speech language models is often summarized with text-output scores and nominal bit widths. Those numbers alone do not establish behavior that depends on information missing from a transcript, or efficiency for a particular runtime. We introduce an evaluation protocol that separately tests lexical output, a transcript-insufficient endpoint, and a measured packed implementation. In a Qwen2-Audio case study, a translation-selected 6-bit allocation improves chrF by 2.36 on a frozen English-to-German replay, with paired 95% bootstrap interval [1.04, 3.62], but loses 3.91 percentage points on speaker-disjoint emotion recognition. At the same 6-bit budget, the uniform structural control reaches higher emotion accuracy than the selected allocation, and the front-layer control is also higher by point estimate on the same frozen set. At 7 bits, chrF improves by 3.28 with interval [2.08, 4.59], the emotion interval against FP16 includes zero, and a same-budget front-layer control still exceeds the selected allocation. A separate matched-budget 4.08-bit study finds roughly 10-point emotion deficits for every tested low-bit allocation and no selected-allocation advantage over frozen controls. Finally, a dequantized average-6-bit simulation retains the FP16 peak memory. This case study identifies a precision-dependent mismatch between lexical output, waveform-dependent behavior, and nominal precision. It does not establish a general failure of low-bit speech models or a deployment benefit for the selected allocation.


17. ArtifactBench: Lineage-Aware Evaluation of AI-Generated Music Detectors under Distribution Shift

Authors: Heewon Oh

Categories: cs.SD, eess.AS

AI-generated music detectors are commonly compared using aggregate scores on benchmarks whose training overlap, generator lineage, source provenance, and audio-transformation history are only partially observable. This paper introduces ArtifactBench, a lineage-aware evaluation suite for measuring detector behavior across generator families and versions, real-music domains, collection-cohort shift, and inference coverage. The benchmark groups source recordings and their derived variants by content identity, separates calibration from final testing, records inference failures independently from classification errors, and reports source-level performance with uncertainty in addition to aggregate metrics. We evaluate multiple publicly available detectors under a version-pinned common protocol and examine how leakage control, cohort availability, threshold policy, and model-specific missingness alter measured performance and model ranking. On the 562-track common-success test intersection, ArtifactNet obtains 0.982 AUROC and 0.918 balanced accuracy, compared with 0.761/0.776 for the public Deezer detector; SpecTTTra and CLAM fall below 0.30 AUROC under this shifted cohort. These results also expose substantial generator- and real-domain shifts that aggregate scores alone conceal.


18. ArtifactNet: Detecting AI-Generated Music via Forensic Residual Physics

Authors: Heewon Oh

Categories: cs.SD, eess.AS

Detecting artificially generated music requires distinguishing synthesis-related traces from musical content and artifacts introduced by audio distribution. We present ArtifactNet, a compact framework based on learned forensic residuals. A pretrained music source separator supplies residual targets during training, but the final system replaces it with ArtifactUNet, a task-directed bounded-mask extractor. Seven channels describe harmonic and percussive residual structure, their balance, and temporal variation, and a lightweight convolutional classifier produces recording scores. The extractor and classifier contain 4.03 million parameters. We evaluate the retained checkpoints on a frozen, recording-level test containing 534 generated and 479 real recordings after conservative source-family and identity exclusions from the reconstructed training lineage. At the pre-established threshold of 0.225, ArtifactNet obtains F1=0.9671, 93.63% recall, and no observed false positive; AUROC and average precision are 0.9982 and 0.9988. Released SpecTTTra and CLAM checkpoints are evaluated on the same files under their model-specific frontends. A diagnostic further shows that the released CLAM head changes substantially with execution batch composition. In a separate paired four-codec study on 100 real and 100 generated recordings, codec-aware extraction reduces the mean score range from 0.276 to 0.033 for real music but increases it from 0.035 to 0.249 for generated music; Opus recall falls from 96% to 74%. The results support compact residual detection while exposing class-dependent codec and implementation sensitivities.


19. Misrecognition or Abstraction? Rethinking Outputs of Sound Event Recognition

Authors: Naoya Tomida, Yuki Okamoto, Keisuke Imoto

Categories: cs.SD, eess.AS

Conventional general sound recognition systems typically output deterministic sound event labels, implicitly assuming that the target sound class can be correctly identified from the input audio. However, in real listening situations, the sound event class is not always clearly identifiable. Human listeners may nevertheless understand their surroundings from an ambiguous sound without identifying its exact sound event class. This motivates a discussion of how the outputs of sound recognition systems should be redesigned under such uncertainty. As a basis for this discussion, this paper proposes an output representation for sound event recognition that combines a sound event class, its confidence score, and an onomatopoeic description of the sound. The proposed representation preserves conventional class-based recognition while providing an additional onomatopoeic description of acoustic characteristics that can remain informative even when the class prediction is uncertain. Experiments using ESC-50 and ESC-50-Onomatopoeia show that the proposed method achieves sound recognition performance comparable to that of a conventional recognition-only system. In addition, an LLM-as-a-judge evaluation and subjective listening experiments indicate that the proposed output is preferred over conventional deterministic outputs based on the sound event label, particularly when used to support understanding of the surrounding environment. These results suggest that such output representations can make sound event recognition more informative and communicative under uncertainty.


20. Speaker Disentanglement of Speech Pre-trained Model Based on Interpretability

Authors: Junhua Li, Xiaoxu Zhu, Aaron J. Li

Categories: cs.SD, eess.AS

Self-supervised speech models learn representations that capture both content and speaker information. Yet this entanglement creates problems: content tasks suffer from speaker bias, and privacy concerns arise when speaker identity leaks through supposedly anonymized representations. We present two contributions to address these challenges. First, we develop InterpTRQE-SptME (Timbre Residual Quantitative Evaluation Benchmark of Speech pre-training Models Encoding via Interpretability), a benchmark that directly measures residual speaker information in content embeddings using SHAP-based interpretability analysis. Unlike existing indirect metrics, our approach quantifies the exact proportion of speaker information remaining after disentanglement. Second, we propose InterpTF-SptME, which uses these interpretability insights to filter speaker information from embeddings. Testing on VCTK with seven models including HuBERT, WavLM, and ContentVec, we find that SHAP Noise filtering reduces speaker residuals from 18.05% to nearly zero while maintaining recognition accuracy (CTC loss increase under 1%). The method is model-agnostic and requires no retraining.


21. Anomalous Sound Detection Meets Noise-Aware Self-Supervised Learning

Authors: Takuya Fujimura, Gordon Wichern, Yoshiki Masuyama, Christoph Boeddeker, Kohei Saijo et al.

Categories: eess.AS

In this paper, we introduce noise-aware self-supervised learning (NA-SSL) models for noise-aware anomalous sound detection (NA-ASD). NA-ASD is an ASD task with two-channel audio recordings, where one microphone is located close to the target machine and the other is located farther away to capture noise. For this task, we simulate two-channel recordings using diverse audio datasets and train NA-SSL models to extract clean SSL representations of the close-microphone signal by using the far-microphone recording dominated by background noise as auxiliary information. The NA-SSL models are then used as frontends in the standard ASD framework. Our experimental evaluation on the DCASE 2026 Challenge Task 2 development dataset demonstrates the effectiveness of the NA-SSL framework across three base SSL models (BEATs, EAT, and Dasheng), both with and without discriminative fine-tuning. Furthermore, the challenge results proved the effectiveness of the proposed approach, where the NA-BEATs system won the challenge by a large margin, achieving an official score of 70.24%, while the second-place system achieved 65.46%.


22. Reducing Speaker Residual by Considering Pinhole Effect in Voice Anonymization

Authors: Zeyan Liu, Weili Jiang, Liping Chen, Kong Aik Lee, Boyu Zhao et al.

Categories: eess.AS

Voice anonymization aims to protect privacy by suppressing speaker identity while preserving linguistic content and prosody. However, residual speaker attributes in non-identity representations may still increase linkability and weaken privacy protection. To this end, this paper proposes a fine-tuning strategy with a pinhole loss for well-trained voice anonymization frameworks to further reduce residual speaker attributes. Inspired by the pinhole effect, the pinhole loss measures the linkability of anonymized utterances from the same source speaker. By minimizing this loss, linkability is reduced, thereby improving privacy protection. Experiments on multiple anonymization frameworks, pseudo-speaker generation methods, and datasets show improved privacy protection while maintaining utility. Audio samples can be found in https://anonymous.4open.science/r/Pinhole-loss-fine-tunning-4628.


Authors: Pedro H. L. Leite, Pedro Benevenuto Valadares, Luiz Wagner Pereira Biscainho

Categories: eess.AS

Leading commercial and open-source Text-to-Speech (TTS) models fail to emulate the regional phonetic diversity of Brazilian Portuguese (pt-BR). By aggregating disparate dialects into a single training distribution, they generate a synthetic “diluted” accent: a phonetic profile attempting to represent all regional distributions simultaneously, but ultimately carrying phonological ambiguity dissociated from natural socio-phonetic realizations. This work introduces a speech deepfake detection methodology combining multilingual phone recognizers with classical signal processing to extract phoneme-level features in consonantal and vocalic realizations with high geographic variance. The analysis reveals that the distributional gap over these features suffices to distinguish natural and synthetic voices through unsupervised Kernel Density Estimation, establishing dialectal inconsistency as a useful and interpretable feature for spoofing detection in pt-BR. Evaluation on pt-BR anti-spoofing datasets shows that these explainable, lightweight, low-dimensional features can boost the performance of foundation models on the task, and show generalization capabilities in a cross-dataset leave-one-out setup.


24. Low-Rank Frequency Convolution and Noise-Range Augmentation for Real-Time Pitch Estimation on Edge Devices

Authors: Venkat Suprabath Bitra, Homayoon Beigi

Categories: eess.AS, cs.LG, cs.SD

Pitch estimation on an edge device is constrained in three ways at once. The model must be small, it must stay accurate when the input is noisy, and one frame must be produced inside the frame period. In this report the Frequency Convolution Network (FrCN) of our earlier work is factored into a low-rank form. The number of parameters is reduced by 35.9%, from 17,787 to 11,397, and accuracy is not reduced, either in domain or on two corpora the model was never trained on. The range of the noise used during training is also shown to dominate the architecture in setting how the model behaves when the interference is severe. When the training noise floor is lowered from +6.02 dB to -20 dB, 0.35 points of clean RPA50 are lost and RPA50 at -20 dB is raised from 2.44 to 22.41. This effect is about two orders of magnitude larger than any architectural effect that was measured. The semi-orthogonal constraint used in TDNN-F is found to be redundant with the normalization inside the bottleneck, and accuracy is reduced when both are applied. Whether the factorization saves time depends on the runtime: in eager PyTorch the factored model is 45% slower, while in a compiled kernel it is 15% faster. For deployment, a small C kernel was written. It needs 83 kB on disk and no runtime library beyond libc and libm. It is faster than OpenBLAS on all four CPUs that were tested, faster than ONNX Runtime by 3.8 times, and faster than PyTorch by 13 times. Its output was checked against PyTorch on 271,893 held-out frames per model, and the same pitch bin was selected on every one of them.


25. LLM can Read Spectrogram: Encoder-free Speech-Language Modeling

Authors: Ruchao Fan, Yiming Wang, Yuxuan Hu, Bo Ren, Yufei Xia et al.

Categories: eess.AS, cs.SD

Recent speech-aware large language models (Speech-LLMs) rely on pre-trained speech encoders to convert audio into semantic/acoustic rich representations consumable by LLM. In this work, instead, we explore: can an LLM learn to read Mel spectrogram directly without a dedicated speech encoder? We propose Mel-LLM, an encoder-free Speech-LLM that feeds lightly pre-processed Mel-spectrogram patches directly into the LLM through a linear projection, allowing the LLM to learn speech-text alignment purely through its own parameters. We focus on speech understanding tasks, including automatic speech recognition (ASR), spoken QA and audio understanding. For ASR, we evaluate on the OpenASR Leaderboard public sets and production-level scaling experiments, demonstrating that the encoder-free solution achieves competitive performance with only limited degradation compared to encoder-initialized counterparts. We find that when data is limited, initialization from a multimodal checkpoint (Phi-4-MM) is crucial for maintaining performance. We also present ablation studies suggesting which LLM layers are most involved in speech adaptation. Beyond ASR, we extend Mel-LLM with general speech/audio understanding tasks, revealing an acoustic-semantic trade-off: directly exposing the LLM to Mel-spectrogram input improves paralinguistic and non-ASR acoustic tasks, while knowledge-intensive spoken QA remains more challenging than encoder-anchored systems. We additionally include a text-to-speech (TTS) proof-of-concept with a next-token VAE decoder, showing that direct Mel generation is possible but still trails stronger latent-diffusion generation.


26. When and why do handcrafted cues help self-supervised anti-spoofing? A causal and faithfulness analysis

Authors: Yugwon Won

Categories: eess.AS, cs.SD

Most spoofing countermeasures now place a light classifier on top of a self-supervised (SSL) speech encoder. A growing line of work adds handcrafted acoustic features and fuses them by cross-attention, partly because the attention weights appear to explain which cues the model relies on. Two things about such fusion remain untested: whether the handcrafted features contribute anything once a strong SSL encoder is already in place, and whether the attention map faithfully reflects what the classifier actually uses. We study both questions with MOSAIC, a single model that handles logical access (LA) and physical access (PA) attacks by projecting a 152-dimensional biophonetic vector into six query tokens attending over thirteen intermediate layers of WavLM-Large. Rather than pursuing state-of-the-art accuracy, we treat the model itself as the object of analysis and apply inference-time interventions that zero out selected components. Removing the handcrafted branch lowers the EER by 0.67 percentage points under LA, so WavLM alone is sufficient there, but raises it by 0.76 points under PA, where the branch supplies replay channel evidence the encoder does not carry. The 6x13 attention map is highly stable under resampling (bootstrap rank correlation 0.99), yet its faithfulness varies by domain: layer attention weights predict the causal importance of each layer under PA (rho = +0.62) but not under LA (rho = -0.28). The handcrafted branch and its attention-based explanation are therefore trustworthy for replay attacks and not for synthetic ones. On standard benchmarks the model is mid-range on LA and ahead of the official baselines on cross-source deepfakes (6.21% EER on ASVspoof 2021 DF). The contribution is not a new architecture but a checking procedure: verify handcrafted cues and attention maps by intervention before trusting either.


27. Generative Learning for Ambisonic Upscaling

Authors: Amit Milstein, Nir Shlezinger, Boaz Rafaely

Categories: eess.AS, eess.SP

Ambisonics Upscaling (AU) aims to enhance the spatial resolution of sound fields by estimating high-order Ambisonics (HOA) components from low-order observations. While deep learning and model-based strategies have been considered for AU, both approaches exhibit significant performance degradation in realistic scenarios, where reverberant sound fields violate the directional sparsity inherent to discriminative mappings. In this work, we address AU as a generative task rather than a deterministic reconstruction, expanding generative modeling to specifically target the recovery of spatial information in reverberant speech. We investigate two dominant continuous-time generative paradigms, adapting both Score-based Generative Model and Flow Matching to these complex acoustic settings. We provide an extensive numerical study comparing our methods against state-of-the-art baselines in various acoustic scenarios. Additionally, we conduct subjective listening tests to evaluate the perceived quality and spatial accuracy of the proposed generative framework across various reverberant scenarios. The studies reveal that Flow Matching consistently outperforms both its discriminative counterparts and Diffusion-based paradigms in all reverberant settings.


28. SocialOmni: Benchmarking Audio-Visual Social Interactivity in Omni Models

Authors: Tianyu Xie, Jinfa Huang, Yuexiao Ma, Rongfang Luo, Yan Yang et al.

Categories: cs.AI

Evaluating omni-modal large language models (OLMs) in multi-party dialogue requires more than answer correctness on pre-segmented inputs. We introduce SocialOmni, an offline diagnostic benchmark that separates three turn-level decisions: identifying who is speaking, deciding when a designated participant should enter at an annotated query time, and determining how that participant should continue the dialogue. SocialOmni contains 2,000 perception items and a quality-controlled core split of 200 interaction-generation items, including naturally occurring speaker-visibility mismatches. Each item is independently checked by three human annotators. Evaluated systems receive only query-time-bounded multimodal evidence, while manually verified reference continuations are reserved for response judging. A complete three-judge ensemble scores every eligible response, with leave-one-judge-out and family-sensitivity audits. Across 11 OLMs, rankings vary substantially by axis, and coverage-adjusted scores reveal when high conditional response quality depends on selective turn entry. The protocol does not measure persistent streaming state or wall-clock latency.


29. Explainable Recommendations at Scale: LLM Rationales for YouTube Music Artist Discovery

Authors: Xiao Liu, Yanwei Song, Srivaths Ranganathan, Yuan Chen, Zheyun Feng et al.

Categories: cs.AI, cs.IR

Modern music streaming platforms face a persistent tradeoff: exploiting familiar content versus driving the exploration of novel items. While users frequently desire discovery, they hesitate to select unknown artists over proven favorites. Providing transparent, natural language rationales that explain why an unexplored item is recommended lowers this barrier. However, while Large Language Models (LLMs) excel at this nuanced explainability, their real-time deployment is severely bottlenecked by prohibitive inference costs and computational overhead. In this paper, we present an industry case study of a decoupled recommendation architecture that successfully scales exploration without compromising latency. Our system isolates LLM inference asynchronously offline, pre-computing personalized candidate pools of undiscovered artists alongside tailored rationales. Large-scale online A/B experiments validate our design. We demonstrate that combining LLM-backed recommendations with these explanatory rationales significantly reduces the trust barrier for new content, yielding statistically significant improvements in both user exploration and overall engagement on the discovery surfaces.


30. Efficient and Adaptive Simultaneous Speech Translation with Fully Unidirectional Architecture

Authors: Biao Fu, Donglei Yu, Minpeng Liao, Chengxi Li, Xinjie Chen et al.

Categories: cs.CL

Simultaneous speech translation (SimulST) produces translations incrementally while processing partial speech input. Although large language models (LLMs) have shown strong capabilities in offline translation tasks, applying them to SimulST poses notable challenges. Existing LLM-based SimulST approaches either incur significant computational overhead due to repeated encoding of bidirectional speech encoder, or they depend on a fixed read/write policy, limiting the efficiency and performance. In this work, we introduce Efficient and Adaptive Simultaneous Speech Translation (EASiST) with fully unidirectional architecture, including both speech encoder and LLM. EASiST includes a multi-latency data curation strategy to generate semantically aligned SimulST training samples and redefines SimulST as an interleaved generation task with explicit read/write tokens. To facilitate adaptive inference, we incorporate a lightweight policy head that dynamically predicts read/write actions. Additionally, we employ a multi-stage training strategy to align speech-text modalities and optimize both translation and policy behavior. Experiments on both in-domain (MuST-C) and out-of-domain (Europarl-ST) En-De and En-Es datasets demonstrate that EASiST offers superior latency-quality trade-offs compared to several strong baselines.


31. Quantifying Consonant Contributions to Word Intelligibility via Acoustic Masking

Authors: Eunjung Yeo, Kwanghee Choi, Krupaben Kothadia, Visar Berisha, Julie M. Liss et al.

Categories: cs.CL

Consonants contribute unequally to whether a word is understood. Given the limited time available for therapy, ranking consonants by contribution to intelligibility helps prioritize intervention targets in motor speech disorders. However, measuring this contribution relies on perceptual studies that are difficult to scale. This paper presents a scalable method that measures consonant contribution using acoustic masking. We silence one consonant at a time in an isolated word and test whether an automatic speech recognition (ASR) model still recognizes the word. We define a consonant’s contribution score as the proportion of its masked instances for which the word becomes misrecognized, which we refer to as the mask-induced misrecognition rate (MMR). We relate MMR to two linguistic factors previously reported to correlate with consonant contribution, namely phoneme frequency and functional load. We apply this analysis across four languages, English, Spanish, German, and Czech, using three ASR architectures, MMS (encoder-only), Whisper (encoder-decoder), and Qwen3-ASR (LLM-based). Using partial Spearman correlations, we find that phoneme frequency correlates negatively with MMR while functional load correlates positively. In other words, more frequent consonants are less disruptive when masked, whereas consonants carrying more lexical contrast are more disruptive. Further cross-language analysis shows that consonant rankings agree only partially across languages, indicating that consonant contribution is language-dependent.


32. Federated Multilingual Speech-LLMs: Architecture and Aggregation Strategy Benchmarking

Authors: Jordi Luque, Aleix Sant, Fernando López

Categories: cs.CL, cs.AI

We present a comprehensive benchmark of Federated Learning (FL) for multilingual Automatic Speech Recognition (ASR), evaluating four Speech-LLM architectures on the Multilingual LibriSpeech dataset. We compare FedAvg and FedProx across frozen and unfrozen encoder configurations, demonstrating that optimized learning rates are critical for performance. Specifically, independently tuning the learning rates for the speech encoder, connector, and decoder yields the lowest error rates, with full three-component adaptation (LoRA for encoder and decoder, full training for the connector) producing the best FL results. We observe that FedProx efficacy is architecture-dependent, providing notable advantages in multilingual pre-trained architectures (e.g., EuroLLM over TinyLlama when keeping the encoder fixed); this indicates that LLM backbone capacity plays a key role in mediating resilience to heterogeneous data distributions. These findings offer concrete design guidance for deploying multilingual Speech-LLMs in privacy-sensitive, distributed environments.


33. Me, Myself, and My Voice: Exploring Cultural and Linguistic Identity in AAC AI-generated Voices

Authors: Tobias Weinberg, Aaleyah Lewis, Ricardo E. Gonzalez Penuela, Weicong Hong, Jennifer Mankoff et al.

Categories: cs.HC

Voice is a central element of identity. We recognize people by their voice, and we uniquely express who we are with it. For people who rely on augmentative and alternative communication (AAC) systems, such as speech-generating devices (SGD), the device’s voice becomes an identity marker others associate with them. Yet, it is hard to find a voice that truly aligns with one’s identity both linguistically and culturally. Although modern AI-generated voices can reproduce diverse accents and speaking styles, AAC users still lack accessible ways to articulate how they want an identity-aligned voice to sound like. We first conducted a survey of AAC users (across eight countries) to characterize current voice representation, finding that non-binary, transgender, and non-US-born respondents rated their current voice support identity alignment consistently lower than other respondents. To examine how AAC users respond to voices designed to reflect their cultural identity, we built a tool that elicits cultural markers through guided questions and generates personalized voice candidates for participants to hear and reflect on. After participants heard the voices, we interviewed them to examine what it means for a voice to feel culturally representative, how they interpreted voices with cultural connotations, and how these voices shaped their sense of identity and agency. Our findings show that cultural voice alignment runs deeper than accent or language alone; it touches on belonging, self-recognition, and what it means to be heard as who you are.