每日 arXiv

音频 · 语音 · 音乐 · 声学

每日新爬取的 arXiv 论文,按评分排序。中文摘要帮你快速判断是否值得深读。

2026-07-14
日期2026-07-14
已评分
均分
最高

Daily Papers — 2026-07-14

21 papers on audio, speech, music, and acoustics.

1. Contrasting statistical patterns in melodic and molecular evolution reveal distinctive constraints in a culturally evolving system

Authors: John M McBride, W Tecumseh Fitch

Categories: q-bio.PE, cs.SD, physics.soc-ph | 13 pages, 3 figures, 12 extra pages of supplementary information Score: 6.95/10 (Obj:8 Id:6 Ind:8 Comp:8 Eff:6 Nov:7 Rep:8)

  • Strength: 首次实现 40,000 旋律变体的大规模自动对齐,四项经典生物信息学分析揭示旋律演化签名与分子演化系统性差异(mutability-prevalence r=-0.95;substitution rate 每半音 log-linear 降 20%;metrical hierarchy r=0.71, p=4.8e-7),并创造性地将协方差分解为 repetition-linked 成分。
  • Weakness: 因果识别存在缺口(key-finding 分析有循环风险,mutability-prevalence 因果方向不确定),”melodic skeleton” 假说提出四个可检验预测但本文未执行任何一个,跨文化验证样本量极小(n=8-9)。
  • Full review: Claude Code 全文七公理审稿

Evolved sequences can be used to infer the rules of evolution. Orally transmitted folk melodies are evolved sequences whose similarity to protein sequences (one-dimensional, drawn from a limited alphabet) invites application of bioinformatics methods to study cultural evolution. A major obstacle is that melodies encode rhythm, which breaks some assumptions of standard sequence-alignment algorithms. We develop a rhythm-aware alignment method and apply it to \num{40000} Irish dance tune variants, enabling the first large-scale automated melodic alignment. Four canonical bioinformatics analyses – mutability, substitution matrices, positional conservation, and covariance – reveal patterns distinct from those of molecular evolution, revealing the forces that shape each domain: biochemical and biophysical constraints for proteins; memory, motor, and social biases for melodies. Together the results show that bioinformatics provides a powerful framework – conceptual as much as algorithmic – for studying cultural evolution. Although the cultural transmission of music has been discussed for centuries, here we show how to analyze it at large scale.


2. Audio-Native Speech Recognition with a Frozen Discrete-Diffusion Language Model

Authors: Harsha Vardhan Khurdula, Abhinav Kumar Singh, Yoeven D Khemlani, Vineet Agarwal

Categories: cs.AI, cs.SD | 10 pages, 2 figures, 6 tables Score: 6.37/10 (Obj:8 Id:5 Ind:8 Comp:8 Eff:5 Nov:7 Rep:5)

  • Strength: 首次实现冻结离散扩散语言模型(DiffusionGemma 26B)的音频原生 ASR,仅训练 42M 参数(0.16%),通过 CTC 损失突破投影器-注意力梯度死锁,在 LibriSpeech test-clean 达到 6.6% WER,解码成本与转录长度解耦(~8 步并行)。
  • Weakness: 绝对 WER 与 SOTA 差距巨大(6.6% vs <2%),多语言结果更弱(FLEURS-en 15.7% vs Whisper 4-5%),CTC 解锁机制缺乏受控消融验证,评测样本量小(n=100-1000),无代码开源声明。
  • Full review: Claude Code 全文七公理审稿

Automatic speech recognition is dominated by autoregressive decoders that emit one token at a time. We ask whether a discrete diffusion language model can transcribe speech instead, refining a whole transcript in parallel over a small number of denoising steps. We train an audio-native interface for DiffusionGemma, a 26B mixture-of-experts model that generates text by uniform, random-token discrete diffusion rather than the absorbing-mask scheme common to recent diffusion language models. A frozen Whisper encoder supplies acoustic features, a lightweight projector maps them into the model embedding space, and low-rank adapters let the frozen backbone attend to the new modality. About 42M parameters are trained, which is 0.16 percent of the backbone. We find that the natural training objectives fail to ground the audio because their gradient reaches the projector only through attention that has already dismissed it. A connectionist temporal classification loss applied through the frozen output head breaks this deadlock. The resulting model reaches 6.6 percent word error rate on LibriSpeech test-clean, transcribes in roughly eight parallel steps regardless of utterance length, and uses a single adapter trained on six languages, which we evaluate here on English, Hindi, and Mandarin.


3. ZipL-Dialog: Memory-Efficient Long-Form Spoken Dialog Synthesis via Latent Flow Matching

Authors: Jihwan Kim, Nam Soo Kim

Categories: eess.AS | Accepted to Interspeech 2026. 5 pages, 1 figure, 3 tables Score: 6.10/10 (Obj:8 Id:6 Ind:8 Comp:5 Eff:6 Nov:5 Rep:6)

  • Strength: 将 ZipVoice-Dialog 的帧级 mel 域 CFM 迁移到 25 Hz 潜在空间,实现最大峰值 GPU 内存 11.22× 降低和推理 2.23× 加速,使分钟级对话 TTS 在单张 40 GB GPU 上单 pass 合成成为可能。
  • Weakness: 核心质量指标全面退化(WER 相对增加 23-51%,cpSIM 下降),消融实验在单说话人设置而非主 dialog 设置下进行,且方法各组件均为已有技术的工程组合。
  • Full review: Claude Code 全文七公理审稿

Zero-shot dialog TTS benefits from flow-matching, but minute-scale generation on dense mel-spectrograms causes severe memory bottlenecks, often forcing unnatural chunked synthesis. We propose ZipL-Dialog, which shifts conditional flow-matching into a 4x time-compressed (25 Hz) latent space. To preserve acoustic fidelity under compression, we employ a deterministic mel autoencoder with auxiliary mel-domain supervision and optimize the ZipFormer’s hierarchical downsampling schedule. Experiments show that ZipL-Dialog reduces maximum peak GPU memory by 11.22x and accelerates inference by 2.23x over the baseline, substantially lowering the memory footprint of single-pass multi-minute dialog synthesis while maintaining perceptual naturalness.


4. Spatial-Frequency Cued Generative Fixed-Filter Active Noise Control Based on Deep Learning in Reverberant Environments

Authors: Boxiang Wang, Haowen Li, Dongyuan Shi, Junwei Ji, Ziyi Yang et al.

Categories: eess.AS, eess.SP Score: 6.05/10 (Obj:8 Id:5 Ind:8 Comp:5 Eff:6 Nov:5 Rep:8)

  • Strength: 通过理论推导(Eq. 22)将最优控制滤波器显式表达为噪声源 3D 位置的函数,并用多任务 CRNN(0.24M 参数)在未见混响环境上实现 >91% 距离分类准确率和 >95% 方位角准确率,实测声路径上达到 >10dB NR。
  • Weakness: 缺乏组件级消融实验(仅空间 vs. 仅频率 vs. 联合),未与级联 baseline(独立定位+GFANC)比较,实测验证仅覆盖 2 个噪声源位置,且未引用/比较 2025 年最新 SOTA 变体 GFANC-RL。
  • Full review: Claude Code 全文七公理审稿

Generative fixed-filter active noise control (GFANC) effectively attenuates noise with diverse frequency characteristics through the combination of sub control filters. However, it does not incorporate the spatial information of the noise source, which limits its performance, particularly in reverberant environments. To address this limitation, this paper proposes a novel spatial-frequency cued GFANC (SF-GFANC) method that exploits both three-dimensional (3D) spatial and frequency information of the noise source. Specifically, a multi-task convolutional recurrent neural network (CRNN) is designed to estimate the source distance, elevation angle, and azimuth angle as spatial cues, while predicting the combination weights of sub control filters as frequency cues. These spatial-frequency cues jointly guide the generation of the appropriate control filter. In addition, a theoretical analysis of the optimal control filter in reverberant environments is presented, highlighting the importance of 3D spatially conditioned control filter design. Evaluations using both simulated and measured acoustic paths demonstrate that the CRNN is robust to unseen acoustic environments and noise types. Furthermore, the results confirm that SF-GFANC outperforms representative ANC algorithms when handling noise sources across diverse 3D locations and frequency characteristics in reverberant environments.


5. What is a Musical Scale? Regularity and Convention in the Organization of Pitch

Authors: John M McBride

Categories: cs.SD | 13 pages, 3 figures, includes a 3-page statistical reporting checklist Score: 6.05/10 (Obj:8 Id:5 Ind:6 Comp:8 Eff:5 Nov:5 Rep:8)

  • Strength: 将”musical scale”从一个monolithic概念系统拆分为empirical core、scale membership、scale typology三层,并用prototype theory配合40,000首Irish dance tunes数据illustrate scale categorization的conventional层面,为跨文化scale研究提供了coherent的definitional framew。
  • Weakness: 论文是conceptual/programmatic的,未在large-scale cross-cultural study中demonstrate该framework比已有approaches(如作者自己前期的McBride et al. 2023或Harasim et al. 2020的axomatic scale theory)产生更好的empirical outcomes,utility主要是prospective而非demon。
  • Full review: Claude Code 全文七公理审稿

Musical scales are near-universal in human music, and most readers will feel they already know what a scale is. On closer inspection, however, the literature lacks a consensus definition: which conditions are necessary and sufficient shifts across disciplines and traditions, and the term turns out to cover several distinct objects. I argue this is less a failure of rigour than a sign that ``scale’’ names several related objects: prescriptive abstractions, instrument tunings, statistical regularities in performed pitch, perceptual categories, social conventions. I adopt an empirical definition – a scale as a statistical regularity in pitch organisation relative to a tonic – that is portable across traditions and computable from recordings, and situate it alongside the other senses of the term. Even this empirical core is not purely observational, as convention enters in deciding which pitches belong to a scale. And a further step of grouping scales into named categories is a separate convention, which I approach through prototype theory and illustrate with examples from Irish music. Separating these layers provides a basis from which scales can be re-examined empirically and cross-culturally.


6. Real-time Generation of Listener Nodding via Prediction of Kinematic Parameters for Avatar Dialogue Systems

Authors: Kazushi Kato, Koji Inoue, Taiga Mori, Divesh Lala, Tatsuya Kawahara

Categories: cs.HC, cs.SD | Accepted by 28th ACM International Conference on Multimodal Interaction (ICMI ‘26), Long paper Score: 6.00/10 (Obj:8 Id:6 Ind:8 Comp:5 Eff:6 Nov:5 Rep:8)

  • Strength: 在 VAP 框架上首次实现实时联合预测点头时序与四个运动学参数(range, speed, repetitions, swing-up),fine-tune 策略使 Corr. 从 0.41 (SVM-Mimi) 提升至 0.52,并开源完整系统 (MaAI, pip install)。
  • Weakness: 核心创新(context-aware 参数预测)在主观评估中 Method E vs D 无显著差异(全部 7 指标 p>0.2),且绝对 Corr.≈0.52 仅解释约 27% 方差,仅单一日语数据集验证。
  • Full review: Claude Code 全文七公理审稿

In human dialogue, we achieve smooth communication by expressing nonverbal cues such as eye contact, nodding, and facial expressions with precise timing. It is expected for conversational avatars to express these cues appropriately to realize natural and human-like interactions. This study focuses on nodding, which is crucial for demonstrating active listening and encouraging further user utterances. We propose a model that predicts both timing and kinematic parameters representing the motion features of listener nodding in real time. The proposed model consists of a timing prediction module and a kinematic parameter prediction module. Each implements a dyadic attention network over the speaker and listener channels based on the technique of Voice Activity Projection (VAP). Unlike conventional models, this approach enables real-time prediction of kinematic parameters based on the specific context of the dialogue rather than just predicting the timing. Furthermore, we demonstrate the effectiveness of fine-tuning the kinematic parameter prediction module initialized from the trained timing prediction module. The proposed model is lightweight and capable of real-time operation, and it has been integrated into an avatar dialogue system. Subjective evaluation experiments shows that our proposed method significantly outperforms both a baseline with stochastic timing and another with fixed-motion nodding. The code and trained models are available at https://github.com/MaAI-Kyoto/MaAI.


7. ChartGenEval: Corruption-Tested Multi-Dimensional Feedback for Rhythm-Game Chart Generation

Authors: Jhen-Ke Lin

Categories: cs.SD, cs.AI Score: 6.00/10 (Obj:8 Id:6 Ind:8 Comp:5 Eff:6 Nov:5 Rep:8)

  • Strength: 七个评测输出轴在 80 个 held-out 歌曲组上全部通过九项 corruption 测试(含 sensitivity 和 invariance 标准),且揭示了 LM perplexity 在 common-pattern rewrite 下下降 37% 和 self-similarity 在 loop collapse 下上升 62% 的 wrong-way incentive,为该领域首次提供 corruption-valid。
  • Weakness: Corruption-tested metric validation 方法论从 text/music generation 已有工作迁移而来,非方法论原创;且框架复杂度较高(6 问题 / 7 轴 / 9 测试 / 14 calibrated scores),关键 proxy diagnostic 结论(perplexity 和 self-similarity 的 wrong-way incentive)仅在 40 首歌的 develo。
  • Full review: Claude Code 全文七公理审稿

A generated rhythm-game chart need not reproduce one official note sequence: many note choices can fit the same song and difficulty. Reference-note agreement therefore measures reconstruction, not the full design problem. We introduce ChartGenEval, a six-question evaluation framework with an automatic, corruption-tested core. It leaves note choice open while anchoring timing to the song: the matched official chart supplies only its authored timing map, never target notes. We test each core output with dose-controlled failures rather than assume that a familiar statistic measures chart quality. Across 80 held-out song groups, seven output axes satisfy prespecified sensitivity and invariance criteria in nine nonredundant tests. Complementary stress tests on the 40-song development panel expose two broader lessons. A chart-wide phase estimate recovers injected shifts of 15, 30, and 60 ms while chart-only outputs remain essentially unchanged. Common-pattern rewriting lowers mean language-model perplexity by 37%, and loop collapse raises mean self-similarity by 62%. ChartGenEval therefore reports separate, role-specific signals instead of one proxy or total score. This profile provides automatic feedback for comparing and iterating generators; selected outputs are candidate optimization targets or constraints after task-specific stress testing.


8. Low-Latency Neural Models for Real-Time Music Enhancement

Authors: Emmanouil Karystinaios, Jonathan Greif, David Nadrchal, Paul Primus, Gerhard Widmer

Categories: cs.SD Score: 5.95/10 (Obj:8 Id:6 Ind:6 Comp:8 Eff:5 Nov:5 Rep:8)

  • Strength: 首个系统性实时音乐增强 benchmark,运行时证据可靠(所有因果模型 RTF<0.114),核心分析结论”无差别增强可在多个客观指标下恶化输入”是反直觉且可靠的经验发现,立体声退化诊断(MM-SNR delta=-0.0725)精确定位了失败模式。
  • Weakness: 作为 benchmark 型论文效用不足——最佳实时模型绝对改善仅约 1 dB SI-SNR,SonicMaster 上所有实时模型 SI-SNR 均为负值(比退化输入更差),缺少非学习基线、实时音乐分离方法、MSR Challenge 系统等关键 baseline,且无任何主观评测验证。
  • Full review: Claude Code 全文七公理审稿

Music recordings and live streams are often affected by noise, reverberation, spectral imbalances, or artifacts that degrade listening quality. While speech enhancement has matured into a well-defined research area, music enhancement is less established because musical signals combine overlapping sources, wide bandwidths, strong dynamics, and intentional production effects. We study real-time music enhancement under strict causal and low-latency constraints. We formulate the task around recovery of the intended produced mix from acoustic and production-oriented degradations, adapt compact causal networks to music, and compare speech-derived real-time baselines, an external music-denoising model, an offline restoration reference, and a music-specific MusicFilterNet-MS variant. On the tested hardware, all causal models run faster than real time, but improvements depend strongly on the dataset, degradation type, and metric family; under several objective criteria, indiscriminate enhancement can worsen the degraded input. The main contribution is therefore a benchmark and an analysis rather than a universal best model: real-time music enhancement is feasible, but robust improvement requires degradation-aware modeling, stereo-aware processing, identity-preserving correction, and evaluation beyond a single objective score.


9. AVSCap: Orchestrating Audio-Visual Synergy for Omni-modal Video Captioning

Authors: Yanghai Wang, Jiahao Wang, Jiafu Tang, Yuanxing Zhang, Zhe Cao et al.

Categories: cs.CV Score: 5.95/10 (Obj:8 Id:5 Ind:5 Comp:5 Eff:8 Nov:4 Rep:6)

  • Strength: AVSCap-7B 在 AVSCapBench 上总分 60.44,超越所有开源 baseline(次优 AVoCaDO 49.31,+11.13),在 synergy recall 上 57.70 远超 baseline(次优 29.13),并在 UGC-VideoCap(82.4)和 Daily-Omni(66.6)上接近甚至超过 Gemini-3-Pro。
  • Weakness: 核心方法(Qwen2.5-Omni + SFT + GRPO)与 AVoCaDO(ICLR 2026)高度重叠,真实新颖性有限;数据生成(Gemini-3-Flash)、训练 reward(GPT-5)、评测 judge(Gemini-3.1-Pro)均依赖闭源 API,关键 pipeline 不可独立复现;训练 reward 与评测范式重叠构成循环风险。
  • Full review: Claude Code 全文七公理审稿

Omni-modal video captioning is not merely combining visual captioning with audio transcription: a useful caption must describe how visual actions, speech, music, and sound effects co-evolve. Existing large multimodal models often fail at this relational step, treating audio and visual streams as loosely coupled observations, relying on automatic speech recognition, and under-specifying non-speech sounds and their links to visual events. We present AVSCap, a framework for audio-visual captioning centered on explicit cross-modal event binding. First, we construct AVSCap-130K, a tri-modal training corpus generated by a decoupled-then-fused pipeline that anchors visual and acoustic evidence before composing grounded omni-modal captions. Second, we train AVSCap-7B, a 7B captioner with a two-stage strategy: supervised fine-tuning establishes baseline capabilities, while sample-efficient reinforcement learning uses hybrid rewards to optimize acoustic completeness and audio-visual synergy. Our scaling analysis shows that reinforcement learning brings larger gains than increasing SFT data. Third, we introduce AVSCapBench, a benchmark that decomposes captions into visual, audio, and synergy events and evaluates them with fine-grained event recall. Experiments on AVSCapBench and external benchmarks show that AVSCap-7B improves non-speech audio coverage and cross-modal binding, delivering the best overall performance among evaluated open-source models.


10. Investigating the Integration of Spatial Information in Foundation-Model-Based Speaker Diarization

Authors: Marc Deegen, Adrian Meise, Reinhold Haeb-Umbach

Categories: eess.AS | Accepted at IWAENC 2026 Score: 5.60/10 (Obj:8 Id:5 Ind:8 Comp:5 Eff:6 Nov:4 Rep:7)

  • Strength: 在统一框架下对比三种多通道-FM集成方案,发现beamformer在overlap区段有害(DER增加1.1pp),conditioning方案最优(macro DER 15.3% vs baseline 16.28%),且错误分析揭示spectral-only系统在overlap区段反直觉地优于spatial-only系统。
  • Weakness: 三个核心方案均直接采用已有工作(尤其conditioning方案来自同一作者团队的前期工作 [8]),论文novelty增量主要限于对比实验和错误分析,且未与pyannote、ECAPA-TDNN、SphereVBx等主流diarization系统对比,baseline覆盖面有限。
  • Full review: Claude Code 全文七公理审稿

Spatial information gleaned from multi-channel input has been shown to lead to improvements in meeting processing tasks like diarization and source separation. At the same time, diarization based on features extracted by large pretrained single-channel foundation models, such as WavLM, achieved state-of-the-art performance. This work compares three approaches to integrate spatial features into foundation model-based diarization systems: the cascade of a beamformer and a single-channel foundation model, a multi-channel foundation model, and the conditioning of the downstream network on explicitly extracted spatial features. Results show that the beamformer front-end is even detrimental to diarization performance in regions of overlapped speech, while best performance is achieved with the conditioning, demonstrating that the incorporation of explicit spatial features is a competitive approach to foundation-model-supported diarization. This approach is further subjected to a detailed error analysis showing that the conditioning system removes errors to a good extent that would occur when either only spectral or only spatial features were used.


11. AutoSIFT: Automatic Style Sifting for Controllable Speech Generation with Arbitrary Style Infilling

Authors: Haowei Lou, Junda Wu, Chengkai Huang, Tong Yu, Hye-young Paik et al.

Categories: cs.SD Score: 5.50/10 (Obj:6 Id:5 Ind:3 Comp:5 Eff:6 Nov:5 Rep:6)

  • Strength: 提出了 Arbitrary Style Infilling 任务框架和 DSA/RSA 评测指标,在分类解耦上大幅超越基线(Emotion B.Acc 67.2% vs ParaMETA 50.1%,Language 98.8% vs 91.1%),ASI 门控机制在单类别编辑时接近 100% DSA。
  • Weakness: S-SIM 评测使用自训练 Style Extractor 构成循环论证(AutoSIFT S-SIM 0.8512 的巨大优势可能被膨胀),缺少 DMP-TTS/ControlSpeech/FleSpeech 等直接相关同期工作的对比,核心 PAL 损失直接继承自前作 ParaMETA 且代码未公开。
  • Full review: Claude Code 全文七公理审稿

State-of-the-art text-to-speech (TTS) models achieve impressive naturalness and expressiveness, yet fine-grained, disentangled control over speaking styles remains challenging. In professional scenarios such as film dubbing, game voice acting, and video content generation, users often need to modify a specific style category, such as emotion, age, or gender, while preserving all others. Existing style-controllable TTS methods typically rely on either text-described styles or speech-reference style transfer, making it difficult to jointly control explicit semantic attributes and preserve subtle, text-undescribed prosodic details. We propose AutoSIFT, a controllable speech generation framework for category-level style editing. AutoSIFT decomposes speaking style into known text-describable categories and unknown residual styles that capture non-verbal prosody and speaker-specific nuances. It consists of a generalized Style Disentangler, which extracts category-aware style prototypes from reference speech, and an Arbitrary Style Infiller, which selectively infills unspecified style categories from the reference. By replacing only text-specified style categories while preserving residual speech-derived styles, AutoSIFT enables natural, expressive, and highly customizable speech generation.


12. Explainable-by-Design Audio Deepfake Detection via Wiener-Hopf Linear Prediction

Authors: Mattia Tamiazzo, Simone Milani, Massimo Iuliani, Marco Fontani

Categories: cs.SD, cs.AI, cs.CR, cs.MM | Accepted at ACM IH&MMSec 2026 Score: 5.50/10 (Obj:8 Id:5 Ind:9 Comp:5 Eff:5 Nov:5 Rep:5)

  • Strength: 将经典 Wiener-Hopf 线性预测系数作为 2D 图像输入轻量 CNN(422K 参数),在三个公开数据集上达到平均 EER 3.16%,并通过 Grad-CAM 揭示模型关注静音与过渡区域,提供了物理可解释的检测路径。
  • Weakness: 在最关键 benchmark ASVspoof 2019 上 EER 5.97% 显著落后于 Wav2Vec2 (2.56%),缺少最简 LP 分类器 baseline 和直接前驱 Borrelli et al. 2021 的实验对比,且”20倍低复杂度”声称无定量证据支持。
  • Full review: Claude Code 全文七公理审稿

The rapid advancement of synthetic speech generation methods has made audio deepfake detection a critical challenge in multimedia forensics. While recent approaches achieve high detection accuracy, they typically rely on black-box architectures that offer limited interpretability and high computational complexity. In this paper, we propose an explainable-by-design audio deepfake detection framework based on Wiener-Hopf linear prediction, processed by a lightweight 2D Convolutional Neural Network (CNN). This design enables a direct and transparent connection between classification outcomes and the acoustic properties of the signal. Experimental results on benchmark datasets demonstrate competitive detection performance while maintaining significantly lower computational complexity compared to state-of-the-art solutions. The interpretability analysis using Grad-CAM reveals that the classifier focuses on low-order predictor coefficients and on silence and transitional regions, suggesting that the Wiener-Hopf predictor captures reverberation characteristics and subtle statistical inconsistencies in synthetic speech. Finally, robustness experiments show that fine-tuning effectively recovers detection performance under common post-processing degradations, including additive noise, MP3 compression, and telephone filtering.


13. Listen first: Output-based multi-microphone speech enhancement

Authors: Panos Apostolidis, Svend Feldt, Zheng-Hua Tan, Jan Østergaard, Jesper Jensen

Categories: eess.AS | Accepted at the International Workshop on Acoustic Signal Enhancement (IWAENC) 2026 Score: 5.50/10 (Obj:6 Id:5 Ind:8 Comp:5 Eff:5 Nov:5 Rep:5)

  • Strength: 在 SNRi=-5dB 下,output-based MPDRF 相对 input-based MVDR 获得 ~4.2dB ΔSNR 提升(11.00 vs 6.78 dB)及 ESTOI/PESQ 显著改善,且在 RTF mismatch 条件下仍保持优势(Wilcoxon p=0.05)。
  • Weakness: 仅对比单一 MVDR baseline,缺少 deep learning beamforming SOTA 对比和关键 ablation(GP 选择 vs. 随机选择的贡献被省略),且 output-based paradigm 的 general claim 远超实验验证的窄仿真场景。
  • Full review: Claude Code 全文七公理审稿

Traditionally, hearing-aid speech enhancement (SE) algorithms rely on input-based feature estimation, often derived by a voice activity detector (VAD), to configure beamformers. Yet features extracted from noisy microphone signals can become unreliable in challenging acoustic scenes where users most need help. We introduce a novel paradigm in which the settings of a sound processing system are determined by evaluating characteristics of its output. To demonstrate this idea, we employ an output-based system that selects among a set of minimum power distortionless response (MPDR) beamformers. Although MPDR beamformers are typically avoided due to their sensitivity to steering errors, we show that they become effective within an output-based framework. We compare the proposed system to a conventional input-based minimum variance distortionless response (MVDR) baseline. Experimental results show that the proposed system consistently outperforms the MVDR baseline, particularly at low SNRs, in terms of SNR, ESTOI and PESQ.


14. UD-ASD: A Unified Diffusion Model for Anomalous Sound Detection

Authors: Pengxiang Gao, Yu Qiu, Yanzhi Song

Categories: cs.SD | 5 pages, 3 figures, Interspeech 2026 Score: 5.47/10 (Obj:8 Id:5 Ind:8 Comp:5 Eff:5 Nov:4 Rep:6)

  • Strength: 统一 diffusion 模型在 DCASE2022 Task 2 上达到 Hmean AUC 77.16%,通过 Condition Projector 实现多机器共享模型,减少部署开销。
  • Weakness: 方法为已有组件(DDPM + ID embedding + GMM)简单组合,缺少最简 baseline(no-CP unified model)消融,仅在一个数据集上评测且未与 DCASE2022 挑战赛 top systems 对比,提升幅度有限。
  • Full review: Claude Code 全文七公理审稿

Anomalous Sound Detection (ASD) aims to determine whether faults have occurred by monitoring sounds. Existing methods detect a limited range of anomalies, exhibit poor generalization, or train a separate model for each machine. Diffusion models possess strong generalization and can generate specific data with condition guidance. We propose a unified diffusion model only with a small module. The audio is first transformed into log-Mel spectrograms. The lightweight module embeds machine IDs into condition embeddings, guiding the model to reconstruct data for specific machines. Then diffusion model reconstructs data with condition, using Gaussian Mixture Models to fit the distributions of reconstruction errors. Our unified model could monitor multiple machine types and learn more fundamental feature spaces with cross-domain learning. Experiments on DCASE2022 Challenge Task 2 show that our model achieves 3.44% AUC and 2.52% pAUC improvements over baseline, validating its effectiveness.


15. An Omnilingual-ASR-Based Speech-LLM System for the 2nd MLC-SLM Challenge

Authors: Shuming Fang, Shuifei Zeng

Categories: cs.SD, cs.AI | Accepted to INTERSPEECH 2026. 4 pages + references. Technical description of our 2nd MLC-SLM Challenge Task 1 submission Score: 5.47/10 (Obj:8 Id:6 Ind:8 Comp:5 Eff:5 Nov:3 Rep:6)

  • Strength: 在 Development set 上实现 macro tcpMER 29.27%(vs 官方 baseline 79.15%,相对降低 63%),并通过 4 组消融实验清晰识别出 embedding-based clustering 和 overlap-disabled segmentation 是性能关键。
  • Weakness: 系统是已有公开组件的纯工程组合(零方法新颖性),仅与极弱的官方 baseline 比较,Evaluation set 上 tcpMER 50.23% 且 Dev-Eval gap 达 21 个百分点,缺少与其他参赛者的横向比较和代码发布。
  • Full review: Claude Code 全文七公理审稿

We describe our submission to Task 1 of the 2nd MLCSLM Challenge: a cascaded diarization-then-recognition system that combines DiariZen-Large-s80 (WavLM-Large) segmentation, CAM++ embedding-based two-speaker clustering, and a LoRA-adapted omniASR LLM 7B v2 recognizer, with no oracle segmentation or speaker labels at test time. On the official Development set (150 conversations, 21 language/accent categories) the system attains a macro tcpMER of 29.27%, versus 79.15% for the official baseline; on the Evaluation set it scores 50.23%. We also analyze two engineering choices that substantially affect tcpMER. First, embedding-based speaker clustering outperforms an end-to-end-style alternative that assigns speakers from ASR <sc> turn markers alone. Second, overlap-aware segmentation, although intended to raise diarization recall, increases tcpMER because overlapped speech is transcribed twice.


16. PolarBM: Complex-valued Boltzmann Machine for Modeling Audio Signals in Polar and Log-polar Coordinates

Authors: Toru Nakashika, Kohei Yatabe

Categories: cs.LG, cs.SD, eess.AS, stat.ML | Submitted to IEEE Trans. ASLP Score: 5.40/10 (Obj:8 Id:6 Ind:8 Comp:5 Eff:5 Nov:5 Rep:4)

  • Strength: LogPolarRBM 在 PESQ=4.00、MOS=4.51 上达到与自然语音无统计显著差异的重建质量,理论上统一了 Rice/Nakagami/noncentral-chi/Gamma 分布为 PW-NCCG 特例。
  • Weakness: 实验仅用 4.2 分钟单一说话人数据,PolarRBM 在 PESQ 上输给 GVM-RBM(3.65 vs 3.75),无代码开源,且核心组件全部来自同一作者前作的自引用链,新颖性增量有限。
  • Full review: Claude Code 全文七公理审稿

Although vast amounts of data, such as audio signal spectra, are naturally represented using complex numbers, conventional machine learning methods often simplify complex-domain problems by employing frameworks designed for real-valued variables. While this simplification offers computational benefits, it discards structural information regarding the inherent relationship between amplitude and phase. In this paper, we propose a novel Boltzmann machine (BM), named PolarBM, capable of naturally handling complex-valued variables in the polar coordinate (i.e., an amplitude-phase representation). PolarBM defines a probability density function for complex variables in which the phase explicitly depends on the amplitude, thereby capturing the physically important relationships of complex-valued signals. Furthermore, to process audio signals in accordance with human auditory perception, we propose LogPolarBM, which models amplitude on a logarithmic scale. This extension yields a flexible conditional probability density function, a power-weighted noncentral complex Gaussian (PW-NCCG) distribution, whose marginal amplitude distribution encompasses the Rice, Nakagami, and noncentral chi distributions as special cases. For practical applications, we also introduce the restricted variants of these proposed models: PolarRBM and LogPolarRBM. Experimental results demonstrate that by explicitly modeling the dependency between amplitude and phase, the proposed RBMs achieve superior modeling accuracy compared to conventional models, including deep neural networks. Although our experiments focus on audio signals, the utility of the proposed BMs is not limited to audio applications; their potential extends widely across various fields of science and engineering that involve complex-valued data, such as wireless communications and quantum mechanics.


17. The Sound of Absence: Audio-Language Embedding Models Struggle with Negation

Authors: Chun-Yi Kuan, Hung-yi Lee

Categories: eess.AS, cs.AI, cs.CL, cs.LG, cs.SD | Manuscript in progress Score: 5.40/10 (Obj:8 Id:6 Ind:5 Comp:8 Eff:5 Nov:5 Rep:4)

  • Strength: 系统性证明音频-语言嵌入模型存在严重否定盲区(Negation MCQ 准确率 0.2-7.1%,远低于 25% 随机基线),受控诊断(ESC-50, 1000 items)定量揭示否定仅产生 30% 的内容替换级嵌入位移,该规律在 CLAP 和 MLLM-based WAVE-7B 中均成立。
  • Weakness: 遗漏直接先验工作 Vasilakis et al. (arXiv 2601.13931, 2026-01) 对音乐 CLAP 否定建模的研究(6 个月前发表,未被引用),benchmark 构建依赖闭源 LLM(Claude Opus 4.8)且代码/数据未开源,MCQ 选项逻辑正确性的人工校验统计缺失。
  • Full review: Claude Code 全文七公理审稿

Audio-language embedding models such as CLAP are widely evaluated on matching present sound events, but rarely on negation. We show this affirmation-only evaluation hides a key limitation: these models fail to encode negated sound concepts, mapping affirmative and negated captions to nearly identical representations. To expose this blind spot, we introduce NegEval-Audio, a framework that converts existing datasets into two negation-aware tasks, Retrieval-Neg and Multiple-Choice Negation (MCQ-Neg), to probe whether models distinguish present from absent events. On AudioCaps and Clotho, performance degrades sharply under negation, with negation-type MCQ accuracy falling far below chance, and the failure persists even for a recent multimodal LLM-based embedding model. While a training-free steering method improves MCQ-Neg, it yields marginal gains for Retrieval-Neg. This indicates that affirmation bias is a fundamental flaw in the representation geometry, necessitating explicit negation-aware training objectives.


18. Audio Diarization: A New Paradigm for Exploring Audio Recordings with Unknown Event Classes

Authors: Alexander Werning, Reinhold Haeb-Umbach

Categories: eess.AS | accepted at IWAENC 2026 Score: 5.30/10 (Obj:6 Id:4 Ind:8 Comp:5 Eff:5 Nov:5 Rep:5)

  • Strength: 提出 audio diarization 任务范式并展示 EEND 适配后可在未知类别场景下以 34.6% DER 显著优于闭集 SED baseline 的 45.3% DER(11pp 提升),同时支持 overlap 检测。
  • Weakness: 核心能力(同类事件 grouping)验证失败率超 55%(DMultiMix both-miss),baseline 比较不公平(SED 用更多数据+fine-tuned encoder+post-processing),且方法为已有组件直接组合无架构创新。
  • Full review: Claude Code 全文七公理审稿

We propose a new task, audio diarization. The motivation is that there are applications, such as audio monitoring in an unknown environment, where initially the sound event classes to be recognized are unknown. For such a scenario, we propose to first localize in time relevant sound events and to classify them, e.g., by comparing with known event classes, in a second step. This contribution is dedicated to the first step, which we call audio diarization, as it is reminiscent of the speaker diarization stage that precedes and simplifies the second stage, speech recognition, in multi-talker conversational speech processing. In this contribution, we define audio diarization as detecting onset and offset times of sound events with overlap for an open set of classes and without user prompts. We show how a speaker diarization system can be adjusted for audio diarization and propose an evaluation setup. Compared to a closed-set sound event detection system, the proposed system achieves similar performance with the additional ability to detect novel sounds.


19. Open-Source Intelligence and Music Information Retrieval for Geographic Attribution of Musical Affect and the Ecological Limits of Population Inference

Authors: Mohammadreza Rashidi

Categories: cs.CR, cs.SD | 16 pages, 12 figures Score: 5.30/10 (Obj:8 Id:6 Ind:8 Comp:5 Eff:5 Nov:4 Rep:9)

  • Strength: 跨国音乐结构差异的统计证据极强(8特征全部 p<0.001,leap ratio p<10⁻⁹⁰,Cliff’s delta 达 0.90),且 fail-closed checker 重新推导论文每一个数字,可复现性远超领域标准。
  • Weakness: 核心负结果(0/6 相关性不显著)基于 n=7-10 的极小样本且未做功效分析,同时遗漏至少三篇直接相关工作(IZA DP 14258、Springer BRM 2021、arXiv:2506.13199),Related Work 未覆盖”音乐-国家指数”文献。
  • Full review: Claude Code 全文七公理审稿

A common intuition holds that a region’s music mirrors the temperament of its people, so that melancholic melodies mark melancholic populations. We test the measurable half of that intuition and reject the inferential half. Using the Essen Folksong Collection, a corpus of thousands of notated folk melodies, we extract real melodic and affect-related features from 2393 deduplicated melodies spanning 16 countries and 7 geographic regions, with the analysis performed on symbolic scores rather than audio. The mode of each melody is computed with a key-finding algorithm rather than read from the file, because the collection’s own documentation warns its major and minor labels are unreliable. Cross-country differences in melodic structure are large and highly significant. All 8 tested features differ across countries at p<0.001, with the leap-related features reaching p<10^-90, and China carries a distinctive wide-leap, high-activity signature (arousal composite +1.24 standard deviations, mean absolute interval 2.77 semitones against Germany’s 2.17). We then test the inferential half. We correlate the regional musical-affect measures with two published, validated national indices, the World Happiness Report ladder score and the Hofstede individualism index. None of the 6 correlations is significant (0 of 6). The geography of musical affect is real and measurable, but it does not predict how happy or how individualist a population is, and any claim that it does is an ecological fallacy. We release the full extraction and analysis pipeline, and a fail-closed checker re-derives every number in this paper from the data.


20. Traceback Translators Against Forgetting in Continual Fake Speech Detection

Authors: Enrico Gottardis, Mattia Tamiazzo, Simone Milani

Categories: cs.CV, cs.AI, cs.MM, cs.SD | Accepted at EUSIPCO 2026 Score: 5.16/10 (Obj:8 Id:5 Ind:8 Comp:5 Eff:5 Nov:4 Rep:5)

  • Strength: 仅训练 21K 参数的 traceback translator 在冻结 backbone 下实现源域精度保持(ASV19: 95.0% AUC / 9.74% EER 与 baseline 持平)并支持跨语言适应(ADD22 中文数据集 98.1% AUC / 6.96% EER),参数效率比 full retraining 高 500 倍。
  • Weakness: 在新数据集上绝对效果不可用(FoR EER 11.37%, ITW EER 10.64%),远逊于 full retraining(3.27%, 0.28%)和 [11] 的 CL Mc 策略(5.00%, 2.80%),方法核心是已有组件(adapter+CORAL+PC)的简单组合,且 “first” 声称被多篇同期 adapter-based 音频 deepfake 工作削弱。
  • Full review: Claude Code 全文七公理审稿

Fake speech detectors are increasingly challenged by the development of new and more accurate generative models. To cope with this problem, continual learning techniques are nowadays widely considered feasible strategies for updating models to new datasets, but they also lead to decreased performance on previously seen samples (catastrophic forgetting). In this work, we propose a forgetting-resilient solution based on the adoption of domain translators within a frozen detector, which remaps the new feature spaces into the original ones by means of a traceback translator network. Experimental results show that this strategy enables the achievement of high detection rates with respect to traditional retraining, while minimizing the computational effort and preserving the detection accuracy on previous data.


21. Neural Morphing: Sequence-Optimized Token-Level Morphing in Neural Audio Codecs

Authors: Emmanouil Karystinaios

Categories: cs.SD | In proceedings of the 29th International Conference on Digital Audio Effects (DAFx) 2026 Score: 3.16/10 (Obj:5 Id:2 Ind:5 Comp:4 Eff:2 Nov:3 Rep:5)

  • Strength: 实现了完整的 VST3/AU 实时插件,beam search 将 palette-index jitter 减半(11.52k vs 24.66k),RVQ group 分层提供了可操作的 structure/detail 控制轴,RTF 0.22-0.24 低于实时。
  • Weakness: 未引用或比较直接竞品 Latent Granular Resynthesis (arXiv:2507.19202, 先发表12个月),无任何外部 baseline 对比,仅测试打击乐单一材料,无人类听感评估,核心新颖性存疑。
  • Full review: Claude Code 全文七公理审稿

Neural audio codecs were originally developed for high-fidelity compression; however, their latent token representations and expressive decoders also constitute a powerful substrate for controllable audio transformation. This work introduces Neural Morphing, a training-free token-domain audio effect that selects residual-vector-quantized (RVQ) token grains from a user palette and decodes the edited stream through a pretrained codec. The method combines an RVQ-group transfer policy that separates coarse, middle, and fine codebook groups with a continuity-constrained sequence matcher that replaces independent greedy selection with bounded beam search. The intended output is a controlled hybrid: the source preserves rhythmic organization while the palette contributes timbral color and residual detail. We focus on the implementation and realtime behavior of a deployable VST3/AU system, including chunked rendering, palette-size scaling, and backend health checks.