Daily Papers — 2026-08-14
12 papers on audio, speech, music, and acoustics.
1. What Makes a Good Layer? Assessing the Layer-Wise Intrinsic Properties of Music Foundation Models
Authors: Angelos-Nikolaos Kanatas, Yuexuan Kong, Pablo Alonso-Jiménez, Xavier Serra, Dmitry Bogdanov
Categories: cs.SD, cs.LG, eess.AS | 11 pages, 2 figures, 2 tables. Accepted at ISMIR 2026. Project page: https://angeloskanatas.github.io/music-fms-layer-eval/ Score: 7.30/10 (Obj:9 Id:7 Ind:8 Comp:8 Eff:8 Nov:8 Rep:8)
- Strength: 覆盖 12 个音乐基础模型与 15 个下游任务,逐层内在度量可作层选择代理:top-3 代理层仅距逐任务 oracle 0.4pp,且比穷举层扫描少约 4-16 倍监督 probe,新提出的 PTE 指标在调性任务上取得全模型一致的相关(均值 Spearman ρ̄=0.80)。
- Weakness: 全部下游数字仅单一 seed 且未报告方差与置信区间,三种预训练范式中对比范式仅含 2 个模型,且分析代码承诺在 camera-ready 后才发布,第三方当前无法复跑指标计算。
- Full review: Claude Code 全文七公理审稿
Music foundation models are commonly used as frozen audio feature extractors, yet selecting which layer to extract from remains largely heuristic. Current practice defaults to fixed depths or multi-layer fusion, with limited understanding of why certain layers transfer better across downstream tasks or how representation quality varies with depth and pre-training paradigm. We conduct a systematic layer-wise analysis of 12 music foundation models spanning three pre-training paradigms (masked modeling, autoregressive modeling, and contrastive learning), characterizing their hidden representations through intrinsic geometric and transformation-based properties. Correlating label-free representation-quality metrics with layer-wise performance across 15 downstream tasks, we find that several metrics track layer quality for genre classification, emotion recognition, automatic tagging, and beat tracking, albeit with varying strength across tasks and pre-training paradigms. However, all metrics fail on tonal tasks such as key estimation and chord recognition, indicating that no single property serves as a general proxy for representation quality across music information retrieval tasks. To address this gap, we introduce a pitch-transposition equivariance measure that captures properties missed by these standard metrics, providing a consistent indicator of tonal quality across model families. Finally, we show that intrinsic metrics can serve as effective proxies for layer selection, matching or outperforming trainable multi-layer fusion methods, particularly in limited-data settings.
2. A Parameter-Free Few-Shot Evaluation for Elephant Vocalisation Classification
Authors: Christiaan M. Geldenhuys, Thomas R. Niesler
Categories: eess.AS, cs.LG, cs.SD, q-bio.QM Score: 6.74/10 (Obj:8 Id:8 Ind:9 Comp:9 Eff:7 Nov:5 Rep:2)
- Strength: 在 EV 小数据集上,Perch v2 嵌入的无参数质心分类器从 k=1 起反超全监督 LR 基线(mAP 0.3853 对 0.3654)、从 k=2 起反超循环基线,k=5 达 mAP 0.5415,并用 bootstrap(R=100 抽取、B=5000 重采样)量化为百分位区间。
- Weakness: 未发布任何代码、数据或权重链接,且优势仅在 EV 数据集、mAP 指标上成立,LDC 大数据集(3985 查询段)各 k 均低于训练基线,AUROC 指标也未反超 AERD。
- Full review: Claude Code 全文七公理审稿
We present a parameter-free episodic evaluation of nearest-centroid classification for elephant vocalisations on fixed pretrained acoustic embeddings, across the Elephant Voices (EV) and Linguistic Data Consortium (LDC) datasets. Rather than asking which embedding yields the best classifier when trained on all available labelled data, we ask how the simplest classifier performs as labelled exemplars per class are varied. Each class is represented by the mean of its support-set embeddings, and each query is assigned to the nearest centroid under squared Euclidean distance. We evaluate this centroid classifier on the Perch (ver. 1), Perch (ver. 2), and HuBERT (base, layer 2) embeddings, together with mel frequency cepstral coefficient (MFCC) features, in an N-way k-shot manner under the same cross-validation protocol as the trained baselines. A bootstrap over 100 resampled support sets quantifies the sampling noise. On the smaller, low-resource EV dataset, the centroid classifier using the stronger Perch (ver. 1) and Perch (ver. 2) embeddings overtakes the fully-trained logistic regression classifier from a single exemplar per class and the stronger recurrent classifier from two. Over the reduced set of call types on which the strongly-supervised end-to-end baseline was trained, the centroid classifier matches and then surpasses that baseline in mean average precision (mAP), from a few exemplars per class. On the larger LDC dataset, where labelled exemplars are abundant, the trained baselines retain their advantage at every k considered. At five exemplars per class, the centroid classifier using the strongest embedding, Perch (ver. 2), attains a mAP of 0.542 on the EV dataset and 0.368 on the LDC dataset. Parameter-free nearest-centroid classification is the stronger choice when labelled exemplars are few and the fixed embedding already encodes the features that separate the call types.
3. The MPB Corpus: A Dataset of Melody, Rhythm, Harmony, and Melody-Harmony Relationships in Brazilian Popular Music
Authors: Carlos de L. Almada, Hugo T. de Carvalho, Felipe D. Martins
Categories: cs.SD, cs.DL, cs.IR, eess.AS | 22 pages, 13 figures Score: 6.60/10 (Obj:8 Id:5 Ind:5 Comp:6 Eff:6 Nov:8 Rep:8)
- Strength: 首个四参数巴西流行音乐语料,500 首作品编码出 8,426 个旋律/节奏 word、17,053 个和弦与 23,447 个旋律-和声音符,置换检验证明 r-字母节奏分布可显著区分作曲家(Robs=1.072,p<10^-4)。
- Weakness: 全库由单一标注者手工标注且无 inter-annotator agreement,统计检验仅覆盖节奏一个维度,和声与旋律维度的效用及与既有语料的定量对比均未提供。
- Full review: Claude Code 全文七公理审稿
This paper presents the MPB Corpus, a collection of 500 musical pieces encoded across four musical parameters: melodic contour, melodic rhythm, harmony, and the relationship between melody and harmony. It constitutes the most comprehensive and detailed dataset to date for computational musicology of Brazilian music. To support the encoding process, we introduce specific analytical models designed to capture rhythmic and melodic information with precision, alongside tailored visualizations and metrics summarizing key musical parameters. Finally, we provide a brief qualitative exploratory analysis of the dataset, illustrating its potential to both formulate and systematically address musicological questions concerning the genre.
4. H2H Music Improv: A Communication Model and Audio-Visual Dataset for Music Improvisation
Authors: Aleksandra Teng Ma, Anthony Cammarota, Jiayi Wang, Alexandria Smith, Cheng-Zhi Anna Huang et al.
Categories: cs.SD, cs.MM, eess.AS | Published in the Proceedings of the Society for Music Information Retrieval Conference (ISMIR) 2026 Score: 6.37/10 (Obj:8 Id:5 Ind:5 Comp:8 Eff:6 Nov:7 Rep:6)
- Strength: 首个带双向意图标注的自由即兴影音数据集(6h8m、37 clips、5 位专家乐手、1,368 条标注),实证发现状态对齐中位数 stability 81.3% 对 proposal 23.6%、negotiation 0%,量化验证了意图与感知的不对称。
- Weakness: 缺少标注一致性度量(kappa/alpha)、外部听者验证与替代标注方案对照,且对齐指标计算代码未发布,效用止于资源价值而未见任何下游算法演示。
- Full review: Claude Code 全文七公理审稿
Current real-time AI improvisation systems lack the communication awareness human musicians rely on: rather than treating communication as a foundational algorithm design concern, most systems layer interaction strategies post-hoc onto generative algorithms through explicit controls and predefined modes. This gap persists in part because no formalized, machine-readable communication model with musicians exists. To address this, we study how expert musicians communicate in free (non-idiomatic) improvisation, unconstrained by prior discussion or agreement. Through a collaborative co-design process with expert improvisers, we derive a communication model that (1) captures how free improvisers negotiate musical ideas and enter stable musical spaces, and (2) is formalized as a machine-readable annotation scheme. We further present the H2H (Human-to-Human) Music Improvisation dataset: six hours of audio-visual expert duo improvisations with clean per-player stems and per-player annotations of both their own intentions and their perception of their partner’s intentions. To our knowledge, this is the first such dataset for free improvisation. Together, the communication model and the dataset offer a new lens and resource for studying musician communication and may in future inform the design of AI musical partners that communicate by design.
5. AT-ADD: All-Type Audio Deepfake Detection Challenge Summary
Authors: Yuankun Xie, Haonan Cheng, Jiayi Zhou, Xiaoxuan Guo, Tao Wang et al.
Categories: cs.SD | Accepted to ACM MM 2026 Score: 6.32/10 (Obj:9 Id:6 Ind:8 Comp:6 Eff:6 Nov:6 Rep:4)
- Strength: 提供了首个跨四类音频(speech/sound/singing/music)的 closed-setting 深度伪造检测基准,26 个 unseen 生成器评测设计扎实,排行榜显示 T1 冠军 Macro-F1 90.71%、T2 冠军 96.10%。
- Weakness: 全篇无官方/minimal baseline 分数、无 T2 分音频类型细分表现、数据集与代码未提供公开下载链接,复现与效用证据均缺失。
- Full review: Claude Code 全文七公理审稿
This paper summarizes the ACM Multimedia 2026 AT-ADD Grand Challenge on all-type audio deepfake detection. AT-ADD contains two tracks: robust speech deepfake detection under realistic acoustic and channel variations, and type-agnostic detection over speech, environmental sound, singing voice, and music. We describe the challenge tasks, dataset and evaluation-set design, official leaderboard results, and common design patterns observed in participating systems. The best Track 1 system achieved 90.71% Macro-F1 on the final evaluation set, while the best Track 2 system achieved 96.10% Macro-F1. The final submissions show that strong systems commonly combine large-scale self-supervised audio representations, data augmentation, multi-crop inference, and structured fusion or routing. The results also reveal remaining challenges in generalization to unseen generators, robustness to realistic speech-domain distortions, and balanced performance across heterogeneous audio types.
6. Separate First, Then Associate: A Two-Stage Approach for Real-World Audio-Visual Speech Enhancement
Authors: Tongtao Ling, Zhong-Qiu Wang
Categories: eess.AS Score: 6.32/10 (Obj:8 Id:6 Ind:8 Comp:7 Eff:6 Nov:5 Rep:6)
- Strength: 解耦的两阶段”分离-再关联”框架在 ISCSLP 2026 Real-World AVSE Challenge 上相对官方 AV-ConvTasNet baseline 取得 Track 1 SI-SDR 从 -5.9 到 10.4 dB(+16.3 dB)、CER 从 102.5% 到 16.4%、SPK-SIM 从 0.328 到 0.737 的显著提升。
- Weakness: 在官方 OVRL 排名中位列非 baseline 参赛系统末位(Track 1 5.43、Track 2 5.29),且未提供 oracle/随机关联或耦合系统的消融对比,无法定位分离器与关联器各自的性能贡献。
- Full review: Claude Code 全文七公理审稿
Audio-visual speech enhancement (AVSE) aims at extracting target speech from multi-speaker mixtures by exploiting visual cues. Although recent studies have reported strong performance on simulated datasets, the performance, however, often drops dramatically when they are applied to real-world audio-visual recordings. To bridge this gap, the Real-World AVSE Challenge held in the ISCSLP 2026 conference calls for participants to design a practical solution for AVSE under real-world conditions, where speaker overlap, acoustic interferences, room reverberation and visual degradations naturally co-exist. In our submission to the challenge, we propose a decoupled separation-then-association approach. It consists of two stages: a separation stage in which a trained, audio-only model (i.e., not using visual cues) is used to separate input multi-speaker mixture to individual speaker signals, followed by an association stage, where an audio-visual CLIP model is used to identify the separated speech signal with the highest similarity with the target speaker’s facial video via cross-modal similarity matching. Evaluation results on the challenge dataset show the effectiveness of our proposed approach.
7. Singer-Informed Vocal Source Separation for Multi-Singer Music Mixtures
Authors: Jocelyn Xu, Minje Kim
Categories: eess.AS, cs.SD | Accepted at IWAENC 2026 Score: 6.26/10 (Obj:8 Id:5 Ind:6 Comp:8 Eff:6 Nov:5 Rep:8)
- Strength: 提出 enrollment 驱动的 singer embedding + FiLM/Concat 条件化方案,在二重唱测试中目标歌手 SI-SDR 从 0.33 dB 提升到 5.58 dB、FAD 从 2.28 降至 0.37,并开源完整代码与预训练权重。
- Weakness: 仅对比未具备目标歌手提取能力的 Open-Unmix 基线,缺少与 UNMIXX、MedleyVox 等 SOTA 多歌手分离方法的公平比较、显著性检验与独立基准验证。
- Full review: Claude Code 全文七公理审稿
Music source separation systems typically extract a single vocal track and do not distinguish between multiple singers. We study singer-informed vocal source separation for multi-singer mixtures. Our framework introduces a short enrollment recording of a target singer to guide separation through a learned embedding. The singer embedding is incorporated using feature concatenation or feature-wise linear modulation (FiLM), enabling the model to focus on the target singer while suppressing interference. We construct a duet dataset based on DAMP-VSEP with quality filtering and non-overlapping enrollment segments. Experiments on solo and duet settings show that while baseline models perform well for single-singer mixtures, the proposed method improves target-singer extraction in multi-singer cases, increasing target-singer SI-SDR from 0.33 dB to 5.58 dB. Fréchet Audio Distance (FAD) further shows improved perceptual quality and better alignment with target audio distributions. Code and checkpoints are available at https://github.com/jocelynxu01/singer-separation-paper.
8. Ambisonics Encoding of Room Impulse Responses using a Device-Agnostic Diffusion Model
Authors: Eloi Moliner, Christoph Hold, Juan Azcarreta Ortiz, Sebastian Prepelita, Ishwarya Ananthabhotla et al.
Categories: eess.AS, cs.SD | IWAENC 2026 Score: 5.79/10 (Obj:8 Id:6 Ind:5 Comp:5 Eff:6 Nov:6 Rep:3)
- Strength: 扩散先验 + DPS 后验采样在 FDTD 与 Treble-10 上以 EDC 0.018/0.019、NPM 0.874/0.803 一致优于线性(0.049/0.106、0.999/1.042)与神经编码器基线,仿真听测中 13 名听者对旋转后渲染仍接近参考。
- Weakness: 全部客观评测由同一 ATF 前向模型仿真闭环生成,真实设备端到端评测缺失,实测 Eigenmike 听测中与设备特定条件扩散无显著差异,且无代码/数据开源支撑可复现。
- Full review: Claude Code 全文七公理审稿
We address the problem of encoding room impulse responses (RIRs) into high-order Ambisonics (HOA) representations from arbitrary and potentially insufficient or incomplete microphone array measurements. This task is fundamentally ill-posed for microphone arrays with limited spatial capture capabilities, such as irregular or sparse arrays, as classical linear methods fail to reconstruct high-order spatial detail. We introduce a diffusion-based generative framework that models the statistical properties of HOA RIRs. This enables device-agnostic encoding from arbitrary microphone arrays, potentially unseen during data measurement. Our approach incorporates a posterior sampling procedure that enforces consistency between the estimated signals and the measurements while plausibly reconstructing spatial information that is unobservable from the limited measurements alone. Experiments on simulated data demonstrate that our method outperforms linear and neural baselines, achieving accurate HOA RIR estimation up to 12th order. A listening test with binaural renderings, including both simulated and measured RIRs, further confirms that the proposed method yields higher perceptual similarity to reference Ambisonics RIRs than all baselines. The flexibility and accuracy of the proposed framework opens new possibilities for scalable acoustics simulations.
9. Beyond the pale: Assessing prevalence and contents of extremist speech in LLM training data
Authors: Dmitry Nikolaev, Ashley A. Mattheis
Categories: cs.CL | Accepted to the CPSS workshop @ KONVENS 2026 Score: 5.74/10 (Obj:8 Id:5 Ind:4 Comp:6 Eff:6 Nov:7 Rep:3)
- Strength: 采用三定义多模型保守流水线与专家编码,给出 Dolma 20 万文档样本中极端主义内容约 1/2000(约 100 篇)的下界估计,按 10 倍高估容错外推至约 50 亿文档语料仍达数十万篇潜在有害文档。
- Weakness: 缺乏基于黄金标注集的精确率/召回率与关键词最小基线对照,独立性仅靠 30 篇人工编码支撑,且抽样脚本、推理流水线与完整抽取结果均未公开发布。
- Full review: Claude Code 全文七公理审稿
Despite a strong interest on the part of the research community in the topic of trustworthy and safe AI, the composition of the text corpora that large language models (LLMs) encounter in pre- and post-training has not yet drawn much attention. In this work, we address the question of whether LLMs are exposed to unfiltered, uncontextualised extremist speech. Using several definitions of extremist speech, stemming from official documents and research literature, and an extraction pipeline combining automated text processing with expert verification, we provide a lower bound on the prevalence of extremist documents in Dolma, an open training corpus underpinning the OLMo series of models. We show that Dolma is likely to include hundreds of thousands of documents containing extremist content and hate speech of several types, including direct calls for violence, and discuss the implications of this for data curation and model pre-training.
10. S2Dialog: Multimodal Dialogue Retrieval with Semantic and Acoustic-Style Modeling
Authors: Xueqi Wang, Zhigang Wang, Runqing Zhang, Zhenqi Jia, Junfeng Zhao
Categories: cs.CL Score: 5.70/10 (Obj:8 Id:5 Ind:6 Comp:6 Eff:6 Nov:5 Rep:5)
- Strength: 在 DailyTalk 上以对话级多模态检索取得 R@10=50.68%(最好 baseline 25.60%)与 R@50=83.56%(best baseline 68.40%),消融证实 Bottom-K 硬负样本(移除后 R@10 降至 21.12%)与声学分支均贡献显著。
- Weakness: 评测 GT 由与模型骨干相同编码器(Sentence-BERT + Wav2Vec2-IEMOCAP)自动构造并经人工重排,标签同源导致 R@K 提升可能被高估,且仅单一数据集、无下游任务验证,代码仓库仅有占位 README 未发布代码。
- Full review: Claude Code 全文七公理审稿
Multimodal dialogue retrieval aims to retrieve dialogues from multimodal dialogue banks that are similar to a target dialogue in terms of both textual semantics and acoustic conversational styles. Such dialogue-level retrieval is crucial for many dialogue-related tasks, including Emotion Recognition in Conversation, Spoken Dialogue Systems, and Conversational Speech Synthesis, where external dialogue examples can provide valuable semantic and stylistic references. However, existing retrieval methods are still largely limited to utterance-level or unimodal matching, and often fail to capture the global semantic coherence and stylistic consistency of an entire dialogue. To address this gap, we propose S2Dialog, a unified framework for dialogue-level semantic-style retrieval from multimodal dialogue banks. Specifically, S2Dialog consists of a Dialogue-level Textual Retriever and a Dialogue-level Acoustic Retriever, which encode the textual and acoustic modalities of a dialogue into dialogue-level representations, respectively. To further enhance multimodal retrieval, we introduce Dialogue-level Textual-Acoustic Contrastive Learning, which aligns semantically and stylistically similar dialogues while distinguishing unrelated ones. Extensive experiments on the multimodal dialogue dataset DailyTalk demonstrate that S2Dialog achieves outstanding retrieval performance.
11. Distinguishing AI-Generated Music from Edited Audio as a Hard-Negative Robustness Task
Authors: Alexandru-Stefan Morosanu, Valerian Cecan, Stefan-Daniel Achirei, Laura Erhan
Categories: cs.SD, cs.AI, cs.LG | Accepted at the RobustifAI 2026 Workshop @IJCAI-ECAI 2026, Bremen Score: 5.63/10 (Obj:7 Id:3 Ind:7 Comp:8 Eff:5 Nov:7 Rep:3)
- Strength: 视频级系统在锚点分组留出测试集上达 0.811 balanced accuracy / 0.813 macro F1,其中 AI 类 clip F1 0.836 显著高于编辑类 0.720,量化证明了编辑音频作为硬负类会造成真实混淆。
- Weakness: 基线仅两个轻量方案(RF 0.52、MobileNetV2 0.601 bal. acc),未与 CLAP/PANNs/AST 或任何现役检测器对比,且数据集与代码均未公开,0.811 的结果无法独立复现。
- Full review: Claude Code 全文七公理审稿
AI-generated music detectors are commonly evaluated against original songs, but real-world uploads are often remixed, re-encoded, pitch-shifted, or otherwise edited. These edited versions form a difficult negative class: they are not generated by AI, yet they may introduce spectral artifacts that resemble synthetic audio fingerprints. We study this problem as a hard-negative robustness setting for AI-generated music detection, focusing on AI-generated and edited variants derived from the same anchor songs. We compile a YouTube-based dataset of AI, edited, and original variants, using the original tracks only as references, and train a binary AI versus edited detector. Audio is processed as 10-second clips and passed as raw waveforms to a pretrained PaSST spectrogram transformer. To reduce leakage, all splits are performed by anchor song. On the held-out test set, the final video-level system achieves 0.811 balanced accuracy. At clip level, AI-generated clips reach an F1-score of 0.836, while edited clips reach a lower F1-score of 0.720. The results suggest that AI-generated music retains detectable fingerprint-like spectral cues beyond ordinary editing, but the lower edited-class F1-score shows that these cues can still overlap with artifacts from edited audio. Grad-CAM visualizations are used to inspect whether high-confidence predictions rely on localized time-frequency regions.
12. Acoustic UAV Detection in Battlefield Scenarios: Handling Noise, Domain Shift, and Weak Labels
Authors: Vadym Vilhurin, Volodymyr Sydorskyi, Andrii Shevtsov
Categories: cs.SD, cs.AI, cs.CV Score: 5.58/10 (Obj:8 Id:6 Ind:6 Comp:6 Eff:6 Nov:5 Rep:2)
- Strength: 在乌克兰战场真实录音数据集上,测试集 Small Drone 检测 F1 从 AST-Drone 的 55.4% 提升至 small 集模型的 81.14% 与 full 集双模态集成的 82.22%(超越 Zvook 专有基线 69.10%),且使用约 5 倍更少的目标域样本。
- Weakness: 战场数据集与全部模型权重/代码均未发布,且缺少最简谱特征/SVM 与跨麦克风对齐(CycleGAN、特征对齐)等基线对比,无法由第三方独立复现或验证。
- Full review: Claude Code 全文七公理审稿
Passive acoustic sensing offers a critical, cost-efficient, and, crucially, passive alternative for detecting small unmanned aerial vehicles. However, the practical deployment of acoustic systems is discouraged by extreme environmental noise and sensor-induced domain shift caused by heterogeneous hardware. This paper addresses these challenges by introducing a robust framework optimized for real-world battlefield conditions. We propose the integration of Per-Channel Energy Normalization (PCEN) and attention-based pooling to enhance feature extraction under low signal-to-noise ratio scenarios. We further propose a domain-aware training strategy that leverages auxiliary classes and multi-microphone data to mitigate cross-domain performance degradation. Evaluated on a unique dataset of combat-zone recordings from the Ukrainian frontlines, our approach significantly outperforms existing baselines, increasing the F1 score from 55.4% to 78.6%. This paper was originally presented at the International Conference on Military Communication and Information Systems (ICMCIS), organized by the Information Systems Technology (IST) Scientific and Technical Committee, IST-224-RSY - the ICMCIS, held in Bath, United Kingdom, 12-13 May 2026.