每日 arXiv

音频 · 语音 · 音乐 · 声学

每日新爬取的 arXiv 论文,按评分排序。中文摘要帮你快速判断是否值得深读。

2026-07-16
日期2026-07-16
已评分
均分
最高

Daily Papers — 2026-07-16

8 papers on audio, speech, music, and acoustics.

1. What does the model actually see? Evaluation protocols and input availability in data-driven prediction of room acoustic parameters

Authors: Akın Oktav

Categories: eess.AS, cs.SD | 12 pages, 4 figures. Submitted to Acta Acustica Score: 7.70/10 (Obj:9 Id:9 Ind:9 Comp:8 Eff:6 Nov:8 Rep:6)

  • Strength: 通过 2×3 因子协议消融在两个厅堂上量化了协议选择对 R² 的决定性影响(跨协议 spread up to 0.75 vs 跨模型 spread ~0.1),并实验识别了 hybrid CNN 的 position fingerprint 机制(K3 sign reversal: +0.11 at 75 positions vs −0.23 at 10 positions)。
  • Weakness: 证据基础限于两个厅堂单 campaign,honest 协议下 6 个核心参数中 4 个 R² < 0.25 且低于最简空间插值,数据非完全开源,混合 CNN 训练超参数未报告。
  • Full review: Claude Code 全文七公理审稿

Machine-learnt models are increasingly used to predict ISO 3382-1 room acoustic parameters from sparse measurements, with reported coefficients of determination frequently above 0.85. This paper shows that such figures are often determined by the evaluation protocol rather than by the model. Using a multi-condition measurement campaign in a 264-seat conference hall and a 180-seat concert hall, three model families were evaluated under a factorial protocol ablation: validation splits either row-based or grouped by receiver position, and input features either including measured-at-test quantities or restricted to source-receiver geometry and environmental state. Row-based splits with measured-at-test inputs reproduce the high reported accuracies (mean $R^2$ 0.81 for the core parameters); grouping the splits by position and restricting inputs to information available at an unmeasured position reduces these to 0.09-0.57, reordering the apparent difficulty of parameter classes. A hybrid CNN evaluated with the target’s own impulse response as input is shown to exploit it as a position fingerprint rather than as transferable acoustic information; training-only signal access yields no gain for any parameter tested, including reverberation time. Under the deployment-consistent protocol, the spread between Random Forest, the hybrid CNN, and inverse-distance weighting is an order of magnitude smaller than the spread between protocols for a fixed model; the learnt models retain a genuine advantage for sound strength and reverberation time, and the high accuracy of the original pipelines re-emerges as condition interpolation at measured positions (band means 0.80-0.88), a distinct and operationally useful task.


2. SLT 2026 REAL-TSE Challenge: Real-world Target Speaker Extraction from Conversational Recordings

Authors: Shuai Wang, Zihan Qian, Ke Zhang, Jiangyu Han, Zikai Liu et al.

Categories: eess.AS, cs.SD | Overview paper of Real-TSE Challenge Score: 6.30/10 (Obj:9 Id:5 Ind:5 Comp:8 Eff:6 Nov:6 Rep:9)

  • Strength: 首次在真实双语(Mandarin+English)对话录音上组织 TSE challenge,24 系统/18 团队的聚合分析揭示 “public neural metric 可被 over-optimize”(DNSMOS-OVRL 与人类 MOS LCC 仅 0.165)这一可迁移经验规律。
  • Weakness: 评测闭环在发布时已被参赛者污染(对抗扰动 inflate 分数、branch selection by automatic score),基线仅用弱 BSRNN 且 DNSMOS over-optimization 只做 post-hoc 人类校验而非全程独立锚点。
  • Full review: Claude Code 全文七公理审稿

We introduce the REAL-TSE Challenge, an IEEE SLT 2026 satellite challenge on target speaker extraction~(TSE) from real conversational recordings. Given a multi-speaker mixture and one or more enrollment utterances from a target speaker, participating systems must recover only the target speech. Unlike simulated read-speech benchmarks, REAL-TSE evaluates Mandarin and English recordings that contain natural overlap, reverberation, noise, channel mismatch, and conversational dynamics. The challenge defines two complementary tracks: an Online track for low-latency streaming extraction and an Offline track for full-context processing. Systems are evaluated with Token Error Rate (TER), Speaker Similarity (SpkSim), DNSMOS, and target-speaker activity F1. This overview paper describes the task definition, datasets, baselines, evaluation protocol, submitted systems, condition-wise findings, and lessons for future real-world TSE benchmarks.


3. SceneBind: Binding What and Where Across Vision, Audio and Language

Authors: Mingfei Chen, Zijun Cui, Ruoke Zhang, Hyeonggon Ryu, Eli Shlizerman

Categories: cs.CV, cs.AI, cs.MM, cs.SD | Project website: https://scenebind.github.io/ Score: 6.20/10 (Obj:8 Id:6 Ind:5 Comp:8 Eff:8 Nov:5 Rep:6)

  • Strength: 提出将全局语义嵌入与 object-centric 语义-空间 slot 结合的全模态场景表示,在 spatial retrieval 上取得 +153% R@1 / +62% mAP 的显著提升,且零样本超越 AVLoc SOTA (cIoU@0.2: 52.65 vs. 38.71)。
  • Weakness: 训练数据空间标注依赖闭源 Gemini + baseline 模型 (ImageBind/CLAP) 做验证(循环风险),object grounding 绝对精度仅 ~20%,各组件 (DETR query, bipartite matching, InfoNCE) 均为已有技术组合,缺乏组件级机制创新。
  • Full review: Claude Code 全文七公理审稿

We present SceneBind, an omni-modal representation of realistic scenes with joint semantic and 3D spatial understanding across vision, audio and language. Existing omni-modal encoders excel at instance-level semantics (i.e., what is present), but often lack explicit spatial structure (i.e., where it is). SceneBind addresses this gap by representing each scene as a semantic-spatial entity, combining a global semantic embedding with object-centric semantic-spatial slots. This representation explicitly captures object-level semantics, spatial attributes, and uncertainty. We further propose SceneBind Matching, a semantic-spatial matching scheme that integrates global scene similarity with object alignment, supporting cross-modal scene retrieval and object grounding. To train and evaluate SceneBind, we curate a novel real-world binaural audio-visual dataset with structured semantic and spatial annotations, and propose a training protocol for aligning semantic and spatial signals across modalities. SceneBind is compatible with large-scale pretrained semantic encoders, adds lightweight spatial modeling with only a few additional tokens. It achieves state-of-the-art scene and spatial retrieval while enabling strong zero-shot transfer to downstream tasks such as audio-visual localization.


4. RW-Voice-EQ Bench: A Real World Benchmark for Evaluating Voice AI Systems

Authors: David Ayllon, Alice Baird, Jeffrey Brooks, Franc Camps-Febrer, Jakub Piotr Cłapa et al.

Categories: cs.SD, cs.AI | Benchmark and leaderboard: https://huggingface.co/spaces/HumeAI/rw-voice-eq Score: 6.00/10 (Obj:8 Id:6 Ind:5 Comp:8 Eff:6 Nov:6 Rep:6)

  • Strength: 跨 TTS/STS/SU/ASR 四领域的统一多维 benchmark(31 TTS + 15 STS + 18 SU + 40 ASR 系统,80 万+人类评分),发现性能高度维度特异性,并设计了 4 种 ASR benchmark contamination 诊断方法。
  • Weakness: 关键分析标注为 preliminary 且无统计显著性检验,SU ground truth 依赖 Hume 自有模型标注存在独立性风险,评测 prompt/rubric/代码未完全开源,STS ablation 仅 19 个 scenario 统计效力不足。
  • Full review: Claude Code 全文七公理审稿

Current voice AI benchmarks typically evaluate isolated capabilities such as speech intelligibility, word error rate, or text-based dialogue quality, but they rarely test whether systems harness the acoustic information that distinguishes spoken language from its textual representation. To this end, we introduce the Real World Voice EQ Bench, a multidimensional benchmark for evaluating voice AI across text-to-speech (TTS), speech-to-speech (STS), speech understanding (SU), and automatic speech recognition (ASR). Our evaluations indicate that performance is highly dimension-specific. For TTS, naturalness, expressiveness, identity stability, and reliability are largely independent evaluation dimensions. For STS, access to audio does not guarantee use of vocal affect, and some agents remain largely transcript-driven. For SU, models perform unevenly across paralinguistic tasks. For ASR, real world accent, emotion, noise, and conversational conditions expose failures that are not captured by established clean-speech benchmarks. Together, these results show that voice AI should be evaluated as a profile of acoustic, expressive, interactional, and robustness capabilities rather than by a single aggregate score.


5. Can Tokens Compete? Token Representations against Supervised CNN Backbones for BirdCLEF+ 2026

Authors: Anthony Miyaguchi, Murilo Gustineli, Adrian Cheung

Categories: cs.SD, cs.AI, cs.LG Score: 4.79/10 (Obj:8 Id:5 Ind:8 Comp:5 Eff:4 Nov:4 Rep:9)

  • Strength: 诚实的负面结果记录(20+ 单变量实验、编解码器失败、SSL 不确定)和完全开源的代码仓库,复现性强。
  • Weakness: 核心发现与已有文献一致(specialist > generalist),rank 1894 不具竞争力,token 失败的因果机制未隔离(同时改变架构、预训练域和参数量)。
  • Full review: Claude Code 全文七公理审稿

This paper details the DS@GT ARC team’s approach to BirdCLEF+ 2026, multi-label detection of animal vocalizations in soundscapes from the Pantanal wetlands. The 2026 edition adds about an hour of labeled soundscapes, shifting the task toward supervised pipelines fit to the labeled set. First, we build a competitive supervised baseline that ensembles a frozen Perch v2 backbone, a trained HGNetV2-B0 sound-event-detection network, and a non-bird prototypical head, reaching a private leaderboard score of 0.936 at rank 1894 within a 90-minute CPU budget. Second, we ask whether token-based representations can compete, contrasting codec representations from neural audio codecs against semantic representations from foundational embeddings. We compare two bioacoustic specialist models against four token-based encoders trained on AudioSet. The repository for this work can be found at https://github.com/dsgt-arc/birdclef-2026.


6. Large Audio Language Models for Spoofing-Aware Speaker Verification

Authors: Sofya Savelyeva, Mariia Perunova, Evgeny Kushnir, Artem Dvirniak, Dmitrii Korzh et al.

Categories: cs.SD, cs.AI Score: 4.74/10 (Obj:6 Id:5 Ind:5 Comp:4 Eff:5 Nov:4 Rep:5)

  • Strength: 首次系统评估 LALM 用于 SASV,LoRA+复合损失适配后 SALMONN-7B 在 ASVspoof5 20K 子集上 Acc 89.30%、a-DCF 0.19 优于重新实现的融合 baseline (a-DCF 0.31-0.41),并提供了 ASV-CM trade-off 的增量消融。
  • Weakness: 方法是已有技术(LoRA/AAM/BCE/GRPO/CoT)的组合无新机制;EER (0.22) 差于 baseline (0.13-0.14) 未解决;评测仅 20K 采样子集非完整集;CoT/GRPO 未超过硬标签 SFT;理由质量未经人类验证;代码未开源、CoT 标注依赖闭源 Step-Audio R1。
  • Full review: Claude Code 全文七公理审稿

Recent advances in text-to-speech and voice cloning make high-quality spoofing inexpensive and scalable, threatening voice authentication systems, especially automatic speaker verification (ASV). Existing defenses mainly address this threat through binary countermeasures (CMs) for deepfake detection or spoofing-aware speaker verification (SASV), where current systems are dominated by modular ASV-CM fusion and cascaded pipelines. Although large audio language models (LALMs) have shown promise on related audio tasks, including CM and ASV, their use for SASV remains unexplored, despite their capacity to produce natural-language rationales for auditing and robustness beyond discriminative predictions. This work systematically evaluates LALMs for SASV against conventional pipelines under zero-shot prompting, supervised adaptation, reasoning-oriented training, and reinforcement-learning-based optimization. Our results show that pretrained LALMs are near chance in the zero-shot setting, confirming that they are not natively suited to SASV, but that task-specific adaptation closes this gap. We further find that competitive SASV performance can be achieved through several distinct routes. These findings position LALMs as a promising and auditable foundation for unified SASV, while clarifying where conventional cascade systems still lead.


7. WanSong v1.0 Technical Report

Authors: Binghui Chen, Pandeng Li, Yu Liu, Jingren Zhou

Categories: eess.AS, cs.CV | Wan Team Score: 4.37/10 (Obj:8 Id:5 Ind:3 Comp:5 Eff:5 Nov:4 Rep:2)

  • Strength: 在 25B 参数 / 6M 小时数据规模上验证了纯扩散端到端歌曲生成路线,PER 7.43% 显著优于 Suno V5 (22.80%) 和 LeVo (27.11%),双音轨设计解决了 CFG 下 vocals/BGM 平衡问题。
  • Weakness: 核心优势指标(musicality 5.49)依赖自训练 judge 模型,存在循环论证;缺少 ACE-Step/DiffRhythm 2 等同期扩散 baseline 对比;消融实验不足以隔离各组件贡献;代码/模型/数据/benchmark 均未开源,几乎完全不可复现。
  • Full review: Claude Code 全文七公理审稿

Music generation foundation models have recently attracted significant industry attention. However, achieving efficient generation and high-fidelity long-form audio while supporting controllability remains challenging. To address these needs, we present \textbf{WanSong}, a simple yet powerful approach for long-form, commercial-grade song generation. Unlike autoregressive (AR) and cascaded multi-stage pipelines (\eg, AR followed by diffusion), \textbf{WanSong} is a pure diffusion-based model that directly generates high-fidelity, multilingual songs up to 5 minutes and outputs dual stems (vocals and background music) in a single run. In addition, our diffusion framework enables faster inference through step-distillation, and offers an efficient pathway for fine-tuning and customization to support downstream editing tasks.


8. MIDI-RAE-JEPA: Hierarchical Representation Learning and Generation for Symbolic Music

Authors: Scott H. Hawley

Categories: cs.SD, cs.LG, eess.AS | 8 pages, 8 figures Score: 3.16/10 (Obj:5 Id:2 Ind:4 Comp:4 Eff:2 Nov:3 Rep:8)

  • Strength: 完整的 encoder-decoder-generation pipeline 在消费级 GPU 上实现,重构 F1=0.995,等变性曲线单调递增验证了 pitch/time equivariance,代码开源可复现。
  • Weakness: 核心 novelty claim 被 OpenReview lieErtGZb6(JEPA for symbolic music)直接挑战;未与 MusicBERT/MidiBERT-Piano/MUPT/Aria 等强 baseline 比较;多组件拼装缺乏充分 ablation 隔离;EMOPIA 4-class accuracy 0.488 未显著超越简单 baseline;生成质量无定量评估。
  • Full review: Claude Code 全文七公理审稿

Rich internal representations of musical structure are essential for music understanding tasks such as machine-assisted music co-writing, yet self-supervised approaches for symbolic music representation remain underexplored, particularly those that encode the hierarchical multiscale nature of musical structures. We present MIDI-RAE-JEPA, combining a pitch- and time-shift equivariance objective with LeJEPA and a Swin Transformer V2 encoder to learn such hierarchical representations of symbolic music encoded as piano roll images. The time-shift equivariance objective encourages the model to internalize temporal musical relationships. The encoder is trained purely on self-supervised objectives – including a masked embedding predictor (MEP) – with collapse prevented via SIGReg. A separate decoder trained on the frozen encoder embeddings achieves reconstruction F1 of 0.995, and a flow matching generative model conditioned on those embeddings produces generations that closely match the pitch register and rhythmic density of the conditioning excerpt, while mismatched conditioning yields unrelated but musically plausible output. Learned representations outperform a Haar scattering transform baseline on a downstream emotion classification task, and embedding distances increase monotonically with pitch and time shift magnitude, confirming measurable equivariance. These results suggest that equivariance-based SSL objectives, combined with sufficient fine-level encoder capacity, provide a viable path toward semantically rich, generatively useful representations of symbolic music.