每日 arXiv

音频 · 语音 · 音乐 · 声学

每日新爬取的 arXiv 论文,按评分排序。中文摘要帮你快速判断是否值得深读。

2026-07-18
日期2026-07-18
已评分
均分
最高

Daily Papers — 2026-07-18

10 papers on audio, speech, music, and acoustics.

1. Falsification-Based Verification of LLM-Generated Optimization Models: Sound Test Batteries and Their Detection Limits

Authors: Haifeng Li, Mo Hai

Categories: cs.SE, cs.AI Score: 7.70/10 (Obj:9 Id:8 Ind:9 Comp:9 Eff:7 Nov:8 Rep:8)

  • Strength: 从 LP 对偶、比较静态和多面体极限导出六类 sound 结构测试,0.0% 假阳性 vs ReLoop 阈值法的 54.9%,并在 326 个 NL4OPT faithful seeds 上复现了理论的 class-by-error detectability matrix 包括其零点。
  • Weakness: 作为选择器在 4 个 benchmark 中 3 个未超过 plain majority voting,且全部实验仅基于单一 Qwen2.5-7B generator,三个 benchmark pool 使用 prefix subset,utility 泛化性证据不足。
  • Full review: Claude Code 全文七公理审稿

Large language models now translate natural-language descriptions of decision problems into solver-ready optimization models, but they fail silently. A generated model often runs and still formulates the wrong problem. This paper develops a theory of falsification-based verification for this setting. Every numeric quantity in the description is a typed slot, and a candidate model is tested only through solver calls on slot-transformed instances; no reference model or label is consulted. From duality, comparative statics, and polyhedral limit arguments we derive a battery of test classes covering directions, curvature, crush probes, prohibitive limits, annihilation, and exchange. Every test is sound, so a violation certifies unfaithfulness and the false-positive rate is zero by design. We characterize what such verification can never see, give conditions under which the canonical error classes are detected with certainty, and prove that no fixed-threshold perturbation tester is simultaneously sound and nontrivial. Experiments on 326 ground-truth models from NL4OPT and four benchmark families confirm the theory. The battery attains a 0.0% false-positive rate against 54.9% for a threshold tester, detects 70.0% of certified conditional-class mutants, convicts 40.4% of the mutants invisible to execution-accuracy scoring, and reproduces the predicted detectability pattern including its zeros.


2. HARP: Harmonic-Aware Residual Partitioning for Neural Audio Codecs

Authors: Qiaoyu Yang, Lixing He, Binyue Deng, Weifeng Zhao

Categories: cs.SD | Accepted to Interspeech 2026 Score: 7.20/10 (Obj:9 Id:9 Ind:9 Comp:8 Eff:6 Nov:6 Rep:9)

  • Strength: HARP 仅修改训练损失即实现 RVQ 频率层级化,同架构下 SI-SDR 提升 +0.25–0.96 dB,MUSHRA 低比特率 +9 分且 IQR 显著收窄,推理零开销。
  • Weakness: 仅比较 2 个 baseline(DAC+BSCodec),缺 HiFi-Codec GRVQ、MUFFIN/MBS-RVQ、WavTokenizer 等直接相关先验的实验对比,且 BSCodec 在 MUSHRA 中因 checkpoint 不稳定被排除。
  • Full review: Claude Code 全文七公理审稿

Neural audio codecs with residual vector quantization (RVQ) normally treat all frequencies uniformly, so their codebooks become spectrally entangled. Truncating stages then removes an unpredictable mix of frequencies. Parallel band decomposition addresses this by splitting audio into independent bands, but fragments the latent space and loses cross-frequency coherence. We introduce HARP (Harmonic-Aware Residual Partitioning), a training strategy that partitions RVQ stages into frequency-ordered groups where each group refines its target band while the decoder retains access to all lower frequencies. Overtones are reconstructed in the context of their fundamentals, preserving coherence that parallel methods lose. HARP requires no architectural changes; it only modifies the training loss, leaving inference identical to standard RVQ. On speech, music, and general audio, HARP outperforms both standard RVQ and parallel decomposition. MUSHRA listening tests also show perceptual improvements.


3. Do Speech Tokens Leak Voiceprints? Speaker Inversion Attacks Against End-to-End Speech Language Models

Authors: Ye Lu, Yihan Yan, Zhaoyang Zhang, Zhitao Ou, Runze Liu et al.

Categories: cs.SD, cs.AI, cs.CR Score: 6.74/10 (Obj:8 Id:5 Ind:6 Comp:5 Eff:8 Nov:6 Rep:6)

  • Strength: 首次将 speech LM frontend tokens 的 voiceprint 泄露形式化为 Speaker Inversion Attack,在 4 个代表性 frontend 上用 3 秒 token 切片实现 CosSim >0.70、EER 0.017-0.038 的 speaker embedding recovery,问题定义清晰且实验覆盖广。
  • Weakness: SpInv 方法是 VICReg+DINO+ArcFace+distillation 的组合,无组件级 ablation 证明各模块必要性,缺少最简 baseline 对比;attack training 数据与 target encoder 训练数据可能重叠未讨论;代码未开源。
  • Full review: Claude Code 全文七公理审稿

End-to-end speech language models increasingly represent user speech with speech tokens rather than relying exclusively on cascaded ASR–LLM–TTS pipelines. Although these tokens support expressive and low-latency spoken interaction, they may also preserve sensitive speaker characteristics. We investigate whether exposed speech tokens leak voiceprints and formulate this risk as a speaker inversion attack. We introduce Audio BERT (AuB), a trainable model that constructs token embeddings from discrete codebooks and aggregates them into speaker-sensitive representations, and propose SpInv, a two-stage inversion method built on AuB to recover embeddings in the space of an attacker-specified speaker encoder. We evaluate Moshi, Higgs3, Kimi-Audio, and Qwen3-Omni using speaker-disjoint protocols on the VoxCeleb dataset. Extensive experiments show that, with only three seconds of frontend output, SpInv achieves cosine similarities above 0.70 in the attacker-specified speaker-encoder space.


4. Pseudo-label distillation for discriminative anomalous sound detection

Authors: Takuya Fujimura, Tomoki Toda

Categories: eess.AS Score: 6.74/10 (Obj:8 Id:6 Ind:8 Comp:6 Eff:8 Nov:5 Rep:8)

  • Strength: 用 <10% 参数/MACs 的紧凑判别前端,通过 SSL 特征聚类伪标签蒸馏,在 DCASE 2020–2025 上达到或超过 Raw-SSL 性能(如 DCASE 2022 eval +6.6pp),NRFT 在 DCASE 2025 上接近详细标签上界。
  • Weakness: 主框架非本文首创(引自作者 ICASSP 2025 前作),NRFT 本质是经典 PCA/GEVD 子空间方法迁移到 SSL 特征空间,缺乏对”为何 GEVD 能在 SSL 特征空间有效分离噪声/机器方向”的机制解释,且未对比朴素 PCA 降维 baseline。
  • Full review: Claude Code 全文七公理审稿

Discriminative anomalous sound detection (ASD) methods train a feature extractor through a classification task using machine-information labels. They then detect anomalies in the resulting feature space based on distances to normal samples. The discriminative feature space effectively captures machine characteristics, leading to high ASD performance. However, this approach benefits from detailed labels, which are costly to obtain. An alternative is a self-supervised learning (SSL)-based label-free approach. This approach directly uses SSL features for ASD and has shown competitive performance. However, SSL models are typically large and computationally expensive. To address these problems, we propose a simple pseudo-label distillation framework. The proposed method generates pseudo labels from SSL features and trains a compact discriminative feature extractor using these pseudo labels. To suppress the effect of noise on pseudo-label generation, we also propose lightweight noise-robust feature transformation (NRFT) methods utilizing a small amount of clean machine-sound data or isolated noise data. We conducted comprehensive evaluations and analyses on the DCASE 2020-2025 Task 2 datasets using four SSL models. The results demonstrate that pseudo-label distillation not only transfers the performance of SSL models to a compact model but also further improves performance by leveraging available coarse labels and data augmentation. Also, our NRFT methods provide further gains.


5. Is One Score Enough? Assessing Singing Quality of Songs with Temporal Score Curves

Authors: Yishan Lv, Jing Luo, Xinyu Yang, Zhizheng Wu

Categories: cs.SD, cs.MM, eess.AS Score: 6.32/10 (Obj:8 Id:6 Ind:5 Comp:5 Eff:8 Nov:5 Rep:5)

  • Strength: 提出 full-length song SQA 新任务并用两阶段伪标签+CLS-style 聚合在两个数据集全部 8 个指标上一致超越 clip-level SOTA(KTAU 相对提升达 13.95%),消融覆盖 teacher/融合/聚合三组变体,baselines 覆盖 VoiceMOS 2024 top systems。
  • Weakness: 方法是标准组件(MuQ + BERT [CLS] + dual-pooling + 伪标签蒸馏)的组合无新机制,且核心卖点数字来自不公开的 Internal Dataset + 自引同团队 teacher model (Lv et al. 2026),代码/checkpoint 开源未声明,独立可复现性不足。
  • Full review: Claude Code 全文七公理审稿

Singing Quality Assessment (SQA) has become increasingly important for practical multimedia applications and Music AI systems, yet existing studies predominantly focus on short singing clips and remain insufficient for full-length songs. Unlike clip-level assessment, full-length song SQA requires modeling how singing quality varies across different audio segments and how these local variations influence the overall evaluation of vocal performance. Moreover, the scarcity of segment-level annotations makes effective supervision challenging, as directly assigning a single overall score label to every segment tends to treat different segment qualities as equivalent. To address these challenges, we propose SongSQA, a two-stage framework for full-length song SQA. In the first stage, a Segment Score Predictor is trained with pseudo labels generated by a pre-trained teacher model, enabling segment-level singing quality prediction without requiring manual segment annotations. In the second stage, a Song Quality Aggregator integrates segment features and predicted segment scores into unified segment embeddings, and employs a learnable song embedding together with self-attention to capture the connection between segment-level vocal performance and overall song quality. In this way, SongSQA dynamically aggregates critical quality cues across the song to produce a holistic quality prediction, while also generating a temporal segment-level quality curve. Experimental results demonstrate the effectiveness of SongSQA for full-length song SQA, achieving up to a 13.95% relative improvement in KTAU over the strongest baseline, while consistently improving other evaluation metrics across all datasets.


6. RealDESED: A Real-World Domestic Sound Event Detection Benchmark

Authors: Florian Schmid, Paul Primus, Alexander Fichtinger, Tara Jadidi, Tobias Morocutti et al.

Categories: eess.AS, cs.AI, cs.SD | Submitted to the DCASE 2026 Workshop (Detection and Classification of Acoustic Scenes and Events). Resources: Dataset (Zenodo): https://zenodo.org/records/20056072; code and baseline implementation (GitHub): https://github.com/fschmid56/RealDESED Score: 5.95/10 (Obj:8 Id:5 Ind:8 Comp:5 Eff:6 Nov:5 Rep:9)

  • Strength: 首个全真实、多标注者(645 人)+ reviewed 评测的家庭 SED 数据集(5,710 录音/37.85h/15 类),数据和代码完全开源,填补 DESED 合成-真实 domain gap。
  • Weakness: 仅单一 backbone(ATST-F)且未对比 DCASE 2024 leaderboard 强系统,方法侧全为已有标准技术组合,分析结论均为领域常识再验证,无新机制或反直觉发现。
  • Full review: Claude Code 全文七公理审稿

This paper presents RealDESED, a real-world domestic sound event detection (SED) benchmark comprising 5,710 audio recordings collected by 652 participants in their homes. Each recording is between 15 and 35 seconds long and contains temporally precise annotations for 15 common domestic sound classes. In contrast to existing SED datasets, which typically rely on simulated soundscapes or broad web-crawled audio, RealDESED consists exclusively of recordings captured in natural domestic environments, reflecting realistic variability in recording devices, device placement, acoustic conditions, background sounds, and naturally occurring event co-occurrences. A distinguishing characteristic of the dataset is its multi-annotator labeling scheme, where each recording is independently annotated by multiple annotators, while the validation and test sets undergo an additional review process to ensure high annotation quality and reliable benchmarking. Furthermore, the dataset provides rich metadata, including recording device, device placement, environment labels, and textual scene descriptions. We establish a strong transformer-based baseline and investigate annotation aggregation strategies, post-processing methods, long-form inference, and the impact of recording metadata on model performance. Our baseline achieves a macro-averaged PSDS1 score of 0.731 on the test set. We believe RealDESED provides a valuable benchmark for developing and evaluating robust SED systems under realistic domestic conditions, helping to bridge the gap between current research benchmarks and real-world deployment.


7. NABEATs: Noise-Aware Audio Representation Learning

Authors: Takuya Fujimura, Yoshiki Masuyama, Gordon Wichern, Christoph Boeddeker, Julius Richter et al.

Categories: eess.AS | Accepted at IWAENC 2026 Score: 5.79/10 (Obj:8 Id:5 Ind:8 Comp:5 Eff:5 Nov:5 Rep:5)

  • Strength: NABEATs 首次将参考噪声条件化引入通用音频 SSL,在 unseen 噪声 (MUSDB18) 条件下仍能改善 6 个下游任务性能,而无参考的 DBEATs 在此条件下退化,验证了参考噪声对泛化的关键作用。
  • Weakness: 缺少代码开源、与最相关先验工作 MMD (denoising distillation) 的明确技术对比、以及多个通用音频 SSL backbone baseline;DCASE 2025 提升幅度有限 (+1.58),训练规模仅用 FSD50K 而非 AudioSet。
  • Full review: Claude Code 全文七公理审稿

We propose the concept of noise-aware audio self-supervised learning (SSL), whose goal is to encode audio mixtures while suppressing undesired noise, and present Noise-Aware BEATs (NABEATs) as a BEATs-based realization of this framework. Audio SSL models are designed to handle a wide range of audio signals. Consequently, under noisy conditions, they cannot effectively focus on the target sounds relevant to a downstream task, resulting in degraded performance. To address this issue, NABEATs is trained to estimate clean BEATs representations from a noisy audio signal with an auxiliary reference noise input. This reference noise enables the model to account for specific noise characteristics at inference time, thereby achieving better generalization across operating environments. Our experimental evaluations demonstrate that NABEATs significantly improves performance of various downstream tasks under noisy conditions and also generalizes well to unseen noise types.


8. Just A Rather Very Intelligent Spoken Agent

Authors: Chen Chen, Zhehuai Chen

Categories: cs.AI Score: 3.79/10 (Obj:5 Id:4 Ind:2 Comp:4 Eff:4 Nov:4 Rep:3)

  • Strength: 在 long-horizon agent 生态中首次系统化定义”user-agent mediation”评测问题,双轨协议(task completion + user interaction)结构清晰,Track 1 在 4 个 frontier worker 上均获 4.86-11.78% 一致提升,说明 mediation 介入对任务完成有可测正向效应。
  • Weakness: 整条评测链路(user simulator / expert user / Jarvis brain / judge / WildClaw 评分)全部由 GPT 家族闭源模型构成且零人类校验,对 benchmark 型论文构成 fatal 级独立性缺陷;仅 34 任务、无方差与显著性、无 ablation、无最简 baseline、无代码开源,不满足 benchmark 论文更严的证据标准。
  • Full review: Claude Code 全文七公理审稿

Long-horizon AI agents are becoming increasingly capable, yet their interaction with users remains surprisingly thin. In most workflows, users give an initial instruction, receive only selective textual updates, and lose a clear sense of what the agent is doing or when to step in. This leaves a missing part in the current agent ecosystem: an always-on Jarvis-style mediator that keeps the agent continuously reachable to the user. Such a mediator should support real-time spoken interaction with the user, answer questions without interrupting the worker, proactively report progress or confusion, and inject user guidance back into the agent’s execution when useful. In this work, we introduce JarvisBench, a benchmark for measuring the dual value of mediation in long-horizon agent workflows. JarvisBench contains two complementary tracks: an agent-collaboration track that measures whether mediation improves downstream task completion, and a user-interaction track that measures whether mediation makes ongoing execution more understandable, responsive, and accessible to users. We instantiate the benchmark with a modular reference Jarvis prototype and evaluate it on 34 text-only WildClaw tasks executed in OpenClaw. Preliminary results with GPT-5.5, Claude Opus 4.7, Gemini-based, and GPT-based worker agents suggest that Jarvis-style mediation can provide trace-grounded responses to user questions and improve task performance when sparse user guidance is injected at appropriate moments. The results also show that effectiveness depends strongly on the mediator’s LLM brain, highlighting both the promise of this missing middle layer and the need for broader community effort. Demo page https://cchen1436.github.io/jarvis


9. Efficient Audio-Visual Event Recognition via Knowledge Distillation and Dynamic INT8 Quantization of a Hybrid Cross-Attention Network

Authors: Parinaz Binandeh Dehaghani, Danilo Pena, A. Pedro Aguiar

Categories: eess.SP, cs.SD | 15 pages, 4 figures Score: 3.58/10 (Obj:5 Id:2 Ind:8 Comp:3 Eff:3 Nov:2 Rep:5)

  • Strength: 在 AVE 数据集上将自家 hybrid cross-attention teacher 压缩 59.06% 参数、INT8 后模型 2.04 MB,精度仅降 2.99%,pipeline 端到端跑通且训练曲线稳定。
  • Weakness: 无任何外部 baseline 与消融,三标准组件(宽度收缩+KD+dynamic INT8 PTQ)无新机制;2.04 MB 不含 VideoMAE/AST backbone,且 INT8 推理 latency 反而更高,”edge efficient deployment” 核心主张缺乏真实端到端证据。
  • Full review: Claude Code 全文七公理审稿

Audio-visual event recognition (AVER) has achieved significant performance improvements through transformer-based multimodal architectures. However, the high computational complexity, large memory footprint, and inference cost of these models hinder their deployment on edge and resource-constrained devices. This paper presents an efficient compression framework for hybrid cross-attention-based audiovisual event recognition by combining architectural model compression, knowledge distillation, and dynamic INT8 quantization. A high-capacity teacher model integrates VideoMAE for visual representation learning, the Audio Spectrogram Transformer (AST) for audio feature extraction, and a hybrid cross-attention fusion network for multimodal feature integration. A lightweight student model is constructed by reducing the hidden feature dimension, the number of attention heads, and the feedforward network size while preserving the overall network architecture. The student model is trained using knowledge distillation to effectively transfer discriminative knowledge from the teacher. Finally, dynamic INT8 post-training quantization is applied to further reduce the model size for efficient deployment. Experimental results on the Audio-Visual Event (AVE) dataset show that the proposed framework reduces the number of trainable parameters in the multimodal fusion module by 59.06%, with only a 2.14% decrease in classification accuracy compared with the teacher model. Furthermore, dynamic INT8 quantization reduces the model size from 10.71 MB to 2.04 MB while maintaining competitive recognition performance. These results demonstrate that the proposed framework provides an effective trade-off between recognition accuracy and computational efficiency, making it a promising solution for deployment on resource-constrained edge AI platforms.


10. Explainable Lightweight Compact Deep Models for Speech Emotion Recognition

Authors: Nelly Elsayed

Categories: cs.SD, cs.AI, cs.ET, cs.LG | Accepted in the IEEE ICMLA 2026 Conference Score: 3.42/10 (Obj:5 Id:2 Ind:5 Comp:2 Eff:5 Nov:2 Rep:5)

  • Strength: 在 SAVEE 上以仅 33,208 参数的 compact CNN+ASP 达到 96.875% accuracy / 0.977 UAR,参数量比对比表中的代表性方法低 1-3 个数量级。
  • Weakness: 无任何消融实验,未与 Light-SERNet/Deep-Net/emotion2vec 等最直接同类或强 baseline 比较,测试集仅 1 个说话人约 120 条语句,无代码开源,三件套(log-Mel+CNN+ASP+Grad-CAM)均为已有方法直接拼装且作者自述不引入新架构。
  • Full review: Claude Code 全文七公理审稿

Speech Emotion Recognition (SER) is an important component in a wide range of human-centered applications, including healthcare, customer service, and human-omputer interaction. In medical and decision-support settings, there is increasing interest in models that not only achieve accurate emotion recognition but also support transparent predictions and efficient deployment. However, many existing SER approaches rely on complex deep learning architectures that limit interpretability and increase computational cost. This paper presents an explainable and lightweight speech emotion recognition framework based on a compact convolutional neural network architecture. The proposed approach utilizes log-Mel spectrogram representations to capture spectro-temporal speech characteristics and employs attentive statistics pooling to emphasize emotionally salient temporal segments. To improve model transparency, gradient-based class activation mapping (Grad-CAM) is incorporated to visualize the time-frequency regions that influence the model’s predictions. Experimental evaluation on the SAVEE emotional speech dataset demonstrates that the proposed framework achieves competitive recognition performance while maintaining a compact architecture with significantly fewer parameters than many existing SER models. The results indicate that efficient convolutional architectures combined with interpretable analysis can provide a practical balance between recognition accuracy, computational efficiency, and model transparency.