每日 arXiv

音频 · 语音 · 音乐 · 声学

每日新爬取的 arXiv 论文,按评分排序。中文摘要帮你快速判断是否值得深读。

2026-08-30
日期2026-08-30
已评分
均分
最高

Daily Papers — 2026-08-30

4 papers on audio, speech, music, and acoustics.

1. TEMPO: Temporally-grounded Multi-task Post-training for Large Audio-Language Models

Authors: Apoorva Kulkarni, Kaousheik Jayakumar, Sreyan Ghosh, Utathya Aich, Ramani Duraiswami et al.

Categories: cs.SD, cs.AI, cs.LG | Accepted at EMNLP 2026 Main Conference. Project page - https://kaousheik-26.github.io/tempo Score: 7.21/10 (Obj:8 Id:7 Ind:6 Comp:8 Eff:6 Nov:8 Rep:8)

  • Strength: 首个统一语音/声音/音乐时间戳的 LALM,五任务单模型:ASR WER 69.7%→43.5%、diarization DER 44.2%→25.4%、mIoU 0.44→0.71,均超过显式做过时间戳训练的 Qwen3-Omni 与 Audio Flamingo Next,且代码、119K 训练集与权重已完整开源(HF 数据集 122,611 行)。
  • Weakness: 评测集与训练语料同源却无去污染核查、grounding 查询依赖闭源 GPT-4o 改写,且与专家系统(Parakeet-TDT、pyannote-3.1)的正面对比缺失——”统一模型可替代专用模型”的核心效用主张缺关键证据,GRPO 增益不足 1.5 点且音乐任务回退。
  • Full review: Claude Code 全文七公理审稿

Large audio-language models (LALMs) describe audio at the clip level but cannot assign timestamps to the events, speakers, or sounds they identify. Despite being essential for downstream tasks like speech recognition and dense audio captioning, timestamping remains a key limitation of most LALMs. We present TEMPO (Temporally-grounded Multi-task Post-training), the first unified model to handle audio, speech, and music timestamping tasks. Our core contribution is a supervised fine-tuning (SFT) stage built on three innovations: atomic timestamp tokens, a time-aware projector that injects sinusoidal wall-clock encodings into audio frame embeddings, and a distance-aware Gaussian loss. Our training is based on a synthetic-to-real curriculum. We further introduce, to our knowledge, the first application of reinforcement learning to unified audio timestamping, using GRPO with verifiable temporal rewards that directly optimize the evaluation objectives. Rather than serving as the primary source of performance gains, GRPO acts as a refinement stage on top of the SFT checkpoint, providing modest additional improvements. To support this work, we build a training dataset containing 119K samples and an evaluation benchmark containing 10K samples, drawn from established corpora across five tasks. On this benchmark, TEMPO outperforms Audio Flamingo Next and Qwen3-Omni, two state-of-the-art LALMs explicitly trained on timestamped data. Experiments confirm that SFT delivers most of these gains, with GRPO providing consistent but moderate refinements.


2. PhysWave: Physics-Guided Latent Diffusion Models for Controllable Spatial Audio Generation

Authors: Lingfeng Yao, Chenpei Huang, Xingke Yang, Ziye Geng, Changqing Luo et al.

Categories: cs.SD, cs.AI, cs.MM | Accepted by EMNLP 2026 Main Conference. Project website: https://lingfengyao.github.io/PhysWave/ Score: 7.00/10 (Obj:9 Id:7 Ind:6 Comp:6 Eff:7 Nov:8 Rep:5)

  • Strength: 文本条件下 FOA 空间一致性大幅领先:静态 DoA 误差 1.73°(ImmerseDiffusion 7.07°)、运动源 5.78°(SonicMotion 31.09°),反平方相关 0.48→0.73,并附 874.5h FOA 数据集与免训练推理 guidance。
  • Weakness: 客观评测全部基于与训练同源的合成 FOA 渲染、无真实录音客观验证,音质劣于最优基线(FD 21.22 vs 13.86),且代码/checkpoint/数据集在官方仓库均未发布(README 三项未勾选、无 license)。
  • Full review: Claude Code 全文七公理审稿

Text-to-spatial audio generation, such as text-to-First-Order Ambisonics (FOA), provides a convenient way to create spatial audio for billion-dollar gaming and film industries. However, existing text-to-FOA methods are largely data-driven and may produce audio that violates acoustic relations between source direction and distance. They also separate descriptive and parametric control, forcing users to trade usability for precision. In this paper, we present PhysWave, a physics-guided latent diffusion model for controllable text-to-FOA generation. PhysWave unifies natural-language and trajectory control through a shared waypoint-caption representation, and augments diffusion training with two differentiable acoustic priors: spherical-harmonic direction consistency and inverse-square distance consistency. To support dynamic spatial generation, we further construct a 300K-clip FOA dataset with diverse sound categories and source trajectories. Extensive results show that the proposed priors help PhysWave generate spatially consistent FOA audio while maintaining competitive audio quality. Further analyses show that these physics priors improve spatial consistency during training and can also be used as inference-time guidance for training-free spatial refinement.


3. What Are You Listening to? Temporal Music Grounding for Audio-to-Text Large Language Models

Authors: Kun Fang, Ziyu Wang, Ichiro Fujinaga

Categories: cs.SD, cs.IR Score: 6.40/10 (Obj:9 Id:6 Ind:6 Comp:5 Eff:6 Nov:8 Rep:4)

  • Strength: 首个具备精确符号-音频对齐的音乐时间定位基准;诊断出 SOTA 音频-LLM 定位近乎失效(Qwen2-Audio IoU@0.5 F1=0.0000、Music Flamingo 0.0342),任务特定训练可将其推至 0.9221(MAE 9.34 ms)。
  • Weakness: 素材仅单声部合成钢琴、外部效度未验证;数据/代码/模型未发布(GitHub 检索 0 结果),LLM judge 对全部微调模型打出 Global Acc 1.0000、grounding→理解结论依赖 51 样本概念组且跨骨干不一致,可复现与证据链均有缺失。
  • Full review: Claude Code 全文七公理审稿

Large audio-language models can produce fluent and musically plausible responses, yet it often remains unclear whether those responses are grounded in the audio input. We introduce temporal music grounding, a task in which a model returns one or more time spans corresponding to a queried musical note, event, or pattern. To evaluate this capability, we present MusicGroundingBench, a controlled benchmark suite built by rendering algorithmically generated piano MIDI to audio, yielding exact symbolic-to-audio alignment. The suite comprises two subsets: MGBench-3N, which evaluates note-level grounding in clips containing up to three notes, and MGBench-2B, which evaluates structured grounding and short-form music understanding in two-bar excerpts. Experiments show that temporal music grounding remains challenging for current audio-language models, whereas task-specific training yields substantial gains. We further report exploratory evidence on the relationship between grounding supervision and music understanding. These results establish MusicGroundingBench as a controlled testbed for assessing whether audio-language models ground their responses in temporally localized musical evidence.


4. How Well Do Generative Music Models Follow Emotion Conditioning?

Authors: Morteza Heydari

Categories: cs.SD, eess.AS | 7 pages, 3 figures, 2 tables Score: 5.79/10 (Obj:8 Id:6 Ind:4 Comp:8 Eff:6 Nov:5 Rep:4)

  • Strength: 在 GTZAN 全量 1,000 首 × 3 系统 × 5 配置上给出首个统一 VA 情感跟随评测,发现文本条件一致优于音频条件(InspireMusic 文本胜率 67.3%,Wilcoxon p<0.001,d=0.502)且效价比唤醒度更易保留(全部系统 MAE-V < MAE-A)。
  • Weakness: 评估闭环不独立——Music2Emotion 同时定义目标、写入提示并对结果打分,文本条件优势部分来自提示与裁判共享表示且无第二 MER 模型或人类听感交叉验证;论文未发布任何代码、提示模板或生成输出。
  • Full review: Claude Code 全文七公理审稿

Recent generative music models offer increasingly fine-grained control through text and audio conditioning, yet how faithfully they follow intended emotional cues remains an open question. We address this gap with a unified evaluation pipeline for emotion-following in generated music. Using all 1000 tracks in GTZAN, we extract semantic audio descriptions with DashengLM, an audio captioning model, and estimate source-track valence and arousal with Music2Emotion, a music emotion recognition model. We construct affect-aware text prompts by combining descriptions with top-ranked emotion tags and generate 30-second outputs with three systems, Stable Audio Open, MusicGen, and InspireMusic, evaluating both text- and audio-conditioned generation. To measure emotion-following, we compute valence and arousal on generated audio and compare them with the source tracks using absolute error and Euclidean distance in valence-arousal space. Text-conditioned generation consistently outperforms audio conditioning, with MusicGen (text) and InspireMusic (text) achieving the best performance, while audio-conditioned variants prove less stable. We further find that valence is preserved more reliably than arousal and that emotion-following varies substantially across genres. These findings underscore the importance of evaluating affective controllability directly rather than relying solely on general quality or prompt-relevance metrics.