每日 arXiv

音频 · 语音 · 音乐 · 声学

每日新爬取的 arXiv 论文,按评分排序。中文摘要帮你快速判断是否值得深读。

2026-09-02
日期2026-09-02
已评分
均分
最高

Daily Papers — 2026-09-02

10 papers on audio, speech, music, and acoustics.

1. A Common Measure of Communication for Speech Brain-Computer Interfaces

Authors: Dulhan Jayalath, Benjamin Ballyk, Oiwi Parker Jones

Categories: cs.LG, q-bio.NC | Code and OVMI Explorer available from the project page at https://neural-processing-lab.github.io/OVMI/ Score: 8.10/10 (Obj:9 Id:8 Ind:7 Comp:8 Eff:8 Nov:8 Rep:9)

  • Strength: 提出语音 BCI 首个跨词表可比的开词表互信息度量 OVMI,暴露 50 词系统准确率 94.0% 实际仅传达参考熵 6.4% 的过估计,并以 OVMI 选词表在三个语音域取得最高 16.3% 相对准确率提升。
  • Weakness: 跨系统比较依赖对称错误标量估计器且未做敏感度分析,非侵入式对照全部出自本组存在自评闭环,词表选择实验未在真实用户数据上验证,参考分布选择引入未规范的自由度。
  • Full review: Claude Code 全文七公理审稿

Speech brain-computer interfaces (speech BCIs) translate neural activity into language, offering a path towards restoring speech for people with paralysis and, more broadly, enabling new forms of natural human-computer interaction. Despite this promise, the field lacks a common measure of progress because systems use different datasets, recording methods, types of speech, and vocabularies, so their reported scores are rarely comparable. Underlying this measurement problem are two unresolved questions: (i) what distribution of words should a speech BCI enable a user to communicate, and (ii) how much information from this distribution can a system convey. We address both by deriving open-vocabulary mutual information (OVMI), an information-theoretic quantity that measures the information conveyed by a decoder relative to a reference distribution over the words a user may wish to communicate. This allows capabilities measured under different conditions, such as distinct vocabularies, to be evaluated on a common communication scale. We show that ordinarily reported accuracy, word error rate (WER), and other metrics computed only over the words a system supports can overstate how much of a user’s intended speech the system can communicate. We then use OVMI to compare existing systems, expose trade-offs between how much of the user’s language a system supports and how accurately it decodes those words, show that these comparisons depend on what the user is expected to communicate, and demonstrate that selecting a vocabulary to maximise OVMI yields up to 16.3% relative improvement in accuracy across three speech domains. OVMI therefore provides the speech BCI community with a principled way to compare heterogeneous systems, improve vocabulary design, and measure progress in the field.


2. ARFT: A Synchronized Multimodal RF-Acoustic Dataset for Positioning in Distributed Environments

Authors: Daan Delabie, Jarne Van Mulders, Bert Pyck, Gustav Nilsson Gisleskog, Gilles Callebaut

Categories: eess.SP, eess.AS Score: 7.70/10 (Obj:9 Id:8 Ind:5 Comp:8 Eff:8 Nov:7 Rep:9)

  • Strength: 首个同步 RF+声学室内定位公开数据集:5011 个三模态对齐样本(5.57×2.89 m、42 天线 920 MHz CSI + 91 麦克风)、1.25 mm 光学真值,声学基线 2D 误差 0.110 m(P90 0.195 m),采集栈、notebook 与合并 CSI 全开源。
  • Weakness: 摘要声称支持 RF-only/联合定位但论文只实现并验证了声学一条管线,RF 定位停留在探索阶段且无任何联合模态实验;单一房间、单频带、纯静态采集,且「无同类数据集」的新颖性声明缺乏全网穷尽验证。
  • Full review: Claude Code 全文七公理审稿

This paper documents the acoustic-radio fusion in Techtile (ARFT) dataset, a synchronized measurement campaign for distributed wireless sensing and positioning in the Techtile testbed. Ultrasonic and radio frequency (RF) signals are simultaneously transmitted and captured at multiple positions in a 2D spatial grid inside the Techtile testbed. Each acquisition cycle corresponds to one rover stop, one position sample, one acoustic recording and one RF snapshot. The RF modality is recorded as per-host pilot measurements and released as channel state information (CSI) tensors for 42 antennas mounted at the ceiling of the room. The acoustic chirp is recorded with 91 synchronized microphones and transmitted by a synchronized quasi omni-directional speaker. The campaign spans 5011 spatial position samples spanning a 5.57 m by 2.89 m area with complete RF and acoustic data. We detail how the measurements are recorded, quantify per-experiment coverage, describe the RF and acoustic data contents, and explain how the aligned modalities support RF-only, acoustic-only, and joint positioning workflows. Furthermore, a static acoustic positioning pipeline that performs anchor selection, pulse-compression ranging, and least squares (LS) localization is elaborated.


3. SonicCaps: Large-Scale Diverse and Fine-Grained Captioning for Improved Audio-Retrieval

Authors: Zineb Lahrichi, Marc Ferras, Gaël Richard, Geoffroy Peeters

Categories: cs.SD, cs.CL, cs.MM, eess.AS Score: 7.58/10 (Obj:9 Id:9 Ind:6 Comp:8 Eff:8 Nov:6 Rep:6)

  • Strength: 15M 描述/700k 音频、每音频约 24 条异质描述的一对多数据集,配合多描述采样使 CLAP 检索 AudioCaps-Val T2A R@10 从 82.4 提至 91.3、FoleyBench 零样本 R@5 从 9.1% 提至 25.6%,并证明多样性驱动收益、质量仅边际贡献。
  • Weakness: 两个发布模型在 HF 暂未检索到、数据许可实为 cc-by-nc-4.0 与摘要声称的 CC BY 4.0 不一致、商业基准不可复现、ESC-50 评估被自认的 FreeSound 泄漏污染,独立泛化证据仅剩低绝对值的 FoleyBench。
  • Full review: Claude Code 全文七公理审稿

Recent advances in audio-language modeling have been driven by large-scale audio captioning datasets. However, existing datasets remain limited by low semantic diversity, generic descriptions lacking acoustic details, and one-to-one audio-caption mappings that poorly reflect the inherent ambiguity of auditory perception. We introduce SonicCaps, a large-scale audio captioning dataset comprising ~15M captions paired with ~700k audio clips, generated using a multi-modal large language model (Qwen3-Omni) conditioned on both audio and text. To explicitly promote diversity, we generate around 24 captions per audio via structured prompt engineering and few- shot generation, spanning main descriptions, rephrased variants (verbosity, style) and semantic tags. Human evaluation shows that SonicCaps is rated significantly higher than existing captioning datasets, with fine-grained analyses indicating that our captions are perceived as more descriptive and precise, which strongly correlates with quality judgments. Finally, training CLAP models on SonicCaps with a multi-caption sampling strategy consistently improves audio retrieval and zero-shot classification, with stronger generalization across public and commercial benchmarks. We release both SonicCaps and two specialized CLAP models on hugging face: https://huggingface.co/datasets/Zineb/SonicCaps.


4. VibeVoice-ASR-Streaming Technical Report

Authors: Yujie Tu, Zhiliang Peng, Jianwei Yu, Li Dong, Songchen Xu et al.

Categories: eess.AS Score: 7.58/10 (Obj:9 Id:8 Ind:6 Comp:7 Eff:8 Nov:7 Rep:8)

  • Strength: 首批 LLM 端到端流式说话人归属 ASR,7B 五集平均 WER/CER 24.66 超过 Gemini 3.5 Transcribe Live(25.23),13 个归属设置 12 个最佳,期望延迟 2.00s vs Azure CT 8.21s,权重与代码已开源。
  • Weakness: 闭源基线数字全部作者自测且公平条件未说明,合成训练数据与评测集无去污证明,归属能力依赖 7B 规模(1.5B→7B 差 12.76 cpWER)而机制未解耦,42 万小时训练数据不公开导致无法独立重训。
  • Full review: Claude Code 全文七公理审稿

Traditional speaker-attributed ASR systems treated ASR and speaker diarization as two separate tasks. Recently, end-to-end models such as VibeVoice-ASR have unified the two tasks within a single model. However, existing unified models still mainly support offline recognition, making it difficult to meet the low-latency requirements of real-time voice assistants and agents. To tackle this issue, we present VibeVoice-ASR-Streaming, one of the first LLM-based end-to-end approaches to streaming speaker-attributed ASR. It interleaves fixed-size audio chunks, a small amount of lookahead audio and previous text. This allows the model to produce ‘‘who said what’’ as speech arrives, without a separate diarization stage. For transcription accuracy, our 7B model achieves the lowest average WER/CER across five evaluation sets. For speaker attribution, it achieves the best or tied-best on 12 of 13 evaluation settings. We release the 1.5B and 7B model weights together with inference code.


5. Sensing Bone-Conducted Speech with Earbuds

Authors: Christoph Weyer, Peter Jax

Categories: eess.AS | 14 pages, 12 figures, journal Score: 7.30/10 (Obj:9 Id:8 Ind:6 Comp:9 Eff:7 Nov:7 Rep:5)

  • Strength: 首次系统测量 17 名受试者两款耳塞的 OV 振动频谱与空间特性:低通 −93 dB/十倍频程、第一主成分(进出耳道口方向)占 79–89% 方差、单轴对准后 100–400 Hz 平均衰减仅 0.7–1.5 dB(±45° 内 <3 dB)。
  • Weakness: 500 Hz–1 kHz 信号没入传感器噪声使 1 kHz 以上结论仅为有偏上界,无下游任务(ASR/增强)验证,无代码/数据发布,且相关工作新颖性检索因 Semantic Scholar/OpenAlex 查询失败未能独立交叉确认。
  • Full review: Claude Code 全文七公理审稿

Clear capture of the wearer’s own voice (OV) is essential when using earbuds for mobile communication. However, OV capture remains challenging in noisy environments. Bone-conducted (BC) speech, which can be sensed as vibrations of the earbud housing, can be used to improve OV capture. However, neither bandwidth nor spatial characteristics of OV-induced earbud vibrations have been analyzed in detail, despite both characteristics being relevant, e.g., for sensor choice and placement. This study investigates both characteristics, based on measurements with two earbud models. Spectrally, results indicate that OV-induced earbud vibrations exhibit a low-pass characteristic, with a steep roll-off of -93 dB per decade above 400 Hz. Thus, sensors with comparatively low noise floors are required to sense the vibrations above \SI{1}{\kilo\hertz}. Spatially, results indicate that the earbuds mainly vibrate in and out of the ear canal entrance, with high consistency between subjects and fits. Simulations confirm that this enables capture of the high-power vibrations below 400 Hz by a single-axis sensor with less than 1.5 dB mean attenuation.


Authors: Yichen Liu, Quanwei Zhang, Haozhe Wang, Donghao Zhou, Xiaojie Li et al.

Categories: cs.MM, cs.CV Score: 7.20/10 (Obj:9 Id:8 Ind:6 Comp:9 Eff:8 Nov:6 Rep:4)

  • Strength: 在 200 条测试脚本上将 Shot Boundary MAE 从 1.11s 降至 0.042s(-96%,约一帧),Dialogue Acc@0.5s 从 28.3% 提升至 84.1%,同时 IQ(0.7032)、WER(8.48%)、Sync-C(2.78)均为全表最佳,28 人盲测五维度全胜,方法零可学习参数仅用 LoRA。
  • Weakness: 基线未获得同等时间精修监督而去精修消融显示数据管线贡献约一个数量级(0.375s→0.042s),收益归因混杂;训练/测试限于两个未公开短剧集,无代码、检查点或数据发布,跨域泛化与独立复现证据均缺失。
  • Full review: Claude Code 全文七公理审稿

Joint audio-video generation models have made substantial progress in visual quality and audio-visual synchronization. However, they still provide limited control over when shot transitions occur and dialogue is spoken. This limitation constrains their application in script-driven content creation, where timing errors can undermine narrative coherence and the viewing experience. Current joint generators align video and audio representations on a shared temporal axis, yet the precise timing of shots and dialogue specified in a structured prompt is encoded only in the prompt’s text representation and remains unaligned with the temporal coordinates of either modality. Consequently, video and audio may remain synchronized with each other while both fail to follow the script timeline. This mismatch motivates us to extend temporal alignment beyond video and audio to include the structured script. We therefore introduce Temporal Context Routing (TCR), which maps the script timing onto the shared temporal axis of video and audio generation and routes each prompt’s guidance to the corresponding positions in both modalities. Compared with the baseline on 200 test scripts, TCR reduces Shot Boundary MAE by 96%, from 1.11 s to 0.042 s, and raises Dialogue Acc@0.5 s from 28.3% to 84.1%. TCR achieves these improvements while maintaining visual quality and audio-visual synchronization comparable to those of the baselines. A user study further shows that participants prefer TCR on all five evaluated dimensions.


7. Auditory Illusion Benchmark for Large Audio Language Models

Authors: Hayoon Kim, Eunice Hong, Kyogu Lee

Categories: cs.SD, cs.AI | Accepted to the 2026 IEEE International Conference on Acoustics, Speech, and Signal Processing (ICASSP 2026) Score: 7.20/10 (Obj:9 Id:7 Ind:6 Comp:7 Eff:7 Nov:8 Rep:6)

  • Strength: 首个 LALM 听幻觉基准(14,829 强制选择试验、10 幻觉 × P/P+K 标注、13 模型原始回答公开),HLA/RA/ISI 三指标配 20 人受控人类听测,揭示域反转现象且自证措辞可移动 ISI 达 0.6。
  • Weakness: 提示词敏感性消融仅覆盖 3/10 模型而其余 ISI 排名未验证协议鲁棒性,人类基线 20 名绝对音感受试者代表性不足,音频不随附且论文数字只能统计重现、无法逐位复现,缺文本先验污染对照。
  • Full review: Claude Code 全文七公理审稿

Perceptual illusions have long served as crucial probes into human cognition, revealing biases and limitations of perception. In the auditory domain, such illusions provide a unique lens for testing whether Large Audio Language Models (LALMs) replicate human perceptual tendencies. Despite their importance, most benchmarks focus on visual illusions or general audio tasks, leaving auditory illusions underexplored. To this end, we present AIB, the first auditory illusion benchmark for LALMs, covering ten representative illusions across music, sound, and speech, each annotated for the presence of knowledge-based priors. Our methodology pairs model evaluation with controlled human listening studies, enabling direct comparison of responses. Results show systematic differences: while most LALMs remain signal-faithful on low-level acoustic illusions, several exhibit more human-like responses when linguistic or musical priors are involved, although no model matches the human perceptual profile. These findings highlight the current limitations of LALMs as cognitive models. By establishing auditory illusions as a rigorous testbed, our work offers a new perspective for probing neural black-box models and advancing understanding of auditory cognition. AIB is publicly available at https://github.com/gillosae/aib.


8. VoxReason: Listener-Free Evaluation of Source-Grounded Speech Planning Before Synthesis

Authors: Mengzhe Geng

Categories: cs.SD, cs.CL, cs.LG, eess.AS Score: 6.90/10 (Obj:8 Id:8 Ind:8 Comp:8 Eff:5 Nov:7 Rep:7)

  • Strength: 把合成前交付规划的源引用合法性变成确定性可校验任务,shortcut 对照证明 slot 精度不安全(oracle 1.000→0.242),7B locality SFT+CF 修复将 plan-slot/局部性从 0.684/0.141 提到 0.919/1.000,源通道效应(+0.488)远超规模效应(+0.023)。
  • Weakness: 评测构造过窄(2 条 target texts、1 个 scene、零公开 context-audio、确定性 key→plan 映射、无独立人工标注),波形质量与人工听感被排除,checkpoints 与 raw audio 不公开,无端到端效用证据。
  • Full review: Claude Code 全文七公理审稿

Expressive speech systems make a decision before any waveform is rendered: how an utterance is delivered. In dialogue agents, narration, and role-conditioned TTS, that hidden planning step sets affect, pitch, energy, rate, pause, emphasis, and stance, yet downstream audio scores rarely reveal whether those choices were licensed by the source record, a source-use failure that occurs before any waveform exists. VoxReason makes that pre-synthesis decision measurable as a listener-free task for source-grounded speech planning. Before synthesis, VoxReason measures whether delivery choices are grounded in cited source records. Systems output a source-cited speaking-plan with evidence citations, and a deterministic verifier checks citation legality, slot agreement, unsupported state, schema validity, and one-cue counterfactual locality. On 1,440 checked source-label cases, shortcut controls show why slot accuracy alone is unsafe: a key-lookup oracle reaches 1.000 plan-slot accuracy on seen keys, while an emotion prior still reaches 0.958 slot accuracy on source-key-disjoint cases without citing intensity or identity. In a separate 100-case learned source-key-disjoint comparison, a 7B locality SFT+CF repair improves plan-slot accuracy/locality from 0.684/0.141 to 0.919/1.000, and removing source records lowers citation-required grounded score by 0.488. Rendered waveform quality remains outside the present evaluation.


9. Understanding Automatic Mixing: A Subtask-Oriented Analysis of Two-Stage Mixing System

Authors: Jinjie Shi, Wei Hua, Kunzhu Xie, Make Li, Yuchen Liu et al.

Categories: cs.SD, eess.SP | Accepted at the International Society for Music Information Retrieval Conference (ISMIR 2026). 6 pages, 5 figures. https://sparrowreivun.github.io/TwoStageMixingAnalysis/ Score: 6.80/10 (Obj:8 Id:6 Ind:8 Comp:8 Eff:6 Nov:7 Rep:6)

  • Strength: 通过三个受控听觉实验(18 名可靠听者、1170 条评分)首次系统归因两阶段自动混音收益,两个两阶段变体均显著优于单阶段基线(2S-MEGAMI 61.57 vs 54.65, q=0.043;2S-Diff-MST 28.69 vs 19.44, q=7.3e-4),并发现免训练 ELL 在组内平衡上不输神经模型。
  • Weakness: 全部结论仅基于 3 个流行/摇滚选段且 Diff-MST 存在地板效应,7 组与 4 组分组在粒度和语义上混淆无法归因,响度分离的必要性未获统计支持(q=0.151),系统代码与 checkpoint 未公开仅统计层可复核。
  • Full review: Claude Code 全文七公理审稿

Automatic mixing transforms multitrack recordings into perceptually coherent, balanced, and aesthetically consistent mixes. In real-world production, this task is challenging due to large track counts, diverse instrumentation, and strong inter-track dependencies. Two-stage systems address this complexity by separating intra-group processing from inter-group mixing, yet it remains unclear whether their gains arise from stronger component models or from explicit task decomposition. We present a subtask-oriented analysis of automatic mixing through three controlled listening experiments. We investigate whether full-mix models transfer to intra-group mixing, whether downstream models compensate for grouping and loudness errors, and whether two-stage decomposition improves full-mix quality. Across three dense pop and rock excerpts, transfer differs between the evaluated models; inappropriate grouping causes clear downstream degradation, while altered loudness relationships have weaker and model-dependent effects. Both two-stage variants significantly outperform their corresponding single-stage baselines. These findings support explicit separation of local balance and global mix coordination as a useful design principle for automatic mixing. Code and audio examples are available online.


10. Scalable Direction-Following TTS via Voice Impression-Guided Pseudo Triplet Construction

Authors: Kenichi Fujita, Yusuke Ijima

Categories: cs.SD, cs.CL, cs.LG | 5 pages,4 figures, Accepted to INTERSPEECH 2026 Score: 6.05/10 (Obj:8 Id:5 Ind:6 Comp:6 Eff:6 Nov:7 Rep:4)

  • Strength: 首次量化验证”仅伪三元组(35 万对,1,600 说话人合成)即可实现稳定保说话人的方向跟随修改”,且伪+真实组合将对齐推至 LLM 指标 3.79±0.01、主观 SMOS 3.35 vs 仅真实数据 2.67,F0 分析(0.05 vs 0.14)诚实解释了伪数据保守性来源。
  • Weakness: 缺乏任何外部系统基线与幅度归一对照,Pseudo/Recorded 的说话人保持与对齐优势分别被修改幅度和数据量级(7.4 万 vs 6,899 对)混杂;印象可控 TTS、估计器与全部数据均未发布,仓库仅含音频 demo。
  • Full review: Claude Code 全文七公理审稿

Voice actors often re-read the same script while modifying their delivery in response to performance directions. We study this setting as direction-following TTS, where a system generates a new utterance that reflects a given direction relative to a reference utterance while preserving speaker identity and linguistic content. A key challenge is the lack of training data capturing such relative modifications. To address this, we propose a scalable pseudo-triplet construction pipeline that generates~(reference utterance, direction text, modified utterance) triplets. It generates controlled style variations using an impression-controllable TTS model and uses an LLM to produce natural language directions from estimated impression differences. Experimental results demonstrate that pseudo-triplets alone enable stable speaker-preserving modification, and that combining pseudo and recorded data further improves direction alignment while maintaining speaker similarity. Audio examples are available on our demo page https://ntt-hilab-gensp.github.io/IS2026pseudo/