每日 arXiv

音频 · 语音 · 音乐 · 声学

每日新爬取的 arXiv 论文,按评分排序。中文摘要帮你快速判断是否值得深读。

2026-08-27
日期2026-08-27
已评分
均分
最高

Daily Papers — 2026-08-27

18 papers on audio, speech, music, and acoustics.

1. Not all generalisation failures can be bought back: four boundaries in affective audio modelling

Authors: Jingyi Zhang, Xiaotong Yao

Categories: eess.AS | 40 pages, 8 figures. Supplementary Information (25 pages, 8 Extended Data figures) included as an ancillary file. Code, run records and per-panel source data: https://github.com/AuraJZ/affective-audio-boundaries Score: 7.84/10 (Obj:9 Id:9 Ind:6 Comp:8 Eff:7 Nov:8 Rep:8)

  • Strength: 用四语料×四表示×三生理库把泛化失效拆成可定价的两类:同域损失 20–31% 且 100 个目标标签回收 64%,跨域损失 84–95% 且四个预训练表示无一回收,生理边界所有信息源低于可达天花板的 1/3。
  • Weakness: 生理”信息受限”结论实质只建立在一个 EEG 语料(31 人)且组级 p 未过多重校正(最小 q=0.16)、目标侧学习曲线未测;跨域结论限音乐↔环境声一对组合与唤醒单维度,无法外推到更宽边界。
  • Full review: Claude Code 全文七公理审稿

Models mapping acoustic properties onto affective response underpin applications from music recommendation to sound design, yet are evaluated almost entirely within the corpus they were fitted on. When one fails outside it, the standard response – more data, or a larger model – assumes every failure is a shortage of resources. We show it is not, and that the alternative calls for the opposite remedy. Using four corpora of rated sound, four pretrained representations and three corpora of physiological recording, we pushed one mapping across four boundaries an application must cross: to new material, to edited audio, to a sensor in place of a self-report, and to an individual listener. At each we report the ceiling the target permits, the fraction surviving the crossing, and the price in target-side observations of closing the gap. Within a corpus, prediction reaches 84% of the ceiling set by inter-listener agreement. A same-domain corpus swap costs a fifth of that, and a hundred target labels return two-thirds of the loss. Crossing between music and environmental sound costs four-fifths to all of it, and four pretrained representations recover none of it. Against physiological response no information source we constructed exceeds a third of the attainable ceiling. “The model does not generalise” is therefore two diagnoses, not one, with mutually exclusive remedies; treating the second as the first is the more expensive mistake.


2. Mapping Written Words to Spoken Words in a Different Language Using Only Visual Grounding

Authors: Gabriel Pirlogeanu, Dan Oneata, Horia Cucu, Herman Kamper

Categories: cs.CL, eess.AS | 9 pages, 5 figures, 5 tables, preprint, submitted to IEEE Transactions on Audio, Speech and Language Processing Score: 7.30/10 (Obj:9 Id:8 Ind:6 Comp:8 Eff:6 Nov:7 Rep:7)

  • Strength: 免训练对比式对齐方法在跨语言关键词定位上大幅超越注意力神经基线(P@10 localization 49.9% vs 10.4%,spotting 63.0% vs 18.8%),负例挖掘带来 111% 相对提升并使结果对 captioner 选择鲁棒。
  • Weakness: groundtruth 依赖 9.4% WER 的印地语 ASR 与 ChatGPT 词典等目标场景缺乏的资源,且弱监督下 localization 仅 49.9%(转写 topline 89.8%),半数检索片段需人工筛选;核心代码尚未完整开源。
  • Full review: Claude Code 全文七公理审稿

In many low-resource settings, even just eliciting speech for data collection is difficult. One promising approach has been to ask speakers to describe images. But how do we build models from such visually grounded speech data? Given a dataset of images with Hindi spoken captions, we consider how we can map a written English keyword to spoken realisations of that word in Hindi. Previous work trained end-to-end multimodal neural models. Instead, we explore a simpler alignment-based approach built on self-supervised speech representations. Written English tags are automatically obtained from images using off-the-shelf image captioning systems. Hindi utterances associated with the same keyword are then aligned (using self-supervised features), and alignment evidence is aggregated to identify recurring speech segments corresponding to the target word. Experiments evaluating keyword spotting and localization show that our alignment-based approach outperforms a previous attention-based neural model. We also show the benefit of incorporating negative examples during alignment. Our work demonstrates that cross-lingual word-to-speech mappings can be learned directly from visual grounding without transcriptions or explicit model training.


3. SURE-Challenge: Evaluating Speech Evidence Before Speech-LLM Generation

Authors: Mengzhe Geng

Categories: eess.AS, cs.CL, cs.SD Score: 7.30/10 (Obj:9 Id:9 Ind:5 Comp:6 Eff:7 Nov:7 Rep:8)

  • Strength: 首个针对 Speech-LLM 生成前准入决策的评测基准,冻结的 energy+Whisper 分数规则使 Qwen2-Audio 非支持拒绝率从 15/204(0.074)升至 196/204(0.961)且支持准确率不变(0.930)、下游调用减少 41%,并在六个骨干上回放一致。
  • Weakness: 跨域效用边界明确但无缓解方案(Common Voice 保留率在 τ=0.90 降至 0.505、重叠语音仅拒 0.092、babble 种子间 18–24/54),474 例小规模单作者评测无第三方复现,且两条近邻引用(AHa-Bench、HalluAudio)经检索未能核实存在性。
  • Full review: Claude Code 全文七公理审稿

Speech LLMs are usually graded after they answer, although an operating system first has to decide whether to send a waveform to the model. We define the Speech-Unsupported Rejection Evaluation Challenge (SURE-Challenge) for this admission step. The benchmark pairs LibriSpeech-derived transcription and first-word question answering with unsupported silence, colored noise, synthetic tones, and source-ambiguous babble under disjoint source splits. Front-end ablations use Qwen2-Audio; the selected energy-plus-Whisper-score rule is then replayed before six speech/audio LLMs. On the leakage-screened 474-example SURE-Extended test set, raw Qwen2-Audio rejects 15/204 unsupported inputs, whereas the fixed rule rejects 196/204 and leaves supported accuracy unchanged. External evaluations qualify this result: Common Voice retention drops as the Whisper-score threshold is tightened, and no-speed babble gives 18 to 24 rejected clips out of 54 across regenerated seeds. The result identifies a pre-generation error mode missed by answer-only scoring.


4. Alias-Free Oscillator Synchronization via Additive Synthesis

Authors: Jonas Roth, Domenic Keller, Oscar Castañeda, Christoph Studer

Categories: eess.AS, eess.SP | To be presented at DAFx26 Score: 7.30/10 (Obj:9 Id:8 Ind:5 Comp:5 Eff:8 Nov:7 Rep:8)

  • Strength: 将振荡器同步统一为作用于任意带限波形傅里叶系数的线性谱重采样变换,支持 hard/mirrored/pulsar 三种模式,并流片 HASY ASIC(6 mm²、65 nm、96 kHz/24 bit、512 谐波、变换 5 个采样周期内完成、250 MHz/242 mW、SINAD>41 dB),Python 参考实现开源。
  • Weakness: 全部评测基线为作者自建,未与 BLIT/BLEP/PTR 等既有方法做任何定量对比且无听感实验;O(N²) 变换无算法层压缩(软件仅 172–370 Hz 更新率),当前流片芯片存在控制逻辑 bug,ASIC RTL 未开源。
  • Full review: Claude Code 全文七公理审稿

Oscillator synchronization is a widely used sound-synthesis technique, but straightforward digital implementations suffer from aliasing artifacts. This paper presents an alias-free method for digital emulation of oscillator synchronization of arbitrary periodic waveforms based on additive synthesis. Starting from a finite set of Fourier-series coefficients representing a bandlimited free-running waveform, we derive linear spectral-resampling transforms that map these coefficients to those of the bandlimited synchronized waveform. Beyond conventional hard synchronization, the proposed approach also supports two additional soft-synchronization modes. To address the high computational complexity of the proposed method, we introduce HASY, a 6 mm^2 application-specific integrated circuit (ASIC) fabricated in 65 nm CMOS technology. HASY generates one 96 kHz, 24 bit alias-free synchronized waveform with up to 512 harmonics and computes the spectral-resampling transform within only five audio-sample periods.


5. A Frequency-Domain Artificial Reverberator Plug-In

Authors: Jonas Roth, Nishanth Kumar, Silvan Krebs, David Wieland, Christoph Studer

Categories: eess.AS | To be presented at DAFx26 Score: 7.30/10 (Obj:9 Id:6 Ind:5 Comp:9 Eff:9 Nov:5 Rep:9)

  • Strength: 五步 STFT 噪声载波混响算法紧凑优雅,交付为 VST3/AU 开源插件(AGPL-3.0,v1.1.0),单实例 REAPER CPU <2.6%(M2 Pro),N=2^5–2^14 全程可调,附 Python 参考实现与 DAFx26 演示。
  • Weakness: 全文无任何评测:无听感测试、无客观指标、无与既有 magnitude-decay/稀疏频域混响器的对比,创作扩展价值仅有作者自述,且核心机制为既有谱衰减概念的参数级组合。
  • Full review: Claude Code 全文七公理审稿

We present FDverb, a frequency-domain artificial reverberator, based on the idea of a vocoder with a noise carrier signal. Using a short-time Fourier transform (STFT) for analysis and synthesis, FDverb generates late reverberation by weighting spectral noise components with envelopes. We extend FDverb with early reflections, nonlinear decay, and pitch shifting. These extensions enable creative sound-design applications. We provide FDverb as an open-source DAW plug-in, using the JUCE framework.


6. Evaluating Loss Functions in Differentiable Out-of-Domain Sound-Matching with Partial Parameter Distance

Authors: Amir Salimi, Daniel Penner, Kalvin Eng, Abram Hindle, Osmar R. Zaïane

Categories: cs.SD, cs.AI Score: 7.20/10 (Obj:9 Id:8 Ind:5 Comp:6 Eff:7 Nov:8 Rep:6)

  • Strength: 首次系统评估 OOD 可微声音匹配中的损失函数:4 损失 × 7 场景 × 每格 ≥300 试验 + 盲听验证,提出 PPD 指标并取得与人类听测 5/7 场景的最优损失一致率(听测信度 ICC3k=0.86)。
  • Weakness: 听测面板全部为 4 位作者本人且关键参数由同一批人选取,独立验证不足;代码未公开、真实录音目标与高维参数空间未验证,PPD 无法泛化为通用指标。
  • Full review: Claude Code 全文七公理审稿

In out-of-domain (OOD) sound-matching, a synthesizer is optimized to mimic a sound it did not generate. OOD evaluation of loss functions is underexplored in part because the standard “parameter loss” metric requires a shared parameter space between target and imitator, which OOD settings lack. We introduce Partial Parameter Distance (PPD), which applies parameter loss only to the critical parameters that mismatched synthesizers share (e.g., filter cutoffs), enabling automatically evaluated OOD experiments; we verify its results with blinded listening tests. Across seven scenarios involving band-pass filtering, amplitude modulation, and pitch-bending, we evaluate four differentiable loss functions (SIMSE_Spec, L1_Spec, JTFS, DTW_Envelope). Loss-function effectiveness remains tightly coupled to the method of synthesis: SIMSE_Spec excels at filter-cutoff recovery, DTW_Envelope at amplitude-modulation recovery, and JTFS at smooth pitch trajectories. Parameter-based evaluation agrees with listening tests on the top-ranked loss function in five of seven scenarios, demonstrating its utility as a diagnostic tool.


7. A Mixed-Behavior Vote Model for Multimedia Subjective Quality Votes, Means, and Variances

Authors: Jaden Pieper, Stephen D. Voran

Categories: cs.MM, cs.SD, eess.AS | Submission venue TBD Score: 7.00/10 (Obj:8 Id:6 Ind:6 Comp:8 Eff:7 Nov:7 Rep:7)

  • Strength: 提出 UVR 与四次多项式方差模型,仅 2 参数使 16 个跨语音/图像/视频数据集中 13 个无加权拟合全程落在可行方差区域内,纠正了旧抛物线模型系统性违反最小方差下界的问题,并给出 a=0.05/a=0.54 两条数据质量诊断阈值。
  • Weakness: 缺少下游任务量化验证——未展示新方差模型对 MOS 预测、样本量设计或异常实验检测的任何可测量改进,且加权拟合权重选择无形式化准则、仓库无 license、α(y) 无统计推断支持。
  • Full review: Claude Code 全文七公理审稿

The relationship between subjective test vote variance and vote mean (or MOS) is well-studied, and the mathematically admissible vote variance region has been previously defined. We propose a reduced admissible variance region called the Unimodal Variance Region (UVR) that better describes real subjective rating behavior of multimedia. Further, subjective vote variance is often modeled as parabolic. We explain that, in practice, the parabolic model often violates the admissible region in the variance vs. MOS plane and we propose alternatives that respect the admissible region. We also present a parametrized random process to model votes that mixes voting processes and produces a realistic range of vote variances within the UVR at any desired MOS. This process was inspired by and comports with voting behavior that is observed in many subjective tests. By modeling vote variance from a subjective experiment, this vote model offers additional interpretable insights into voting behavior observed in a given experiment. We present example results from 16 datasets spanning speech, image, and video subjective quality experiments.


8. Direct or Mediated? Task-Dependent Audio Information Routing in Large Audio Language Models

Authors: Yizhou Zhang, Wangjin Zhou, Xin Gu, Yichi Wang, Wei Tan et al.

Categories: cs.SD, cs.MM | 8 pages, 2 figures Score: 6.84/10 (Obj:8 Id:8 Ind:8 Comp:8 Eff:6 Nov:7 Rep:3)

  • Strength: 跨 4 个 LALM 的受控拼接实验揭示任务依赖鲁棒性差距(ASR WER 仅升 0.10–1.30pt,AQA 准确率最高从 93.83% 跌至 35.33%),并以敲除归因+线性探测的双重证据链将失败定位于 prompt 中介信息的下游检索瓶颈(AQA 崩溃时音频属性仍以 67.0–90.1% 准确率线性可解码)。
  • Weakness: 无代码无数据发布(GitHub 搜索零结果),诊断止步于瓶颈发现而未验证任何缓解手段,且波形直接拼接与真实多音频场景的差距未量化、缺等长单段对照排除长度混淆。
  • Full review: Claude Code 全文七公理审稿

Large Audio Language Models (LALMs) have demonstrated strong performance across a wide range of audio understanding tasks. However, they are typically evaluated on single, coherent audio segments, leaving their behavior under less familiar input configurations underexplored. We study this issue through a controlled setting in which two audio segments are concatenated into a single input. Across multiple LALMs, we observe a striking task-dependent robustness gap: automatic speech recognition (ASR) remains comparatively stable, whereas audio question answering (AQA) degrades substantially. To investigate the mechanisms underlying this disparity, we analyze how audio information is routed through LALM decoders using layer-wise attention knockout. The results reveal distinct task-dependent pathways. ASR relies primarily on direct retrieval from audio tokens by answer tokens, whereas AQA depends more strongly on a mediated route in which audio information is first integrated into prompt tokens and subsequently accessed during generation. We further probe prompt-token representations under audio concatenation and find that task-relevant audio attributes remain readily decodable, particularly in middle and later decoder layers, even when AQA performance deteriorates sharply. This dissociation indicates that the failure cannot be explained by complete loss of audio information from the decoder states and is instead consistent with a downstream bottleneck in retrieving or utilizing prompt-mediated information during answer generation. Together, our findings reveal task-dependent audio information routing in LALMs and highlight information utilization as a potential limitation on their generalization.


9. Topic Matching in the Wild: Benchmark and Lessons from Real-World ASR Transcripts

Authors: Saman Rahbar, Xiliang Zhu, Irvin Cardoza, David Rossouw

Categories: cs.CL, cs.AI, cs.LG | Accepted at the 11th Workshop on Natural User-generated Text (W-NUT 2026), EMNLP 2026. Camera-ready version. 9 pages, 2 figures, 3 tables Score: 6.84/10 (Obj:8 Id:8 Ind:6 Comp:8 Eff:7 Nov:6 Rep:5)

  • Strength: 首个基于真实生产 ASR 的 topic matching 基准(11 家公司、173 话题、2,655 对三人标注、Fleiss’ κ = 0.660),证明轻量 LLM + 自然语言描述达 F1 = 0.847,超 regex(0.721)12.6 点(McNemar p < 10⁻¹⁴),且 Flash-Lite 以 16 倍速度、500 倍成本优势 Pareto 支配 Pro 级模型。
  • Weakness: 评测数据因隐私不可发布且无配套代码仓库(GitHub 搜索零结果),核心结果与标注质量无法被外部检验;缺干净文本对照使噪声鲁棒性与一般语义能力无法解耦,表示×匹配器设计非正交、单一 prompt 模板无敏感性分析。
  • Full review: Claude Code 全文七公理审稿

In contact centers, real-time agent-assist tools determine, for each of many predefined topics, whether a live customer utterance is relevant and display a coaching card to the agent when it is. The input is noisy and challenging: ASR(Automatic Speech Recognition) transcripts of spontaneous phone conversations, which can be unclear, repetitive, and mostly lack punctuation. To systematically study this real-world task, we curate a human-annotated topic-utterance judgments dataset sourced from real call-center transcripts. We compare three types of matchers: a regex-based baseline, zero-shot sentence embedding encoders, and Gemini-based LLM matchers. In addition, two types of topic representations are studied in our benchmark:keyphrases and natural language description. Our empirical experiments highlight the superior performance of lightweight LLM matchers over embedding and regex models when equipped with natural language descriptions.


10. RoboGesture: Real-Time Semantic-aligned Co-Speech Gestures Generation for Humanoid Interaction

Authors: Zifan Wang, Ziang Ren, Pengyang Shi, Zirui Wang, Chenghuai Lin et al.

Categories: cs.RO, cs.CV | Accepted at ECCV 2026. Project page: https://RoboGesture.github.io Score: 6.80/10 (Obj:8 Id:10 Ind:6 Comp:6 Eff:16 Nov:14 Rep:4)

  • Strength: 在 Unitree G1 上实现全链路实时语义手势交互:FGD 0.8452 达最强基线 DiffSHEG 2.2316 的 2.6 倍好,MPC 自碰撞率 4.16%→0.13%,动作生成 120 FPS,人类评估四维全胜。
  • Weakness: 官方代码仓库创建当日仍为空(0 文件、无权重、数据集未确认公开),SemanticBEAT 自建无真值动作、碰撞指标与自家 MPC 循环依赖,关键超参无敏感性分析且无统计显著性报告。
  • Full review: Claude Code 全文七公理审稿

Enabling humanoid robots to respond to human speech with synchronized and semantically meaningful gestures is fundamental to natural human-robot interaction. However, this task faces three critical barriers: the scarcity of semantically rich datasets, the “modality eclipse” where models ignore audio cues in favor of kinematic inertia, and the sim-to-real gap regarding physical safety. We propose RoboGesture, a robot-centric framework that co-designs data, modeling, and control to power a complete interactive human-humanoid system in which the robot listens, responds, and gestures in real time. We first establish the RoboGesture dataset featuring over 300 gesture categories and develop an automated pipeline to synthesize large-scale collision-free, robot-specific audio-motion pairs. Our architecture features a Hierarchical Semantic-Acoustic Aligner that extracts multi-granular prosodic and semantic cues directly from raw audio tokens. These cues drive a Streaming Conditional Motion Generator based on a diffusion transformer with conditional flow matching. To ensure high responsiveness, we introduce Anti-Inertia CFG Masking, which prevents the model from collapsing into repetitive historical patterns by compelling it to proactively mine control signals from the audio modality. Finally, an MPC-based safety filter ensures real-time, collision-free execution on physical hardware. Experiments on a Unitree G1 humanoid demonstrate that RoboGesture generates safer, more rhythmic, and more semantically appropriate responses compared to state-of-the-art baselines.


11. Multi2AV-Safety: Benchmarking Safety in Multimodal-to-Audio-Video Generation

Authors: Kaichao Jiang, Changtao Miao, Baiqi Wu, Zhiyuan Lu, Kang Yang et al.

Categories: cs.AI Score: 6.60/10 (Obj:8 Id:6 Ind:6 Comp:8 Eff:6 Nov:8 Rep:4)

  • Strength: 首个覆盖全部 11 种 T/I/A/V 条件组合的音视频生成安全基准(11,024 实例,4 机制 × 5 危害类),用 e=(C,k,ρ) 框架量化出 Composed Harm 下全模态护栏 recall 仅 39.2–46.2%、Diluted Harm 下 k=1 时 recall 低至 58.0%、RHR 随模态数从 32.8% 升至 40.4% 的组合式风险缺口。
  • Weakness: 数据集与代码截至 2026-09 完全未发布(承诺 2026-10 且高危部分限权访问),GitHub 0 结果;全模态 guard 仅 2 个、无与既有多模态安全基准的 head-to-head;标注依赖 Claude/GPT 生成并同族质检且未报告人工一致率;未测良性输入误杀率,无法校准安全-可用性权衡。
  • Full review: Claude Code 全文七公理审稿

Audio-video generation is rapidly moving from prompt-driven synthesis toward multimodal conditioning, where text, images, audio, and video can jointly shape the generated output. This shift changes the nature of safety evaluation: harmful intent may no longer reside in any single input, but instead emerge from how otherwise benign or weakly harmful conditions interact across modalities and time. Existing safety benchmarks, however, remain largely prompt-centric or tied to fixed conditioning interfaces, leaving such compositional risks difficult to study systematically. To bridge this gap, we introduce Multi2AV-Safety, the first safety benchmark, to the best of our knowledge, to cover all 11 non-singleton T/I/A/V conditioning configurations for audio-video generation, comprising 11,024 attack instances. Evaluation on Multi2AV-Safety reveals systematic weaknesses in representative multimodal safety guards across attack mechanisms and harm-evidence structures. Our evaluation reveals two complementary failure modes: harmful semantics can emerge from the combination of individually benign inputs, while explicit harmful cues can become harder to detect when mixed with benign multimodal context. Together, these results identify \emph{compositional risk perception} as a central capability gap in safeguarding multimodal-conditioned audio-video generation: current safety guards fail to reliably integrate safety evidence across modalities and time, even when all conditioning inputs are observable. The dataset will be publicly released in October 2026.


12. Your Voice Cloning System is Secretly a Voice Anonymizer

Authors: Romolo Muletta, Felix Matthias Saaro, Mark Cieliebak, Jan Deriu

Categories: cs.CL Score: 6.47/10 (Obj:8 Id:5 Ind:4 Comp:6 Eff:8 Nov:6 Rep:8)

  • Strength: 零训练复用 XTTSv2 实现七欧洲语言匿名化,EER 0.49(近最优 0.5)、CV WER 0.16 相比 VoicePrivacy 2022 冠军 SALT 相对降约 38%,ΔUTMOS +0.17 显著优于基线 −0.74/−0.62,CMOS 人类听测亦最佳。
  • Weakness: ECAPA2 既构造伪说话人又当 ASV 评测器存在循环依赖、无任何第二攻击模型交叉验证,全文零消融使三个自研组件的增益无法归因,人类听测仅 30 条英语语句且避开方法脆弱情形。
  • Full review: Claude Code 全文七公理审稿

Speaker anonymization suppresses speaker-identifying attributes from speech while preserving linguistic content and quality. We propose repurposing XTTSv2, a multilingual voice cloning model trained on 27k hours of speech, for speaker anonymization without retraining. Our key insight is that XTTSv2’s voice cloning capabilities preserve prosodic structure independently of speaker identity, enabling voice conversion by conditioning on a pseudo-speaker. We introduce an iterative refinement strategy that balances privacy and utility by maximizing a harmonic mean of speaker dissimilarity and intelligibility. Evaluated on seven European languages across CommonVoice and Multilingual LibriSpeech, our system achieves near-optimal privacy (EER $\approx$ 0.49), competitive intelligibility, and substantially better speech quality than dedicated anonymization baselines, while requiring no language-specific training. We release the code here: https://github.com/rm00cr/coqui-tts.


13. Recovering Expert Critic-Sourced Network Adjacency between Musical Artists from Acoustic Distributions: A Construct-Validity Approach

Authors: Elena Badillo-Goicoechea, Fengfeng He

Categories: stat.ML, cs.LG | Accepted workshop paper at USR Workshop, RecSys, 2026, Minneapolis, MN, USA Score: 6.40/10 (Obj:8 Id:5 Ind:7 Comp:8 Eff:5 Nov:8 Rep:4)

  • Strength: 在 artist-disjoint 冷启动划分下用 80 维 Essentia 分布 + 边际 Wasserstein 距离以 AUC 0.767 (CI 0.761–0.775) 外部验证了乐评邻接信号,可恢复性随乐评共识单调升至 0.865,并与流派社会学类型学分层对齐。
  • Weakness: 全文无任何对照基线(朴素声学/流派/流行度均未比较),所有 AUC 均在 3:1 人为负采样下取得而真实图密度仅 8.6×10⁻⁴,未做 top-k 排序评测,且无代码或语料开源。
  • Full review: Claude Code 全文七公理审稿

Music recommendation relies primarily on two signals: user-item interactions, which fail in the cold-start regime, and intrinsic musical content, available for any recording. We argue that a third, largely untapped signal is both richer and more principled: critical adjacency, the pairwise relation established when an expert critic explicitly links two artists in long-form prose. It encodes deliberate judgments about which artists belong together. Prior work established its internal validity, showing it recovers coherent, interpretable communities and can match collaborative filtering in user-satisfaction simulations, with no user data. What has been missing is external validation: whether this critic-sourced relation is grounded in the music itself versus sociological context. We test it against acoustic content, reframing the question as one of construct validity. Representing artists as empirical distributions over 80 low-level Essentia acoustic descriptors and modeling pairwise proximity via marginal optimal-transport (Wasserstein) distances, we evaluate how far critical adjacency is sonically recoverable under a cold-start, artist-disjoint split. Our ensemble recovers these edges at out-of-sample AUC of 0.767 (95% CI 0.761-0.775). Recoverability rises monotonically with critical consensus, reaching 0.865 on multi-source attested edges. Stratified evaluations align with sociological models of genre: tightly bounded, scene-based genres show higher recoverability than broad industry umbrella terms. Critical discourse is thus a rich source of information for recommendation, decomposing into a reproducible “sonic core” and a “sociological remainder” driven by narrative positioning, subcultural context, and canonical placement. The work offers both a scalable cold-start discovery mechanism and a sociologically grounded approach to MIR and MRS research.


14. PolyMap: A 64-Channel Polyphonic Guitar Pickup System

Authors: David Wieland, Jonas Roth, Christoph Studer

Categories: eess.AS, eess.SP | To be presented at DAFx26 Score: 6.26/10 (Obj:8 Id:5 Ind:6 Comp:7 Eff:6 Nov:6 Rep:7)

  • Strength: 64 通道逐弦多点拾音系统完成度高:端到端延迟 5.4 ms(低于 Presonus 基线 6.1 ms,硬件仅 5 采样),噪声归因与逐位置频谱验证扎实,硬件设计文件(CERN-OHL-P-2.0)+ 插件 + 示例音频全开源。
  • Weakness: 核心功能虚拟拾音位置仅有线性插值描述、无任何音色保真度验证,−67~−73 dBFS 拾音器嘶声与仅 Reaper 可用的 64 通道限制未经听感与跨平台评估,摘要声称的下游研究应用(逐弦转录/tone transfer)零实际演示。
  • Full review: Claude Code 全文七公理审稿

In electric guitars, the vibrations of the strings are typically sensed by coils of wire combined with a magnet, called pickups. The pickups and their position along the strings contribute strongly to the instrument’s sound. Most guitars feature one to three pickups, each spanning across all strings with fixed positions and generating a single mono output. The work of this Master’s Thesis at ETH Zürich introduces a new pickup system called PolyMap, which senses each string individually and at multiple locations. The system is demonstrated with a custom-made eight-string guitar that contains eight pickups per string for a total of 64 pickups. The signals from these 64 pickups are individually digitized inside the guitar and transmitted over a multichannel audio digital interface (MADI), a low-latency digital audio interface, to a computer for further processing. PolyMap enables high-resolution sensing of an electric guitar’s strings and enables extensive post-processing capabilities for musicians, audio engineers, and researchers. To the best of our knowledge, this is the first polyphonic guitar pickup system with such a complete feature set.


15. When Text Misleads: Inconsistent-Aware Reasoning for Audio-Grounded Dialogue

Authors: Yen-Ju Lu, Yuzhe Wang, Yaohan Guan, Xiluo He, Jiarui Hai et al.

Categories: cs.CL, cs.AI, cs.LG, eess.AS | 24 pages, 4 figures Score: 6.21/10 (Obj:8 Id:6 Ind:7 Comp:6 Eff:5 Nov:8 Rep:3)

  • Strength: 把转录捷径形式化为跨模态分歧并构建 501 题冲突/一致双集基准 ContraTalk,实证文本-only LLM 在一致题 91.7–98.2% 而冲突题跌至 33.0–47.7%,直接 Audio-LLM 仍有 29.7–39.9% 掉入转录陷阱,并有 Audio Twin 无训练证据接口(冲突题最高 50.5%、mislead 降至 29.4%)。
  • Weakness: 方法改进缺乏统计稳健性(Sonnet+AT 冲突题 +2.8pp 且 CI [44.9, 55.9] 与纯文本重叠、一致题骨干依赖退化最高 11.4pp)、agentic 管线与证据阈值无组件级消融,且无代码/数据发布、构建依赖受限的 Seamless Interaction Dataset,基准不可复现。
  • Full review: Claude Code 全文七公理审稿

Understanding spoken dialogue requires joint reasoning over lexical content and paralinguistic acoustic signals such as emotion and conversational intent. However, existing evaluations often allow shortcuts based on transcripts or single-modality solutions, obscuring whether models genuinely ground predictions in speech. We formalize this failure mode as cross-modal disagreement, where transcripts suggest plausible but incorrect surface interpretations while acoustic cues such as prosody or speaking style support different answers. We develop a scalable framework that identifies text-biased surface interpretations and converts disagreement regions into conflict QA examples. We also include consistent cases where transcript-based and speech-grounded interpretations agree, enabling evaluation beyond adversarial audio dependence. This results in ContraTalk, a controlled benchmark containing 501 questions across five discourse dimensions: interaction behavior, emotion state, dialogue act, social stance, and conversational intent. We further develop an agentic-style reasoning framework that converts speech into an Audio Twin, a text-readable representation of localized acoustic cues that exposes acoustic evidence to the reasoning model. Experiments show that strong text-only LLMs exceed 90% accuracy in consistent cases but drop to 33-48% in conflict cases. Direct AudioLLMs provide only partial grounding, still selecting the transcript-biased trap in roughly 30-40% of conflict cases. Our Audio Twin framework improves conflict-case accuracy while reducing trap selection, but its consistent-case behavior remains backbone-dependent. These results identify transcript-based shortcuts as an important failure mode in spoken dialogue understanding and show that explicit acoustic evidence aggregation provides a more controllable interface for diagnosing and improving speech-grounded reasoning.


16. Soft Active Electromyography Interface for Machine Learning-Enabled Silent Speech Recognition

Authors: Yuta Kurotaki, Shusuke Yamakoshi, Reitaro Yoshida, Yutaka Isoda, Tamami Takano et al.

Categories: cs.LG, cs.HC, cs.SD | 17 pages, 5 figures, supplementary information Score: 5.90/10 (Obj:8 Id:5 Ind:5 Comp:5 Eff:6 Nov:7 Rep:5)

  • Strength: 手戴软体按需 EMG 接口实现 30 词无声语音分类 97.2±1.3%(3 名受试者)并完成四命令无人机实时控制,LM 互连在 1000 次循环/50% 应变下电阻波动 <±0.5%、SNR 31.0-35.5 dB,物理层按需采集提供无需算法的隐私机制。
  • Weakness: ML 流水线(MFCC+DNN)沿用既有工作且未与硬件贡献做因果分离,评测仅 3 人单会话结构化逐词输入、无跨会话重贴附验证、无人机演示无成功率/延迟数字,数据与代码在正式发表后仍未公开。
  • Full review: Claude Code 全文七公理审稿

Silent speech recognition (SSR) provides an alternative communication pathway in the absence of audible speech. However, conventional approaches are limited by the need for constant facial attachment, privacy concerns, and unstable signal acquisition. Here, we propose a soft, active electromyography (EMG) interface that enables word-level SSR using machine learning. Worn on the hand, the device uses a fingertip electrode that can be positioned near the lips to acquire EMG signals only when needed. The interface integrates liquid metal (LM) interconnects, transparent flexible printed circuit (FPC) electrodes, and elastomer encapsulation to ensure high mechanical stability during finger motion. A deep neural network trained on these stable signals achieved a mean accuracy of 97.2 $\pm$ 1.3% across three subjects in classifying a 30-word vocabulary, demonstrating robust linguistic discrimination. Furthermore, real-time drone control validates the practicality of this approach in noisy and privacy-sensitive environments where conventional voice recognition fails. This study highlights the potential of soft, wearable EMG systems as secure and intuitive human-machine interfaces.


17. Scaling phoneme-based TTS augmentation for ASR: A unified pipeline and controlled study

Authors: Zhen Wang, TianRui Wu, RongQi Han, Hao Wu, Wei Liang

Categories: cs.CL Score: 5.89/10 (Obj:8 Id:6 Ind:6 Comp:7 Eff:6 Nov:5 Rep:4)

  • Strength: 4 语言 13 测试集受控拆分证实合成规模/文本选择/参考质量为独立控制变量,低资源葡萄牙语(25.7h 真实数据)获 6.3–9.0 绝对 WER 下降,PFGS 零成本实现并达最大 19.3% 相对降低。
  • Weakness: 无代码/语料开源且阿拉伯语候选文本非公开,全部结果为单次运行无方差报告,且未与 SpecAugment、现成开源 TTS 或 Whisper 微调做任何成本效益对比。
  • Full review: Claude Code 全文七公理审稿

Synthetic speech provides scalable supervision for automatic speech recognition (ASR), but its benefit depends on the selected texts, reference speech, and amount of synthesized data. We present a unified phoneme-based TTS-to-ASR augmentation pipeline built around a multilingual TTS model trained from scratch using the F5-TTS architecture with language-ID conditioning. The pipeline combines language-specific grapheme-to-phoneme conversion, reference-speech filtering, candidate-text selection, synthesis, and matched ASR continuation. We further propose phoneme-frequency-guided selection (PFGS), which ranks candidate sentences using phoneme frequencies estimated from real ASR training labels. Experiments with separate monolingual ASR systems for Arabic, French, Italian, and Portuguese span 13 test sets. Across the synthesis-scale sweep, random augmentation improves over matched real-only continuation on 11 test sets. Under a nominal 60% synthesis budget, PFGS improves over real-only training on 12 test sets and over random selection on 9. Its largest relative word error rate (WER) reduction against random selection is 19.3%. With target texts and synthesis counts fixed, reference-speech filtering reduces absolute WER by 0.29 and 0.59 points on Italian and French Common Voice, respectively. These results identify synthesis scale, candidate-text content, and reference quality as important control variables in TTS-based ASR augmentation.


18. Compositional Failure in Audio-Visual LLMs: Late-Layer Prior Dominance Under Cross-modal Conflict

Authors: Adarsh Sudheer, David Li, Omar Elbanna, Ishaan Kodarapu, Arjun Bahuguna et al.

Categories: cs.CL, cs.AI | Accepted to the 2nd Workshop on Compositional Learning at ICML 2026. 7 pages, 4 figures Score: 5.74/10 (Obj:8 Id:5 Ind:6 Comp:5 Eff:5 Nov:6 Rep:6)

  • Strength: 在 VideoLLaMA 2-7B-AV 上证明三种对齐配置(49.0-50.2%)均无法超过 base 的 51.7%,InternVideo2 在冲突子集掉 32.3 点,且对齐只移动预测分布(Yes 占比 56.9%→35.0%)不改冲突准确率,机制审计定位先验承诺稳定在 25.5±1 层、21 个冲突解决头在层 15-18。
  • Weakness: 核心”层 15-18 冲突计算被层 25.5 先验覆盖”的因果链仅相关性支撑,steering/damping 等干预全部留作 future work,全文无任何冲突解决改善方案,且机制证据限于单架构、审计阈值无统计校准、开源 repo 无 license 无第三方复现。
  • Full review: Claude Code 全文七公理审稿

We study audio-visual conflict as a compositional generalization test for AV-LLMs: the model must combine synchronized but semantically incompatible audio and video evidence and decide whether the pair matches. On VideoLLaMA 2-7B-AV, three alignment configurations remain nearchance on the scored exact-string Yes/No subset of AVHBench, even though their output priors shift substantially. Similarly, off-the-shelf InternVideo2 experienced a 32.3% accuracy decrease specifically under cross-modal conflict, accompanied by a 17.3% instruction-following failure. We call this failure mode prior dominance: late-layer commitment to an internally preferred answer pattern that is weakly grounded in the conflicting inputs. To explain this behavior, we conduct a mechanistic interpretability analysis and find that commitment remains concentrated at 25.5 $\pm$ 1 layers. We show that stronger temporal alignment changes answer bias, but do not improve compositional conflict resolution. Code and data to reproduce our mechanistic audit and behavioral evaluations are available at https://github.com/AdarshSudheer09/AVHBench-dmai.