每日 arXiv

音频 · 语音 · 音乐 · 声学

每日新爬取的 arXiv 论文,按评分排序。中文摘要帮你快速判断是否值得深读。

2026-08-31
日期2026-08-31
已评分
均分
最高

Daily Papers — 2026-08-31

32 papers on audio, speech, music, and acoustics.

1. Perceptually Better, Semantically Worse: Measuring Speech Enhancement Impact on LLM-Based Voice Systems

Authors: Randy Frans Fela, Pejman Mowlaee

Categories: cs.SD, eess.AS | Accepted at EMNLP 2026. 12 pages, 4 figures, 8 tables Score: 8.10/10 (Obj:8 Id:9 Ind:6 Comp:8 Eff:8 Nov:7 Rep:9)

  • Strength: 首个量化 SE 前端对下游 LLM 语义影响的评测:2,974 条 SLURP 上 MetricGAN+ 使 ODR 从 0.135 翻倍至 0.318(PESQ 反而 1.76→1.94),模拟回声 ODR 达 0.836(说话人替换,WER 无法刻画),结论跨 Whisper/wav2vec2 排序完全一致(ρ=1.0),代码与 clip 级数据 MIT 开源且被 EMNLP 2026 Industry Track 接收。
  • Weakness: ODR 依赖单一商业 LLM(Gemini 2.5 Flash Lite)闭集意图分类,Pro 模型仅部分复现且倍率缩水(2.36×→1.91×),全部增强条件为 DNS 模拟、每类 SE 只测一个代表模型,开放词表任务与免参考监控代理均缺证据。
  • Full review: Claude Code 全文七公理审稿

Speech enhancement (SE) is commonly applied as a preprocessing step in spoken AI pipelines under the assumption that better audio quality improves downstream task performance. Whether SE-induced distortions propagate to downstream LLM task performance remains an open question. We introduce Output Divergence Rate (ODR), which measures how often SE changes an LLM’s intent classification relative to clean speech, and benchmark five conditions on 2,974 SLURP clips using Whisper large-v3 and wav2vec2-large cascades. Every condition produces ODR significantly above zero ($p < 0.001$, binomial test). MetricGAN{+} more than doubles ODR versus unenhanced noisy speech (0.318 vs. 0.135) despite improving PESQ, and unmitigated echo reaches an ODR of 0.836 through speaker substitution, a failure WER cannot capture. Audio quality metrics range from near-zero to moderate correlation with ODR (SQUIM-MOS $ρ=-0.068$, PESQ $ρ=-0.467$). The MetricGAN{+} and echo results replicate across ASR architectures, indicating that standard audio quality metrics are insufficient for LLM pipeline quality.


2. An Optical Pathway to Movable Rydberg Atomic Quantum Receivers

Authors: Qihao Peng, Qu Luo, Lixia Xiao, De Mi, Neng Ye et al.

Categories: eess.AS Score: 7.50/10 (Obj:9 Id:9 Ind:6 Comp:9 Eff:7 Nov:7 Rep:6)

  • Strength: 提出光学可移动 RAQR 的三因子闭式基带模型(误差 < −20 dB,经 Lindblad 数值解验证),仿真中换能波束整形与光学位移优化合计带来 14 dB + 87 dB 干扰抑制,解析梯度 AO 算法 4–5 次迭代收敛且优于 6 个基准。
  • Weakness: 全部结论基于纯仿真:无硬件实验,87 dB 抑制依赖理想光学位移与完美 CSI,未量化导向误差与重配时延的影响;噪声建模忽略散粒/量子投影噪声的设计依赖;无代码发布,且相对同社区 2026-07 的池级信道整形工作(arXiv:2607.05979)新颖性属于概念组合。
  • Full review: Claude Code 全文七公理审稿

This paper develops an optically movable Rydberg atomic quantum receiver (RAQR), in which the probe and coupling beams are steered within each vapor cell to dynamically reconfigure the effective radio-frequency (RF) sensing position without mechanical actuation. A closed-form equivalent baseband model is derived by separating the atomic transduction coefficient, optical steering phase, and cell-center array response into distinct factors and the accuracy of the resulting model is validated against numerical solutions of the Lindblad master equation. Based on the derived model, we reveal two complementary channel-shaping mechanisms, including intrinsic beam-pattern shaping through RF-to-optical transduction and per-cell phase control enabled by optical displacement. To further exploit these capabilities, a non-convex sum-rate maximization problem is formulated over the optical positions and local oscillator design and solved via an alternating optimization framework with analytical gradients. Simulation results validate the derived model and demonstrate substantial performance gains enabled by optical movability, highlighting its potential as a programmable receiver architecture for future wireless networks.


3. SPHERE: Automatic Music Upmixing via Audio Language Model Post-Training with Spatial Heuristic Rewards

Authors: Zixun Guo, Calvin Murdock, Sanjeel Parekh, W Owen Brimijoin, Simon Dixon et al.

Categories: cs.SD Score: 7.40/10 (Obj:9 Id:8 Ind:7 Comp:9 Eff:6 Nov:8 Rep:5)

  • Strength: 3B 开源学生模型经 Sphere 奖励(6 个感知驱动确定性子奖励)SFT+RL 后训练,客观总奖励 4.981±0.013 超越教师 Gemini 2.5 Pro(4.548)及全部 8 个零样本基线,主观专家评分 5.73 亦反超教师 5.55,且奖励与人类偏好强对齐(440 次 A/B 中 76.8% 高奖励混音被偏好,p=6.65×10⁻³¹)。
  • Weakness: 客观测评全部使用模型自身被优化去最大化的 RSphere,存在循环性且无独立空间真值或专用混音系统基线;RL 增益未转化为感知提升(主观 5.71→5.73),且代码、权重未开源(GitHub 检索 0 结果),教师数据依赖专有 Gemini 2.5 Pro 导致重建成本高。
  • Full review: Claude Code 全文七公理审稿

In this paper, we study the task of automatic music upmixing, wherein a system predicts spatial mixing parameters from a multi-stem recording. Different from existing methods that rely on task-specific music encoders, we approach this task via audio language model (ALM) post-training, leveraging rich representations from existing ALMs, which encode both music semantics and mixing knowledge. Specifically, we propose a post-training recipe that first employs rejection sampling SFT, followed by reinforcement learning (RL) with verifiable rewards (RLVR) via GRPO. We propose Sphere (Spatial Heuristic Rewards), a deterministic reward suite inspired by music mixing conventions, to guide our post-training. It consists of 6 perceptually-motivated sub-rewards and encourages the output mix to be centered, balanced and spacious. More broadly, our results suggest that expert domain knowledge can be encoded as verifiable rewards and distilled into language models, without task-specific architectures.


4. VIBE: Video Instruction-aligned Background music gEneration

Authors: Aryan Vijay Bhosale, Vaibhavi Lokegaonkar, Vishnu Raj, Gouthaman KV, Sreyan Ghosh et al.

Categories: cs.SD, cs.AI, cs.CL, cs.CV, cs.LG Score: 7.40/10 (Obj:9 Id:8 Ind:6 Comp:7 Eff:7 Nov:7 Rep:8)

  • Strength: 提出 Conditioning Connection(逐层加权跨层条件路由,单独加入使 FAD 从 2.538 降至 1.724、约降 32%)与硬可验证奖励(tempo Exact 32.0、MAE 2.4,key Circle-of-Fifths + Krumhansl-Schmuckler)+ 软主观奖励的 5 阶段 RL 课程,主结果 FD 10.128、IS 2.358 全基线最佳,人类 A/B 总体 win rate 65%。
  • Weakness: 评测循环性未讨论(Gemini-2.5-Flash 同源承担训练数据、奖励与终评)、ImageBind 作音画代理、风格奖励寄生于 CMI-RM;指令跟随增益边际(Key Exact 11.0 vs 10.1)、八度等价错误仍在(150 BPM→133.9),且仅支持 10 秒器乐、无演唱。
  • Full review: Claude Code 全文七公理审稿

Current video-to-music (V2M) models lack semantic control and fail to penalize instruction violations, largely due to their reliance on reconstruction objectives and the representational bottleneck of static cross-modal conditioning in Diffusion Autoregressive (DAR) architectures. To resolve this, we introduce VIBE, a novel text-and-video-to-music (T+V2M) generation model that leverages: (1) Conditioning Connection, a depth-wise cross-layer conditioning mechanism that dynamically bridges the planning and diffusion refinement heads and (2) a comprehensive reward modeling taxonomy, optimizing for both hard, verifiable constraints (e.g., tempo, key) and soft, subjective qualities (e.g., musicality, multimodal alignment) with a structured 5-stage training curriculum. Upon evaluation using audio-visual alignment, instruction following, and audio quality metrics, along with a subjective human evaluation study, we observe that VIBE demonstrates enhanced controllability and instruction adherence while performing comparably to most evaluated baselines on generation fidelity and multimodal alignment.


5. Vocal Music under Phoneme-Conditional Analysis

Authors: Hayoon Kim, Kyogu Lee

Categories: cs.SD, cs.CL | Accepted to the 27th International Society for Music Information Retrieval Conference (ISMIR 2026) Score: 7.40/10 (Obj:8 Id:8 Ind:6 Comp:8 Eff:6 Nov:8 Rep:8)

  • Strength: 提出同曲内标记-对照匹配的音位条件分析,覆盖 9 语言 6 语系 5,024 首歌,对齐误差 23ms(NUS-48E 验证),歌曲级特征在歌手分组 9 分类达 85.5% 平衡准确率(随机 11.1%),最大单效应普通话卷舌 MFCC-1 d=−0.80(25,899 对),已获 ISMIR 2026 接收并同日开源全套管线与验证脚本(MIT)。
  • Weakness: 语言排序主要由音色通道驱动而该通道携带人声分离后残余伴奏的未控线索(去音色列后所有语言行重排),H2 累积性预测在 n=9 上未获支持(ρ=+0.27/+0.50,p=0.49/0.17),多数单格效应d<0.15 且无感知可听度实验,语言可分性与音位局部效应之间的因果链条留白。
  • Full review: Claude Code 全文七公理审稿

The vocal music of each language carries a distinctive sonic identity, even without instrumental accompaniment. We ask whether these differences are measurable and traceable to specific phonemes. To tackle this question, we introduce phoneme-conditional analysis, which isolates the acoustic effect of typologically distinctive phonemes by comparing marker syllables against matched non-marker controls within the same song, holding singer, melody, and genre constant. Across nine typologically diverse languages and thousands of songs, we measure effects along five acoustic dimensions. Song-level profiles built from these effects identify the language of an unaccompanied vocal at 85.5% balanced accuracy in a nine-way classification with folds grouped by artist; whether the separability arises by accumulation of the phoneme-local effects themselves is left open. Our findings suggest that phonological structure leaves systematic and measurable traces in how each language is sung.


6. When Does Predictor-Based RL Align with Human Perception? A Study of Subjective Rewards in Codec-Based Speech Language Models

Authors: Joonyong Park, Jerry Li

Categories: cs.CL, cs.SD, eess.AS | Submitted to EMNLP 2026 Score: 7.40/10 (Obj:9 Id:7 Ind:5 Comp:8 Eff:8 Nov:7 Rep:8)

  • Strength: 在同一 CER-zone 奖励门下首次系统对照 GRPO 与 Best-of-8 的人类感知差异(HWR 52.0/46.0/48.0,接近随机),并用 1000 票奖励差距逻辑回归(OR=1.93, 95% CI [1.38, 2.70], p<.001)证明主观奖励平均迁移存在且与残留 CER 混淆无关,给出四条可操作的多奖励选型诊断。
  • Weakness: 所有结论绑定单一 Llasa-1B/XCodec2 日文动漫域且无多 seed 与跨模型验证,AnimeScore 轴在奖励差距分析中不显著、Likability 出现机器 +0.17 与人类 HWR 36% 的脱钩,核心「预测器-人类对齐」命题在最关键轴上未获完全识别。
  • Full review: Claude Code 全文七公理审稿

Codec-based text-to-speech (TTS) models make language-model post-training applicable to speech generation, but it remains unclear when learned perceptual predictors can serve as reinforcement learning rewards without losing alignment with human listeners. We study this question with Group Relative Policy Optimization (GRPO) using learned rewards for anime-like speaking style, naturalness, likability, and arousal. To prevent perceptual rewards from being optimized through transcript drift, we introduce a character error rate (CER) zone constraint and compare policy optimization with Best-of-$N$ reranking under the same reward gate. Across single-reward runs, each reward primarily improves its own target metric, showing that subjective predictors are not interchangeable quality surrogates. Multi-rater A/B tests further show uneven human transfer, while a reward-gap analysis separates average transfer from within-axis calibration: signed reward gaps significantly predict listener choices in the pooled analysis, whereas residual CER gaps do not, but per-axis calibration remains heterogeneous. Best-of-8 is a strong human-level baseline and is not clearly worse than GRPO perceptually, suggesting that GRPO should be viewed as amortizing reward-selected behavior into the policy rather than uniformly outperforming reranking. These results support analyzing subjective speech rewards as predictor-axis-base tuples and provide practical diagnostics for selecting rewards before multi-reward speech post-training.


7. Ouroboros: Self-Referential Backdoor Attacks on Speech Enhancement via Clean Audio Triggers

Authors: Yunjie Zhou, Yuheng Huang, Diqun Yan

Categories: cs.SD, cs.CR | Accepted at INTERSPEECH 2026. This is the author-accepted manuscript, not the ISCA proceedings camera-ready publisher version. 5 pages, 2 figures Score: 7.26/10 (Obj:9 Id:8 Ind:8 Comp:8 Eff:6 Nov:7 Rep:6)

  • Strength: 自指 CleanTrigger 机制在 4 模型 × 2 数据集上实现近完美 ASR(99.88%–100%,2% 投毒率即生效),60 条真实手机录音物理激活率 96.7%–100%,且低通/Wiener/微调三类防御全部失效。
  • Weakness: 未提出任何防御设计或缓解分析(防守方零可操作结论),无代码开源(GitHub 检索 0 结果),”首个语音增强后门攻击”的新颖性 claim 未经独立文献索引验证(Semantic Scholar/OpenAlex 查询均失败)。
  • Full review: Claude Code 全文七公理审稿

Speech enhancement models are widely deployed as frontend modules in real-time speech services, yet their vulnerability to backdoor attacks remains unexplored. Existing backdoor methods are confined to classification tasks and rely on active trigger injection, an assumption incompatible with the passive processing nature of speech enhancement models. In this paper, we propose Ouroboros, a novel backdoor attack framework that leverages the ideal clean outputs of speech enhancement models as natural triggers, enabling inference-time activation without any external trigger injection. Extensive evaluations show Ouroboros achieves near-perfect attack success rates with minimal performance degradation on diverse models and datasets. Physical-world validations confirm that naturally recorded, unaltered clean audio can reliably activate the backdoor. Moreover, Ouroboros generalizes to targeted content-tampering attacks and remains effective against common filtering and finetuning defenses.


8. Likelihood-Constrained Acoustic Reranking for Training-Free Hallucination Mitigation in LLM-Based ASR

Authors: Jiasheng Kuang, Linru Zheng, Hongjin Song, Zhaoqi Cui, Song Li

Categories: eess.AS | 5 pages, 4 figures Score: 7.20/10 (Obj:8 Id:6 Ind:6 Comp:8 Eff:8 Nov:7 Rep:7)

  • Strength: 免训练解码干预在 δ=0.60 下消除 38.8–57.1% 的 LLM-ASR 检测器识别幻觉(OpenSpeech 套件 NetHIR +52.8–70.9 点),LibriSpeech/AISHELL2 上 WER/CER 变化 ≤0.16%,并发布 8800 条冻结记录、人工审核精度 93.3% 的挑战基准(HF 实测可访问)。
  • Weakness: δ 超参由主评测模型 Qwen3 的先导实验选出且数据划分未说明,「无需外部检测器」卖点无对比实验支撑,LCAR 本体无代码仓库(GitHub 实测 0 结果),且 3997 个修复事件中仅 28.8% 忠实恢复;S2/OpenAlex 独立索引核查均失败。
  • Full review: Claude Code 全文七公理审稿

Large language model (LLM)-based automatic speech recognition (ASR) systems achieve strong performance on conventional speech data by leveraging powerful linguistic priors and multilingual capabilities. However, under challenging conditions, these priors can override acoustic evidence, resulting in unintended translation, instruction execution, repetition, or catastrophic deletion. We propose Likelihood-Constrained Acoustic Reranking (LCAR), a training-free decoding method that improves acoustic grounding while preserving support from the base model. At each decoding step, LCAR first retains tokens whose base-model likelihood falls within a margin of the greedy token, then reranks them using an acoustic compatibility score computed from attention-pooled audio embeddings and the existing LM head. By restricting acoustic intervention to plausible, model-supported alternatives, LCAR requires no additional training, external detector, reference transcript, or auxiliary model at inference. We evaluate LCAR on four LLM-based ASR systems using human-audited TTS and open-source speech challenge suites. At $δ=0.60$, LCAR removes 38.8–57.1\% of detector-identified hallucination failures while largely maintaining WER/CER on standard open-source test sets.


9. Closing the Verification Loop: Self-Check Captioning for Long-Paragraph Detailed Audio Captioning

Authors: Fengji Ma, Yan Rong, Xu Li, Chen Zhang, Pengfei Wan et al.

Categories: cs.SD | EMNLP2026 Score: 7.00/10 (Obj:9 Id:7 Ind:6 Comp:7 Eff:8 Nov:8 Rep:2)

  • Strength: 7B 开源模型在 Omni-Cloze 长段落音频描述上达 58.9%(较基座 +44.8pp、超 Omni-Captioner-7B 4.4pp、距 Gemini3.1-Pro 仅 0.6pp),MMAU/MMAR/MMSU 全开源最佳且 MMAU 超 GPT-4o Audio 5.9pp;LACap+SCC-Verifier 作为免训练插件为 Gemini3.1-Pro 提升 +8.6pp。
  • Weakness: 宣称已发布的 50,222 条 LACap-50k 在 GitHub/HuggingFace 实测 0 命中且论文无任何发布链接,数据全链路依赖 Gemini3.1-Pro/Qwen 系模型自审、无新增人工评测,且 LC-SFT 主打的可验证奖励项消融仅 +0.18pp、SCC-Verifier 推理开销无量化报告。
  • Full review: Claude Code 全文七公理审稿

Long-paragraph detailed audio captioning, which requires dense and transcript-faithful descriptions of fine-grained audio content, remains unsolved for current audio-visual multimodal language models. We attribute this failure to two structural problems. The first is data poverty, as no public corpus jointly provides long clips, paragraph captions, and verbatim-transcript fidelity. The second is generation-mode failure, evidenced by a 44.8 to 46.4 percentage-point gap between right-audio and shuffled-audio multiple-choice question (MCQ) accuracy. We address both within Self-Check Captioning (SCC), a unified framework that instantiates audio-grounded question answering as the verification primitive at every lifecycle stage. SCC yields three artifacts. Long-paragraph Audio Caption 50k (LACap-50k) is a 50,222-clip audio-visual corpus with 491.5-word captions and a post-hoc automatic speech recognition (ASR) audit. Layer-Curvature Supervised Fine-Tuning (LC-SFT) is the first on-policy supervised fine-tuning method to weight tokens by intermediate-layer evidence, motivated by our identification of Late-Layer Semantic-Entropy Collapse (SEC). SCC-Verifier arbitrates among caption rollouts via audio-grounded self-answering at inference. Across multiple benchmarks, our system attains state-of-the-art among open-source captioners and is competitive with proprietary baselines. We release LACap-50k to fill the resource gap for long-paragraph detailed audio captioning research.


10. Weakly Supervised Tabla Stroke Transcription via an Adaptive Dynamic Rhythm Language Model (ADRM)

Authors: Rahul Bapusaheb Kodag, Vipul Arora

Categories: eess.AS Score: 7.00/10 (Obj:9 Id:7 Ind:7 Comp:8 Eff:8 Nov:6 Rep:4)

  • Strength: 弱监督塔布拉转录框架 ADRM 在四个数据集上将 stroke error rate 相对 acoustic-only 解码一致降低 18.9%–35.4%(HMR 43.8%→29.6%,TID 24.9%→16.4%),并发布首个序列级弱标注真实即兴独奏数据集 TID(133 分钟、6 种 tāla、30 类鼓点)。
  • Weakness: 代码完全未随论文发布(GitHub 仓库仅含 README,重打分核心管线无法核验,无 seed/显著性检验/调参集声明),且与作者 2026-01 的 TI-SDRM 预印本(arXiv:2601.08537)框架高度重合,onset 对齐在 50ms 容差下 F1 仅 8.9%。
  • Full review: Claude Code 全文七公理审稿

Tabla Stroke Transcription (TST) is central to the analysis of rhythmic structure in Hindustani music, yet it remains challenging due to complex and dynamic rhythmic organization and the scarcity of strongly annotated data. Existing approaches largely rely on fully supervised learning with onset-level annotations, which are costly and impractical at scale. This work addresses TST in a weakly supervised setting, using only symbolic stroke sequences without temporal alignment of onsets. We propose a framework that combines a Connectionist Temporal Classification (CTC)-based acoustic model with a sequence-level rhythmic language model for rescoring, similar to that used in automatic speech recognition. The acoustic model produces a decoding lattice, which is refined using an Adaptive Dynamic Rhythm Language Model (ADRM) that combines $t\bar{a}la$-conditioned symbolic rhythmic regularities with local stroke dynamics. Moreover, we release a new performance-recorded tabla dataset, named \emph{Tabla Improvisation Dataset}, along with a complementary synthetic dataset for sequence-level weakly supervised TST. Experiments demonstrate consistent and substantial reductions in stroke error rates with ADRM compared to those with acoustic-only decoding, confirming the benefit of incorporating symbolic rhythmic regularities during lattice rescoring for accurate transcription.


11. Textual Acoustic Grounding for Generalizable LLM-Based Deepfake Voice Detection

Authors: Yassine El Kheir, Xin Wang, Wanqing Ge, Tim Polzehl, Sebastian Moeller et al.

Categories: cs.SD, eess.AS | 6pages + 1ref, SLT Submission Score: 6.95/10 (Obj:8 Id:6 Ind:6 Comp:8 Eff:8 Nov:6 Rep:7)

  • Strength: 1.1B 参数系统在 ITW/MLAAD 取 83.96/93.42 Macro-F1(较 ALLM 基线 +16.2% 绝对),冻结 LLM 结论跨全部配置复现,openSMILE 文本注入使 LoRA 从 68.18% 恢复至 91.17%,全部模型 HuggingFace 公开可验证。
  • Weakness: 模型选择疑似按评测集最优种子(best42)挑选且无方差报告,+16.2% 声明依赖未解释的失效基线(FT-GRPO ITW 28.54%),openSMILE 文本序列化模板无代码仓库、无法精确复现,第三方索引(S2/OpenAlex)核查均失败。
  • Full review: Claude Code 全文七公理审稿

Deepfake voice detection suffers from poor generalization across unseen domains. While Audio Large Language Models (ALLMs) show promise, the modality gap between continuous audio embeddings which capture the subtle acoustic details necessary for deepfake detection and the semantic space of LLMs remains a critical, underexplored bottleneck. We address this by benchmarking diverse audio encoders integrated with Qwen LLMs (0.5B to 7B parameters). First, we demonstrate that fine-tuning the LLM alone risks out-of-domain overfitting, making a frozen LLM a stronger, resource-efficient baseline. Second, to explicitly bridge the modality gap, we introduce a cross-modal prompting strategy that injects linguistic-knowledge-driven acoustic features (via openSMILE) as structured text tokens. This explicit textual grounding not only enhances the frozen baseline but also makes LLM fine-tuning more effective. Ultimately, our approach demonstrates state-of-the-art resilience on the out-of-domain ITW and MLAAD benchmarks, yielding over \textbf{16.2\%} absolute improvement in Macro-F1 over existing ALLM baselines while maintaining competitive in-domain performance. All models reported in this work are \href{https://huggingface.co/01Yassine/AudioLLM-Deepfake-Detection}{publicly available}.


12. Playability-Aware Audio-to-Tablature Guitar Transcription via Diffusion Models

Authors: Riccardo Simionato, Louis Bigo

Categories: cs.SD, cs.IR | Accepted for ISMIR 2026 Score: 6.90/10 (Obj:9 Id:8 Ind:5 Comp:6 Eff:7 Nov:7 Rep:6)

  • Strength: 首个(据论文自述)面向吉他 tab 转录的扩散模型 Noise2Fret,五个可玩性辅助损失经单损失消融逐一验证,GOAT 上 tab F 从 0.740 提至 0.795、FPR 最低 0.022,且每前向仅 341M FLOPs(TabCNN 约 1/10、FretNet 约 1/51),RTF 0.67 接近实时。
  • Weakness: 基线仅 2019/2023 两个旧模型、无显著性检验与人类可玩性评测,GOAT 的 Let Ring 标注缺口未量化影响;GitHub 承诺的预训练权重实测 404 且仓库无 license,”首个扩散 tab 模型”的新颖性主张因 Semantic Scholar/OpenAlex 查询失败(429/欠费)未能独立核验。
  • Full review: Claude Code 全文七公理审稿

Guitar tablature transcription requires not only accurate pitch detection but also assigning each note to a specific string-fret position, as the same pitch can be played at multiple fretboard positions. Existing approaches treat this as a standard classification problem, ignoring the musical and physical constraints that govern playable fingering sequences. We propose Noise2Fret, a diffusion model for audio-to-tablature transcription that generates tablature through a continuous latent representation of discrete fret and string targets, conditioned on spectral and audio features. To bridge the gap between pitch accuracy and physical playability, we introduce five auxiliary losses encoding Pitch-Class Distance, Positional Distance, Circle-of-Fifths Distance, String Similarity, and Hand-Span Feasibility directly into the training objective. Experiments on GuitarSet and GOAT datasets demonstrate that the model outperforms baselines while remaining computationally more efficient, and that the auxiliary losses yield consistent gains over the standard training objective.


13. Stride-k Subsampling: Train-Free Audio Token Reduction for Whisper

Authors: Chanhee Cho, Junhyuk Choi, Bugeun Kim

Categories: cs.SD, cs.AI | Accepted EMNLP 2026 Main Score: 6.89/10 (Obj:9 Id:7 Ind:9 Comp:9 Eff:6 Nov:5 Rep:6)

  • Strength: 零训练零参数的 stride-2 子采样在五个 Whisper 规模上保持 k=2 基线 WER,同时切 75% 音频 token、52-58% GFLOPs,在三个 SpeechLM 上端到端延迟降 19.6-27.4%(AF3 MMAU 仅 −1.7),并以 R ≥ kS 判据统一解释跨前端适用性,head-to-head 胜过 SpeechPrune(74.38 vs 62.73)。
  • Weakness: 干净语音结论在噪声下崩溃(0dB 时 Large-v3 +28.6 / Small +57.5 WER),Common Voice 绝对代价 +9.56 至 +13.40 WER,ASR 延迟几乎无改善;零代码发布、SpeechLM 集成细节不足、无统计显著性检验,CKA 机制归因缺乏干预证据。
  • Full review: Claude Code 全文七公理审稿

Whisper exposes speech through a fixed 1500-token encoder interface, now a default representation for ASR decoders and Whisper-based speech language models (SpeechLMs), yet its redundancy remains largely unexamined. We propose stride-k subsampling, a deterministic indexing operation that retains every k-th token after the convolutional stem or encoder transformer. Across five Whisper scales, k=2 preserves baseline WER at both positions, with CKA attributing this stability to acoustic overlap at the stem and attention-induced redistribution at the encoder output. Applying stride-2 at both positions cuts audio tokens by 75% and total GFLOPs by 52-58%, with small WER costs on most ASR benchmarks and larger costs on harder ones. The same configuration extends to three Whisper-based SpeechLMs, yielding modest accuracy drops on stronger baselines and larger drops on weaker ones, while reducing end-to-end latency by 19.6-27.4%. Requiring no training or auxiliary computation, stride-k subsampling exploits Whisper’s preprocessing redundancy, indicating that its audio-token interface carries more capacity than downstream tasks require.


14. Not All Fallbacks Are Failures: Understanding and Recovering from Fallbacks in Mobile Voice Assistants

Authors: Phillip Schneider, Alexandre Mercier, Joshua Oehms, Kristiina Jokinen, Florian Matthes

Categories: cs.CL | Accepted to EMNLP 2026 Score: 6.80/10 (Obj:9 Id:6 Ind:5 Comp:8 Eff:8 Nov:6 Rep:4)

  • Strength: 基于 500+ 用户六个月真实部署数据构建首个 fallback 话语标注数据集 VoxFallbacks(3,030 条,Fleiss’ κ=0.759),证明轻量 embedding 分类器以 <0.05 ms 延迟和 macro F1=0.810 胜过最优 LLM 的 0.727(LLM 延迟 184–1,661 ms),全部实验在一台 M3 Pro 上完成。
  • Weakness: 数据集与代码承诺的 GitHub 仓库实测为空(可复现证据缺失),全部数据来自单一系统/语言/老年人群且标签为无对话上下文的研究者推断(作者自认非 ground truth),错误代价不对称与推荐的 LLM 澄清生成环节均未量化评估。
  • Full review: Claude Code 全文七公理审稿

Robust understanding of user input is a core requirement for voice assistants deployed in real-world environments. In practice, these systems encounter heterogeneous fallback situations caused by noisy audio input, transcription errors, ambiguous requests, incomplete utterances, or unintended activations. Existing systems typically respond with generic fallback messages, which do not resolve the underlying interaction failure and can degrade user experience. We study fallback handling in a deployed smartwatch-based voice assistant for general health support in everyday environments. Our analysis is based on six months of real-world usage data from more than 500 users, yielding a dataset of 3,030 anonymized, naturally occurring fallback-triggering utterances. We contribute (1) an operational taxonomy and the annotated VoxFallbacks dataset of these interactions, (2) a comparative evaluation of different models within a classification pipeline under practical deployment constraints, and (3) practical lessons for designing robust and cost-efficient fallback mechanisms. Results show that lightweight embedding-based classifiers outperform larger generative models on most classification tasks while requiring substantially fewer computational resources.


15. Using Prosody to Predict Syntactic Structure

Authors: Junghyun Min, Alex Warstadt, Tamar I. Regev, Tiago Pimentel, Ethan Gotlieb Wilcox

Categories: cs.CL, cs.AI, cs.LG | 15 pages, 4 figures. EMNLP 2026 camera-ready Score: 6.80/10 (Obj:8 Id:6 Ind:6 Comp:6 Eff:8 Nov:6 Rep:6)

  • Strength: 以互信息框架量化韵律句法信息,发现韵律最高削减句法不确定性 10.2%(时长 3.45 nats/停顿 2.06 nats),自发语音高于朗读语音(长度匹配后仍显著),短语边界占信息量 50-60%,为 IRT/韵律自举假说给出首个大语料定量裁决(EMNLP 2026)。
  • Weakness: 条件互信息 I(S,PW) 出现 −1.2 nats 负值暴露估计不对称、RQ2 核心问题结论不可靠,句法树为 ASR 转写上的 silver 解析使语体对比混入同源误差,且全文无代码、无权重、无数据发布链接。
  • Full review: Claude Code 全文七公理审稿

While it is well-established that prosody carries crucial cues for syntactic structure, the degree and nature of correspondence between these two domains remains contested. We investigate the syntax-prosody interface through an information-theoretic lens, quantifying the interaction between prosodic features and syntactic representations as their mutual information. We provide a general-purpose framework for estimating this quantity over large speech-text corpora using multimodal language models. Our framework is structure-agnostic and modular, insofar as it can be used to measure the contributions of individual prosodic features or components of structure. We evaluate the syntax-prosody relationship for two features (word duration and inter-word pauses) across two domains–read audiobooks and spontaneous conversations–both in English. Our results demonstrate that prosody contains measurable syntactic information, with prosodic features reducing syntactic uncertainty in spontaneous conversations by up to 10.2%. Our findings offer new empirical support for several theoretical accounts of the syntax-prosody interface.


16. Puppeteer: Object-Grounded Posture-Aware Co-Speech Gesture Generation

Authors: Vida Adeli, Soroush Mehraban, Jacob Rommann, Harrison Sanborn, Cole Clifford et al.

Categories: cs.CV Score: 6.80/10 (Obj:8 Id:7 Ind:6 Comp:6 Eff:8 Nov:7 Rep:4)

  • Strength: 首个物体接地+姿态感知的伴语手势扩散框架,BEAT2 上 FGD 3.436 领先 GestureLSM 3.692 且 Div 14.131 比次优高 24%,姿态条件把 LL1 从 16.606 压到 2.446,物体模块使 MeanPen 降低 2.8 倍。
  • Weakness: 物体接地优势无同任务基线且用自家 SDF 指标自评自证,附录 G.1 模型人类评估为空标题,代码/数据集均为 “Soon” 占位、SceneGes 依赖不可复现的美术逐帧精修(28 分钟、仅椅桌两类物体)。
  • Full review: Claude Code 全文七公理审稿

Generating co-speech gestures that are temporally coherent, semantically aligned with speech, and grounded with surrounding objects remains challenging. Prior speech-driven gesture models emphasize audio-gesture alignment but do not explicitly account for posture constraints or surrounding objects, failing to capture the inherent correlation between body gestures and the physical space. We present Puppeteer, a posture-aware, object-grounded co-speech gesture diffusion model operating in a causal latent space. We decompose long gestures into structured primitives and learn a causal variational autoencoder that encodes them into temporally ordered latent tokens, each depending only on the past. We then perform conditional diffusion directly in the causal latent space, conditioning on speech signals, motion history, an initial posture reference, and object geometry to synthesize physically consistent gestures. This temporally ordered latent formulation enables explicit temporal control and supports tasks such as gesture in-betweening and gesture completion. To better assess co-speech gesture synthesis beyond existing measures, we introduce new evaluation metrics tailored to this task. We also created SceneGes, the first curated synthetic 3D dataset of embodied co-speech gestures and corresponding 3D objects, enabling object-grounded gesture generation. Experiments show that Puppeteer generates more diverse and temporally synchronized gestures than prior methods, while enabling object-grounded gesture synthesis.


17. Read the Room, Read the Image: Understanding Indirect Speech Acts in Multimodal Visual Contexts

Authors: Jaehee Kim, Ji Hoon Chung, Seoyoon Park, Unsol Kim, Kyungwon Park et al.

Categories: cs.CL | Accepted to Findings of ACL 2026 Score: 6.70/10 (Obj:8 Id:8 Ind:7 Comp:8 Eff:6 Nov:7 Rep:3)

  • Strength: 提出 READI 多模态 ISA 评测基准(EN 45 + KR 57 项,CCSARP 三级间接性分级),8 个 VLM 实验显示随间接性升高准确率显著下降(英语 Lv2 0.85→Lv3 0.45;韩语 Lv2 0.58),图像打乱消融证实联合推理必要性(Gemini 3 EN 0.93→0.80)。
  • Weakness: 数据集仅 102 项且完全未开源(GitHub 检索 0 结果、全文无发布链接),AI 生成图像无生成流程公开导致不可复现,韩语侧核心相关性结论不显著(r=−0.22, p=.109)而样本量不足以支撑该缺失证据。
  • Full review: Claude Code 全文七公理审稿

Indirect speech acts (ISAs) require pragmatic reasoning over context, as directive intent can- not be inferred from surface form alone. Prior text-based studies and existing multimodal benchmarks largely overlook this requirement, focusing instead on explicitly encoded context or perceptual recognition, and thus underex- plore context-dependent pragmatic understand- ing, particularly in high-context languages such as Korean. We introduce READI, a multimodal benchmark for evaluating ISA understanding through integrated reasoning over visual con- text and dialogue. READI models graded in- directness grounded in pragmatic theory and formulates the task as vision-based pragmatic question answering (V-PQA), supporting cross- lingual evaluation in English and Korean. Ex- periments show that even state-of-the-art multi- modal models struggle with visually grounded indirect speech acts, with performance declin- ing as indirectness increases, underscoring the need for benchmarks that explicitly target con- textual pragmatic reasoning.


18. Language-Statistical Analysis of Neural Audio Codec Tokens Across Architectures, Corpora, and Noise Conditions

Authors: Joonyong Park, Shinnosuke Takamichi, David M. Chan, Shunsuke Kando, Yuki Saito et al.

Categories: cs.CL, cs.SD | Submitted to IEEE Transactions on Audio, Speech, and Language Processing (TASLP) Score: 6.60/10 (Obj:9 Id:6 Ind:5 Comp:6 Eff:7 Nov:7 Rep:6)

  • Strength: 首次将 NAC token 语言统计扩展到 13 编解码器 × 3 噪声条件(117 cell),发现语料几乎不解释方差(η²≤0.005)而架构-噪声按指标主导,单阶 JSD 在 DEMAND 下与 UTMOS 相关 r=−0.76(逐点剔除后系数移动 ≤0.10),collapse/explosion 签名按架构聚集(RVQ×白噪声 14/21、RVQ×DEMAND 13/21)。
  • Weakness: 词表规模与架构族/码率完全混杂未解耦,JSD-质量关联在白噪声下失效(UTMOS Pearson 0.15、MCD 系数剔除非 VQ 后由 0.62 反号为 −0.27)且仅 12 个编解码器均值点,仅 0 dB 单一 SNR,全文无代码发布(GitHub 检索 0 命中)。
  • Full review: Claude Code 全文七公理审稿

Neural audio codecs (NACs) convert speech into discrete token sequences, and prior work has reported that these sequences follow language-like statistical laws. This paper analyzes the token statistics of 13 NACs spanning multi-codebook residual vector quantization (RVQ), single-codebook VQ, and non-VQ designs, evaluated on three corpora under clean, white-noise, and real-world DEMAND-noise conditions. Zipf and Heaps parameters, unigram entropy, codebook occupancy, and Jensen-Shannon divergence (JSD) are estimated from matched token samples with explicit fit-validity safeguards and family-conditional $n$-gram orders. Corpus identity explains little variance in any metric, whereas acoustic condition and quantizer meta-category dominate in a metric-dependent way, and unigram entropy is the metric most strongly associated with meta-category. Clean-to-noise JSD computed at a common unigram order is associated with mel-cepstral distortion most clearly under DEMAND noise. The collapse and explosion degradation signatures previously reported for RVQ codecs concentrate in RVQ cells under white and DEMAND noise, respectively; explosion also occurs in non-VQ codecs, and single-codebook VQ codecs shift in occupancy and distribution shape without either signature. These results provide architecture-conditioned conventions for applying language-statistical analysis to NAC tokens.


19. Sequential Trajectories and Simultaneous Blending: Multi-Emotion Modeling for Instruction-Following TTS

Authors: Yan Zhou, Yun Hong, Yang Feng

Categories: cs.CL, cs.SD | Code is available at https://github.com/ictnlp/HybridEmo. Demo page: https://zhouyan19.github.io/HybridEmo-demo/ Score: 6.55/10 (Obj:8 Id:6 Ind:5 Comp:6 Eff:6 Nov:7 Rep:6)

  • Strength: 首个同时建模情绪轨迹与情绪混合的多情绪 TTS 后训练框架,混合强度从 3.47 提升至 3.71(追平 Qwen3-TTS)且 WER 仅 0.37%,人类评估对 CosyVoice 3 净胜 32 个百分点,权重与评测集已开源。
  • Weakness: 主观评测由与 SFT 标注同家族的 Qwen3-Omni 担任唯一裁判且测试集为自建(混合分支仅 120 条件),轨迹增益仅 +0.09 且被单任务变体反超(3.35 vs 3.33),训练与奖励代码未开源导致核心贡献不可复现。
  • Full review: Claude Code 全文七公理审稿

Natural-language instructions enable flexible control of synthesized speech, yet emotional TTS systems primarily model a single utterance-level affect, leaving multi-emotion control underexplored. We study two complementary multi-emotion TTS tasks: emotion trajectory, which spans several ordered affective stages, and emotion blending, in which multiple emotions coexist throughout an utterance. These tasks expose a supervision mismatch: supervised fine-tuning (SFT) does not explicitly evaluate emotion features, while single-emotion rewards provide neither structure-aware feedback for trajectory completion nor pair-aware feedback for blending. We introduce HybridEmo, a post-training framework that initializes both tasks with SFT and then aligns the speech-token policy through Group Relative Policy Optimization using a sample-aware hybrid reward. For trajectory samples, segment-aligned consistency combines average and weakest-stage evidence to preserve the correctness and completeness of prescribed stages. For blending samples, a GMM-based reward combines frame-level support from the union of target-emotion anchors in an offline emotion space with an utterance-level weaker-target margin. Both branches share an ASR reward and are routed within a unified policy. On MultiEmo-Test, HybridEmo significantly improves trajectory correctness and blending intensity, without a noticeable degradation in speaker similarity. Human evaluation prefers HybridEmo to CosyVoice 3 and EmoVoice-0.5B, with nearly balanced preferences against Qwen3-TTS.


20. Parallel Time-Band Mixing with Learned Observation-Adding for Robust ASR Front-Ends

Authors: Xingyu Shen, Runze Wang, Wei-Ping Zhu, Benoit Champagne

Categories: cs.SD, cs.AI | Accept by Interspeech 2026 Score: 6.50/10 (Obj:8 Id:8 Ind:6 Comp:8 Eff:6 Nov:6 Rep:4)

  • Strength: 提出全并行带分前端 PTBM(0.96M 参数/0.58 GMAC/s,约为 BSRNN 1/3 计算量)+ 免调参学习式观测叠加 LOA,在 DNS 与 CHiME-4 上用冻结 Whisper 全系后端 WER 一致下降(CHiME-4 Eval Large 6.69→6.24,Tiny 35.03→26.07),已获 Interspeech 2026 接收。
  • Weakness: 全文未提供任何真实运行时/延迟/RTF 数据佐证并行加速这一核心卖点,LOA 为 utterance-level 非流式且训练依赖 oracle 网格搜索,方法无代码开源(GitHub 无官方仓库)导致微小 WER 差异(0.01 级)不可核验。
  • Full review: Claude Code 全文七公理审稿

Speech enhancement is often used as a front-end for robust ASR, yet recurrent temporal and cross-band modules introduce sequential dependencies that reduce parallel efficiency. In this paper, we present a sequence-parallel band-split enhancement front-end built on a Parallel Time-Band Mixer (PTBM) block that eliminates within-block recurrent unrolling. PTBM integrates intra-band temporal mixing and per-frame cross-band attention within a unified parallel architecture, enabling efficient contextual modeling across both time and frequency dimensions. The system retains the mask-plus-residual reconstruction interface and introduces learned Observation-Adding (LOA) to suppress ASR-sensitive artifacts without development-set tuning. Experiments on DNS Challenge and CHiME-4 with frozen Whisper back-ends show that the proposed front-end consistently reduces word error rate relative to recurrent band-split baselines while requiring only 0.96 M parameters and 0.58 GMAC/s for the front-end network.


21. Listen to the Latents: Self-Correcting Speech Recognition in Large Audio Language Models Through Hidden-State Interactions

Authors: Chan-Jan Hsu, Jaeyeon Kim, Chao-Han Huck Yang, Shinji Watanabe, Hung-yi Lee et al.

Categories: cs.CL, cs.AI | 24 pages, 8 figures, 2 tables. Accepted to Findings of EMNLP 2026 Score: 6.40/10 (Obj:8 Id:8 Ind:6 Comp:8 Eff:5 Nov:6 Rep:5)

  • Strength: 提出训练自由、零参数的 Hybrid Search:用 ASR-LLM 与 base LLM 的中间层隐状态交互(余弦相似度+范数比)定位语义依赖 token 做定向 LLM 校正,摊销 beam 1.4 恢复 beam search NE-ER 增益的 96.4%(26.57→24.71),错误校正模式 NE-ER 18.29→17.69(McNemar p=3.56×10⁻³⁸),已获 Findings of EMNLP 2026。
  • Weakness: WER 无实质收益(beam 11.11→11.12),收益集中于 NE-ER 单指标且 τ=0 消融显示不用隐状态特征仅目标化多 token 词即得更好 NE-ER(23.78 vs 24.32),核心特征模块边际贡献存疑;全文无代码开源,端到端验证仅 Phi-4-Multimodal,外部新颖性核查因 Semantic Scholar/OpenAlex 429 限流失败。
  • Full review: Claude Code 全文七公理审稿

Recent automatic speech recognition (ASR) systems increasingly integrate large language models (LLMs) to leverage their semantic knowledge, either externally through logit fusion or internally through warm initialization. However, how to effectively combine these two strategies remains underexplored. In this work, we refine warm-initialized LLM-based ASR models by leveraging their own pre-adaptation base LLMs, focusing on LoRA-adapted settings where the base LLM is preserved. To achieve this, we propose Hybrid Search, a targeted correction strategy motivated by two observations. First, interaction features that characterize the relationship between LLM-based ASR hidden states and base-LLM hidden states provide informative signals about a token’s degree of semantic dependence. Second, selectively refining targeted tokens with high semantic dependence improves ASR performance far beyond naive global LLM-correction methods including rescoring and late fusion. Our analysis suggests that, even after semantic knowledge transfer through warm initialization, LLM-based ASR models can still leverage their base LLM to further improve inference-time performance.


22. Linguistic Distance Segregates Latent Representations in Automatic Speech Recognition Systems

Authors: Ting-Hui Cheng, Line Katrine Harder Clemmensen, Sneha Das

Categories: cs.CL, cs.LG | to appear in EMNLP finding 2026 Score: 6.40/10 (Obj:8 Id:7 Ind:7 Comp:8 Eff:5 Nov:6 Rep:5)

  • Strength: 在 6 个语料库约 79,335 样本、4 个模型上证实 L1 语言学距离系统性影响英语 ASR 错误率(Tweedie GLMM 四模型均 p<0.001,Whisper-Small β=0.410/z=25.20,LD 每增一单位期望 WER 上升约 52%),并首次系统报告 L2/L1 潜层分离随深度增强且保留至最终层。
  • Weakness: 纯相关性诊断、无任何缓解或干预实验,L2 熟练度等混杂变量未做代理控制,结论限于英语/开源模型/欧亚样本,且无代码发布(arXiv 页与 GitHub 均无随文仓库)、逐层嵌入提取细节不足以无歧义复现、标签模糊反例仅由 4 人感知实验支撑。
  • Full review: Claude Code 全文七公理审稿

While automatic speech recognition (ASR) models have achieved remarkable improvements in recent years, performance disparities persist across different speaker populations. One such disparity is for speakers whose first languages (L1) are from families distant from English. This paper investigates the relationship between first language background and English ASR performance. Through empirical analysis, we observe that the correlation between speakers’ L1 distance and ASR error rates yields a systematic effect on English Speech, with its strength varying across datasets and models. This association is statistically significant in a follow-up analysis accounting for dataset-level variation in Tweedie mixed-effects models ($p<0.001$ across evaluated models). In addition, analysis of the latent space reveals a L1-based spatial segregation across deeper acoustic layers in the majority of evaluated architectures


23. Conjoint Audio-to-Spikes Encoding and Processing for Efficient Neuromorphic Speech Recognition

Authors: Valentin M. Meunier, Amélie Gruel, Pierre Lewden, Adrien F. Vincent, Sylvain Saïghi

Categories: cs.NE, cs.AI, cs.LG, eess.AS | Under review. 15 pages, 9 figures and 4 tables. This work is supported by a public grant overseen by the French ANR as part of the “Chaires IA” programme (GrAI project ANR 19 CHIA 0003) and as part of the “PEPR IA France 2030” programme (Emergences project ANR 23 PEIA 0002). This research is part of the programme DesCartes and is supported by the NRF Singapore under its CREATE programme Score: 6.37/10 (Obj:8 Id:5 Ind:6 Comp:8 Eff:7 Nov:6 Rep:5)

  • Strength: 13 通道非可学习耳蜗编码器 + 100.5k 参数前馈 SNN 在 Heidelberg Digits 达 99.77%(超此前脉冲 SOTA 至少 1.47 pp),并给出 Zynq-7000 FPGA 完整实现(Q9.22 定点,HD-English 子集 99.63%)。
  • Weakness: 缺少实测能耗证据(全篇依赖自定义的精度/脉冲代理指标,FPGA 亦无功耗测量)与代码发布,且 NHD 采用 80/20 随机切分与基线可能的官方切分不对齐、GSC 落后 SOTA 8.5 pp,核心声明可比性存疑。
  • Full review: Claude Code 全文七公理审稿

Obtaining data from neuromorphic sensors and processing it with Spiking Neural Networks is a promising solution to lower the energy cost of artificial intelligence. The current rarity of natively neuromorphic datasets promotes the development of software tools to translate input sensory data into spikes. However, highly bio-mimetic simulators can be challenging to implement on digital hardware. In this work, we evaluate the neuromorphic encoding and subsequent classification of audio into spikes using a non-learnable, high-level, programmable encoder targeting hardware implementation on FPGA. We quantify the pipeline’s efficiency with hardware-agnostic metrics based on the quantitative spiking activity. Our study focuses on the simultaneous optimisation of encoder and classifier: the first provides efficient and informative data so that the latter achieves a better performance with an overall lower energy cost at learning and inference. This work introduces the first end-to-end neuromorphic spike-encoding and evaluation of the TIMIT dataset. Our simple feedforward network reaches a classification accuracy of 99.77% on a spike-encoded Heidelberg Digits, overcoming the neuromorphic state of the art on this benchmark dataset.


24. Towards Balanced Spectral Reconstruction: Spectrally Adaptive Loss for Streaming Speech Enhancement

Authors: Haixin Zhao, Nilesh Madhu

Categories: eess.AS | Accepted by IWAENC 2026 Score: 6.20/10 (Obj:8 Id:6 Ind:5 Comp:7 Eff:6 Nov:6 Rep:4)

  • Strength: 提出两种频谱加权 STFT 损失修复流式语音增强中高频过度衰减,HF LSD 提升 1.18 dB(8.67→7.48)、HF SI-SDR +0.43 dB、M-RMSE 降 15.2%,信号依赖版本在 MF 区 LSD 7.46 对基线 7.78 全面占优;配套 HyST-Net 以 0.11 M 参数、RTF 0.22(对 FTF-Net 快 4.77 倍)实现持平性能。
  • Weakness: 全局指标持平或略降(PESQ 2.86→2.81、DNSMOS 3.23→3.21)且无主观听测支撑感知主张,评估限于单一 DNS 合成域,代码零发布(GitHub 与演示页均无仓库),超参数全为经验值且 demo 自曝高频嘶嘶声伪影未做分析。
  • Full review: Claude Code 全文七公理审稿

This paper proposes two spectrally weighted STFT loss functions for lightweight streaming speech enhancement, addressing the magnitude over-attenuation in mid-to-high frequency regions caused by the magnitude-phase compensation effect. The proposed sigmoid-weighted loss applies a smooth frequency-dependent modulation to the phase-aware contribution, while the signal-dependent spectrally adaptive loss further conditions the modulation on the ground-truth log-magnitude spectrogram. To evaluate the proposed objectives, we additionally design HyST-Net, a lightweight and competitive backbone with hybrid MHA-GRU spectral-temporal modelling for low-latency streaming scenarios. Experimental results exhibit consistent improvements in high-frequency spectral reconstruction for both losses. The spectrally adaptive loss further enhances the mid-frequency region, resulting in a more balanced spectral reconstruction across the full frequency range.


25. Conversation Coach: A Voice-enabled AI System that Helps Practice Difficult Workplace Conversations

Authors: Fanyou Wu, Suraj Maharjan, Ainur Yessenalina, Dennis Xu Chen, Rahul Srivastava et al.

Categories: cs.AI Score: 6.20/10 (Obj:8 Id:6 Ind:5 Comp:8 Eff:7 Nov:6 Rep:3)

  • Strength: 端到端 vs. 级联语音架构在生产级配对对比中给出清晰量化边界(TTFA P50 3 倍差距、成本 8 倍差距、级联在人格一致性 4 维度占优且 Haiku 消融证明为架构效应),6 个月 40,000+ 经理、54,000+ 会话部署,评估截止期前使用量达 4 倍峰值、最低职级员工占被辅导对 43.3%。
  • Weakness: 教练效果无因果证据(工作满意度差 4.3pp 未达显著,MDE≈13pp),评测全为合成数据且 LLM 裁判含 Claude 系模型与被评级联系统存在家族关联,零代码零数据公开、全部专有服务、生产数据 30 天删除致第三方无法复现;S2(429)/OpenAlex(额度耗尽)独立索引核查均失败。
  • Full review: Claude Code 全文七公理审稿

Effective manager-employee communication is critical for retaining high performers and developing underperformers, yet training managers in these skills remains costly. Text-based chatbots offer a scalable approach but cannot provide realistic rehearsal: managers need to practice speaking aloud to build confidence before high-stakes conversations. In this paper, we propose Conversation Coach, a voice-first AI system that enables managers to rehearse difficult workplace conversations in a realistic spoken format. The system addresses three challenges: achieving low-latency interactions with strong language understanding, enabling adaptive conversations through configurable bot personalities that simulate different employee types, and generating personalized feedback on content and policy compliance. We compare an end-to-end speech-to-speech model with a cascaded approach combining automatic speech recognition, a large language model, and text-to-speech synthesis. The end-to-end approach achieves 3$\times$ lower median (P50) latency with native barge-in capability at an estimated 8$\times$ lower cost, while the cascaded approach offers superior reasoning essential for coaching quality. We deployed the cascaded architecture in production, where 40,000+ managers used it over six months, with adoption patterns indicating selective use for difficult conversations.


26. MusGU+: Toward a Musician-Centered Evaluation Framework and Discovery Tool for Generative Music AI

Authors: Laura Ibáñez-Martínez, Roser Batlle-Roca, Xavier Serra, Martín Rocamora

Categories: cs.SD, cs.AI, cs.CY, eess.AS | Accepted at AIMC 2026 Score: 6.11/10 (Obj:8 Id:6 Ind:4 Comp:8 Eff:6 Nov:5 Rep:7)

  • Strength: 首个面向音乐人适用性的结构化横评:15 条细则 × 三级评分覆盖 10 个学术/工业/商业系统,发现仅 3/10(DDSP-VST、Neutone Morpho、RAVE)满足个人适配约束,Udio 禁止下载输出,三维度不对称挑战民主化叙事,工具与代码(Apache-2.0)均已公开上线。
  • Weakness: 框架作者本人即全部评分者,无第三方评分、无评分者间一致性、无用户研究,Controllability 仅凭文档评估未经体验验证,专有系统(Suno/Udio)评分无快照时间戳且聚合权重未论证,实用性与效度均缺独立证据。
  • Full review: Claude Code 全文七公理审稿

Generative music systems are increasingly presented as tools that democratize music creation, yet their practical suitability for musicians remains underexplored. Prior work includes openness-focused evaluation frameworks, such as MusGO (Music-Generative Open AI), as well as qualitative studies of musicians’ experiences with generative systems. However, these approaches do not support systematic comparison or early-stage discovery of models for creative use. Motivated by such limitations, we introduce MusGU+, a musician-centered evaluation framework organized around three dimensions: Adaptability, Usability, and Controllability. Together, these capture whether a model can be feasibly trained or fine-tuned on personal data, integrated into real-world music workflows, and controlled in musically meaningful ways. We evaluate 10 representative generative music systems and present an interactive discovery tool that enables musicians to explore and filter models according to these criteria. While MusGO remains valuable for promoting responsible research practices, MusGU+ supports informed selection and practical adoption of generative systems by musicians.


27. CoJEPA: Combining Contrastive Learning and JEPA for Global-Local Music Representations

Authors: Gabriel Meseguer-Brocal, Yuexuan Kong, Romain Hennequin

Categories: cs.SD, cs.AI, cs.LG, eess.AS, eess.SP Score: 6.10/10 (Obj:8 Id:6 Ind:6 Comp:7 Eff:6 Nov:6 Rep:4)

  • Strength: 单一 7.1M 共享骨干联合 JEPA+对比损失、完全去除 EMA teacher,在多数 MIR 任务上达到或超过 330M 的 MusicFM(key .640 vs .558、chords L4 .399 vs .258),并首次在音乐域观察到联合目标特有的”中间层达峰”表征深度现象。
  • Weakness: 全部结果单次运行无显著性检验,损失权重/掩码率/预测器均无消融,MusicFM 基线仅用末层探测削弱跨规模比较公平性,且零代码发布、预训练依赖 2200 万曲目私有数据,方法当前不可复现。
  • Full review: Claude Code 全文七公理审稿

Joint-Embedding Predictive Architecture (JEPA) has shown strong performance in learning rich representations through self-supervised prediction in latent space. However, it typically relies on teacher–student architecture with an EMA to stabilise training, and can tend to yield uninformative representations. Contrastive learning is stable to train and produces strong global representations, but remains limited on local tasks by the global nature of its objective. In this work, we combine both into CoJEPA: a single shared backbone jointly trained with a JEPA objective on masked sequence tokens and a contrastive objective on the class token. The contrastive gradient provides stability, removing the need for an EMA teacher entirely, while JEPA enriches the sequence tokens via local predictions that contrastive learning alone cannot provide. Crucially, no extra parameters are added to the backbone: the same model is guided towards richer representations purely through the design of its training signal. CoJEPA takes the best of both worlds, outperforming or matching both individual methods across global and local MIR tasks, with a particularly strong advantage on tonal and harmonic understanding, and without any task-specific architectural changes. CoJEPA shows that combining objectives with complementary inductive biases can substitute for scale, encouraging future work to invest in smarter training objectives over ever-larger models.


28. U-PAST: A Phase-Aware Audio Spectrogram Transformer-U-Net for Single-Channel Speech Enhancement

Authors: Cao Duong Ly, Jörn Anemüller

Categories: eess.AS, cs.SD Score: 5.70/10 (Obj:9 Id:6 Ind:5 Comp:7 Eff:6 Nov:4 Rep:4)

  • Strength: 混响失配下唯一超过未处理输入(9.03 dB)的模型族——U-PAST-H 以 2.40M 参数、4.71 GMACs 取得全场最佳 SI-SDR 9.76 dB(五个基线全部跌破未处理,TF-GridNet 仅 7.45),并在匹配条件取得全场最佳 PESQ-wb 3.41。
  • Weakness: “首个混合 transformer-U-Net 音频增强”的”首个”主张被 arXiv 实测的 MUSE(2024)/U-Former(2022)/DCUC-Net(2023,复数域)等先行工作否证,四条件中三个 SI-SDR 落后 TF-GridNet 1.62–2.44 dB 且混响鲁棒性无机制解释;核心基线为非官方实现、无显著性检验、论文与 GitHub(实测 0 仓库)均无代码发布。
  • Full review: Claude Code 全文七公理审稿

Convolutional neural networks (CNNs), used widely and successfully in audio enhancement, capture long-range time-frequency dependencies only indirectly, through successive convolution and pooling. Here, we present U-PAST, a hybrid transformer-U-Net architecture that addresses this limitation through self-attention dependency-modeling in the complex spectrogram domain. U-PAST tokenizes a complex STFT representation, similarly to the magnitude spectrogram tokenization of the Audio Spectrogram Transformer (AST), applies a multi-layer transformer encoder, and reconstructs the enhanced complex spectrogram with a U-Net-style decoder. We evaluate four architectural variants with between 1.17M and 2.40M parameters on the DNS Challenge, VoiceBank-DEMAND, and LibriMix corpora under matched, acoustic mismatch, and two-dataset mismatch conditions. U-PAST attains the best SI-SDR of any evaluated model under acoustic mismatch and closely trails substantially larger convolutional and time-domain baselines by 0.26 dB to 0.63 dB SI-SDR under the remaining three conditions while achieving the strongest perceptual (DNSMOS) quality under dataset mismatch. The largest evaluated configuration, U-PAST-H (2.40M parameters), is consistently the strongest variant of the family, offering an attractive performance-to-cost trade-off at a small parameter footprint.


29. SISER: Speaker-Invariant Speech Emotion Recognition with Entropy-Based Adversarial Training

Authors: Eunseo Choi, Hyunku Kang, Chanwoo Kim

Categories: cs.SD, cs.CL | Accepted to INTERSPEECH 2026 Score: 5.68/10 (Obj:8 Id:6 Ind:8 Comp:8 Eff:4 Nov:4 Rep:5)

  • Strength: 在 IEMOCAP 说话人无关 4 分类上,SISER 无需数据增强即达 UA 60.63%/WA 58.53%,较无增强基线(51.15%)提升 9.48 且折间方差最低(±1.72 vs ±3.55),2×2 消融证明将说话人判别器从 FC 换为 ECAPA-TDNN 贡献 +6.49~+7.13 UA、超过换编码器的 +1.09~+1.73。
  • Weakness: 仅与 2020 年单一基线比较,对带速度扰动增强的基线净增益仅 +0.72 UA 且无显著性检验,未触及 2021 年后 IEMOCAP 上 65%-75% 的 SSL 系方法区间;声称开源的 GitHub 仓库实测为空壳(仅一行 README,0 star),且判别器对比未控制参数量。
  • Full review: Claude Code 全文七公理审稿

Speech emotion recognition (SER) faces two fundamental challenges: scarcity of labeled data and inter-speaker variability, both of which hinder generalization of emotion recognition systems. While prior adversarial approaches address speaker variability, they fall short in leveraging powerful pre-trained representations. We propose SISER (Speaker-Invariant Speech Emotion Recognition), integrating wav2vec 2.0 as a feature encoder and ECAPA-TDNN as a speaker discriminator within an entropy-based adversarial training scheme. wav2vec 2.0 provides rich self-supervised representations that alleviate dependency on large labeled datasets, while ECAPA-TDNN enables suppression of speaker identity via a stronger adversarial signal than shallow classifiers. Evaluated on IEMOCAP, SISER achieves a UA of 60.63%, outperforming the baseline (51.15%) and wav2vec 2.0 without speaker suppression (56.46%), with ablation emphasizing that the choice of speaker classifier architecture is a key factor.


30. DreamX-Creator: Democratizing Native Audio-Video Generation at 2K Resolution

Authors: Jiashu Zhu, Yanhao Zheng, Ruitian Tian, Rujing Dang, Shen Zhang et al.

Categories: cs.CV, cs.SD Score: 5.63/10 (Obj:7 Id:7 Ind:5 Comp:3 Eff:6 Nov:5 Rep:6)

  • Strength: 7B 原生联合音视频生成器为表 1 披露的最小开源原生模型,RL 后 DeSync 0.1902→0.1351 与 2K Refiner 后 VQ 0.6568→0.6930 均为同量级评测全场最佳,MUSIQ 0.7073/MANIQA 0.4382 领先 FlashVSR/SeedVR,且权重与推理代码已实际开源(GitHub 165 stars、HF 权重齐备、Apache-2.0)。
  • Weakness: 六项核心机制(双级门控、方向掩码、模态路由奖励等)零独立消融,训练数据与 Verse-Bench 的重叠排除协议未声明且 RL 评估器过优化验证被作者自认未完成,训练代码、奖励模型名称与内部数据均未披露,音频美学(CE 4.7463 vs 5.2776)与唇同步(LSE-C 7.8361 vs 8.7354)落后于 22B/33B 系统。
  • Full review: Claude Code 全文七公理审稿

Recent video generators often omit audio or synthesize it in a separate stage, limiting reciprocal modeling of visual dynamics and acoustic events. We present DreamX-Creator 1.0, a compact native joint audio-video generation system centered on a 7B generator. Conditioned on a first frame and a text prompt, the generator jointly denoises modality-specialized audio and video streams. The streams are processed independently in the first half of the network and coupled in the latter half through Gated Cross-Modal Attention, whose token- and head-wise output gates modulate each active cross-modal attention-head output. A unified Audio-Video Data System constructs and filters temporally coherent clips, produces structured multimodal annotations, and organizes clips into capability-oriented data pools. Progressive Joint Training comprises two audio-video pre-training stages followed by High-Quality Finetuning. Audio-Video Reinforcement Learning further post-trains the generator with Modality-Aware Multimodal Feedback that routes video-, audio-, and cross-modal feedback to the corresponding streams. For high-resolution output, our Autoregressive 1-Step 2K Refinement pipeline adapts a bidirectional multi-step teacher into an autoregressive multi-step refiner and distills it into a student requiring one denoising evaluation per temporal chunk. Overall, DreamX-Creator 1.0 achieves native, synchronized audio-video generation with performance competitive with state-of-the-art open-source systems. By releasing our compact 7B generator and 2K Refiner, we seek to democratize native audio-video generation and provide an accessible foundation for future research in unified audio-video generative modeling.


31. Arctic Dispersion Interruption Phenomenon and Sound Source Depth Estimation

Authors: Weng Jinbao, Yang Yanming, Chen Benqing, Xu Dewei, Zhou Hongtao

Categories: physics.app-ph, cs.SD Score: 5.42/10 (Obj:8 Id:5 Ind:6 Comp:7 Eff:5 Nov:5 Rep:3)

  • Strength: 在 272 km 深北极实测中用单水听器时频图上的频散中断频率(对应特征函数一阶波节,mode 7 波节深度随频率从 1428 m 降至 225 m)反演 300 m 标定源深,预测中断位置与 10 通道实测吻合,摆脱了对大孔径垂直阵的依赖。
  • Weakness: 全文无任何量化深度估计误差、无与模态上限法等基线的对比实验,验证仅覆盖单一源深/距离/航次的目视判读,且声速剖面与实验数据未公开导致不可复现。
  • Full review: Claude Code 全文七公理审稿

The sound speed profile in the deep Arctic Ocean causes the surface layer to form a normal mode waveguide. When the depth of the sound source or receiver is near a node of the eigenfunction, the source cannot excite the mode, or the receiver cannot detect it, resulting in the modal amplitude at the receiver being approximately zero. This manifests as a dispersion interruption in the dispersion structure, which can be observed through time-frequency analysis of the received acoustic signal. Based on the interruption frequency identified from the received signal, combined with the relationship between node depth and modal frequency calculated from the ocean sound speed profile, the source depth can be estimated if the receiver depth is known. The phenomenon of dispersion interruption and the method of estimating source depth have been validated through simulations and experiments.


32. Decoupled Latent Flow Matching for Few-Step Joint Vocal-Accompaniment Separation

Authors: Lishi Zuo, Youzhi Tu, Lu Yi, Zezhong Jin, Chongxin Gan et al.

Categories: cs.SD, cs.MM Score: 4.79/10 (Obj:6 Id:5 Ind:5 Comp:5 Eff:4 Nov:5 Rep:4)

  • Strength: 1 步潜域对抗后训练相对 20 步 flow matching 将人声 SDR 从 4.693 提至 5.895 dB、ViSQOL 3.717→3.826,且单一共享 1D 潜判别器在全部六指标上优于波形域多判别器方案(伴奏 ViSQOL 3.676 vs 3.438)。
  • Weakness: 全文零外部基线(判别式与同类少步方案均未对比),50 个评测片段全部取自训练集分布且语料与划分协议未指名,无代码仓库且冻结 VAE 来源与潜空间维度缺失,解耦设计(Separation Encoder+REPA)无独立消融,人声 SDR 距 VAE 重构天花板仍差 2.9 dB。
  • Full review: Claude Code 全文七公理审稿

Generative modeling provides a flexible way to model mixture-conditioned source distributions, but iterative diffusion and flow matching models are costly for long music signals. This paper studies joint vocal-accompaniment separation through latent flow matching, where a pretrained variational autoencoder (VAE) maps mixtures and sources into a compact latent space and a flow matching model generates vocal and accompaniment latents jointly. The proposed framework decouples semantic separation from acoustic velocity prediction through a Separation Encoder and a Velocity Decoder. To reduce sampling cost, we further apply latent adversarial post-training inspired by Flow2GAN for few-step generation. Experiments show that latent adversarial refinement can improve perceptual and separation metrics under a reduced sampling budget.