Daily Papers — 2026-07-19
4 papers on audio, speech, music, and acoustics.
1. SALMONN-2: Advancing General-Purpose Hearing Abilities with Self-Supervised Representations
Authors: Xiaoyu Yang, Xuenan Xu, Wenyi Yu, Siyin Wang, Changli Tang et al.
Categories: eess.AS | Disclaimer: This work has been submitted to the IEEE for possible publication Score: 6.70/10 (Obj:8 Id:6 Ind:9 Comp:5 Eff:9 Nov:5 Rep:8)
- Strength: SALMONN-2 用单一 SSL 编码器(SPEAR)+ MLF adapter 在仅 18.2k 小时训练数据下,在 MMAU-Pro(58.5)、MMAR(64.5)、MMSU(69.5)三个 ALLM benchmark 上达到 SOTA,超过使用 55k-13M 小时数据的 Kimi-Audio/MOSS-Audio/AF-3,并在 SED(70.15)、spoofing(73.65)、SQA(0.661)上超过专门模型。
- Weakness: MLF adapter 本质是 concat + down-projection 的标准操作,与 weighted-sum 的参数量差异未控制,且 timestamp injection 和 SSL encoder 均直接采用已有工作,单个组件新颖性有限。
- Full review: Claude Code 全文七公理审稿
Recent audio large language models (ALLMs) are typically built upon audio encoders trained with large amounts of supervised data. Since self-supervised learning (SSL) audio encoder models are known to learn general-purpose and transferable representations, we investigate whether general-purpose SSL audio representations can serve as an effective foundation for ALLMs. We present SALMONN-2, an ALLM built upon a unified SSL encoder. To better exploit the hierarchical representations learned by SSL encoders, we propose a multi-layer feature fusion (MLF) adapter that aggregates information from all encoder layers before projecting them into the language model. Beyond conventional audio understanding tasks, we further explore multimodal in-context learning (MICL) in ALLMs and study how this capability can be acquired through contextual biasing training. Experimental results show that a general-purpose SSL encoder achieves performance comparable to, or better than, specialised supervised audio encoders while providing a more balanced capability across speech, audio, music and paralinguistic tasks. SALMONN-2 further achieves state-of-the-art performance among comparable-scale open-weight models on ALLM understanding benchmarks, obtaining the best results on MMAU-Pro, MMAR and MMSU. We also show that MICL does not emerge naturally in ALLMs, but can be effectively acquired through targeted contextual biasing training.
2. Adaptive Momentum Enhanced Distributed Multichannel Active Noise Control for Faster Convergence under Communication Delays
Authors: Junwei Ji, Woon-Seng Gan, Boxiang Wang, Ziyi Yang, Haowen Li
Categories: eess.AS, eess.SP Score: 5.10/10 (Obj:8 Id:5 Ind:8 Comp:4 Eff:4 Nov:4 Rep:5)
- Strength: 针对真实 DMCANC 延迟-收敛 trade-off,将自适应动量(cosine similarity + power-law scaling)以低额外开销(~4Lw+2 mult/node)嵌入 ASSS-MGDFxLMS,在突发和波动延迟下均保持稳定且目测快于 ASSS-MGDFxLMS。
- Weakness: 全文无任何定量结果表格,仅 2 个仿真场景、6 节点单一声学配置,baseline 仅限作者自家算法谱系;三个核心组件(MCANC 动量、自适应动量系数、power-law scaling)均有先前工作,组合必要性未通过组件级 ablation 证明。
- Full review: Claude Code 全文七公理审稿
Distributed multichannel active noise control (DMCANC) reduces the computational burden of centralized ANC systems by distributing processing tasks across multiple nodes, while requiring information exchange to achieve satisfactory global noise reduction. To improve robustness under communication delays, the auto-shrink step size mixed-gradients filtered reference LMS (ASSS-MGDFxLMS) algorithm has been proposed. However, the reduced step size inevitably slows convergence. In this work, an adaptive momentum term is introduced to accelerate convergence, where cosine similarity is used to evaluate the alignment between the instantaneous gradient and the momentum component and dynamically adjust the momentum parameter. This design accelerates convergence when the directions are consistent while preserving stability under delayed communication. Simulation results demonstrate that the proposed adaptive momentum ASSS-MGDFxLMS (AMAS-MGDFxLMS) algorithm achieves faster convergence than ASSS-MGDFxLMS while maintaining stable and effective noise reduction performance.
3. AI_LectureNote: A Retrospective Pilot Study of a Post-ASR Workflow for English-Script Rendering and Semantic Drift in Korean-English Medical Lectures
Authors: Kyeongeon Lee, Donghoon Chang, Seungryeol Baek, Taehong Kim, Wonjun Yang
Categories: cs.CL | 12 pages, 4 figures Score: 3.58/10 (Obj:4 Id:4 Ind:2 Comp:5 Eff:3 Nov:2 Rep:9)
- Strength: 代码与标注全部开源、manifest 完整、复现性强;诚实承认单标注人与 4 讲座局限,明确不主张 population rate。
- Weakness: 标注人即系统开发者且无第二标注人/IAA,使核心 taxonomy 存在循环论证;4 讲座/2 说话人/polarity 集中在 1 讲座使现象不可靠;核心”readability vs faithfulness trade-off”在 LLM post-ASR correction 文献中已是预期结果,新颖性增量极小。
- Full review: Claude Code 全文七公理审稿
AI_LectureNote is a historical, readability-oriented post-ASR workflow for Korean-English medical lectures. It rewrites speech-to-text output into study transcripts while restoring Latin-script medical terms rather than Korean phonetic transliterations. We retrospectively evaluate the workflow on four author-recorded lectures across five conditions. In this pilot, post-processing raised the macro English-script rendering rate from 0.39 to 0.71 on the whisper-1 path and from 0.26 to 0.65 when applied to 3-minute chunked gpt-4o-transcribe output. However, English-script rendering did not imply semantic faithfulness: the two post-processed conditions showed semantic drift in 34 and 36 of 282 reference sentences and polarity failures in 11 and 13 of 101 polarity-cue rows. A descriptive cross-input comparison suggested different candidate failure patterns: polarity-failure sets overlapped more strongly across front-ends (Jaccard 0.60; 9 shared of 15 unioned failures) than general semantic-drift sets (Jaccard 0.23; 13 shared of 57 unioned drifts). This single-annotator pilot documents concrete failure modes rather than population rates and supports evaluating surface accuracy, term-script rendering, chunk-level script consistency, and medical-meaning preservation separately.
4. Robust Assamese Speech Recognition through Controlled Fine-Tuning of Whisper Models
Authors: Ganapati Das, Dwipen Laskar, Hasin Afzal Ahmed, Sanjib Kr Kalita, Kshirod Sarmah et al.
Categories: cs.LG | 31 pages, 11 figures Score: 3.47/10 (Obj:6 Id:2 Ind:5 Comp:3 Eff:4 Nov:2 Rep:5)
- Strength: 在 7.66 小时极低资源数据上完成 Whisper-Small 阿萨姆语微调,CER 从 1.9091 降至 13.18%,HER 从 0.5552 降至 0.0183,定性错误分析对正字法/形态学错误类型分类有描述价值。
- Weakness: 无任何消融实验隔离变量贡献,未与 HuggingFace 已有的阿萨姆语 Whisper 微调模型及 Chen et al. 2025 (WER=29.4) 对比,”first” 声明不成立,方法为标准技术组合无机制创新,绝对 WER=43.75% 未达可用水平。
- Full review: Claude Code 全文七公理审稿
Developing Automatic Speech Recognition (ASR) for morphologically rich, low-resource languages such as Assamese is challenging due to insufficient annotated speech data. The pretrained Whisper model performs poorly on Assamese speech recognition tasks. This paper presents a controlled, fine-tuned Whisper-based Assamese ASR system trained on the Mozilla Common Voice 24.0-Assamese corpus. A hardware-aware optimized training pipeline is implemented for resource-constrained environments, employing mixed-precision training and gradient accumulation on Tesla 4 Graphics Processing Units (T4 GPUs). The proposed fine-tuned model significantly outperformed the Zero-shot baseline, yielding Word Error Rate (WER), Character Error Rate (CER), Match Error Rate (MER), and Word Infomation Loss (WIL) of 43.17\%, 13.18\%, 43\%, and 64.81\%, respectively, achieving significant relative improvements of 78.26\%, 93.10\%, 57.0\%, and 35.19\% over the baseline. Semantic evaluation of the fine-tuned model also demonstrates notable improvement over a zero baseline, attaining Bilingual Evaluation Understudy (BLEU) and Metric for Evaluation of Translation with Explicit ORdering (METEOR) scores of 30.81 and 0.5262, respectively. Additionally, the predicted hallucination rate and Real-Time Factor (RTF) are substantially improved by 96.70\% and 32.38\%, compared to the zero-shot baseline.