Daily Papers — 2026-08-29
10 papers on audio, speech, music, and acoustics.
1. Anchoring Speech with Semantics: A Multimodal Adapter Mechanism for Automatic Speech Recognition in Low-Resource Languages
Authors: Kuan-Tang Huang, Cheng-Yeh Yang, Chien-Chun Wang, Hung-Shin Lee, Hsin-Min Wang et al.
Categories: cs.CL, cs.SD, eess.AS | Accepted to EMNLP 2026 (Main Conference) Score: 7.70/10 (Obj:9 Id:9 Ind:6 Comp:7 Eff:9 Nov:7 Rep:6)
- Strength: 在两个约 30 小时的低资源汉语变体上,SAMA-ASR 以仅 +8.1% 参数将 CER 从最佳基线 23.60/19.30 进一步降至 20.64/16.82(与 LoRA 叠加相对再降 12.54%/12.85%),且实用自动锚点版(22.48)反超 oracle 版纯语义基线 TG-ASR(24.28)。
- Weakness: 全部结果为单次贪心解码 CER、无显著性检验与多种子方差,实验仅限两个 Sinitic 变体且多语言诊断为 MT 派生,对其他语系泛化与第三方复现(代码库 1 star)均缺证据。
- Full review: Claude Code 全文七公理审稿
Low-resource ASR remains difficult because scarce transcripts provide limited supervised evidence for target-side generation. To address this gap, we propose SAMA-ASR, a lightweight adapter mechanism that augments the decoder with semantic anchors from auxiliary translations and an acoustic anchor from speech; in principle, the mechanism can be applied to similar encoder–decoder multitask speech models. Through cross-modal adaptation, SAMA-ASR conditions decoder states on translation-derived semantic embeddings and a speech embedding, combining utterance-level meaning with speech-grounded evidence before token prediction. At evaluation time, these semantic anchors can be generated automatically by an upstream speech-to-text translator rather than supplied as oracle translations. Experiments on two 30-hour datasets covering the low-resource Sinitic varieties Taiwanese Hokkien and Hakka show that SAMA-ASR improves over acoustic, prior prompt-based, and semantic-only translation-guided baselines and remains effective in practical automatic semantic-anchor settings; translator-capacity analyses show that useful semantic anchors can be produced by a compact ST model.
2. HEAR Who Said What: Unlocking Speaker-Attributed Reasoning via Counterfactual Voice Grounding
Authors: Dongwook Lee, Sangkwon Park, Eunwoo Song, Che Hyun Lee, Youngho Cho et al.
Categories: cs.CL, cs.AI, cs.SD | EMNLP2026 Main Conference Score: 7.53/10 (Obj:8 Id:7 Ind:7 Comp:8 Eff:7 Nov:8 Rep:8)
- Strength: HEAR 基准以反事实配对设计量化了 20 个 SLM 的语义幻觉(Qwen3-Omni-Thinking 配对精度 71.4%→17.3%),训练出的 A2R-30B 在 HEAR 上提升 31.7 个点(35.4→67.1)并在 WDYL 零样本迁移 +36.4、真人重录音 +48.0,通用能力几乎无损(BBH 93.90 vs 94.00)。
- Weakness: 缺乏与最强闭源模型的差距分析(A2R 67.1 仍落后 Gemini-3-Flash 73.1 与 Gemini-3.1-Pro 88.8)、判别任务近乎未解决(VCD 38.7/VL 48.4)、关键消融仅单种子且标注者间一致性未报告,TTS 伪影混淆路径仅靠真人子集部分排除。
- Full review: Claude Code 全文七公理审稿
Speech Language Models (SLMs) are increasingly deployed in multi-speaker environments, yet their ability to attribute speech to the correct speaker and reason over speaker identities remains unclear. Hence, we introduce HEAR, a conceptually hierarchical benchmark diagnosing the foundational capabilities of speaker-attributed reasoning, comprising 2.4K human-verified samples from 887 diverse multi-party audio clips. Evaluating 20 leading SLMs on HEAR reveals they struggle with these foundational tasks, often relying on semantic priors rather than actual vocal cues. To address this, we present A2R, a 30B model optimized on Counterfactual Audio with Speaker-level Hard negatives (CASH), a dataset designed to guide the model to prioritize acoustic vocal cues over linguistic signals. A2R achieves strong performance on HEAR and exhibits zero-shot generalization to diverse multi-speaker downstream tasks, demonstrating that learned speaker attribution unlocks the model’s latent capacity for speaker-aware reasoning. All resources are available at https://attributetoreason.github.io/AttributeToReason/
3. Beyond Speech: Dual-Domain SSL Fusion for Unified All-Type Audio Deepfake Detection
Authors: Cunhang Fan, Junqin Cao, Tian Gao, Zhipeng Xie, Jun Xue et al.
Categories: eess.AS | 8 pages, 5 figures; accepted to the AT-ADD Grand Challenge at ACM Multimedia 2026 (MM ‘26) Score: 6.84/10 (Obj:8 Id:8 Ind:7 Comp:6 Eff:8 Nov:6 Rep:4)
- Strength: 双域 SSL 融合统一检测器在 AT-ADD Track 2 评测集达 95.58% Macro-F1(第二名),比官方基线高 20.30 点、比自身单融合模型高 1.14 点且增益均匀分布于四类音频,硬路由失败实验(91.46%)为统一路径提供了数据论证。
- Weakness: 方法组件全部为既有技术组合而代码与推理成本均未披露(3 个 300M 级 SSL 前向 + Whisper-large-v3 × 5 裁剪的开销无任何参数量/时延报告),评测集被用于逐步确认集成成员而无 bootstrap 置信区间量化累积偏差。
- Full review: Claude Code 全文七公理审稿
Unified all-type audio deepfake detection aims to determine whether an input clip is real or fake when its audio type may be speech, environmental sound, singing voice, or music. Existing speech-centric or type-dependent solutions are insufficient for this setting because the test-time audio type is unknown, while the required output is still a single binary decision. To address these issues, this paper proposes a dual-domain SSL fusion method that maps heterogeneous audio into a shared binary authenticity space. EAT-large and wav2vec 2.0 XLS-R-300M are used as complementary SSL feature sources, providing broad acoustic and event-level representations as well as waveform-level, vocal, and speech-sensitive representations. Layer-wise weighted fusion integrates multi-level artifacts from different transformer depths, while token-level fusion forms a unified feature pool without enforcing frame-level alignment between the two SSL streams. The fused tokens are summarized by multi-head attentive statistics pooling and classified with a binary MLP head. With conservative speech refinement applied on top of this unified core detector, the submitted system achieves 95.58% Macro-F1 on the AT-ADD Track 2 evaluation set and ranks second in the challenge.
4. V2TATC: A Joint Voice-Trajectory Embedding Framework and Dataset for Air Traffic Controller Situational Awareness
Authors: Louis Brusset, Mathurin Petit, Jordan Kam, Alexandre Bayen
Categories: cs.LG, eess.AS | 41 pages, 21 figures, 9 tables Score: 6.70/10 (Obj:8 Id:8 Ind:7 Comp:5 Eff:5 Nov:8 Rep:6)
- Strength: 首个公开配对语音-轨迹数据集(湾区 8 频率、52,030 对、25.5 GB、CC BY 4.0,已核实可下载),对比对齐将跨模态检索 R@1 从随机 0.02% 提升至 23.5%(×1175),消融证明对比监督使联合空间有效秩从 806 降至 203、方差显著集中。
- Weakness: 模型层效用未达标——R@1 绝对值仅 23.5%、温度触 clamp 下限 0.04 暴露数据多样性不足(训练 0.45 vs 验证 1.585),语音条件轨迹预测误差约 3.1 km/86° 远超管制可用精度,且代码仓库未实际可查(GitHub 检索 0 结果)、逆向检索无量化指标。
- Full review: Claude Code 全文七公理审稿
As air traffic volumes in the National Airspace System continue to expand, in particular in the low altitude airspaces, the need for scalable decision support tools used by air traffic controllers will also require more development. This article introduces Voice-to-Trajectory for Air Traffic Control, a joint voice communication-flight trajectory data embedding framework, that can be a component of situational awareness in congested airspaces, and assist the development of tools for ATC as they reason in real-time over Automatic Dependent Surveillance-Broadcast trajectories, or the intent expressed by pilots in natural language. We show that these data modalities are not independent and represent a common physical referent: an aircraft flying through the airspace. V2TATC maps a voice instruction and the trajectory of the addressed aircraft to nearby points in a single latent space that can be queried in both directions. It combines a self-supervised trajectory encoder, a frozen large-scale speech encoder, a contrastive joint embedding, and a bijective lifting via normalizing flows. We demonstrate V2TATC’s effectiveness on the San Francisco Bay Area, for its concentration of major airports, and its mix of commercial and general aviation low altitude traffic. Lastly, we release a novel paired voice-trajectory dataset, and report experiments on cross-modal retrieval, ablations, and latent-space analysis.
5. Diagnose, Then Refine: A Closed-Loop TTS System with AudioLLM-Guided Correction
Authors: Zeyang Song, Tianchi Liu, Tianrui Wang, Chenglin Xu, Steven Y. Guo et al.
Categories: cs.SD, cs.AI | Accepted by EMNLP 2026 main conference Score: 6.47/10 (Obj:8 Id:5 Ind:6 Comp:6 Eff:8 Nov:7 Rep:4)
- Strength: LoopTTS 用 Filter-Judge-Refiner 闭环修复 TTS 局部韵律缺陷,被标记语音盲评 MOS 从 3.21/3.01 提升到 4.17/4.01,成对偏好 85.67% 胜原始音频、72.3% 胜 EmoVoice 重生成,词级指令跟随 Stress MOS-I 达 4.57。
- Weakness: 缺少与同期指令式语音编辑系统(CosyEdit、dots.tts.edit、FireRedTTS3 等)的直接对比,训练标注/诊断/评测三方共用 Gemini-3-pro 存在循环性,人类评测仅覆盖被标记子集,且代码与 Refiner-DB 数据截至审稿日全部未发布。
- Full review: Claude Code 全文七公理审稿
Current TTS systems typically rely on open-loop, single-pass generation and can produce sporadic local prosodic defects, such as misplaced stress, unnatural pauses, or flattened intonation, that utterance-level metrics often fail to expose. We present LoopTTS, a judge-guided Filter-Judge-Refiner framework for recovering low-quality TTS outputs diagnosed by an AudioLLM. Given an initial utterance from a base TTS model, an AudioLLM Judge identifies salient prosodic issues and generates structured refine instructions; a Refiner, our fine-grained instruction-following TTS model, then performs guided expressive re-synthesis conditioned on the initial utterance, target text, and instruction. To train the Refiner, we construct Refiner-DB, a 42K-example AudioLLM-annotated dataset with word-level prosodic weak supervision. Human evaluation on diagnosed low-quality utterances shows that LoopTTS can detect perceptually salient errors and correct them with the Refiner, outperforming raw generated audio and practical open-loop re-generation baselines in recovery quality. The Refiner also demonstrates stronger instruction-following ability for stress and pause control in targeted prosody modification.
6. When Vocal Tone and Literal Meaning Diverge: An Acoustic-Semantic Incongruity Study for Large Audio-Language Models
Authors: Yu-Wen Chen, William Ho, Maxim Topaz, Zoran Kostic, Julia Hirschberg
Categories: eess.AS | EMNLP 2026 Findings Score: 6.42/10 (Obj:6 Id:6 Ind:5 Comp:8 Eff:7 Nov:7 Rep:5)
- Strength: 受控 TTS 解耦构建 77k 条声学-语义错位基准,首次给出 LALM 错位盲区的层级机制证据(基线 Accdual≤0.320、Qwen2-Audio 语义-声学差随深度从 0.023 增至 0.242),rsLoRA SFT 将域内 Accacou 提升 +0.427 且转录 WER 保持 ~0.07 不退化。
- Weakness: 声学标签仅 150 样本人类校验(64.6%、α=0.54、2/3 评估者为共同作者),Whisper-large-v3 同源 AER 过滤与 Qwen2-Audio 编码器构成未声明的评测偏差,域外 LISTEN 三模型结果方向不一且数据集音频本体在接收后仍未公开。
- Full review: Claude Code 全文七公理审稿
Affective cues across modalities may be incongruous (e.g., sarcasm or mocking praise), potentially leading to misinterpretation when relying on a single modality. Large Audio-Language Models (LALMs) have recently gained popularity and been applied to multimodal emotion recognition, but their ability to disentangle acoustic and semantic cues, especially in incongruent cases, remains underexplored. To address this gap, we introduce CREMA-ASIS, a dataset specifically created to investigate incongruence between acoustic emotion and semantic sentiment cues. It pairs acoustic emotion labels with semantic sentiment polarities. Using this dataset, we evaluate LALM biases within a multitask framework and conduct a layer-wise analysis to identify modality dominance across layers. Our findings reveal that LALMs struggle with semantic-acoustic incongruent cases, rarely predicting incongruity, and that LALMs are predominantly influenced by semantic information. However, supervised fine-tuning significantly improves LALM performance on our CREMA-ASIS test set while preserving transcription accuracy and joint emotion recognition. Results demonstrate potential for enhancing both acoustic and semantic understanding on out-of-domain data.
7. Acoustically Grounded Cost Learning for Open-Vocabulary Audio-Visual Semantic Segmentation
Authors: Tianrui Hui, Shaofei Huang, Qisong Han, Yaxiong Wang, Lechao Cheng et al.
Categories: cs.CV | Accepted by ACM MM 2026 Score: 6.26/10 (Obj:8 Id:5 Ind:7 Comp:5 Eff:8 Nov:6 Rep:4)
- Strength: 类别特定代价学习在 AVSBench-OV 上将 unseen 类 mIoU 从 29.14 提升至 45.59(+16.45)、Overall 从 44.81 提升至 55.20,并在 VPO zero-shot OOD 上达 MIoU 57.72、超过有监督基线 AVSegFormer(56.46)。
- Weakness: 主基准直接基线仅 OV-AVSS 一家、同主题 Pattern Recognition 2026 论文未被对比,且承诺开源的 GitHub 仓库截至审稿日为空仓库(无任何文件与 license),关键实现无从验证。
- Full review: Claude Code 全文七公理审稿
Open-Vocabulary Audio-Visual Semantic Segmentation (OV-AVSS) aims to perform pixel-level segmentation of sound-emitting objects from an open set of categories. The previous method relies on a class-agnostic foreground definition, which groups semantically diverse objects into a heterogeneous positive set, causing the model to learn unstable sounding patterns and produce unreliable proposals. To address this, we reformulate the objective to be category-specific and propose a novel Acoustically Grounded Cost Learning (AGCL) framework to transform the static, audio-agnostic visual-text priors into dynamic, audio-grounded cost representations. For intra-category soundingness discovery, we devise Audio-Modulated Cost Generation (AMCG) and Audio-Guided Temporal Aggregation (AGTA) modules to enable both frame-level sounding region highlighting and video-level temporal refinement with a low-intrusive audio injection mechanism. For inter-category distractor discrimination, we introduce a Synergistic Distractor Mining (SDM) strategy, which selectively penalizes acoustically and semantically confusing negative categories to learn more discriminative decision boundaries. Extensive experiments on the AVSBench-OV dataset demonstrate that our method significantly outperforms previous state-of-the-art approaches, particularly on unseen categories. Code is available at https://github.com/spyflying/AGCL.
8. Sound Analysis for Speed Estimation of Induction Motors Under Non-Stationary Conditions
Authors: Tomas A. Garcia-Calva, Daniel Morinigo-Sotelo, Konstantinos N. Gyftakis
Categories: eess.SP Score: 6.00/10 (Obj:8 Id:6 Ind:6 Comp:8 Eff:6 Nov:5 Rep:4)
- Strength: 低成本麦克风声学法在负载阶跃/振荡等非平稳工况下实现瞬时转速跟踪,MAPE 0.0776%(RMSE 0.0309 r.p.s.)、时延约 1 ms,并以素数阶 BPF 谐波 + sinc 包络论证给出谐波选择的物理解释。
- Weakness: 无任何数值对比基线(FFT/电流法均未实测),验证范围仅 [23.5, 25] r.p.s. 窄区间且启动暂态被排除,无代码与数据发布、噪声鲁棒性未检验,核心 ST-MUSIC 技术承袭同团队 2019 年工作且本文为 SDEMPED 2025 会议版扩展。
- Full review: Claude Code 全文七公理审稿
This paper presents a novel methodology for the estimation of the rotational speed of induction motors operating under non stationary conditions, using acoustic signals acquired by a low cost microphone. The proposed approach integrates multi rate digital signal processing with advanced time frequency analysis to extract speed dependent harmonic components from motor acoustic emissions, while effectively suppressing electrical interference and electromagnetic noise. Numerical simulations and experimental validations are conducted under both steady state and transient operating conditions, considering a wide variety of speed profiles, including slow ramps and abrupt speed variations. The experimental results demonstrate that the proposed method is capable of accurately tracking rapid and dynamic speed changes caused by load disturbances and mechanical irregularities. The estimated instantaneous speed exhibits a high degree of agreement with reference measurements obtained from a conventional speed sensor. The results confirm that the proposed acoustic based technique provides a reliable, noninvasive, and cost effective solution suitable for speed monitoring of induction motors.
9. Free Speech and Artificial Intelligence
Authors: Etienne Brown
Categories: cs.CY, cs.AI | Preprint. Forthcoming in Mark Satta, Étienne Brown, and JP Messina (eds.), Philosophy and Free Speech: An Introduction. London: Routledge, 2027 Score: 5.90/10 (Obj:8 Id:6 Ind:7 Comp:7 Eff:5 Nov:5 Rep:5)
- Strength: 系统组织了推荐算法与 LLM 聊天机器人的言论自由辩论,覆盖 2024-2025 最新案例法(Anderson v. TikTok)与约 40 条文献,并提出「shadowbanning 与删除在传播性利益框架下道德等价」等有启发性的论证角度。
- Weakness: 全章自陈「提出而非解决」,核心概念(采纳、表达性)被承认模糊后搁置,监管理由的因果前提未被证实,且分析仅锚定美国第一修正案而完全缺席欧盟 DSA/AI Act 等现行监管语境,无任何可检验命题或经验证据(仅一条二手引用)。
- Full review: Claude Code 全文七公理审稿
Philosophers and legal scholars are engaged in debates about the implications of artificial intelligence for freedom of expression. This paper analyzes the free speech issues raised by two distinct AI technologies: social media recommendation algorithms and conversational AI (i.e., chatbots powered by large language models). The first part shows that, through their recommendation algorithms, social media platforms control the dynamics of speech visibility in the digital public sphere, making algorithmic recommendation relevant to the philosophy of free speech. The second part turns to conversational AI. It discusses both the reasons for granting or withholding speech rights to artificial agents and users’ right to receive information, which may render specific forms of chatbot regulation illegitimate. Throughout, the chapter also considers whether social media platforms or AI developers hold corporate speech rights. Its general aim is to raise rather than settle questions that arise from the rapid development of AI technologies.
10. Disentangling Representation using Attributes-based Gaussian Estimation for Medical Sound Diagnosis
Authors: Ke Zhao
Categories: cs.AI, cs.CV, eess.AS | 4 figures, 2 tables, code is available at: https://github.com/ZhaoKe1024/DisentangledRepr Score: 5.20/10 (Obj:6 Id:5 Ind:5 Comp:7 Eff:5 Nov:4 Rep:6)
- Strength: 在 COUGHVID 咳嗽音频 COVID 检测上验证集 AUC 达 97.06%(SVM),比最强基线 IDB-SR 的 93.32% 高 3.74 个百分点,消融显示去掉属性分支后 Recall 从 99.00% 降至 75.87%、AUC 降至 84.02%,代码已随论文开源。
- Weakness: 公平性主张无任何量化证据——未纳入性别/年龄等敏感属性、未报告分组公平性指标,全部指标来自仅 720 个原始样本(2,850 段)的验证集且未说明按受试者切分,99.00% Recall 存在说话人泄漏虚高风险,方法与 DIVA 类弱监督解耦先行工作缺少正面对比。
- Full review: Claude Code 全文七公理审稿
Deep learning has a powerful capability of feature extraction. However, the lack of fairness and interpretability in deep neural networks poses limitations to their adoption in the medical domain. This paper proposes a disentangled representation learning (DisenRL) framework, named the Attributes-based Gaussian Estimation for Disentangled Representation (AGEDR), which incorporates Attribute Mapping Embedding (AME) modules designed to map attributes into vectors and align them with a subset of the latent vectors in a Variational AutoEncoder (VAE). This part of the latent vector will be disentangled from the remaining latent vectors by minimizing mutual information. A classifier is then trained using the mean parameters of the latent vectors from the VAE. Extensive experiments demonstrate that AGEDR outperforms both conventional classification models and existing disentangled representation learning methods. The ablation experiments also indicate the disentangling capability and fairness of AGEDR. The source code is publicly available at https://github.com/ZhaoKe1024/DisentangledRepr.