Daily Papers — 2026-08-23
7 papers on audio, speech, music, and acoustics.
1. Pre-Decoding Acoustic Triage for Budgeted Vision-Language Captioning of Untrimmed Egocentric Video
Authors: Masoud Jalayer, Changyi Li, Yu Xiao
Categories: cs.CV, cs.AI, cs.SD | 18 pages, 5 figures, 9 tables. Under review at IEEE BigData 2026, Industrial and Government Track. Code: https://github.com/masjalayer/PreDecoding-AcousticTriage Score: 7.60/10 (Obj:8 Id:8 Ind:8 Comp:6 Eff:8 Nov:6 Rep:8)
- Strength: 用 span-level MIL 目标替换帧级交叉熵,在冻结 BEATs 特征下于 EK-100 24 个录制中 22 个胜过均匀采样(25% 调用时平均 +5.8 点,Wilcoxon p<10⁻⁴),并主动量化 occupancy 上限与视觉前线的算力反例(光学流整账 209 > 全扫 185 GPU 小时)。
- Weakness: 新颖性限于已知组件的目标层重构,VLM 描述质量仅 200 个刺激且无人工评测佐证,核心主张「覆盖率代理即下游效用」缺乏直接证据。
- Full review: Claude Code 全文七公理审稿
Automatically analyzing hours-long egocentric video is increasingly essential for progress monitoring, quality control, and safety in logistics, construction, and manufacturing. Yet current pipelines that process short, fixed-size windows with a vision-language model (VLM) are prohibitively expensive because cost scales with the number of model calls. To reduce this cost, prior work proposes triage policies to select which windows merit a VLM invocation. However, these policies either sample uniformly or rank windows using visual features, which ironically requires the video decoding that the budget constraints are meant to avoid. We propose audio-first triage: select windows using the lightest modality, scored before any video frame is decoded, so the approach composes naturally with token compression or quantization. The novelty lies in the objective, not the representation: rather than a per-frame sound-event detector, we train the selector to trigger once per action. This objective shift improves action coverage by 4.0-10.8 percentage points across all evaluated call rates, using frozen AudioSet-pretrained features without domain-specific sound-event labels. Using fewer than half of the available calls, the triage cuts 9-20% of VLM calls at matched coverage on EPIC-KITCHENS-100 (EK-100), surpasses uniform sampling through the mid-range on Ego4D over 247 clips, and outperforms two recent visual keyframe selectors. Code, the reference implementation and every results file this manuscript reads are at https://github.com/masjalayer/PreDecoding-AcousticTriage.
2. Mitigating Speaker Leakage in Cascaded Multi-talker ASR with Diarization-based Transcript Correction
Authors: Hermann Yepdjio Nkouanga, Minwei Luo, Maggie Wigness, Suresh Singh
Categories: eess.AS, cs.CL | Accepted to INTERSPEECH 2026 Score: 6.80/10 (Obj:8 Id:7 Ind:8 Comp:6 Eff:8 Nov:6 Rep:4)
- Strength: 高泄漏子集上相对 cpWER 最高降 29.26%(Mossformer 在 AMI IHM 65.58→46.39),两骨干、全重叠/部分重叠/远场会议三类数据上全量集一致改善,且方案零训练、仅靠 pyannote 分离验证器。
- Weakness: 未与任何重标注基线(LSEC/AG-LSEC/SEAL)做数值对比,核心”剪除优于重标注”主张无直接证据,且全量集增益多不足 3%、无显著性检验、无代码发布。
- Full review: Claude Code 全文七公理审稿
While cascaded multi-talker ASR (MT-ASR) leverages state-of-the-art foundation models, its performance is often capped by speaker leakage during separation. Prior correction strategies primarily focus on lexical re-labeling for speaker attribution. We propose a complementary pruning-based paradigm that robustly identifies and removes leakage artifacts. Our method utilizes a pre-trained speaker diarization model as a multimodal verifier to prune transcribed segments satisfying a tripartite consensus of temporal containment, lexical cross-validation, and temporal alignment. Results on LibriMix, LibriSpeechMix, and the AMI Meeting corpus show our algorithm consistently reduces cpW ER across diverse overlap conditions. Specifically, on subsets with high speaker leakage, our method achieves relative cpW ER reductions of up to 29%, highlighting its effectiveness in enhancing the reliability of cascaded MT-ASR transcripts in complex acoustic environments.
3. MRMAD: A Multi-Round Multi-Audio Benchmark for Evaluating Acoustic Degradation Perception in Large Audio-Language Models
Authors: Yize Li, Ningyuan Yang, Sile Yin, Sindhuja Thogarrati, Sung-En Chang et al.
Categories: cs.SD, cs.LG, eess.AS | EMNLP 2026 Score: 6.40/10 (Obj:1 Id:2 Ind:1 Comp:1 Eff:2 Nov:2 Rep:1)
- Strength: 首个多轮多音频 LALM 劣化感知基准(8400 题、3 领域、9 类劣化),系统性评测 18 个模型并发现开放模型贴近随机基线(最佳 Qwen3-Omni-Thinking DTI 仅 36.58%),闭源 Gemini 3.1 Pro 领先 49.64 个绝对点。
- Weakness: 评测框架与数据未发布(GitHub Bose/MRMAD 为空仓库)、缺人类基线、诊断仅基于 900 样本无方差报告,当前版本不可复现且绝对分数缺乏锚定。
- Full review: Claude Code 全文七公理审稿
Large audio-language models (LALMs) have shown promising progress in understanding speech, music, and general sound events, yet their ability to reason about how audio signals are degraded remains underexplored. Existing benchmarks primarily evaluate semantic understanding, event recognition, or high-level audio reasoning, leaving a basic question unanswered: Do LALMs understand the differences in audio quality? We introduce MRMAD, a Multi-Round Multi-Audio Degradation benchmark for evaluating audio degradation perception and understanding in LALMs. MRMAD spans speech, music, and sound, and frames evaluation as multi-turn dialogues over multiple audio inputs, requiring models to identify degradation types, compare severity, and perceive corruption changes across turns. Unlike current single-turn audio-language benchmarks, MRMAD evaluates whether LALMs can maintain consistent degradation hypotheses with new evidence and explain low-level acoustic phenomena in natural language. Through a systematic evaluation of 18 representative LALMs from non-thinking to reasoning and Omni models, we find that current models often recognize coarse content while failing to diagnose, compare, or reason about degradations reliably. MRMAD reveals an important yet overlooked aspect of audio-language understanding and provides a diagnostic foundation for building future LALMs that are robust to real-world acoustic conditions.
4. Dual-Scale State-Space Modeling with Speaker-Wise Dynamic CRF for Speech Emotion Recognition in Conversation
Authors: Guan-Hua Wen, Kuan-Yu Chen, Hou-Chiang Tseng
Categories: cs.LG | 5 pages, 2 figures, 4 tables. Submitted to ICASSP 2027 Score: 6.40/10 (Obj:8 Id:8 Ind:8 Comp:5 Eff:8 Nov:5 Rep:2)
- Strength: DSSM-CRF 在 IEMOCAP 上取得 75.81% UA / 74.90% WA(领先第二名 1.91/2.60 个百分点),并用三组匹配对照(说话人分解 +0.99 UA、CRF +1.17 UA、残差转移 +1.72 UA)干净归因了各组件贡献。
- Weakness: 全文未发布任何代码/权重(GitHub 零命中),方法基本由已有组件组合而成,且 MELD 上仅领先第二名 0.98 WA,缺乏显著性检验与效率对比。
- Full review: Claude Code 全文七公理审稿
Conversational speech emotion recognition must reconcile acoustic evidence across temporal scales with two interaction processes: cross-speaker contextual influence and within-speaker emotion evolution. We propose DSSM-CRF, an audio-only architecture that explicitly separates these processes. Bidirectional state-space models encode fused self-supervised speech representations at frame and dialogue scales, so each utterance representation captures local prosody and context from all speakers. The decoder then orders each speaker’s utterances into an independent dynamic conditional random field chain. Consecutive utterances in a speaker’s chain form a transition pair whose score combines a corpus-level transition matrix with a residual predicted from the two contextualized utterances. An auxiliary objective supervises whether each pair changes emotion but does not participate in Viterbi inference. Thus, interlocutor turns affect contextual emotion scores without being treated as transitions in another speaker’s emotion trajectory. DSSM-CRF achieves 75.81% UA and 74.90% WA on IEMOCAP, and 54.72% WA and 49.31% WF1 on MELD. Matched controls demonstrate complementary gains from speaker-wise factorization and CRF modeling.
5. AudioNoisePrints: Model-free audio watermarking using spatial correlation in flow matching TTS
Authors: Timothy Tin-Long, Jian Zhu, Aidan Pine, Mengzhe Geng
Categories: cs.SD, cs.CL Score: 6.30/10 (Obj:8 Id:6 Ind:6 Comp:8 Eff:6 Nov:6 Rep:5)
- Strength: 提出零生成开销、零重训练、零质量折损的音频归属检测方案,在 speed 与 cropping 等强增强下显著优于 AudioSeal(speed 1% 拉伸即令 AudioSeal 失效),并验证空间相关性在 F5-TTS、Matcha-TTS、DiffWave 上普遍成立。
- Weakness: 核心卖点”不损害生成质量”缺乏任何客观质量指标支撑,Echo 增强下准确率仅约 0.5,且未与 GROOT/latent watermarking 等同类别基线对比,也未发布代码。
- Full review: Claude Code 全文七公理审稿
We present AudioNoisePrints, a training-free watermarking pipeline for flow matching and diffusion TTS models, which requires minimal extra computation during inference and does not require retraining the TTS model or reducing the generation quality. We exploited the fact that there are strong correlations between the initial Gaussian noises and the generated outputs in diffusion and flow matching models, such that a simple cosine correlation between the initial noise and the generated output can be used to perform watermaking. Moreover, we train a lightweight detector on top for more aggressive augmentations. Our method outperforms AudioSeal, a strong baseline for audio watermarking under strong augmentations. We experimented on F5TTS and other TTS and vocoder models, and concluded that they all exhibit similar spatial correlation properties, suggesting our watermarking scheme can be used for more flow-matching TTS models and even vocoders in the future.
6. Motion-Aware Reasoning from Speech to Mask Tracks: Runner-up Solution for the MeViS-Audio Track of the 8th LSVOS Challenge 2026
Authors: Jinxing Zhou, Suiyi Zhao, Yanghao Zhou, Ruohao Guo
Categories: cs.MM, cs.CV, cs.SD Score: 5.79/10 (Obj:7 Id:6 Ind:9 Comp:4 Eff:7 Nov:5 Rep:2)
- Strength: 官方 MeViS-Audio 测试集亚军(Final 0.7057,J&F 0.5374),相对季军 J&F +0.0787,验证集定性案例 J&F 达 0.9561。
- Weakness: 六个大模型串联的流水线无组件级消融,Final 落后冠军方案 0.064(0.7057 vs 0.7696),且代码权重未发布、GPT 依赖不可复现。
- Full review: Claude Code 全文七公理审稿
Speech-guided referring video object segmentation aims to recover the mask tracks of objects specified by a spoken motion description. Here, speech carries a linguistic instruction rather than acoustic evidence from a sounding object, so a solution must connect speech recognition, motion-centric temporal grounding, mask tracking, and explicit no-target handling. We introduce Speech2MaskTrack, our approach for the MeViS-Audio track of the 8th LSVOS Challenge. Speech2MaskTrack transcribes the spoken query and compiles it into structured constraints over category, count, direction, interaction role, and temporal phase. SAM3.1 enumerates multiple instance tracks, which TRACE ranks using complete-trajectory motion and relation evidence. A frozen lexical presence gate may suppress the ranked SAM3.1 base prediction. When the gate predicts that a target is present, an available full-expression-conditioned SaSaSa2VA track replaces the SAM3.1 mask. Only outputs that remain empty enter GPT-assisted recovery, which invokes SaSaSa2VA again under query- and mask-level verification. Speech2MaskTrack achieved second place in the official challenge ranking.
7. Cross-Subject Generalization in Decoding Perceived Speech from Non-Invasive Brain Recordings
Authors: Aoke Zhang, Bo Wang, Xihong Wu, Heping Cheng, Jing Chen
Categories: cs.SD, cs.AI | Submitted to IEEE Transactions on Audio, Speech, and Language Processing (TASLP) Score: 5.30/10 (Obj:7 Id:5 Ind:8 Comp:4 Eff:6 Nov:5 Rep:2)
- Strength: CPSD 在三个跨语言/跨模态的 MEG/EEG 数据集上,跨被试 Top-10 相对最强基线提升 6.8%/15.4%/15.8%,且特化阶段仅需多被试训练 7.5%-15.6% 的训练步数。
- Weakness: 无代码、权重或 PKUEEG 数据发布,且大部分增益归因于标准的源预训练+微调流程,PESA 模块的独立贡献伴随巨幅方差、证据偏弱。
- Full review: Claude Code 全文七公理审稿
Decoding perceived speech from non-invasive brain recordings has garnered significant attention in recent years due to its wide range of potential applications. However, existing methods face considerable challenges in cross-subject decoding, primarily due to limited generalizability and the absence of explicit mechanisms for extracting subject-consistent information. These limitations result in high training costs and suboptimal decoding performance. To address these challenges, we propose an innovative Cross-Subject Perceived Speech Decoding (CPSD) framework, which comprises two training stages: source model pre-training and personal specialization. In the source model pre-training stage, contrastive learning is employed to capture shared representations across multiple source subjects. Subsequently, personal specialization initializes the model for the target subject by extracting consistent components from the source model and fine-tuning it using target subject data. Additionally, we introduce the Positional Encoding-based Spatial Attention (PESA) module, which remaps MEG/EEG data into a standardized reference space, thereby enhancing cross-subject consistency and facilitating model training. We evaluate the proposed CPSD framework on three perceived speech neural datasets encompassing different modalities and languages. The results demonstrate that our framework outperforms baseline methods by more than 6.8%, 15.4%, and 15.8% in Top-10 accuracy on the Armeni 2022, PKUEEG 2025, and Broderick 2018 datasets, respectively. Further analyses confirm the effectiveness, efficiency, and robustness of the proposed approach.