Daily Papers — 2026-08-17
22 papers on audio, speech, music, and acoustics.
1. Numerical and perceptual validity of synthetic Head-Related Transfer Functions at scale
Authors: Katarina C. Poole, Lorenzo Picinali
Categories: eess.AS | Version 2 clarifies in the pdf this is a preprint submitted to JASA and awaiting review Score: 8.42/10 (Obj:9 Id:8 Ind:8 Comp:8 Eff:9 Nov:8 Rep:9)
- Strength: 在 200 人规模上系统评测合成 HRTF,行为定位中合成与个体测量在所有极角指标上无显著差异(如极角绝对精度 37.57° vs 32.64°),KEMAR 全面显著更差,为 BEM 合成 HRTF 提供规模化有效性的实证支持。
- Weakness: SRM 任务以 -3 dB 信掩比运行且正确率偏低(合作位 25.3%),任务难度可能掩盖 HRTF 效应,使 RQ4 的零结果难以支撑结论。
- Full review: Claude Code 全文七公理审稿
Individually measuring head-related transfer functions (HRTFs) at scale remains a central challenge for personalised spatial audio, motivating growing interest in synthetic HRTFs. We evaluated the numerical, computational, and behavioural validity of synthetic HRTFs, generated through the boundary element method simulation using Mesh2HRTF, against measured and KEMAR HRTFs using the Extended SONICOM dataset. Across 200 subjects, synthetic HRTFs deviated less from measured than KEMAR in interaural time and level differences, but residual errors, together with elevated spectral distortion, concentrated at low, rear elevations. This is consistent with the omission of torso geometry from the synthesis pipeline. Two computational models revealed a corresponding pattern of predicted localisation errors, with synthetic HRTFs positioned between measured and KEMAR. In a virtual reality localisation task (N = 20), synthetic HRTFs matched measured on every polar metric, while KEMAR was significantly worse. However, behavioural error clustered around the front-back midline regardless of condition, not at the low elevations implicated numerically or by the models. A separate spatial release from masking task (N = 18) showed no effect of HRTF type. Together, these results indicate that high-resolution synthetic HRTFs preserve behavioural localisation performance, despite discrepancies between the numerical/model-predicted bias and the spatial pattern of behavioural error.
2. INSPIRE: A Benchmark for Instruction-Aware Speech Retrieval
Authors: Chen-An Li, Hung-yi Lee
Categories: cs.SD, cs.CL, eess.AS | Interspeech 2026 long paper Score: 7.53/10 (Obj:8 Id:7 Ind:7 Comp:6 Eff:8 Nov:8 Rep:8)
- Strength: 首个指令感知语音检索基准,统一评测四大范式 12+ 模型,量化出语义检索级联最优(DailyTalk Qwen3-Embedding R@10=62.0 vs 随机 0.24)、副语言属性自监督最优(VCTK WavLM R@10=17.5)、多属性合成子集无方法超过 R@100=11.1 的清晰范式分水岭。
- Weakness: 正负样本标签仅源自数据集元数据且无人工抽检,VCTK 仅 80 查询单一属性,基线均无指令感知检索训练,任务难度与基线上限的混淆未完全解除。
- Full review: Claude Code 全文七公理审稿
Existing speech retrieval systems rely on fixed similarity matching and cannot adapt to diverse user intents. We introduce INSPIRE, the first benchmark for instruction-aware speech retrieval, in which natural-language instructions dynamically specify relevance criteria, including semantic content, speaker identity, speaking style, environmental sounds, and their combinations. We evaluate four retrieval paradigms: large audio-language models, cascaded pipelines, self-supervised speech models, and contrastive audio-language models. Our results reveal that no current method robustly handles all retrieval intents. Text-based approaches perform relatively better at semantic retrieval but struggle with paralinguistic attributes, while speech-based models are moderately better at capturing acoustic properties but falter at following instructions. These findings highlight the need for unified architectures capable of instruction-aware speech retrieval.
3. AnyTalk: Speech Animation for Arbitrary Characters Leveraging a Video Generation Model
Authors: Kwan Yun, Serin Yoon, Sunjin Jung, Jung Eun Yoo, Inyup Lee et al.
Categories: cs.GR, cs.CV, cs.MM, cs.SD | accepted to TVCG, Project page at https://serin-yoon.github.io/projects/anytalk/ Score: 7.21/10 (Obj:9 Id:6 Ind:6 Comp:6 Eff:8 Nov:8 Rep:4)
- Strength: 在 5 个顶点 4,542–241,981 的异构角色上,AnyTalk 以零动画数据训练获得 LSE-D 11.304 / LSE-C 3.155(优于 ScanTalk 12.152/2.395),用户研究 vs ScanTalk 自然度 78.6%(p<0.001),蒸馏变体 AnyTalkRT 达 110 FPS 实时推理。
- Weakness: 论文声明可获取的代码在项目页为死链,GitHub 仅发布覆盖 CsF 单点的 AnyTalkCsF 仓库,blendshape 优化与实时蒸馏代码缺失,且 w/o CsF 消融在 LSE-C 上反而最优(3.451 vs 3.397),外观一致性提升口型精度的主张缺少独立量化证据。
- Full review: Claude Code 全文七公理审稿
We present AnyTalk, a novel method for generating 3D speech animations for arbitrary characters without requiring any animation data. While existing audio-driven 3D speech animation methods rely on character-specific training data or laborious rigging/re-meshing, AnyTalk circumvents these limitations by leveraging recent video diffusion models trained on extensive video datasets. We first adapt a pre-trained video diffusion model to a target character through our Character-specific Fine-tuning (\textit{CsF}) technique. By fine-tuning on rendered images of the 3D character paired with zeroed-out audio embeddings (representing “no motion”), we eliminate the need for animation data while preserving the motion prior of large-scale video diffusion model. We then uplift the resulting talking-head video into a 3D speech animation by estimating blendshape parameters through a proposed optimization process. AnyTalk enables lip-synced animations across diverse face meshes and blendshape configurations, significantly reducing manual effort and data requirements. We further enhance usability by distilling AnyTalk into a streamlined network, $\text{AnyTalk}_{RT}$, thereby enabling real-time performance. By leveraging talking-head video generation, our method broadens access to audio-driven speech animation technology for arbitrary characters. The code is publicly available at https://serin-yoon.github.io/projects/anytalk/.
4. Navigating Speech Enhancement for Real-Time MRI: A Systematic Assessment of Signal Quality, Source Preservation, and Downstream Tasks
Authors: Huang-Cheng Chou, Sean Foley, Haley Hsu, Kevin Huang, Szu-Jui Chen et al.
Categories: eess.AS, eess.SP | Submitted to the Journal of the Acoustical Society of America (JASA). 18 pages, 3 figures, 11 tabels Score: 7.21/10 (Obj:9 Id:7 Ind:8 Comp:4 Eff:8 Nov:8 Rep:5)
- Strength: 在 5 个 rtMRI 语料 × 3 个 ASR 的 15 组对比中,RE-USE 在 11 组降低 WER(均值 ΔWER −1.28 个百分点),Denoiser 在 13 组升高(+3.04),系统证明了增强效果具有端点依赖性。
- Weakness: 存档加性噪声探针的混合生成代码与增强模型间直接对比区间均缺失,RE-USE 优势停留于描述性证据,且分析代码仅按需提供、无法独立复现。
- Full review: Claude Code 全文七公理审稿
Audio recorded during real-time magnetic resonance imaging (rtMRI) is heavily contaminated by scanner noise, but it remains unclear whether general-purpose speech enhancement improves the signal for speech research and downstream processing. Three off-the-shelf systems—Denoiser, PASE, and RE-USE—are evaluated across five rtMRI corpora using naturally recorded inputs, a clean-input probe, and an archived paired additive-noise probe. The multi-task evaluation spans learned quality predictors, speaker and phone representations, reference-based intelligibility and quality measures, acoustic–phonetic probes, automatic speech recognition (ASR), and paralinguistic tasks. The central result is that enhancement effects are endpoint dependent: higher predicted-quality scores do not reliably imply better ASR performance or greater source fidelity. Across 15 corpus–recognizer comparisons using corpus-provided processed inputs, RE-USE yielded lower word-error-rate point estimates in 11, whereas Denoiser yielded higher estimates in 13. In the paired additive-noise probe, PASE and RE-USE improved recognized-phone agreement, intelligibility, and perceptual-quality point estimates. Denoiser improved recognized-phone agreement and short-time objective intelligibility (STOI) but reduced speaker-embedding similarity. No system was uniformly best across corpora, recognizers, and endpoints. Enhanced rtMRI audio should therefore be treated as a task-specific transformed derivative rather than a universally improved replacement for the original or DSP-processed waveform.
5. How Fragile Is Your Watermark? Training-Free Structural Removal of Neural Audio Watermarks
Authors: Likhith Kumara
Categories: cs.SD | Accepted at APSIPA ASC 2026 Score: 7.16/10 (Obj:9 Id:6 Ind:6 Comp:8 Eff:8 Nov:7 Rep:6)
- Strength: 用四个 pair-only 结构探针以 17 MFLOP/片段(较 HarmonicAttack 低约 1.4×10^4 倍)实现训练无关的匹配攻击,将 WavMark/SilentCipher/audiowmark 的 payload 在 PESQ≥3.84 下擦除至 5.9-48.4%、AudioSeal 检测标志透明移除,fragility 分数把十种方案分成脆弱(0.80-0.87)与鲁棒(0.27-0.39)两组。
- Weakness: 未与 HarmonicAttack 在同方案同数据上对比成功率,也缺少”匹配攻击优于盲扫电池”的定量对照;探针→攻击原型在已知十方案上指派并验证、未见方案上的预测自述仅部分可靠,且 fragility 仅为下界。
- Full review: Claude Code 全文七公理审稿
Neural audio watermarks are increasingly used to attribute and detect AI-generated speech, so their practical value rests on how cheaply an adversary can remove them. Robustness is usually measured by running a fixed battery of distortions blindly against every scheme. We instead make removal diagnostic: from a few clean/watermarked pairs we compute cheap structural probes that reveal where a watermark sits in the signal (its embedding domain), then apply a single domain-matched attack rather than a blind sweep. We further summarize each scheme with one threshold-free fragility score, the area under its accuracy-versus-quality trade-off, which an accuracy-only benchmark cannot provide. Across ten watermarking schemes the probes separate fragile from robust marks: for magnitude and carrier-domain watermarks a single matched attack erases the payload (WavMark, SilentCipher, audiowmark) or removes the detection flag (AudioSeal) at high objective quality (PESQ >= 3.6), whereas latent-domain marks (VoiceMark, WMCodec, AlignMark, AWARE) resist every training-free attack we apply. The same pair-only probe signatures also identify which watermarking scheme is present (84% over ten schemes).
6. Geometry-adaptive Ambisonic encoding for sparse microphone arrays of variable topology using physics-informed diffusion
Authors: Xiang Zhou, Zhengqiao Zhao, Zhengding Luo, Wen Zhang
Categories: eess.AS, cs.SD Score: 7.05/10 (Obj:9 Id:6 Ind:7 Comp:6 Eff:8 Nov:8 Rep:4)
- Strength: 模拟 FOA 单声源 SI-SDR 达 15.42 dB(优于最佳基线 2.99 dB)、LOCATA 真实录音 SOA 双声源 SI-SDR 5.62 dB 且 -3 dB 相干截止频率从基线 2133 Hz 提升至 3106 Hz,跨未见阵列拓扑保持 9.78 dB 平均 SI-SDR。
- Weakness: 全文未发布代码、权重或数据,且未与最相近的生成式基线 Flow-HOA 做实验对比,无法定位相对生成式 SOTA 的位置;FOA 单声源空间指标(Coh/Mag.Err/ILD.Err)仍被 Parametric 基线反超。
- Full review: Claude Code 全文七公理审稿
Ambisonics delivers compact scene based spatial audio representation, yet higher order Ambisonic encoding poses difficulties for wearables and embedded hardware. Their microphone arrays are often sparse, irregular, and constrained by device specific boundary conditions. These factors make the spherical-harmonic (SH) domain encoding ill conditioned: inverse filtering amplifies noise, while deterministic neural encoders may overfit to array-specific responses or smooth ambiguous higher-order components. This paper presents DiffM2A, a geometry-adaptive conditional diffusion framework for robust Ambisonic encoding from sparse MAs with variable topologies. Its Geometry-Adaptive Spherical Harmonic Projection (GASHP) front-end constructs boundary-aware SH steering functions and applies an energy-normalized modal projection, mapping array-dependent observations to a common modal representation without explicit pseudo-inverse computation. A dual-branch Elucidated Diffusion Model then estimates complex Ambisonic coefficients, conditioned on both the raw microphone spectra and GASHP features. Sound intensity and rotational equivariance losses further enhance inter-channel phase consistency and structured behavior across SH subspaces. Evaluations on both first- and second-order Ambisonic encoding tasks, using simulated room-acoustics and real-world LOCATA recordings, demonstrate that DiffM2A outperforms conventional and neural baseline methods on signal fidelity, spectral accuracy, spatial coherence, and binaural cue preservation. Additional experiments show that these gains are largely retained across unseen five-microphone layouts and under mismatched open-array and rigid-sphere boundary models.
7. Listen, Reason, and Segment: Aligning LALMs with Editorial Judgment for Media Chapterization
Authors: Tony Alex, Wish Suharitdamrong, Sara Atito, Armin Mustafa, Muhammad Awais et al.
Categories: cs.SD, cs.AI, cs.CL, eess.AS | 19 pages, 9 figures, 8 tables Score: 6.90/10 (Obj:9 Id:8 Ind:8 Comp:5 Eff:8 Nov:6 Rep:4)
- Strength: AudioChaps-R1-8B 在首个音频章节化基准 AudioChaps-Eval 上将宏 F1 从主干 AF3-Think-8B 的 28.6 提升到 77.8(+49.2,bootstrap p<1e-4),以 1/4 参数量超越 32B 的 Step-Audio-R1,并将全长度边界检测 F1 从 6.5 提升到 37.6、中位定位偏差从 38 秒降到 10 秒。
- Weakness: 代码、模型权重与三个数据集仅在论文接收后发布,GitHub 仓库当前为空,且 CoT 数据生成依赖闭源的 Step-Audio-R1 与 Gemini 2.5 Pro,逐字复现不可行。
- Full review: Claude Code 全文七公理审稿
Large Audio Language Models (LALMs) have made rapid progress on standardized benchmarks, yet their deployment in practical media workflows, curation, archival indexing, and content distribution remains largely unrealized. We identify automated audio chapterization, the task of segmenting continuous audio streams into thematically coherent chapters, as a demanding and commercially consequential setting that exposes this gap. Chapterization is challenging because boundaries are defined less by objective acoustic events than by subjective editorial judgment, requiring models to reason sequentially over long acoustic contexts and approximate creator-authored boundary decisions. We present AudioChaps, a post-training framework for aligning end-to-end LALMs for this task via Group Relative Policy Optimization (GRPO) guided by Chain-of-Thought (CoT) reasoning. To support training and evaluation, we curate three datasets: AudioChaps-Alignment, derived from creator-annotated chapter boundaries on YouTube; AudioChaps-CoT, which provides structured supervision for well-formatted, high-quality, and evidence-grounded boundary reasoning; and AudioChaps-Eval, a held-out benchmark for audio chapterization. Applying GRPO directly without a Supervised Fine-Tuning (SFT) cold start, AudioChaps-R1-Zero already improves average F1 by 33 points over the state-of-the-art LALM Audio-Flamingo-3-Think. The AudioChaps framework produces our final aligned LALM, AudioChaps-R1, which improves average F1 by 49 points. These results demonstrate that GRPO-trained LALMs can reliably transform unstructured auditory streams into navigable, structured media. Our code, models, and dataset resources will be released upon acceptance at https://github.com/ta012/AudioChaps.
8. Speaker-Normalized Semantic Speech Tokens via Iterative S2U-T2U Refinement
Authors: Hanlin Zhang, Daxin Tan, Dehua Tao, Chengxi Deng, Xiao Chen et al.
Categories: eess.AS Score: 6.84/10 (Obj:9 Id:6 Ind:7 Comp:6 Eff:7 Nov:8 Rep:4)
- Strength: 平行语句 token 一致性大幅提升(UED 344.61→59.44、SelfBLEU-4 27.04→94.17,均超过 R-Spin 等全部基线),VC 双语种 speaker similarity 达最优(EN 0.71、ZH 0.77),且 S2U-T2U 对齐 WER 相对下降 72.0-86.7%。
- Weakness: 未发布代码与模型权重且缺失最相近基线 DC-Spin 的对比,TTS 英文 WER 从 Iter.0 的 1.80 回退至 1.97,说话人探测精度改进仅 0.10%→0.09%(0.01pp),说话人信息减少的证据偏弱。
- Full review: Claude Code 全文七公理审稿
Semantic speech tokens should preserve linguistic content while suppressing speaker- and duration-dependent variation inherited from acoustic inputs. We propose Iterative Semantic Token Purification (ISTP), an alternating speech-to-unit (S2U) and text-to-unit (T2U) training procedure guided by text predictability. Starting from an initial S2U tokenizer, each iteration trains a T2U model on its deduplicated token sequences. The decoded T2U predictions then serve as connectionist temporal classification targets for a newly initialized S2U model, whose outputs supervise the next T2U model. This cycle progressively aligns the two token generators and biases the token space toward information recoverable from text. Experiments on Mandarin and English show substantially improved S2U–T2U agreement. Independently trained de-tokenizers further show that the refined S2U and T2U tokens retain sufficient content for high-intelligibility voice conversion and text-to-speech synthesis. In voice conversion, the generated speaking rate follows the reference more closely. The refined tokens also exhibit substantially improved cross-speaker consistency and reduced probe-recoverable speaker information.
9. ACE-Cap: Active Evidence Acquisition via Agentic Co-Evolution for Long-Paragraph Fine-Grained Audio Captioning
Authors: Fengji Ma, Yan Rong, Xu Li, Xuenan Xu, Chen Zhang et al.
Categories: cs.SD | 9 pages, 5 figures Score: 6.70/10 (Obj:8 Id:6 Ind:5 Comp:6 Eff:8 Nov:8 Rep:4)
- Strength: 7B 级三角色交互框架在 MMAU/MMAR/MMSU 上分别达 73.0/64.1/70.4(超最强开源基线 5.1/8.8/2.2pp)且 Omni-Cloze 64.4% 为全部评测系统最高,同时超过 GPT-4o Audio 并在三基准上与 Gemini 3.1 Pro 互有胜负。
- Weakness: 核心方法 LOOP-GRPO 缺少同一骨干上相对标准 GRPO 的受控消融,训练 reward judge 与评测协议共用 Qwen3.6-27B 构成循环依赖风险,且全文无代码/数据 release,复现闭环不完整。
- Full review: Claude Code 全文七公理审稿
Long-paragraph fine-grained audio captioning requires models to recover diverse acoustic facts while avoiding omissions and unsupported details. However, prevailing captioners remain passive one-shot generators: once a detail is overlooked, they cannot identify the evidence gap, query the audio for targeted information, or decide when sufficient evidence has been collected. We formulate this task as active evidence acquisition and introduce Agentic Co-Evolution for Captioning (ACE-Cap). The framework uses multi-turn interaction between a Composer and an Instruct model to form a closed evidence-acquisition loop. A Captioner first produces an initial description. Conditioned on this description and the interaction history, a text-only Composer asks targeted questions about unresolved acoustic attributes, while an audio-conditioned Instruct model provides grounded answers. The Composer then decides when to terminate and synthesizes the accumulated evidence into a final caption. ACE-Cap trains these roles through a unified gold-to-prediction reward derived from fixed, gold-grounded multiple-choice questions and a frozen caption-only judge. For credit assignment in variable-length interactions, LOOP-GRPO replaces the trajectory-wide scalar advantage with span-aligned signals: leave-one-out contributions of individual questions to the accumulated evidence, a quality-cost utility for stopping, and an evidence-preservation utility for final synthesis. Role-wise warm-up followed by alternating Composer and Instruct optimization keeps each update a well-defined single-policy problem while allowing the roles to co-evolve. ACE-Cap thus turns captioning from passive one-shot generation into an adaptive process that learns what evidence to acquire, when to stop, and how to preserve it in a long-paragraph caption.
10. Emotion Across Speech and Faces: Shared Affective Mechanisms in Multimodal Foundation Models
Authors: Xiutian Zhao, Luqi Sun, Björn Schuller, Berrak Sisman
Categories: cs.CL, eess.AS, eess.IV | 9 pages, 4 figures Score: 6.70/10 (Obj:8 Id:7 Ind:8 Comp:6 Eff:7 Nov:6 Rep:5)
- Strength: 首次在三个多模态基础模型上实现语音与面部情感神经元的双向因果转移,Gemma-4-12B-it 视觉 ESN 引导将 FER 自类情绪提升 +16.46 pp(Self-Cross Gap +21.01),跨 12 组设置随机掩码对照均保持近零效应。
- Weakness: 未发布代码/激活钩子实现与具体随机种子,跨模态转移效应量偏小(引导 +1.8~+6.9 pp)且 neutral 情绪干预模式较弱,方法对前作 ConAct 框架为纯增量复用。
- Full review: Claude Code 全文七公理审稿
Modern multimodal foundation models (MFMs) have made rapid progress on tasks requiring integrated perception across speech, vision, and language, including emotion recognition. However, it remains unclear whether they recognize speech and facial emotion through shared affective functional units or modality-specific pathways. We explore emotion-sensitive neurons (ESNs), sparse decoder neurons selectively associated with emotion categories, in three MFMs: Gemma-4-12B-it, MiniCPM-o-4.5, and Qwen2.5-Omni-7B. Using speech emotion recognition and facial expression recognition as complementary probes, we identify acoustic and visual ESNs. Visual ESNs are causally meaningful: deactivating them selectively impairs recognition of the associated facial emotion, whereas steering their activations selectively enhances recognition of that emotion relative to other emotion categories. Acoustic and visual ESNs further show emotion-matched overlap and similar layer-wise distributions, indicating partial structural alignment between affective representations across speech and faces. Finally, cross-modal interventions reveal bidirectional causal transfer: ESNs identified from one modality produce emotion-specific effects when applied to the other. Our findings provide one of the first cross-modality activation-level analyses of affective functional units in MFMs, suggesting that speech and facial emotion recognition partially converge onto sparse decoder-level components that can be localized and manipulated without training.
11. SingDance: Compositional Zero-Shot Singing-and-Dancing Video Generation with Role-Aware Audio Conditioning
Authors: Tao Feng, Xu Li, Xiangyang Luo, Ming Wen, Huadai Liu et al.
Categories: cs.SD, cs.CV | 9 pages, 5 figures Score: 6.53/10 (Obj:8 Id:6 Ind:7 Comp:5 Eff:7 Nov:7 Rep:5)
- Strength: SingDance 在从未训练的 held-out Song/Source 配置上以角色状态翻转实现组合式零样本边唱边舞,SingDance-50 上 BeatAlign 0.2723 逼近 GT 0.2795、MBCR 0.2558 优于 Wan-S2V 0.1983,且以 5.63B 参数(约 InfiniteTalk 的三分之一)达到 LSE-C 7.995 的竞争性口型同步。
- Weakness: 未对比最强音乐条件舞蹈基线 OmniDance 与 Wan-Dancer,且代码、权重与训练/测试数据均未发布,边唱边舞评测仅基于 50 例自建测试集,核心组合主张无法独立复算。
- Full review: Claude Code 全文七公理审稿
Generating personalized dance videos from a reference image, text prompt, and audio track requires music-conditioned body motion. Singing-and-dancing adds a second requirement: the visible subject must also articulate the vocals. Existing music-conditioned methods focus primarily on choreography, while speech-driven models generally assume that the visible subject produces the input voice, leaving this combined setting largely underexplored. We introduce SingDance, a unified video diffusion framework that formulates controllable vocal articulation as a semantic role: the visible subject is either the source, who produces the vocal signal, or the listener, who receives it from an off-screen performer. Hard-compact routing selects task-relevant speech, music, and role conditions, which are composed through frame-wise joint audio injection; source and listener retain the same speech pathway. Training uses asymmetric supervision: on-screen speaking and curated off-screen conversational-response videos establish role control, while instrumental and song-based dancing-only videos establish music-conditioned body motion. The target Song/Source configuration is never observed during training. At inference, assigning the source role to a song composes separately learned articulation and song-conditioned dance capabilities, enabling compositional zero-shot singing-and-dancing. Experiments demonstrate strong motion–beat alignment and visual fidelity, reliable paired switching of vocal articulation while preserving music-aligned body motion, and highly competitive lip synchronization with substantially fewer generation-time parameters than the strongest speech-driven baseline evaluated.
12. A Multiplication-Free Feature Extractor for Signal Classification: Keyword Spotting Case Study
Authors: Radu Dogaru, Ioana Dogaru
Categories: cs.SD, cs.HC, eess.AS | 5 pages, 3 figures, 2 tables, 1 algorithm, submitted to IEEE Signal Processing Letters Score: 6.26/10 (Obj:8 Id:5 Ind:5 Comp:9 Eff:6 Nov:5 Rep:8)
- Strength: iRDT 作为乘法自由的特征提取器,在 Google KWS 12 类上以 VRES-CNN 分类器达到 94.7% 验证精度,CPU 处理时间较 MFCC 至少快 10×,且仅需加/减/abs/移位即可计算,硬件复杂度约 480 kFlops vs MFCC 3 MFlops。
- Weakness: 缺少在嵌入式平台(FPGA/Cortex-M4)上的实测 LUT/功耗/端到端延迟数据,且未与 TC-ResNet/EdgeSpeechNet 等 SOTA 紧凑 KWS 基线做并列对比,94.7% 来自单次训练/验证划分而无方差报告,难以判断在等硬件预算下的实际竞争力。
- Full review: Claude Code 全文七公理审稿
A very low complexity feature extractor called next iRDT is proposed and evaluated for the problem of keyword spotting (KWS). Unlike any other types of feature extractors including the widely used MFCC, or adaptive, CNN-based ones, our algorithm is multiplier-free and it employs only simple, energy-efficient arithmetic operators. Since keyword-spotting of speech commands (KWS) is a typical application for TinyML platforms requiring low complexity for the signal classification chain, we consider it as a case study to evaluate complexity and functional performance. If properly tuned, iRDT demonstrates similar accuracy to solutions based on MFCC or CNN-based extractors using baseline classifiers on Google’s KWS 12-classes dataset. With a different classifier the system achieved 94.7% validation accuracy. Processing times on CPU for the proposed feature extractor, are at least one order of magnitude smaller than for the MFCC. The proposed algorithm has a very low hardware footprint, making it ideal for ultra-low power edge devices. Code and demo are available [18].
13. A Novel Binaural Cue Preservation Loss for DNN-Based Binaural Speech Enhancement
Authors: Jayteerth Amble, Thomas Haubner, Hendrik Schröter, Christoph Hoog Antink, Henning Puder
Categories: eess.AS, cs.SD | Accepted at the International Workshop on Acoustic Signal Enhancement (IWAENC) 2026 Score: 6.16/10 (Obj:8 Id:7 Ind:5 Comp:7 Eff:6 Nov:6 Rep:4)
- Strength: 联合 ILD-IPD 复平面损失 ℒBC 在双耳噪声抑制模型中取得低于基线 cue loss 与纯去噪模型的 ILD 误差,同时 SI-SDR/MBSTOI 保持接近无约束基线(图 2 曲线,正文无数字表格)。
- Weakness: 全文结果仅以图呈现、无数字表格与统计显著性,且自创 ℒBRE 指标同时用于训练与评测,缺乏主观空间定位听测、混响条件与公开代码。
- Full review: Claude Code 全文七公理审稿
Binaural speech enhancement for hearing aids aims to reduce noise while preserving the interaural cues needed for spatial localization. Although deep neural network-based methods achieve strong noise reduction, they often distort the rela- tionship between the left and right signals. In this paper, we propose two novel binaural cue preservation losses. First, a binaural reconstruction error loss that directly penalizes masking-induced distortion in the relationship between the left and right spectra, providing a more direct measure of the binaural consistency than conventional separate interaural level differences (ILD) and interaural phase differences (IPD) errors as in prior work. Second, a binaural cue loss that jointly models ILD and IPD to better preserve the binaural structure. Experimental results show that both proposed losses maintain strong noise reduction performance and reduce masking- induced distortion compared to the state-of-the-art baseline cue loss, while the second proposed joint binaural cue loss also outperforms the baseline in ILD preservation.
14. Audio-Visual Segmentation via Depth-Guided Collaborative Modeling
Authors: Zhaojin Fu, Yuyang Hong, Qi Yang, Zili Wang, Kun Ding et al.
Categories: cs.CV, cs.AI Score: 6.11/10 (Obj:8 Id:6 Ind:8 Comp:5 Eff:8 Nov:5 Rep:5)
- Strength: DGCM-AVS 在 AVSS 基准上相对 SOTA 取得 10.2% MJ / 8.7% MF 相对提升(R50 绝对 +6.3/+6.2),三基准、两骨干全消融验证了深度模态的效用。
- Weakness: 未发布代码与权重(GitHub 检索 0 仓库)、对比表未纳入引言中提及的近期强基线、参数增至 224.3M 且 S4 增益仅 0.5 MJ,深度先验”首次引入”的新颖性主张缺少系统论证。
- Full review: Claude Code 全文七公理审稿
Audio-Visual Segmentation (AVS) is a fundamental task in multimodal perception that performs pixel-level segmentation of sounding objects in videos by leveraging both visual and audio cues. It has broad applications in video understanding, human-computer interaction, and autonomous driving. However, most existing AVS methods do not explicitly model geometric cues such as relative distance and occlusion, thereby limiting the robustness of cross-modal alignment. In human perception, spatial structure is naturally integrated with audio-visual evidence to accurately localize sounding objects. Motivated by this, we incorporate estimated depth as a spatial structural cue for AVS and propose DGCM-AVS, a tri-modal framework that jointly models audio, visual, and depth information. Specifically, we design a Depth-Aware Dynamic Modulator to improve the separation of adjacent objects while preserving intra-object feature consistency. Furthermore, we propose Depth-Guided Progressive Fusion, which uses depth as an intermediate bridge to progressively align audio cues with visual features. Compared to state-of-the-art methods, DGCM-AVS achieves relative improvements of 10.2 percent in M_J and 8.7 percent in M_F on the AVSS dataset. We believe our study highlights depth as a promising yet underexplored modality for AVS and may encourage further research in this direction.
15. Cached LLM Probability Retrieval for Speech Recognition
Authors: Sheng Li, Takahiro Shinozaki, Tatsuya Kawahara
Categories: eess.AS | under review Score: 6.05/10 (Obj:8 Id:5 Ind:7 Comp:6 Eff:6 Nov:6 Rep:5)
- Strength: Cached LLM probability retrieval improves 1-pass ASR in 28 of 39 settings and cuts Whisper-small WER/CER by an average of 8.13% absolute (up to 13.01% WER on AMI IHM) with no parameter training and no online LLM calls un。
- Weakness: The paper releases no code, data, or cache, omits the most relevant lightweight baselines (classical n-gram LM and kNN-LM rescoring) from the main comparison, and derives the 28/39 win rate from best-of selection across。
- Full review: Claude Code 全文七公理审稿
Large language models (LLMs) enhance automatic speech recognition (ASR) by providing linguistic priors; however, their direct rescoring is costly because it requires evaluating every N-best hypothesis. This paper introduces “cached LLM probability retrieval,” which involves querying a local teacher LLM offline to obtain next-token probabilities for ASR-relevant context-target pairs. These probabilities are then utilized during recognition via cache lookups, backoff strategies, and optional scoring for significant misses. The method is training-free and can integrate with existing recognizers without requiring modifications to acoustic models. Evaluations across various ASR models reveal that cached retrieval outperforms 1-pass ASR in 28 of 39 settings and achieves lower non-oracle errors. Context length analysis indicates that benefits peak at a context length of 8, suggesting that cached probability retrieval is an effective and lightweight ASR adaptation method, in contrast to the heavy training required for Generative Error Correction (GER) or knowledge distillation (KD).
16. Sonifying I2S Transport Signals to Detect Transmission Faults
Authors: Stephen Roddy
Categories: eess.AS, cs.SD | 7 pages, 3 figures, 7 equations Score: 5.90/10 (Obj:8 Id:4 Ind:6 Comp:6 Eff:4 Nov:7 Rep:8)
- Strength: 论文公开完整代码与合成数据集(GitHub I2SSon),并以 300 个样本、重复测量 ANOVA 与聚类分析证明联合立体声表示带来 modest 但一致的 jitter 可分性提升(silhouette 0.090→0.127,jitter 每样本均值 0.619)。
- Weakness: 缺少监听者实验与任何外部基线对比,且分离结果主要由 jitter 单类驱动(clean/bit-slip/word-length 三类每样本 silhouette 约 −0.07 仍大幅重叠)。
- Full review: Claude Code 全文七公理审稿
This paper outlines a sonification design to support fault detection in the transmission of I2S transport signals. I2S is a protocol for communicating real-time digital audio between integrated circuits that, while in wide and general use, does not include built-in error detection. Moreover, given the nature of the protocol transmission faults affecting timing, framing and alignment can be difficult to identify using conventional visual methods. The proposed design addresses this with an approach informed by Audification, wherein oversampling controls temporal rescaling to render protocol structure (SCK and WS) and payload data (SD) across separate stereo channels. A preliminary computational feasibility study was carried out to measure feature-space separability of I2S faults in the generated auditory representations as opposed to listener performance. It evaluates the design across several payload types and error conditions including jitter, bit-slip, and word-length errors. Class separability was assessed through clustering analyses of extracted features. The evaluation results show that while oversampling produces systematic changes in feature values, it does not meaningfully improve separability between error classes. However, a modest but consistent improvement in separability is observed as a function of the joint representation of structural and payload information across channels. The findings suggest that feature-space separability in sonified communication protocol data may be dependent on the integration of complementary information streams, rather than on signal scaling alone.
17. A Data-Efficient Analytical Prior Machine Learning Framework for Sound Reduction Frequency Prediction in Helmholtz Resonators
Authors: Jiaming Li
Categories: cs.LG | 13 pages, 5 figures, 1 table Score: 5.84/10 (Obj:8 Id:7 Ind:6 Comp:5 Eff:7 Nov:5 Rep:2)
- Strength: 解析先验框架在 86 个 COMSOL 标注下把消声频率 MAE 从直接 SVR 的 3.375 Hz 与直接 MLP 的 1.109 Hz 分别压到 0.426 Hz 与 0.371 Hz(较直接学习降 87.4%/66.6%,R²=0.998),并在 20–70 例小数据预算下一致优于直接学习。
- Weakness: 数据与代码仅”按合理请求提供”、无公开仓库或权重,且高保真标签继承自解析模型的 COMSOL 验证集并缺乏任何实验或独立前瞻验证,使数据效率结论难以独立复现与推广。
- Full review: Claude Code 全文七公理审稿
High-fidelity finite-element simulations can provide accurate numerical predictions for side-branch resonators, but large simulation datasets are expensive to generate and purely data-driven surrogates may become unreliable when simulation-labelled data are scarce. This study develops an analytical-prior learning framework that reuses a low-cost analytical model to improve data efficiency under limited high-fidelity simulation budgets. Two complementary routes are considered. When the analytical model remains available at inference, it is retained as an explicit baseline and the simulation data are used to learn only the analytical-to-simulation discrepancy. When a self-contained predictor is required, the analytical mapping is first distilled from abundant low-cost evaluations into a learned prior and then calibrated with the limited simulation data. The framework is evaluated on rectangular side-branch Helmholtz resonators using 86 simulation-labelled geometries and 8,998 non-overlapping analytical-only geometries. The analytical model achieved a mean absolute error (MAE) of 1.333 Hz. Direct support vector regression (SVR) achieved 3.375 Hz, while residual SVR reduced the MAE to 0.426 Hz. A direct multilayer perceptron (MLP) achieved 1.109 Hz, whereas analytical-prior pretraining reduced the error to 0.556 Hz with frozen-prior residual adaptation and 0.371 Hz with full-model fine-tuning. Across training budgets of 20 to 70 simulation-labelled cases, both analytical correction and analytical-prior pretraining consistently improved data efficiency relative to direct learning. These results show that analytical prior information can substantially improve high-fidelity prediction when simulation data are scarce, with explicit correction and prior distillation serving complementary deployment needs.
18. Feedforward Active Speech Suppression Based on Time Series Prediction of Speech Signals Using Neural Networks
Authors: Manami Nishikata, Shoichi Koyama
Categories: eess.AS | Accepted to APSIPA Annual Summit and Conference 2026 Score: 5.74/10 (Obj:8 Id:5 Ind:6 Comp:5 Eff:6 Nov:6 Rep:4)
- Strength: SP-FxLMS 用过去+当前+未来梯度联合更新控制滤波器,在 NN 预测信号下平均 NPR 达 -42.12 dB,比标准 FxLMS 提升 5.03 dB。
- Weakness: 缺少 RLS 与现代深度 ANC(PFANC/Deep ANC/Mamba-masking)基线对比,且 MLP 预测器推理延迟 1.19×10⁻⁴ s 超过采样间隔 6.25×10⁻⁵ s,未满足实时处理。
- Full review: Claude Code 全文七公理审稿
A feedforward active noise control (ANC) method based on time-series prediction for speech signals is proposed. Although current ANC techniques are highly effective against stationary noise, suppressing highly non-stationary speech signals remains a challenging task. We propose an adaptive filtering algorithm for active speech suppression based on neural-network-based time-series prediction of future signals. The update value for the linear control filter is calculated based on the predicted signal, as well as the current and past signals. Numerical experiments indicated that the noise reduction can be improved in both cases: when using the true predicted signal and when using a signal predicted by neural networks.
19. Unadapted Multilingual ASR on a Garrusi Kurdish Evaluation Set: A Common-Reference Staged Normalization Analysis
Authors: Hiwa Asadpour
Categories: cs.CL, cs.SD | 12 pages A4, 4 tables, 2 figures, pilot study Score: 5.70/10 (Obj:8 Id:4 Ind:8 Comp:3 Eff:5 Nov:6 Rep:7)
- Strength: 提出的 common-reference staged normalization 把跨文字 WER 的 13.85 个百分点与 CER 的 49.72 个百分点变化分解到三个假设侧变换上,并首次为 Garrusi 给出 97.85% WER / 51.20% CER 的零样本测量。
- Weakness: 缺失把”正确假设”经同一转写+折叠管线打分的控制实验,导致 ~97.85% WER 中识别误差与打分管线伪迹(613+12,330 未覆盖字符)无法分离,且唯一外部基线 Southern Kurdish 模型未在相同 segmentation 上重测。
- Full review: Claude Code 全文七公理审稿
Evaluating speech recognition for a Kurdish variety written in a Latin field orthography, using a model that outputs Arabic script, creates a measurement problem before a modelling one: direct scoring treats writing-system differences as recognition errors. Jointly normalizing reference and hypothesis avoids this, but also changes reference tokenization, mixing agreement gains with a change in the scoring denominator. I evaluate MMS-1B-all with the Central Kurdish (ckb) adapter, used as released without adaptation, on 1,722 Garrusi questionnaire segments from five speakers (9,763 reference word tokens; 117.9 minutes). I use a common-reference design: the reference is folded once and fixed at 9,763 tokens, while only the hypothesis representation varies. The raw Arabic-script hypothesis scores 111.70% WER and 100.92% CER, with zero exact word matches. Latin transliteration gives 102.36% WER and 57.89% CER; folding it into the reference’s reduced orthography gives 97.85% and 51.20%. Thus RAW-to-FOLDED reduces measured WER by 13.85 points and CER by 49.72 points; folding alone accounts for 4.51 and 6.69 points. Substantial error remains: 14.53% of reference tokens are exact matches, edits are substitution-dominated, and per-segment WER is higher for shorter segments. A Southern Kurdish fine-tuned system (aranemini/southern-kurdish-asr), scored under the same design, performs worse on every speaker (1,703 segments), with 109.56% WER and 55.85% CER. However, 12,330 output characters fall outside the folding table, so these rates must be recomputed against the corrected fixed reference. The MMS output also contains 613 unconverted or unmapped characters, showing that part of the residual error reflects scoring-pipeline limits rather than recognition alone. I will release the fixed reference and segment-level results, subject to source-corpus sharing terms, to support independent checking.
20. Contrastive Learning with Variational Regularization for Multi-Session EEG-to-Speech Decoding
Authors: Tomoaki Mizuno, Toru Nakashika
Categories: eess.AS, cs.SD, eess.SP | Accepted to APSIPA ASC 2026 Score: 5.68/10 (Obj:8 Id:6 Ind:6 Comp:5 Eff:5 Nov:6 Rep:4)
- Strength: 利用跨会话重复 EEG 响应对做对比学习并配合变分正则,CER 从 0.968 显著降至 0.948(相对下降约 2.1%)且 PCC 从 0.262 恢复至 0.274(p=0.046),会话探测精度降至随机水平,会话不变性获得直接证据。
- Weakness: 全部效果仅在单受试者、作者自家基线上验证,CER 绝对值 0.948(约 95% 字符错误)远离实用,SPA 未达显著(p=0.087),且无外部方法横向对比与代码/数据集公开。
- Full review: Claude Code 全文七公理审稿
Reconstructing heard speech from non-invasive electroencephalography (EEG) is challenging due to a low signal-to-noise ratio (SNR) and inter-session variability. While trial averaging improves the SNR, it is difficult to apply to continuous speech. We instead use repeated EEG responses to the same stimulus across different sessions as positive pairs for contrastive learning, and introduce variational regularization that, combined with this contrastive objective, keeps the encoder representation space broad. Experiments on a Japanese EEG dataset show that combining the session-invariant strategy with variational regularization improves the character error rate (CER) while maintaining mel-spectrogram reconstruction fidelity. Session probing confirms that the encoder representations achieve session-invariance.
21. Automatic Transcription of Microtonal Free-Rhythm Vocal Music: A Case Study in Iranian Classical Music
Authors: Sepideh Shafiei, Shapour Hakam, Harsh Dange, Joel Rodriguez Caraballo
Categories: cs.SD Score: 5.32/10 (Obj:8 Id:2 Ind:2 Comp:8 Eff:2 Nov:5 Rep:5)
- Strength: 完整可解释的微音阶自由节奏人声转录工作流(多级音高直方图 + DTW 对齐 + tahrir 专用记谱符号),应用于 Karimi 语料 165 首(145 首 IRMA + 20 首 Tsuge),公开仓库 https://github.com/SepiSha/microtonal-music-autotranscriber 已建。
- Weakness: 全文零量化评测——没有与 Masoudieh ground truth 的音符级比对指标、无基线对比、无逐曲结果(§5 自述定量评测留待未来),且核心代码截至审查日(2026-08-23)尚未发布。
- Full review: Claude Code 全文七公理审稿
This paper introduces a computational workflow for automatically transcribing microtonal, free-rhythm vocal music, with Iranian classical music as a case study. Our approach is based on performances by the renowned vocalist Karimi and ground truth transcriptions by the prominent ethnomusicologist Masoudieh [14], which were subsequently incorporated into the IRMA Audio-MIDI dataset [20]. To accurately extract melodies, we employ pitch histograms in conjunction with Dynamic Time Warping (DTW). Additionally, we introduce specialized musical notations to capture the intricate ornamentations characteristic of the genre, with particular emphasis on the vocal technique tahrir. The transcription process is implemented in Python using the music21 library for symbolic music representation [5]. This study not only advances the field of computational ethnomusicology but also highlights the potential of computational methods in preserving and analyzing complex musical traditions. The transcription system also generates a combined visualization of the audio pitch contour and the DTW-aligned MIDI representation, enabling users to inspect the correspondence between the performance and the generated transcription. A companion visual editor supports expert-in-the-loop correction of the resulting notation.
22. Closing the Affective Loop: Multimodal Speaker-Listener Emotion-Dynamics-Aware Empathetic Social Robots
Authors: Zi Haur Pang, Casey Kennington, Tatsuya Kawahara
Categories: cs.HC, cs.CL, cs.RO | This paper has been accepted for presentation at APSIPA ASC 2026 Score: 4.89/10 (Obj:8 Id:5 Ind:5 Comp:4 Eff:5 Nov:5 Rep:2)
- Strength: AffectLoop 通过把说话人与听话人双情感流接入 GPT-4.1-nano 回复生成,在 5 人 within-subject 试点中把共情回复评分从 4.80 提升到 5.35、valence 正恢复率从 72.2% 提升到 100.0%。
- Weakness: n=5 且无任何显著性检验或效应量,过程指标(对齐、痛苦恢复)由驱动生成的同一条 NRC-VAD/EmoNet 管线自评,且代码、数据、权重均未开放、GitHub 无仓库,系统不可复现。
- Full review: Claude Code 全文七公理审稿
Empathetic social robots should respond not only to what users say, but also to how their emotions dynamically evolve during interaction. However, existing empathetic dialogue systems are often text-centered and primarily model empathy as a one-way mapping from the user’s emotion to the system response, limiting their ability to capture embodied speaker–listener affective exchange. We present AffectLoop, a multimodal speaker-listener emotion-dynamics-aware spoken dialogue system implemented on the Misty II robot. The system tracks the speaker’s verbal and facial affective dynamics, estimates the robot listener’s own verbal and behavioral affective state, and conditions LLM-based response generation on both affective streams. The robot then generates a short spoken empathetic response together with emotionally congruent embodied behavior, forming a closed speaker–listener affective loop. We evaluate the system in a pilot within-subject study with five participants, comparing it with an otherwise identical utterance-conditioned baseline that omits the speaker- and listener-affective-state inputs. The proposed system received higher overall impression ratings, especially for empathetic response and user satisfaction. Post-hoc log analysis further showed higher speaker-listener affective alignment and stronger valence-based distress recovery. These preliminary results suggest that explicitly modeling both speaker emotional dynamics and listener affective state can improve embodied empathetic interaction.