每日 arXiv

音频 · 语音 · 音乐 · 声学

每日新爬取的 arXiv 论文,按评分排序。中文摘要帮你快速判断是否值得深读。

2026-08-26
日期2026-08-26
已评分
均分
最高

Daily Papers — 2026-08-26

17 papers on audio, speech, music, and acoustics.

1. VoiceMem: Streaming Dual-Brain Memory for Real-Time Interaction

Authors: Zhifei Xie, Jiaqi Lang, Ze An, Yifan Zhao, Dongchao Yang et al.

Categories: eess.AS, cs.AI, cs.IR, cs.MM, cs.SD | 18 pages, 9 figures, 6 tables Score: 8.10/10 (Obj:1 Id:2 Ind:1 Comp:1 Eff:2 Nov:2 Rep:1)

  • Strength: 流式双脑记忆系统在 LoCoMo 上以 430 记忆 token 达 91.2 分(134ms 检索、K=5),超自身后端 Mem0 达 24.12 分,并打通 GitHub/HF 权重/数据集/pip 包全链路开源。
  • Weakness: 主要评测依赖 GPT-4o-mini 的 LLM-judge 且响应模型多为同源,ChatMem-Bench 与训练数据出自同一蒸馏管线形成部分循环评测,基准明细与人工标注被推迟到后续报告。
  • Full review: Claude Code 全文七公理审稿

Conversational systems, such as duplex speech language models (SLMs), still lack a streaming, accurate, and empathetic memory system as their soul. We introduce VoiceMem, a simple memory architecture with a parallel informational left brain, an emotional right brain, and streaming memory I/O mechanisms. We further build a complete pipeline for memory-aware SLM training, long-horizon evaluation, and decoupled deployment with interchangeable memory backends. Experiments and real-world deployment show three advantages: i) Accuracy: under top-5 retrieval, the left brain outperforms classical systems such as Mem0 at top-200 by nearly 30 points; ii) Emotional & Personal: the right brain, with short- and long-horizon affective attribution and dual-node persona modeling, achieves state-of-the-art performance across three persona benchmarks and improves the aggregate score by 4.29 points over the previous best system; and iii) Real-Time & Cheap: VoiceMem completes retrieval in 134 ms, well within standard VAD latency, adding no extra conversational delay while maintaining high accuracy and low cost. These results show that VoiceMem provides a practical memory foundation for real-time, personalized, and emotionally aware speech interaction.


2. Lost but not erased: Finding traces of a forgotten language in neural speech models

Authors: Peter Plantinga, Charlotte Moore, Peter W. Donhauser, Krista Byers-Heinlein, Denise Klein

Categories: cs.CL, cs.LG Score: 7.47/10 (Obj:9 Id:8 Ind:8 Comp:8 Eff:8 Nov:6 Rep:6)

  • Strength: 用 Conformer ASR 在 CommonVoice 法/英/德上做突然语言切换,行为上 Ma vs Mpost 收敛到 93.6% vs 93.7%(<0.1% 差),但 pre-phonemic 层 RSA 仍高于跨语言基线 8.3%(CI 2.9-13.7%),且 Ma 比 Mpost/Mctrl 重学 Lpre 快 14.3%/12.8% 达 70% 阈值,spliced-layer 实验定位 savings 于 pre-。
  • Weakness: 框架继承自 Achille 2019/Constantinescu 2025/Moore 2025,方法学增量限于把视觉域 critical-period 框架搬到语音 ASR 并细化三层定位;GitHub 仓库 pplantinga/bilingualnetworks 在公开索引中 0 命中,未声明权重 release;English→French Ma 的 traces 随 pre-training 单调上升(异于其他三对的非线性峰。
  • Full review: Claude Code 全文七公理审稿

International adoptees retain phonological traces of a birth language they can no longer speak or comprehend, a persistence typically attributed to a biologically-timed critical period. We asked whether it could instead reflect the ordinary dynamics of learning, using automatic speech recognition models that simulate the international adoptee experience without maturational confounds. Models were trained on one language and then abruptly switched to a second. We found that traces of the first language persisted throughout second-language training, but mainly in the lowest, pre-phonemic layers. These traces were functional, as models with early exposure re-learned their lost first language 14% faster than naive models; this advantage held even against models adopted early from a related language and disappeared when the earliest layers were substituted from a non-adopted model. We argue that these critical-period effects reflect entrenchment of foundational representations rather than a maturational loss of plasticity, and that experience plays a central role in critical periods in language acquisition.


3. Knowledge Distillation for Efficient Acoustic Echo Control

Authors: Ernst Seidel, Pejman Mowlaee, Tim Fingscheidt

Categories: eess.AS | 5 pages, accepted to EUSIPCO 2026 Score: 7.42/10 (Obj:8 Id:7 Ind:7 Comp:7 Eff:8 Nov:7 Rep:8)

  • Strength: 首次将知识蒸馏用于 AEC,学生 CGGN16-S(0.2M 参数、0.25G FLOPS,teacher 的约 2% 计算量)经两阶段 KD 后达到六倍复杂度模型(1.87G FLOPS)的整体性能,PESQ 2.07 vs 2.04、LPS 0.74,双讲下 ERLEBB 从 10.68 提升到 11.19,主观 MOS 优于 DLAC-Kalman。
  • Weakness: 新颖性限于 KD 技术的任务迁移与 loss 配方消融,缺少理论解释;所有消融与选择基于单一 dev 集、α=0.5 固定,主观听测未报告置信区间与显著性,且全文无 GitHub/数据链接,实验仅在仿真数据上验证。
  • Full review: Claude Code 全文七公理审稿

In recent years, many efforts have been made to supersede classical acoustic echo control (AEC) algorithms with more powerful machine-learned approaches. While surpassing the performance of well-established adaptive filters is very much possible, a remaining challenge is computational complexity. Popular architectures, such as convolutional recurrent networks (CRNs), are by multiple orders of magnitude computationally more expensive than classical signal processing solutions. Scaling down such models is usually straight-forward, but it comes at the cost of a notably reduced performance. We show - to the author’s knowledge for the first time in AEC - how these performance drops can be successfully alleviated to a large degree by employing an effective knowledge distillation (KD) process, enabling more potent efficient AEC. Our proposed CGGN16 student AEC models show significantly less near-end speech distortion at only 2% of its teacher’s computational complexity, surpass the overall performance of a six times more complex model trained on ground-truth labels, and outperform other AEC-focused architectures from recent literature.


4. AllMusicCaps: Album Reviews as Complementary Supervision for Music CLAP

Authors: Pablo Alonso-Jiménez, Xavier Lizarraga-Seijas, Xavier Serra, Dmitry Bogdanov

Categories: cs.SD | Accepted at ISMIR 2026. Code, weights, and data available at https://github.com/MTG/allmusiccaps Score: 7.05/10 (Obj:8 Id:8 Ind:6 Comp:6 Eff:7 Nov:6 Rep:9)

  • Strength: 从 AllMusic 专家评论蒸馏出 245,346 轨道级 caption 数据集,与既有语料组合后在人类书写基准 Song Describer 上 MRR 15.1→18.8(+3.7),且增益集中于叙事/评价/场景类复杂查询(复杂度≥3 档),盲区经注册诊断量化实证。
  • Weakness: 评论数据单独训练全面落后四语料基线(Song Describer 15.1→14.7),增益依赖组合且集中在单一基准;乐器/制作细节评测(MTG-Jamendo Instrument)无一超越外部基线,复杂度标注与数据生成共用同一 LLM 构成部分评测循环。
  • Full review: Claude Code 全文七公理审稿

Recent open text-audio contrastive models (CLAPs) are typically trained with LLM-generated captions derived from tag datasets or web search results, which tend to be accurate but expressively narrow. As a complementary source, we explore human-written album reviews, specifically expert reviews from AllMusic: they exist at scale and carry narrative cues, evaluative adjectives, and scene framing that other sources lack. Since raw reviews are too noisy for direct use as captions, we first build a caption corpus with 24,5346 samples via an LLM preprocessing pipeline that identifies descriptive musical quotes and rewrites them into training-ready captions. We find that album review supervision yields the largest retrieval gains on a human-written caption benchmark (Song Describer), particularly for complex queries that other existing caption datasets leave uncovered. In addition, we revisit the training recipe and show that SigReg regularization, which encourages an isotropic Gaussian distribution in the embedding space, improves MLP probing across classification tasks, as well as text-to-music retrieval. The resulting model outperforms open CLAP-style baselines on text-to-music retrieval, zero-shot classification, and most MLP probing tasks. We release the review-derived caption dataset and model weights to support future research.


5. Formal, Executable and Explainable Runtime Monitoring of Spoken Air Traffic Control Operational Procedures

Authors: Roberto Luvini, Giacomo Longo, Alessandro Armando, Enrico Russo

Categories: cs.AI, cs.CL, eess.AS Score: 6.70/10 (Obj:1 Id:2 Ind:1 Comp:1 Eff:2 Nov:2 Rep:1)

  • Strength: 首个将 ICAO 口语管制义务编码为五类时序公式并端到端落地到真实流量监控的系统,F1=0.85(盲评 12 条违规),1,495 个合成情况全中,并在 Überlingen/Comair 两起事故中比撞击提前 19-31 秒触发对应程序违规。
  • Weakness: 真实评测仅三小时 12 条人工标注、无端到端基线定量对比、无代码/数据/标注发布,且 37.6% 的例行 schema 覆盖率意味着多数程序域未被真实数据验证。
  • Full review: Claude Code 全文七公理审稿

Air traffic control procedures are executed through spoken exchanges between controllers and pilots. These interactions are essential to the safety of air transportation: failures in their execution can create severe operational hazards, as evidenced by past fatal accidents. Assessing whether an instruction has been followed requires relating what was said to the aircraft concerned, its state, and the obligations that pilots must meet. We present a runtime verification framework that monitors such procedures by checking controller-pilot exchanges, surveillance data, and onboard observations. The framework parses radio communications into events linked to the entities they concern and merges them with surveillance and onboard observations into a time-stamped trace. The ICAO-derived obligations as formalized as temporal formulas with explicit time bounds and evaluated over execution traces. Every violation is reported along with the breached obligations and the observations that support the verdict. With real traffic, the complete pipeline reaches an F1 of 0.85 against blind human-annotated violations; in 1,495 synthetic situations derived from two public corpora, the monitor logic returns the expected verdict in every case. In two historical accidents reconstructed from official investigation reports, the monitor identifies the same procedural deviations documented by the investigators.


6. PAGS: Autofocusing Photoacoustic Tomography via Speed-of-Sound-Adaptive Gaussian Splatting

Authors: Jiarui Ge, Jintao Ma, Bangxu Fan, Jinyan Zhang, Xiaokang Yang et al.

Categories: cs.CV, physics.med-ph | 13 pages, 6 figures Score: 6.63/10 (Obj:8 Id:6 Ind:6 Comp:8 Eff:7 Nov:6 Rep:6)

  • Strength: 把 PACT 盲自动聚焦重述为 SH 探针参数化的各向异性路径平均声速场学习,解析高斯声学投影使每迭代耗时降约 2.8×(6.7s→2.4s)、参数量压缩 12.5K×,in silico 上相对 vanilla SlingBAG 提升 SSIM 0.331→0.537、PSNR +1.3dB。
  • Weakness: 定量指标全部计算在均匀 SoS UBP 参考上而非真实初始压力场,且未与迭代 TV/时间反演/学习法重建等实用基线对比;物理幻影与稀疏采样实验仅有定性展示,Dual-SoS oracle 基线领先 UBP 5.0dB 而 PAGS 仅领先 oracle 2.1dB,效用论证的量化强度不足。
  • Full review: Claude Code 全文七公理审稿

Photoacoustic computed tomography (PACT) combines optical absorption contrast with acoustic detection for high-resolution deep-tissue imaging. A persistent challenge is that unknown speed-of-sound (SoS) heterogeneity changes acoustic time-of-flight, causing defocusing artifacts when reconstruction assumes a uniform SoS. Existing SoS-adaptive methods either rely on calibrated acoustic priors or optimize dense physical medium models, which becomes expensive and difficult to scale in 3D. We propose PAGS, a differentiable framework for blind autofocusing PACT via speed-of-sound-adaptive Gaussian splatting. PAGS represents the initial pressure field with sparse Gaussian photoacoustic (PA) sources and replaces explicit medium recovery with a compact anisotropic path-averaged SoS (ASoS) field parameterized by spherical harmonic probes. This latent propagation field directly controls source-to-transducer arrival-time alignment, while an analytic Gaussian acoustic projection maps the source representation to transducer signals efficiently. The resulting closed-loop signal-domain optimization jointly updates the Gaussian PA source parameters and the ASoS field from measured data, without calibrated SoS priors. Experiments on simulated and physical phantom data demonstrate improved reconstruction sharpness under heterogeneous acoustic media, robustness to sparse-view sampling, and computational benefits from the analytic Gaussian projection.


7. Dissonance Spectrum explicitly models perceptual frequency interactions for better music understanding

Authors: Tianle Wang, Xinyi Tong, Liangke Zhao, Jishang Chen, Sirui Zhang et al.

Categories: cs.SD, cs.AI Score: 6.63/10 (Obj:8 Id:8 Ind:8 Comp:7 Eff:6 Nov:5 Rep:6)

  • Strength: 6 个配对种子下 DS 在 MusicQA(BERTScore-R .9024 vs 基线 .8952)与 Music2Emo(R̄²VA .6589 vs 基线 .6473)所有端点上取得最高均值,且受控乐理验证音程 ρ=.951、调式 ρ=.893。
  • Weakness: 相对架构匹配的 CQT 分支增益极小(+.0028/+.0047)且符号检验 Holm p=.09375 不显著,同时全篇无听感实验,代码仓库无公开链接、GitHub 检索零命中。
  • Full review: Claude Code 全文七公理审稿

Conventional music representations describe acoustic energy over time and frequency but do not explicitly expose relations among simultaneous frequency components. We introduce the \emph{Dissonance Spectrum} (DS), a nonnegative time–frequency representation that applies a tolerance-based rational pitch-relation kernel with logarithmic harmonic distance to a constant-Q spectrum and attributes aggregate pairwise interactions back to individual frequency bins. Controlled music-theory tests show strong ordinal agreement for intervals, harmonic-function connections, and church modes, and weaker but significant agreement across diverse chord voicings. DS is then encoded by a lightweight parallel branch whose zero-initialized residual projection preserves the baseline function at initialization. Across six paired training seeds in open-ended music question answering and categorical and dimensional music emotion recognition, DS obtains the highest mean on every reported endpoint relative to the unchanged baseline, a parameter-matched Gaussian-input branch, and an architecture-matched magnitude-CQT branch. These results support DS as an interpretable, complementary representation, while listener-specific perception and broader task coverage remain open problems.


8. Generative vs. Encoder Large Language Models for ASR Evaluation: A Comparative Study

Authors: Thibault Bañeras-Roux, Shashi Kumar, Driss Khalil, Sergio Burdisso, Petr Motlicek et al.

Categories: cs.CL Score: 6.58/10 (Obj:8 Id:7 Ind:8 Comp:6 Eff:6 Nov:5 Rep:8)

  • Strength: 在 HATS 上系统性对比 24 个编码器/解码器 LLM,最佳配置 SemDist 与人类判断一致率达 90%(编码器 Sentence-CamemBERT-large)对 89%(Qwen3-Embedding-8B),均大幅超过 WER 63%;生成式 judge 成对选择最高 94%(GPT-4.1)、开源 Qwen3.5-35B 达 92%。
  • Weakness: 全部实证基于法语单一基准 HATS,跨语种外推未验证;定性分类与 SemDist 相关最高仅 -0.66(作者自认不足以取代连续指标);LLM-as-judge 的偏好/偏差既有文献未援引,且全文未提供代码链接,最优层配置仍需用户自行网格搜索。
  • Full review: Claude Code 全文七公理审稿

Automatic Speech Recognition (ASR) is typically evaluated using Word Error Rate (WER), which poorly reflects semantic similarity. While embedding-based metrics correlate better with human judgments, the respective roles of encoder and decoder-based Large Language Models (LLMs) remain underexplored. This paper presents a comparative study of both families for ASR evaluation. We analyze BERTScore and SemDist across different LLMs, layers, and pooling strategies, showing that both metrics can achieve strong correlation with human judgments when properly configured. For decoder models, we investigate generative LLMs in two settings: pairwise hypothesis selection via prompting and direct qualitative error classification. Our results show that encoder-based metrics remain highly competitive, while generative LLMs perform strongly in hypothesis comparison and improve the interpretability of ASR evaluation.


9. InteractGesture: Progressive Chunk Guidance for Continuous Streaming Co-Speech Gesture Control

Authors: Ekkasit Pinyoanuntapong, Ajinkya Deogade, Paul Streli, Wenjing Zhang, Joanna Materzynska et al.

Categories: cs.CV | ECCV 2026 Workshop - Interactive Social Avatars Score: 6.50/10 (Obj:1 Id:2 Ind:1 Comp:1 Eff:2 Nov:2 Rep:1)

  • Strength: 纯推理期空间控制把平均关节误差压到 6.34cm(vs Sequential 11.70cm、IK 19.97cm),且流式 Progressive 的 FGD 0.431 优于离线 Synchronous 的 0.442,模型完全冻结。
  • Weakness: 代码仓库提交当日即为空(0 文件)、无权重与种子报告,且核心机制是 OmniControl 式训练无关引导的既有路线的流式化组合,”实时”卖点缺任何延迟测量。
  • Full review: Claude Code 全文七公理审稿

Co-speech gesture generation has made significant progress toward realistic full-body motion from speaker audio, yet existing models lack fine-grained spatial controllability of individual joints. To address this, we introduce \emph{InteractGesture}, a model-agnostic, inference-time method for spatially controllable gesture generation. \emph{InteractGesture} guides target latent estimates of a diffusion sampler through a differentiable RVQ-VAE decoder, backpropagating spatial control gradients to adjust motion latents during sampling. A primary challenge in streaming co-speech generation is chunk-wise dependency: standard sequential inference freezes prior chunks, preventing spatial constraints in future chunks from adjusting preceding trajectories and causing boundary inconsistencies. To overcome this limitation, we propose \emph{Progressive Chunk Guidance}, a chunk-window strategy that maintains an active set of editable chunk latents with staggered delays, enabling spatial constraints to propagate gradients backward across chunk boundaries during streaming generation. Experiments on the BEAT2 dataset show that \emph{InteractGesture} improves multi-joint spatial control while preserving overall gesture quality. Furthermore, our approach supports diverse applications, including sparse joint positioning, dense joint trajectory control, and directional pointing. Our project page is available at https://exitudio.github.io/interactgesture-page .


10. CSAVocoder: A Causal Spatial Audio Vocoder Towards Real-Time Spatial Audio Generation

Authors: Zhiyuan Zhu, Han Wang, Wenxiang Guo, Yu Zhang, Changhao Pan et al.

Categories: eess.AS Score: 6.30/10 (Obj:8 Id:7 Ind:7 Comp:5 Eff:7 Nov:5 Rep:5)

  • Strength: 在约 1500 小时双耳/FOA 数据上,CSAVocoder 的空间一致性(ANG/DIS COS 62.11/77.05,MUSHRA-P 88.0 对次优 72.0)全面超越 8 个声码器基线及 DSP/BinauralGrad 空间化器,且 RTF 0.1587 满足实时流式。
  • Weakness: 因果基线与空间基线对比不公平(逐通道推理 + 缺标准化因果套件)、音频质量弱于 Vocos/WaveFM 的代价归因于因果约束但因果/非因果变体对照并不支持,且代码与权重未发布。
  • Full review: Claude Code 全文七公理审稿

Spatial audio vocoders are able to convert mel-spectrograms produced by generative models into spatial audio waveforms. Most neural vocoders are designed for monaural audio, and direct extensions to spatial audio can degrade spatial quality by ignoring inter-channel cues. We present CSAVocoder, a causal GAN-based spatial audio vocoder that jointly optimizes waveform fidelity and spatial rendering. Our framework introduces a Spatial Adaptor that fuses multi-channel mel-spectrograms with dynamic source-listener pose information, together with a spatial consistency discriminator that supervises inter-channel cues. To meet real-time requirements, we design a strictly causal, stateful generator that supports efficient streaming inference with constant memory overhead. Experiments on large-scale spatial audio datasets show that CSAVocoder improves spatial fidelity at competitive audio quality and real-time performance.


11. Leveraging Speech Acts for Low-Data and Cross-Domain Conversation Derailment Forecasting

Authors: Angela Yifei Yuan, Christine De Kock, Christopher Leckie

Categories: cs.CL Score: 6.20/10 (Obj:1 Id:2 Ind:1 Comp:1 Eff:2 Nov:2 Rep:1)

  • Strength: 在 3 数据集×5 档数据规模×双向跨域的实验矩阵中,Hparallel 以言语行为辅助任务在全部域内 AULC 上取得第一(如 GITHUB AUPRC 66.84 vs PLM 57.16),300 样本即达到纯文本模型 500-1000 样本的水平。
  • Weakness: 全量数据下 SA 优势消失(PLM 反超)、GITHUB→CMV 转移近乎失效(MacroF1 峰值 36.19),且核心 SA 标签与匿名代码均不可获取,数值复现存疑。
  • Full review: Claude Code 全文七公理审稿

Conversational derailment forecasting aims to predict when online discussions will escalate into hostility, enabling proactive moderation. Existing approaches often struggle in low-data settings and to generalize across domains. This poses a challenge for new platforms and smaller communities where annotated data is limited. We propose modeling pragmatic representations of conversations to reduce lexical noise and improve generalizability. Specifically, speech act information is used as an auxiliary learning signal alongside textual semantics. Experimental results show improved performance across three datasets, particularly in low-data and cross-domain settings.


12. Mandarin Humorous Homophone Recognition and Disambiguation in Automatic Speech Recognition

Authors: Sicheng Jin, Jinghao Chen, Mostafa Shahin, Beena Ahmed, Aditya Joshi

Categories: eess.AS | Accepted for publication at ISCSLP 2026 Score: 5.89/10 (Obj:8 Id:6 Ind:6 Comp:7 Eff:6 Nov:6 Rep:2)

  • Strength: 首次把普通话刻意谐音修辞(HumourPhone)作为 ASR 评测对象,提出的适配器在弱基线 Whisper-v3 上把低频同音词 T-CER 从 24.74% 降到 20.39%、近音 HumourPhone T-CER 相对降 16.5%,并以 target-aware 指标(T-CER/MUR/TSS)限定标注 span 完成评估。
  • Weakness: 增益仅对弱基线与近音组兑现,强基线 Qwen3-ASR 普通同音词 T-CER 反升(2.19%→3.42%)、全同音组几乎无改进;HumourPhone 评测集仅 80 条 TTS 合成音频,且代码、数据、权重均未发布,核心实验依赖闭源 GPT-5.5 路由无法复现。
  • Full review: Claude Code 全文七公理审稿

Automatic mispronunciation detection and diagnosis (MDD) plays a crucial role in L2 Mandarin pronunciation learning. While end-to-end (E2E) based MDD methods have substantially improved phoneme-level detection accuracy, diagnostic feedback remains limited, as segmental and tonal errors are not explicitly separated. In this paper, we propose a phonological feature-based MDD framework that models both segmental and tonal attributes within a unified Wav2Vec2-CTC architecture. Experimental results show that the proposed method reduces the False Acceptance Rate (FAR) by 10.1% and the Diagnostic Error Rate (DER) by 23.6% compared with the phoneme-only baseline system. By decomposing phonemes into low-level phonological components, the proposed approach enables more detailed and interpretable diagnostic feedback for L2 learners.


13. A Training-Free Proactive Defense Against Partial Speech Manipulation via Self-Embedding Steganography

Authors: Yigitcan Özer, Zhe Zhang, Wanying Ge, Xin Wang, Junichi Yamagishi

Categories: cs.SD, cs.AI | 6 pages; 4 figures; 1 tables; accepted at Interspeech 2026 Score: 5.63/10 (Obj:7 Id:5 Ind:6 Comp:7 Eff:6 Nov:5 Rep:4)

  • Strength: 免训练主动防御框架在词级换词攻击下检测 EER 达 8.9-10.0%(单词)/4.4-5.1%(双词),远优于近随机的被动基线,且经 15 次重复嵌入实现 100% 位精确自载荷恢复。
  • Weakness: 承诺的「恢复」无任何量化指标(无 PESQ/词错误率),未测 LSB 对有损压缩的鲁棒性,未对比同类主动防御 AudioSeal,且无代码发布。
  • Full review: Claude Code 全文七公理审稿

Partial deepfake speech, where only limited segments of an utterance are synthesized or manipulated, poses a significant challenge to existing deepfake detection systems. As the proportion of spoofed regions decreases, passive detectors become increasingly unreliable, and accurate detection and restoration remain challenging. In this paper, we revisit audio steganography from a new perspective and propose its use as a proactive defense against partially deepfaked audio. In particular, we consider a self-embedding strategy in which a clean speech signal embeds a compressed representation of itself, enabling post-hoc extraction of reference content. We demonstrate how existing audio steganography methods can be repurposed to support detection of partial deepfakes through codec-based restoration. Experiments on a benchmark dataset show that the proposed approach complements passive defenses. Remarkably, the proposed method operates without any training, providing a robust and data-efficient alternative for partial deepfake detection.


14. Fine-Tuning Whisper for Automatic Speech Recognition in Baniwa: A Preliminary Study

Authors: Leonardo Duart, Tiago Fonseca, Thiago Chacón

Categories: cs.CL, stat.ML | 12 pages, 3 tables. Preliminary study Score: 5.60/10 (Obj:1 Id:2 Ind:1 Comp:1 Eff:2 Nov:2 Rep:1)

  • Strength: 0.54 小时(1,373 条孤立词/短句录音)微调 Whisper Small 取得 37.5% WER、7.45% CER,为巴尼瓦语建立首个 ASR 基线。
  • Weakness: 无零样本基线或对比方法、无错误分解,语料/代码/权重均不公开,结果既无边际增益证据也无法复现。
  • Full review: Claude Code 全文七公理审稿

Automatic Speech Recognition (ASR) technologies have achieved remarkable performance in recent years through the use of large multilingual foundation models. However, most advances remain concentrated on high-resource languages, while indigenous languages continue to suffer from a lack of speech resources and language technologies. This work presents a preliminary study on the adaptation of Whisper for Automatic Speech Recognition in Baniwa, an indigenous Arawakan language spoken in Brazil, Colombia, and Venezuela. The experiments were conducted using a corpus of 1,373 manually transcribed recordings obtained from a linguistic documentation project. The corpus contains approximately 0.54 hours of speech and consists primarily of isolated words and short elicited utterances. The Whisper Small model was fine-tuned using supervised learning and evaluated using Word Error Rate (WER) and Character Error Rate (CER). The best model achieved a WER of 37.5% and a CER of 7.45%, demonstrating that multilingual foundation models can be successfully adapted to extremely low-resource indigenous languages. The results establish an initial baseline for Baniwa Automatic Speech Recognition and provide a foundation for future research involving larger datasets, language-specific adaptation strategies, and post-processing techniques.


15. Cooperative Multi-Agent Reinforcement Learning for Adaptive Aggregation in Semi-Supervised Federated Learning with non-IID Data

Authors: Rene Glitza, Luca Becker, Rainer Martin

Categories: cs.LG, cs.DC, cs.SD, eess.AS, eess.SP | 5 pages, 4 figures, ICASSP 2026 - 2026 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) Score: 5.58/10 (Obj:8 Id:6 Ind:6 Comp:4 Eff:5 Nov:4 Rep:8)

  • Strength: 在 DCASE 2020 Task 2 的 14 客户端半监督音频 AST 联邦任务上,pFedMARL 本地 F1 全面超过 Ditto(LS 0.87 vs 0.69)并显著压制 FedAvg,把对抗客户端聚合权重从 ~0.5 压至 ~0.1,代码仓库(NexuFed/pFedMARL,78 个 Python 文件)可复跑全部三个非 IID 场景。
  • Weakness: 纯 local 训练在本地指标上几乎全胜 pFedMARL(QS F1 0.84 vs 0.77、CS 0.98 vs 0.96),摘要「outperforming local」无本地指标支撑;实验仅单一数据集/14 客户端/无种子方差、无非对抗对照,且无智能体间通信的「合作多智能体」命名与 HAPFL/FedMRL 等前置双智能体工作未作对比。
  • Full review: Claude Code 全文七公理审稿

Federated Learning (FL) enables distributed training of machine learning models while preserving data privacy. However, FL struggles with heterogeneous, non-IID client data distributions, resulting in sub-optimal and biased global models. In this paper, we propose pFedMARL, a novel approach leveraging Multi-Agent Reinforcement Learning (MARL) with Twin Delayed Deep Deterministic Policy Gradient (TD3) to dynamically adapt aggregation strategies in FL settings. Our method employs a server-side agent adjusting client contributions to optimize global model robustness and client-side agents balancing global and local updates to personalize models effectively without pre-training. We demonstrate superior performance of pFedMARL for training a semi-supervised audio spectrogram transformer, matching or outperforming FedAvg, Ditto, and local training approaches across multiple non-IID scenarios and in the presence of adversarial clients. Our results indicate that pFedMARL actively improves accuracy, robustness, and fairness, making it suitable for real-world deployments.


16. Combining Self-Embedding Audio Watermarking with Ultra-Low-Bitrate Neural Codecs

Authors: Yigitcan Özer, Xin Wang, Zhe Zhang, Junichi Yamagishi

Categories: cs.SD, cs.AI | 6 pages; 1 figure; 2 tables; submitted to WIFS 2026 Score: 5.53/10 (Obj:6 Id:5 Ind:8 Comp:6 Eff:5 Nov:5 Rep:5)

  • Strength: 在 AV-Deepfake1M 验证集(1,480 条)上,所有编解码器×LSB 位深配置下负载均无错恢复且重建语音可懂,最佳配置 SemantiCodec 平均检测 EER 14.47%、帧级定位 AUC 0.840,并提供哈希上界(EER≈0、AUC>0.999)作为对照。
  • Weakness: 理想信道下检测 EER 仍高达 14.47–32.11%(TAAE)、删除操作定位 AUC 仅 0.614–0.665(接近随机),全程无任何基线水印方法横向对比,无代码/数据发布,且「重建质量与检测性能解耦」等机制结论缺乏定量验证。
  • Full review: Claude Code 全文七公理审稿

Partial manipulation of speech recordings, where only localized segments of an utterance are altered, poses a significant challenge for content integrity verification, as reliable detection and localization of such edits becomes harder as the manipulated proportion decreases. Watermarking offers a proactive defense alternative by embedding auxiliary information prior to distribution; classical hash-based schemes achieve near-perfect detection and localization under ideal conditions, but the original content cannot be recovered once a segment is manipulated. Building on a prior self-embedding audio steganography framework, this work presents an initial exploration of proactive defense performance under ideal conditions, extending the investigation along three axes: frame-level localization, multi-bit least significant bit variants, and evaluation across multiple ultra-low-bitrate neural codec representations. By embedding a compact neural codec representation rather than a cryptographic hash, the framework additionally enables recovery of the manipulated regions, while supporting training-free detection and localization without spoofed examples. Experiments across four controlled manipulation types under ideal channel conditions show that the embedded payload, and hence an approximate reconstruction of the authentic content, is always fully recovered without bit errors. The results also indicate that the choice of neural codec is the dominant factor for detection and localization performance.


17. Acoustic Echo Control Based on Sound Object Identification for Suppressing Howling Caused by Complicated Acoustic Paths

Authors: Osamu Hoshuyama

Categories: eess.AS, cs.SD | 5 pages, 3 figures. Submitted to IEEE Signal Processing Letters Score: 5.10/10 (Obj:1 Id:2 Ind:1 Comp:1 Eff:2 Nov:2 Rep:1)

  • Strength: 提出以声音对象识别 + 默认静音 + 条件半双工替代路径估计式 AEC,在多免提终端仿真(T60=500ms,双房间三终端)中于 AEC 收敛后稳定抑制啸叫,概念框架清晰且挑战梳理系统。
  • Weakness: 全篇唯一证据为频谱图主观观察,无 ERLE/PESQ/STOI 等任何客观量化,无最小基线与实用方案对比,双讲下过度静音损伤语音而缺缓解设计,且无代码无数据。
  • Full review: Claude Code 全文七公理审稿

This paper proposes acoustic echo control based on sound object identification for suppressing acoustic echo and howling in conferencing environments with complicated acoustic paths, where multiple hands-free terminals coexist in the same room. Conventional acoustic echo cancellers target fixed intra-device echo paths; however, unintended paths, for example, those formed via inter-terminal communication, are difficult to control and can lead to howling. Instead of estimating echo paths, the proposed approach identifies sound objects and keeps channels muted by default, allowing pass/playback only when the signal is judged not dentical to recently observed objects. This breaks echo loops caused by repeated reproduction of the same sound object and can be viewed as an extension of classical voice switching toward conditional half-duplex operation. Technical challenges and connections to related techniques are discussed, and a simulation shows howling suppression together with a trade-off against speech quality.