每日 arXiv

音频 · 语音 · 音乐 · 声学

每日新爬取的 arXiv 论文,按评分排序。中文摘要帮你快速判断是否值得深读。

2026-08-03
日期2026-08-03
已评分
均分
最高

Daily Papers — 2026-08-03

21 papers on audio, speech, music, and acoustics.

1. Discriminative Axis, Not Data Volume: What a Contrastive Corpus Teaches an Audio Embedding

Authors: Abdul Basit Tonmoy

Categories: cs.CL, cs.SD | 10 pages, 4 figures Score: 7.40/10 (Obj:9 Id:9 Ind:6 Comp:9 Eff:7 Nov:8 Rep:6)

  • Strength: 在同一音频上通过 caption collision 干预因果性地恢复情感识别 +0.089(3 seed 非重叠),同时用 4x 更大的 emotion-captioned 语料在 matched exposure 下产生 -0.0007 的 null 效果,干净证明语料库结构而非数据量决定对比嵌入编码什么属性。
  • Weakness: 核心效果在 acted emotion paradigm(CREMA-D/Ravdess 共享 lexical-matched 结构)内最强,在非 acted IEMOCAP 上衰减到 ~1/4,且 Section 4-8 的 fine-tune 实验为单次运行无 seed replicate,限制了可迁移性和统计可信度。
  • Full review: Claude Code 全文七公理审稿

Scaling the corpus is the default remedy when a contrastive representation lacks an attribute. We report a case where it does nothing, and identify what does: adding a lexical-speech round to a frozen-base multimodal embedding model raises zero-shot keyword spotting by 76 points while reducing speech-emotion recognition by 14. The loss is not a capacity limit: fine-tuning on 7,442 clips from a prosody-controlled corpus recovers emotion past its pre-speech level at a five-point keyword cost. Nor is it data volume: 29,428 mined clips whose captions explicitly name emotions, at matched exposure, move emotion by -0.0007. The difference is structural: a contrastive objective encodes an attribute only when the in-batch negatives cannot be separated without it; the controlled corpus holds sentence content fixed, so prosody is the only separating signal, whereas mined captions name emotion yet remain separable by scene content. Intervention on the same audio confirms causality: raising caption similarity does not recover emotion, but collapsing caption diversity so that emotion becomes the only separating axis recovers it by 8.9 points across three seeds, with a smaller, same-signed gain on a non-acted corpus, while keyword accuracy trades back. Corpus structure, not size or caption vocabulary, controls what a contrastive audio embedding encodes.


2. Allocation Before Ranking: Decoupled Token Compression for OmniLLMs

Authors: Zhenghui Guo, Yilin Yang, Yuanbin Man, Miao Yin, Weidong Shi et al.

Categories: cs.AI, cs.SD Score: 7.26/10 (Obj:8 Id:8 Ind:9 Comp:8 Eff:7 Nov:6 Rep:6)

  • Strength: 通过 attention decomposition 精确识别了 shared top-K 中”一个 score 做两个决策”的结构性缺陷,并提出 allocation-before-ranking 解耦方案,25% retention 下在 Qwen2.5-Omni-7B 上保留 98.7% 性能且 Pareto 优于 OmniZip@45%。
  • Weakness: Baseline 覆盖偏薄(主表仅 3 个),且核心 modality-decoupled budget idea 与同期工作 OmniScope 高度重叠,代码尚未开源。
  • Full review: Claude Code 全文七公理审稿

Token compression in OmniLLMs is typically posed as a single saliency-ranking problem: score each multimodal token, keep the top-K. We argue this abstraction is mis-specified. The same attention score simultaneously decides two things: how much retained capacity each modality receives, and which tokens within a modality are kept. A shared top-K rule therefore inherits this audio-favoring allocation prior, spending retained capacity on audio before video tokens have a chance to compete. We propose Macer, a training-free compressor that first assigns explicit audio and video budgets, then performs allocation-normalized ranking within each modality at modality-specific shallow layers. Macer significantly reduces token cost while preserving accuracy across audio-grounded, audio–video joint, visual-dominant, and video-centric benchmarks. At 25 % retention, Macer preserves 98.7 % of full-token performance on Qwen2.5-Omni-7B and 97.3 % on Qwen2.5-Omni-3B. On Qwen2.5-Omni-7B, this 25 % setting reaches OmniZip-level performance at 45 % retention while using lower FLOPs. On OmniVinci-9B, the same allocation-before-ranking principle improves over shared top-K ranking by up to 12.9 points.


3. AcoustiTrace: When Plausible Sound Violates Physics

Authors: Shiyang Li, Yuewen Cao, Yihao Liu, Yuandong Pu, Baochang Zhang et al.

Categories: cs.MM, cs.SD Score: 6.74/10 (Obj:8 Id:6 Ind:5 Comp:8 Eff:8 Nov:6 Rep:6)

  • Strength: 首个将音视频生成的声学物理真实性按声学过程三阶段八维度进行 per-output 诊断的 benchmark,9 个模型评测显示 RT60 Consistency 和 Range Attenuation 是普遍弱项(最低 17.22/100),range-guided sampling 将 R² 从 0.7138 提升至 0.8580(80.16% win rate)。
  • Weakness: 评测器依赖多个外部模块(OV-AVEL, FlexSED, Grounded-SAM, VDA, Qwen MLLM)但未做组件误差消融;RT60 视觉估计器训练/测试集场景重叠(非 scene-disjoint);数据构造与评测器验证共享同一声学关系闭环,独立性有隐患。
  • Full review: Claude Code 全文七公理审稿

Recent audio-video generators can produce semantically plausible and apparently synchronized sound, yet may still violate the acoustic processes implied by visible events and environments. Existing benchmarks provide limited support for attributing such violations to particular acoustic processes and quantifying their severity. We introduce AcoustiTrace, a diagnostic benchmark that formalizes acoustic physical realism in audio-video generation. AcoustiTrace organizes text-to-audio-video (T2AV) and image-to-audio-video (I2AV) evaluation around the acoustic process, covering sound generation, propagation environment, and acoustic reception through eight dimensions grounded in measurable acoustic quantities. Based on these evaluation dimensions, we construct a large-scale dataset organized around acoustic mechanisms, comprising real-world audio-video recordings and acoustically annotated RGB-D observations, and use it to develop targeted prompt suites and validated evaluators. Experiments reveal that even leading generators still struggle with fundamental acoustic processes despite producing plausible sound events. Finally, we show that the diagnostics AcoustiTrace provides for specific acoustic relations can guide model refinement toward more physically faithful audio and open new directions for incorporating acoustic principles into training objectives, reward modeling, and candidate selection.


4. Music Restoration via Latent Operator Optimization and Diffusion Model Priors

Authors: Michal Švento, Eloi Moliner, Valtteri Kallinen, Lauri Juvela, Vesa Välimäki et al.

Categories: eess.AS, cs.AI | Accepted to the the 27th International Society for Music Information Retrieval Conference (ISMIR 2026) Score: 6.60/10 (Obj:8 Id:5 Ind:8 Comp:6 Eff:5 Nov:6 Rep:8)

  • Strength: LOUDAR 首次将盲逆问题求解迁移到音频潜在空间,用 9.4k 参数的 causal Conv1D latent operator 配合 diffusion prior 实现 test-time 自适应恢复,在 MUSHRA 听力测试中显著优于监督基线 Apollo (p=0.002) 且在激进失真下表现最稳定。
  • Weakness: 客观指标上 LOUDAR 在多数 embedding 指标上输给 Apollo(AFxRep CD: 0.184 vs 0.146),guitar restoration 绝对效果远离实用水平(AFxRep CD: 0.547 vs AE 上界 0.203),且缺少对 latent operator 架构和关键超参数的系统消融实验。
  • Full review: Claude Code 全文七公理审稿

Music restoration seeks to recover a clean signal from an observed recording degraded by an unknown effect, distortion, or corruption. Existing systems often rely on paired training data and distortion-specific supervision, which limits their use when the forward process is not known in advance. We propose LOUDAR (Latent-space Optimization of Unknown Distortion for Audio Restoration) a general-purpose restoration method that operates in the latent space of a pretrained audio autoencoder and models the unknown distortion as a learnable latent operator. At inference time, LOUDAR alternates between estimating the clean latent variable and updating the latent operator parameters. An unconditional latent diffusion model provides a prior over clean audio and regularizes this inference by steering the latent estimate toward the manifold of clean recordings. Because the degradation model is adapted per input, the approach is broadly applicable across diverse restoration problems. We evaluate LOUDAR on singing voice effect removal and restoration, as well as guitar distortion removal, and show that it consistently improves over degraded inputs and is competitive with supervised and unsupervised baselines in waveform and latent domains.


5. The Role of Disfluencies in Speech Translation

Authors: Maike Züfle, Maria Teleki, Fabian Retkowski, Vilém Zouhar, Oliver Grabner et al.

Categories: cs.CL Score: 6.30/10 (Obj:8 Id:6 Ind:5 Comp:8 Eff:6 Nov:5 Rep:8)

  • Strength: 构建了首个8语言disfluency-annotated语音翻译benchmark(Uh-Mazing),通过消融实验识别出EDITED类disfluency(false starts/self-repairs)是翻译质量损失的主要驱动(-3到-4 COMET点),并发现模型倾向省略而非误译disfluency(deletion占1,098/3,154 error spans)。
  • Weakness: benchmark仅含80条utterances×8语言=640条翻译,规模偏小;LLM-as-judge使用GPT-5.2同时评估GPT-5.2自身的翻译输出,存在self-evaluation循环风险;inference-time干预(ICL提升chrF≤5点)未报告统计显著性或人类评估确认。
  • Full review: Claude Code 全文七公理审稿

Current speech translation systems, including SpeechLLMs, are trained on cleaned text and tend to strip disfluencies like filled pauses and false starts rather than translate them. We show this comes at a cost: disfluencies carry meaning that gets lost when speech is cleaned up. To study this systematically, we introduce Uh-Mazing, a benchmark of human-translated, disfluency-annotated Switchboard speech covering English into eight target languages. Across these languages and several architectures, we find that false starts and self-repairs, not filled pauses or discourse markers, drive most of the translation-quality loss, and that models which fail to preserve a disfluency tend to omit it rather than mistranslate it. We show inference-time decoding can mitigate this without retraining, and release the benchmark and code.


6. StreamTalk: Streaming Co-Speech Gesture Generation with Key-Pose Anchoring

Authors: Xiangyue Zhang, Jianfang Li, Jiaxu Zhang, Kaixing Yang, Steven Hoi

Categories: cs.CV, cs.AI, cs.GR | ECCV 2026 Score: 6.20/10 (Obj:8 Id:6 Ind:5 Comp:8 Eff:6 Nov:7 Rep:6)

  • Strength: 闭环 generate-retrieve-refine 机制将流式手势生成的长程漂移问题转化为可操作的方向纠偏问题,在 BEAT2 上 FGD 达到 SOTA (0.383/0.293),自交叉帧数从 387 降至 39,76 FPS 实时运行。
  • Weakness: 评测仅限 BEAT2 单数据集且运动数据库从训练集构建导致训练-评测闭环依赖;遗漏 LiveGesture (CVPR 2026) 和 DuoGesture 两个重要同期/近期基线。
  • Full review: Claude Code 全文七公理审稿

Real-time co-speech gesture generation must produce 3D motion clip by clip as speech arrives. Existing streaming methods are open-loop: each clip depends on past context, but the model cannot check or correct its trajectory. Small errors therefore accumulate and cause drift over long sequences. We observe that this failure is mainly caused by the lack of a forward constraint rather than poor short-clip quality. A plausible key pose at the end of each clip provides a destination anchor that limits drift. Based on this observation, we propose StreamTalk, a closed-loop framework with a periodic generate-retrieve-refine cycle. Streaming Pose-Guided Generation first predicts a coarse clip, retrieves a plausible tail pose from a speaker-specific motion database, and refines the clip using this pose before continuing to the next window. During training, Stochastic Anchor Masking randomly masks pose and translation frames, teaching the model to recover complete motion from sparse boundary conditions. A part-aware DiT separates hand, body, and translation streams to reduce interference between global displacement and local articulation. On BEAT2, StreamTalk achieves state-of-the-art FGD, reduces long-horizon drift relative to open-loop baselines, and runs in real time at 76 FPS. Project page: https://xiangyue-zhang.github.io/StreamTalk/.


7. Can Foundation Models Hear What Made That Sound? A Tiered Benchmark of Audio-Language Models and Traditional Classifiers for Closed-Set Sound Source Identification

Authors: Sajjad Abdoli, Ghassan Al-Sumaidaee, Ahmad ElShiekh, Ahmed Rashad

Categories: cs.SD, cs.AI Score: 6.00/10 (Obj:8 Id:6 Ind:5 Comp:8 Eff:6 Nov:6 Rep:3)

  • Strength: 提出 tiered evaluation framework 将 11 种跨范式音频分类方法在统一 2,242 样本/23 类数据上分层比较,并通过 within-class 分析正确识别出”holistic beats detailed”实为 difficulty confound,结构性错误模式(Water Running distance、Doorbell wireless、Siren default bias)跨方法一致可复现。
  • Weakness: 数据集私有未释放且无代码开源,Tier A 仅覆盖 Gemini 系列窄且重,未引用已有的 AudioBench 等同类 benchmark,CoT 分析与 confidence calibration 发现与多篇已有工作重叠,单次运行无方差估计。
  • Full review: Claude Code 全文七公理审稿

We benchmark eleven audio classification methods: five task-aware closed-set LLMs (four Gemini models plus open-weight Kimi-Audio-7B-Instruct), four fixed-vocabulary taggers (YAMNet, PANNs, Whisper-AT, and SSLAM), a zero-shot audio-text model (CLAP), and an audio-grounded LLM (BAT). We evaluate them on a closed-set sound-source identification task over 2,242 clips spanning 23 fine-grained classes and 11 categories. Since these methods differ fundamentally in how they receive the task and how outputs are scored, we group them into four evaluation tiers rather than one leaderboard, reporting macro Precision, Recall, F1, and false-negative rate per tier. The best model, Gemini-3.1-Pro-Preview, reaches 85.6 percent category-level F1 and 56.7 percent fine-grained F1. Kimi-Audio is competitive for its size, reaching 67.5 percent category-level F1 and 32.9 percent fine-grained F1, but fails to answer 1.6 percent of samples. SSLAM and CLAP match or exceed the best closed-set model at the category level without seeing the candidate list, but fall behind at the fine-grained level. Analyzing the Gemini models’ chain-of-thought across 8,968 responses, we find that response length does not predict accuracy, an apparent “holistic judgment beats detailed analysis” effect is better explained as a difficulty confound, and wrong answers are stated confidently 92 to 100 percent of the time. We report full per-class confusion matrices and metrics for all eleven methods, identify the structural error modes behind most of the accuracy loss between granularities, and give practical guidance for choosing among these method families.


8. EchoCache: Energy-Guided Cross-Modal Caching for Efficient Audio-Driven Video Generation

Authors: Jiayu Chen, Xiaoyu Wu, Rongshan Gao, Maoliang Li, Zihao Zheng et al.

Categories: cs.CV | EchoCache is honored to be accepted by ACM MM 2026 Score: 6.00/10 (Obj:8 Id:6 Ind:8 Comp:5 Eff:6 Nov:5 Rep:8)

  • Strength: 在 Wan2.2-S2V/EMTD 上实现 2.46x 加速且在 FVD/Sync-C/CSIM 上优于 TeaCache、MagCache、TaylorSeer 三个缓存 baseline,音频能量引导的 latent 级缓存策略经消融验证有效(Random anchor 导致 FVD 从 162.66 升至 192.26)。
  • Weakness: 未对比直接竞品 SyncCache(同期工作,在 Wan-S2V 上声称 3.75x 加速),且加速后的绝对质量仍显著低于原始模型(FID +16%、FVD +25%),三个组件均为已有技术的组合,独占新颖性不足。
  • Full review: Claude Code 全文七公理审稿

Audio-driven video generation (A2V) has achieved promising progress in synthesizing temporally coherent and audio-visually aligned videos, yet its inference remains expensive due to the iterative denoising process of diffusion models. Existing caching methods mainly exploit temporal redundancy in visual features while overlooking the cross-modal alignment of A2V, where audio drives visual generation with highly non-uniform temporal importance. In this paper, we identify two levels of misalignment in existing A2V caching methods: temporal-semantic and computation-storage misalignment. To address them, we propose EchoCache, an energy-guided cross-modal caching framework for efficient A2V generation. EchoCache leverages audio time-frequency energy as a saliency anchor to guide latent-level cache updates and further introduces a dynamic timestep-latent caching mechanism with quantized cache management for joint efficiency and memory optimization. Extensive experiments on mainstream A2V models show that EchoCache consistently improves the latency-quality trade-off while preserving generation quality and audio-visual consistency. In particular, on Wan2.2-S2V over the EMTD benchmark, EchoCache achieves a 2.46x speedup with the best overall performance. Code is available at https://github.com/IF-LAB-PKU/EchoCache.


9. Hear, Invoke, and Understand: A Skill-Calling Multimodal Agent for Large Audio Language Models

Authors: Yuwen Wang, Tian-Hao Zhang, Minghao Cai, Yilin Ren, Ziyang Jiang et al.

Categories: cs.MM, cs.AI Score: 5.95/10 (Obj:8 Id:6 Ind:5 Comp:5 Eff:6 Nov:5 Rep:6)

  • Strength: 在 56 个任务上构建了完整的音频智能体训练+评测闭环(HIU-Corpus 65K 轨迹 + HIU-Bench 1,395 样本),SpeechAgent-R 通过 SFT+GRPO 达到 80.05 overall,较同基座 agent harness 提升 15.04 分,并设计了 ID/OOD 划分评估工具组合泛化能力。
  • Weakness: 评测指标与训练 reward 使用相同权重结构(0.05/0.25/0.70)和相同任务指标构成评测-训练闭环;缺少规则路由器等最简 baseline 对比;方法为标准 SFT+GRPO 的领域迁移,无算法级创新;截至 review 时代码/数据尚未公开。
  • Full review: Claude Code 全文七公理审稿

Complex acoustic problems may require models to perform acoustic operations, interact with external tools and reason over the resulting textual or processed-audio observations rather than answer directly from a fixed audio input. We study such problems as tool-interactive audio reasoning and develop SpeechAgent-R, an audio agent that coordinates its intrinsic multimodal understanding with external skills and tools. To support this capability, we construct HIU-Corpus, comprising 65,492 interaction trajectories and 507.6 hours of audio across 24 tasks, 8 skills and 9 tools. SpeechAgent-R first learns structured interaction behaviors through trajectory-based supervised fine-tuning and then improves its decisions through multi-turn reinforcement learning. We further introduce HIU-Bench to jointly evaluate task performance, interaction quality and generalization to diverse task settings. It contains 1,395 samples across 56 tasks, including in-distribution (ID) and out-of-distribution (OOD) splits with substantial shifts in tool usage and workflow composition. SpeechAgent-R achieves 84.17 on ID tasks and 70.94 on OOD tasks, improving over the base model under the same agent harness by 15.40 and 14.23 points. These results demonstrate that learning skill and tool coordination improves audio agents’ ability to handle diverse task settings and adaptive tool interactions.


10. P-MUSE: Prompt-MIDI-Optional Model for Unified Instrumental Music Synthesis and Editing

Authors: Chong Jing, Junan Zhang, Jing Yang, Yulun Wu, Fan Fan et al.

Categories: cs.SD Score: 5.90/10 (Obj:8 Id:6 Ind:5 Comp:5 Eff:8 Nov:5 Rep:6)

  • Strength: P-MUSE 在所有可比基线上全面显著超越现有 MTM 系统(如 FAD 从 7.655 降至 1.238 vs MIDI-VALLE),首次统一了 paired/style/mixed 三种提示模式与生成/编辑任务,Tail-Drop 策略在 MTM 和 TTS 两个任务上均验证有效。
  • Weakness: 核心组件(FIM、课程学习、有限区间 CFG)均为已有技术的跨领域迁移,组合的 synergy 缺乏独立证明;基线比较未报告参数量和训练数据规模,公平性存疑;评测依赖自动转录模型 YourMT3 和 VGGish 特征,独立性不足;模型代码/checkpoint 开源状态未明确。
  • Full review: Claude Code 全文七公理审稿

MIDI-to-Music system renders the melody and rhythm of a target MIDI sequence into musical segment while cloning instrument timbre from a prompt recording. Existing systems typically adopt one of two distinct paradigms: conditional generation with prompt audio alone, which remains applicable when aligned prompt MIDI is unavailable, and In-Context Learning with paired prompt audio and MIDI, which exploits cross-modal alignment for stronger control on MIDI following and timbre similarity. We introduce P-MUSE, an instrumental MIDI-to-Music framework that unifies both paradigms via a multi-stage Curriculum-Learning supporting prompt-MIDI-optional inputs. P-MUSE further unifies music generation and local editing through a shared fill-in-the-middle formulation. Grounded in theoretical analysis and empirical study, we propose a phase-aware classifier-free guidance scheduling principle for Transcription-to-Audio systems, alongside a Tail-Drop strategy. Finally, to advance research in this field, we establish the first comprehensive benchmark, covering various prompt modes, generation/editing tasks, and four representative instruments: piano, guitar, bass, and drums. Demos are available at https://p-muse.github.io/.


11. Uncertainty-Aware Crossmodal Fusion for Classification of Animal Behavior

Authors: Ehsan Yaghoubi, Florian Haselbeck

Categories: cs.SD, cs.AI Score: 5.79/10 (Obj:8 Id:6 Ind:8 Comp:5 Eff:5 Nov:5 Rep:8)

  • Strength: UAF 在 identity-disjoint 评估协议下通过维度级不确定性加权融合实现了 vs. 静态拼接 15.7%/20.4% relative macro-F1 提升,且消融实验清晰证明不确定性融合(而非时间建模深度)是性能驱动因素。
  • Weakness: 缺少与 PANNs/AST/AVES 等已验证 bioacoustic backbone 的直接对比,identity-disjoint 协议下无外部 baseline 可比,且核心融合机制(高斯逆方差加权)和 loss 设计均直接改编自 COLD Fusion (TPAMI 2023),新颖性有限。
  • Full review: Claude Code 全文七公理审稿

Artificial intelligence offers substantial potential for acoustic monitoring of animals, from welfare assessment in precision livestock farming to wildlife conservation and ecological research, where vocalizations can indicate health, stress, and social states earlier and at lower cost than manual observation. However, recordings in these settings are obtained under uncontrolled conditions, including environmental noise, reverberation, overlapping calls, and sensors that degrade without notice. As a consequence, automated classification of animal vocalizations remains challenging, and the two dominant acoustic representations show complementary limitations: raw waveforms preserve temporal microstructure but degrade under clipping and reverberation, while log-Mel spectrograms capture harmonic organization but lose phase information and are sensitive to broadband noise. To address these challenges, we propose Uncertainty-Aware Fusion (UAF), a dual-stream framework that estimates Gaussian uncertainty for each representation and fuses them via uncertainty weighting. This mechanism assigns greater weight to the more confident representation with no reliability labels required. In a cross-species, identity-based evaluation excluding all individuals seen during training, UAF (mean pooling) achieves 59.4\% accuracy / 39.7\% macro F1 on the 17-class SoundWel pig vocalization benchmark and 73.1\% accuracy / 71.5\% macro F1 on the 3-class DogBark dataset, outperforming static-concatenation fusion by 15.7\% and 20.4\% relative macro F1, respectively. Ablations over four temporal aggregation strategies show that uncertainty fusion, rather than the temporal characteristics of animal calls, is the primary driver of the performance gain.


12. MEMS Microphones as Ultrasonic Transducers: Nonlinear Electrostatic Actuation and a Parametric Array Prototype

Authors: Xiaoyu Niu, Zihuan Liu, Ehsan Vatankhah, Yuqi Meng, Neal A. Hall

Categories: eess.AS, physics.app-ph | 14 pages, 9 figures. Expanded article based on work presented at the 183rd Meeting of the Acoustical Society of America. This version adds a 28-die parametric-array prototype, analytical and finite-element modeling, and extended discussion. An earlier meeting abstract covering part of this work appeared in J. Acoust. Soc. Am. 152, A50-A51 (2022), DOI: 10.1121/10.0015506 Score: 5.60/10 (Obj:8 Id:5 Ind:8 Comp:6 Eff:5 Nov:5 Rep:6)

  • Strength: 28-die MEMS麦克风阵列在collapse-snapback模式下成功产生83/93 kHz双频超声,生成10 kHz差频方向性声波,FEA(Westervelt方程)与测量在10 dB伪声校正后吻合。
  • Weakness: PA差频SPL仅35 dB(远低于60 dB对话水平),填充因子仅3.25%,未与Wygant 2009或Ahn 2019的MEMS PA原型做定量性能比较,collapse-snapback模式相对线性模式的PA增益未经消融验证。
  • Full review: Claude Code 全文七公理审稿

This paper investigates commercial-style capacitive MEMS microphone dies as air-coupled ultrasonic transmitters under nonlinear pull-in and snap-back actuation and demonstrates a compact parametric-array prototype. A single die produces large diaphragm displacement and measurable ultrasonic pressure in air. A 28-die array driven at 83 and 93 kHz generates a directional component at the 10 kHz difference frequency. Measurements are compared with analytical radiation theory and finite-element modeling, and the effects of aperture, fill factor, device uniformity, and receiver nonlinearity are discussed.


13. Domain-Specific Evaluation of Text-to-Speech Systems: A Multi-Metric Benchmarking Study

Authors: Ali Jafar, Amal Sarmad, Shifa Yousaf, Maryam Bashir

Categories: cs.CL, cs.LG | 17 pages, 1 figure. Submitted to Computer Speech & Language (Elsevier) Score: 5.53/10 (Obj:6 Id:5 Ind:8 Comp:5 Eff:5 Nov:4 Rep:9)

  • Strength: 跨 4 个 speaking style × 4 个 TTS 系统的 960 对音频多指标评测,发现 subjective-objective inversion(Gemini TTS MUSHRA 最高 3.040 但客观指标全面最差)和 Emotional domain 五指标一致最难(MCD 12.03 dB, F0 RMSE 889 cents),为 Urdu TTS 评测提供了可复现开源框架。
  • Weakness: 所有评测指标和 domain-stratified 设计均来自已有文献,未引用或对比已有的 MANGO/Rethinking MUSHRA (arXiv:2411.12719, TMLR 2025) 和 Preferences of a Voice-First Nation (arXiv:2604.21481) 等直接相关工作;主观评测样本量(20 listeners, 10 per domain)远低于 ITU-R 标准;4 个系统在。
  • Full review: Claude Code 全文七公理审稿

Recent advances in neural text-to-speech (TTS) systems have substantially improved speech naturalness and intelligibility across many languages. However, comprehensive evaluation methodologies that jointly assess perceptual quality, speaker similarity, and acoustic fidelity across diverse speech domains remain limited, particularly for low-resource and underrepresented languages. This paper presents a reproducible, multi-metric benchmarking framework for systematic evaluation of modern TTS systems through domain-specific analysis. The proposed framework integrates complementary subjective and objective evaluation protocols and is demonstrated through a comprehensive case study on a representative low-resource language spanning four speech domains: Formal, Conversational, Literary/Storytelling, and Emotional. Four state-of-the-art TTS systems – Indic-Parler-TTS, MMS-TTS, Microsoft Edge TTS, and Google Gemini TTS – are evaluated using MUSHRA listening tests, ABX discrimination tests, speaker similarity scoring with Resemblyzer, and acoustic analyses based on mel-cepstral distortion (MCD) and F0 RMSE over 960 audio pairs. Results reveal substantial variation in TTS performance across speech domains, with emotional speech consistently presenting the greatest synthesis challenge (mean MCD 12.03 dB; mean F0 RMSE 889 cents), while conversational speech achieves the highest overall acoustic fidelity. Beyond the empirical findings, this work provides a reproducible evaluation framework, publicly releasing evaluation scripts, result tables, and executable Colab notebooks to support standardized benchmarking and future research on TTS evaluation for low-resource languages.


14. Multi-Backbone Self-Supervised Ensembles for Audio Deepfake Detection and a Cross-Track Analysis of Generation-Detection Asymmetry

Authors: Seunghyun Kim, Junghyun Kim, Jiyoung Woo

Categories: cs.SD, cs.CR | Accepted at CLEF 2026, ImageCLEF-Deepfake task. Published in CEUR-WS CLEF 2026 Working Notes Score: 5.50/10 (Obj:8 Id:5 Ind:6 Comp:8 Eff:5 Nov:5 Rep:5)

  • Strength: 同一团队同时参与攻防双轨,提出三区域骨干几何和 OR/AND 不对称性框架,通过四证伪实验+PCA 维度消融提供可迁移的经验规律,检测 Final Score 0.9522 在 ImageCLEF 2026 竞赛中具有竞争力。
  • Weakness: 缺少与 AASIST/RawNet2 端到端基线的公平实验比较,所有 LOSO AUC 受限于团队内部 fake 分布,11.25% 假阳性率未获解决,top-960 transductive 阈值使用测试集信息,代码未开源。
  • Full review: Claude Code 全文七公理审稿

This paper describes the participation of team “Go-To-Germany” in the ImageCLEF 2026 Audio Deepfake Detection and Generation task. Our detection system, built on a four-backbone self-supervised learning (SSL) ensemble combining WavLM-Large, Wav2Vec2-XLS-R-300M, ECAPA-TDNN, and x-vector representations, achieved a final score of 0.9522 on the official ImageCLEF 2026 evaluation, with perfect accuracy (1.0000) on participant-generated deepfakes and 0.8875 on the held-out organizer ground-truth real data. For the Generation sub-task, our official team submission, an F5-TTS v1 baseline processed with a uniform reverberation pass and submitted as a deliberate anti-forensic probe, ranked first with a final score of 0.4304 (word error rate (WER) 4.99%, character error rate (CER) 2.07%); details of our four-model program (GLM-TTS, F5-TTS, XTTS v2, CosyVoice3), from which the official entry was drawn, appear in the paper. We present a cross-track analysis revealing a pronounced asymmetry: our detection system identifies 100% of participant-generated deepfakes, while our official generation entry, despite ranking first in the Audio Generation sub-task and evading 61.4% and 56.2% of participant and organizer detectors, attains a Final Score of 0.4304 against 0.9522 on the Detection side. We further report falsification-based ablation experiments (LOSO 56-speaker cross-validation, three-region backbone geometry, bootstrap confidence intervals, and PCA analysis) that motivate our architectural-insurance hypothesis for multi-backbone SSL ensembling. We complement these results with five cross-track insights and five pre-registered falsification experiments connecting generation-side evasion to detection-side design decisions, and we openly report an 11.25% false-positive gap on held-out organizer real recordings as the principal open challenge for deployment.


15. SAGE: Switch-Aware EEG-Guided Soft Gating for Target Speaker Extraction with In-Trial Switching

Authors: Xuefei Wang, Ximin Chen, Yuting Ding, Chunlin Li, Fei Chen

Categories: eess.SP, cs.SD Score: 5.47/10 (Obj:8 Id:5 Ind:5 Comp:5 Eff:6 Nov:5 Rep:5)

  • Strength: SAGE 将 in-trial attention switching 重构为 dynamic soft selection 问题,结合 dual-stream separator、switch-aware temperature-scaled gating、learnable latency compensation 和 MC-dropout uncertainty strategy,在自建数据集上实现 8.67 dB SI-SDR。
  • Weakness: 所有组件均为已有技术组合(ConvTasNet + temperature scaling + learnable shift + MC dropout),缺少最简 hard-switching baseline 和 EEG 因果贡献验证,仅在 18 人自建数据集上 within-subject 评测,关键超参数和代码未公开,且未引用/比较同期 SAASDNet 和 shortcut learning 风险研究。
  • Full review: Claude Code 全文七公理审稿

EEG-guided target speaker extraction is challenging under in-trial auditory attention switching, where neural noise and intrinsic latency can delay or destabilize attention tracking. Conventional methods struggle with dynamic switches and often cause discontinuities at switching points. Therefore, we propose SAGE, a switch-aware EEG-guided soft gating framework that treats in-trial switching as dynamic selection. SAGE generates two candidate speech streams with a robust separator and uses an EEG-guided switch-aware gating module to produce smooth fusion weights and suppress transition artifacts. We further integrate latency-compensated alignment and an uncertainty-driven conservative strategy to handle latency discrepancies and fluctuating EEG reliability. SAGE outperforms baselines, achieving 8.67 dB SI-SDR and 88.24% STOI while reducing average switching latency to 2.04 s. By coupling neural decoding with speech separation, it enables robust target extraction in dynamic scenarios.


16. An End-to-End Workflow for Fin Whale Song Detection, Note Characterization, and Localization with Distributed Acoustic Sensing

Authors: Dídac Diego-Tortosa, Miriam Romagosa, Arantza Ugalde, Hugo Latorre, Sergi Ventosa et al.

Categories: cs.SD, physics.geo-ph | 14 pages, 8 figures, 2 tables. Manuscript in preparation for journal submission Score: 5.40/10 (Obj:8 Id:5 Ind:8 Comp:5 Eff:5 Nov:4 Rep:8)

  • Strength: 完整的无监督 DAS 长须鲸监测流水线,代码开源,pick-level precision 达 0.990,并首次在 DAS 上展示了 type-A/type-B 音符表征和重叠发声分离能力。
  • Weakness: 仅 6 首歌声评测且无任何 baseline 定量比较,核心组件均为已有方法直接应用,Goestchel et al. (2026) 已覆盖检测+关联+定位的核心功能,消融实验完全缺失。
  • Full review: Claude Code 全文七公理审稿

Submarine fiber-optic cables instrumented with distributed acoustic sensing (DAS) provide an effective approach for large-scale monitoring of fin whales. We present an end-to-end workflow for detecting, characterizing, and localizing fin whale notes, tested on two submarine telecom cables in the Strait of Gibraltar and western Alboran Sea. The workflow applies a kurtosis-value picker adapted to narrow-band fin whale notes. Channel-wise detections are grouped into individual notes using density-based spatio-temporal clustering, cluster agglomeration, and hyperbolic fitting to reject incoherent picks. The retained clusters are characterized through temporal, spectral, and energy-related descriptors that support note-type discrimination and estimation of inter-note intervals. Relative arrival times across DAS channels are then used in a grid-search procedure to estimate candidate source locations. Evaluation against manually annotated detections from six fin whale songs yielded median pick-level precision of 0.990 and recall of 0.744, and median cluster-level precision of 0.880 and recall of 0.806. Representative applications demonstrate separation of overlapping vocalizations, characterization of type-A and type-B notes, and the inference of apparent source movement. By transforming dense DAS recordings into compact note-level bioacoustic information, the workflow provides an integrated framework for fin whale monitoring and a basis for adaptation to other synchronized acoustic receiver arrays.


17. SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks

Authors: Yu Zhang, Ruiqi Li, Changhao Pan, Ke Lei, Xiang Yin et al.

Categories: eess.AS, cs.SD | Technical Report by ByteDance Score: 5.05/10 (Obj:8 Id:4 Ind:5 Comp:5 Eff:6 Nov:5 Rep:2)

  • Strength: 统一 instruct 与 zero-shot 多音频模态生成的完整系统设计,在 expressiveness 指标上多个 benchmark 领先(SwanBench-Speech Richness 3.90/Hierarchy 3.70 monologue, 3.66/3.85 dialogue; SwanBench-Scene Mean MOS 4.22),Unified MoE ablation 显示对 Instruction。
  • Weakness: 核心组件(Engram、reward-conditioned QC、GRPO、curriculum 各阶段)均无独立 ablation,且系统完全依赖 ByteDance 闭源数据(23M+ 小时内部语音)和闭源工具(Seed-ASR、SwanVerifier),代码和模型未开源,外部无法独立验证。
  • Full review: Claude Code 全文七公理审稿

Speech and audio generation is often needed in animation dubbing, audio drama, movies, advertising, games, podcasts, and short-video production. In these scenarios, creators may need to design voices without reference recordings, control speaker styles with natural language, support acoustic scenes with environments and audio effects, and later reuse the designed voices. Therefore, it is important to support multi-speaker speech and audio generation for both instruct and zero-shot tasks. The instruct task requires a caption of the environment, speaker styles, and fine-grained content, while the zero-shot task uses reference audio together with the same fine-grained content. We address these tasks from both the data and model sides. First, we propose SwanData-Caption, which cleans raw speech and audio data, adds targeted synthetic coverage, and annotates diverse and accurate multi-level captions. Then, we propose SwanTale, a multi-speaker expressive speech and audio generation model that supports both zero-shot and instruct tasks. We introduce SwanVAE to support high-quality multi-audio-modality generation. Then, we adopt reward-conditioned quality control and Engram conditioning, along with Unified MoE for multi-task and multi-audio-modality modeling. In addition, we use curriculum learning and GRPO post-training to let the model progressively learn and strengthen its capabilities. Experimental results show that SwanTale leads on multiple key zero-shot and instruct metrics, achieves the best expressiveness scores in both tasks, and supports complex instruct generation involving multi-speaker speech and audio. Demos can be found at https://swanaigc.github.io/#swantale.


18. Comparative Validation of GPT-4o-mini and Teacher Mean Scores for Automated Scoring of Music Analysis Responses: Single-Pass Deployment, Repeatability, and Strategy-Specific Bias

Authors: Baicheng Lin, Lingxi Jin, Kyung-Seok Min

Categories: cs.SD, cs.HC, stat.AP | Proceedings of the International Meeting of the Psychometric Society: The 91st Annual Meeting, Seoul, Republic of Korea, 2026 Score: 3.89/10 (Obj:5 Id:4 Ind:8 Comp:2 Eff:5 Nov:2 Rep:2)

  • Strength: 在 300 份真实学生答卷上系统比较了三种已知提示策略,发现 Fs+CoT 与教师均值一致性最强(ICC=0.657)、SC 重复性最高(intra-LLM ICC=0.990)但个体一致性弱、RAG 系统性过评分(bias=+1.787),维度层面 Terminology 一致性最弱。
  • Weakness: 无方法新颖性(三种策略均为已发表方法的直接应用),无最简基线或组件 ablation,绝对效用未达可用水平(最佳 ICC=0.657 属 moderate),核心结果完全依赖闭源 GPT-4o-mini 且代码/prompt/数据/评分标准均未开源,无法独立验证。
  • Full review: Claude Code 全文七公理审稿

Scoring open-ended music analysis responses is time-consuming and requires nuanced judgments of harmonic knowledge and formal understanding. This study evaluates the validity and repeatability of GPT-4o-mini for rubric-based scoring of music analysis essays, using teacher mean scores as the benchmark. A dataset of 300 university-level student responses was scored by teachers on four dimensions: Harmony, Form, Reasoning, and Terminology. GPT-4o-mini scored the same responses using three prompting strategies: few-shot prompting with chain-of-thought reasoning (Fs+CoT), retrieval-augmented generation (RAG), and self-consistency based on five internal generations per administration (SC). Each strategy was administered three times with the model, prompt, rubric, and response held constant. Single-pass scores represented an operational scoring condition, whereas median aggregation across three runs was used to examine robustness. Agreement with teacher mean scores was evaluated using correlation, intraclass correlation, Krippendorff’s alpha, quadratic weighted kappa, and scoring error indices. Fs+CoT showed the strongest agreement with teacher mean scores in both single-pass scoring and median aggregation. RAG showed systematic over-scoring, whereas SC produced highly repeatable scores but weaker individual-level agreement. Dimension-level analyses showed that scoring performance varied across rubric components, with Terminology generally showing weaker agreement than Reasoning. These findings indicate that GPT-4o-mini can generate stable scores for complex music analysis responses, but prompting strategies produce distinct scoring profiles. Operational use therefore requires strategy-specific calibration, dimension-level validation, and continued human oversight.


19. SGAD: A State-Guided Adaptive Decision Framework for Robust EEG-Based Auditory Attention Switch Decoding

Authors: Yuting Ding, Xuefei Wang, Ximin Chen, Chunlin Li, Fei Chen

Categories: eess.SP, cs.SD Score: 3.89/10 (Obj:8 Id:2 Ind:5 Comp:4 Eff:5 Nov:3 Rep:4)

  • Strength: 提出 6 种分层评估协议以揭示数据划分偏差;SGAD 在 Sw-F1 上较 EBD 基线平均提升 13.6pp,同时降低 SDL 0.20s,显示出在切换检测上的一定有效性。
  • Weakness: 未引用或比较两篇直接先前工作(Heintz 2025 HMM 后处理, Yao 2026 Markov switching model),这两者用更简洁的概率模型解决相同的状态引导平滑问题;仅与简单启发式基线(WD/PBD/EBD)比较,且仅使用了作者自建数据集。
  • Full review: Claude Code 全文七公理审稿

Achieving robust EEG-based auditory attention switch decoding (AASD) is crucial for intelligent hearing aids. However, its application is limited as EEG non-stationarity complicates sequential decision-making, and insufficient control of potential confounding factors may overestimate performance. Therefore, we propose a state-guided adaptive decision (SGAD) framework that infers attention transition states via causal state detection and dynamically modulates temporal smoothing through state-guided adaptive gating. We further introduce six hierarchical evaluation protocols to assess generalization across audio, speaker, and subject dimensions. Experimental results show that SGAD improves decoding accuracy and stability while maintaining low response latency across evaluation scenarios. Performance variations across protocols further suggest data partition-related biases. Together, these findings advance robust AASD for neuro-steered hearing applications.


20. Deep Learning-Based Active Trim Panels for Enhanced Aircraft Interior Noise Control

Authors: Boxiang Wang, Malte Misol, Zhengding Luo, Junwei Ji, Xiaoyi Shen et al.

Categories: eess.AS Score: 3.79/10 (Obj:6 Id:2 Ind:2 Comp:5 Eff:4 Nov:5 Rep:4)

  • Strength: 提出温度感知 SFANC 方法,用 0.12M 参数 1D CNN 在频率变化场景下实现 17.5 dB NR(1 秒内),优于 FxLMS 的 5 dB,且温度感知 ANC 在文献中属首次。
  • Weakness: 评测在单一合成模型上闭环完成(训练/测试同源),无物理验证;温度变化场景下 FxLMS 在 3 分钟后反超 TP-SFANC,核心场景优势不持续;未与作者自己的 prior SFANC-CNN 系统公平对比,无消融实验隔离温度感知的贡献。
  • Full review: Claude Code 全文七公理审稿

Active noise control (ANC) trim panels offer an effective solution to suppress multi-tonal noise in aircraft. The selective fixed-filter ANC (SFANC) method, characterized by low computational complexity, high robustness and rapid response, is suitable to handle multi-tonal engine noise that varies in frequency due to changes in the rotational speed of the engine shaft. However, real-world conditions introduce variations in lining temperature, altering acoustic and structural paths and degrading noise reduction performance. To address this challenge, a temperature-perceptive SFANC (TP-SFANC) approach is proposed that employs a lightweight one-dimensional convolutional neural network (1D CNN) trained using a multi-task learning strategy. By processing both reference and error signals, the 1D CNN learns frequency and temperature characteristics to dynamically select the optimal control filter. Numerical simulations demonstrate the effectiveness of the proposed method in attenuating multi-tonal noise across varying frequencies and lining temperatures.


21. Sounding Canvas: Embedding Algorithms in Networked, Sensorial Sound Art

Authors: Luciano Ciamarone, Dora Motèque, Marco Giordano

Categories: cs.SD | 23 pages, 7 figures Score: 3.68/10 (Obj:5 Id:2 Ind:4 Comp:5 Eff:2 Nov:4 Rep:8)

  • Strength: 完整的软硬件系统实现并开源代码(GitHub),将电容传感、CNN视觉-音频映射、LSTM在线事件管理和WebSocket网络通信集成到物理画布中,在威尼斯等地实际展出。
  • Weakness: 无任何定量评估或对照实验——论文自述”no formal analysis has been conducted”且”no definitive claims are made”,各算法组件(CNN映射、HOMM vs LSTM、网络化效果)均未与baseline比较,无法判断系统复杂度是否换来真实能力提升。
  • Full review: Claude Code 全文七公理审稿

Sounding Canvas turns painting into a touch-responsive multimodal installation by embedding capacitive sensors, real-time decision models, and networking inside the canvas. Touches trigger spatialised sounds that appear to emanate from the painting itself. The work embeds algorithms physically, as sensing and computation concealed behind the artwork; perceptually, through an offline visual-to-sonic mapping that aligns a painting’s features with sound descriptors; and performatively, through online models that shape live interaction with visitors and with remote canvases over a network. We describe the artistic rationale and technical implementation, combining a CNN-based offline mapping that defines the sound vocabulary with two online event managers, a higher-order Markov model and an LSTM-based policy, that balance responsiveness with guided exploration. We discuss how these layers make algorithms perceptible through behaviour rather than code, how networking transforms solitary touch into distributed co-authorship, and how the system raises questions of authorship, agency, and evaluation in embedded algorithmic artworks.