每日 arXiv

音频 · 语音 · 音乐 · 声学

每日新爬取的 arXiv 论文,按评分排序。中文摘要帮你快速判断是否值得深读。

2026-08-11
日期2026-08-11
已评分
均分
最高

Daily Papers — 2026-08-11

15 papers on audio, speech, music, and acoustics.

1. Seeds Before Objectives: Rethinking Evaluation for Low-Resource Garhwali ASR

Authors: Karamvir Singh Batra, Prathamjyot Singh, Ashima Sood, Jasmeet Singh, Sahil Sharma

Categories: cs.CL | 19 pages, 3 figures. Accepted for oral presentation at ICNLSP 2026, Trento, Italy, September 2026 Score: 7.74/10 (Obj:8 Id:9 Ind:9 Comp:8 Eff:7 Nov:6 Rep:9)

  • Strength: 在低资源方言 ASR 中首次建立多 seed 可复现 benchmark(5 seed × 3 目标函数 + 配对 Wilcoxon 检验 + 功效分析),可靠证明 Focal CTC、matra-weighted CTC 和 Hindi→Garhwali 迁移均不优于标准 CTC(47.0% WER),并统一解释为表征瓶颈。
  • Weakness: 核心方法论工具非原创(Bouthillier et al. 2021 等),仅在一个方言/一个语料库上验证,5 seed 经作者自己证明可能不足(需 11–20 seed 达 80% power),47% WER 远未达到部署水平。
  • Full review: Claude Code 全文七公理审稿

At corpus sizes typical of low-resource dialects, single-run comparisons can yield gains that do not replicate. We show this for Garhwali, an under-resourced Indo-Aryan language of the central Himalaya, building the first reproducible multi-seed ASR benchmark on the official VAANI splits, with per-seed outputs and significance testing. Re-examining plausible gains, we find them fragile: neither Focal CTC nor a matra-weighted objective beats standard CTC under seed-level testing, the matra objective fails to cut even its targeted errors, and Hindi-to-Garhwali transfer gives no gain over direct fine-tuning. What holds up is mundane: w2v-BERT 2.0 with standard CTC reaches 47.0% WER over five seeds, beating the larger MMS-1B and comparable models; pretraining design, not parameter count, drives performance, and speed augmentation gives a small, largely consistent gain. Multi-seed evaluation on official splits separates real gains from seed noise.


Authors: Dasol Lee, Minhee Lee, Seonguk Ju, Daewoong Kim, Harin Lee et al.

Categories: cs.SD | 8 pages, 6 figures, 4 tables. Accepted at ISMIR 2026 Score: 7.68/10 (Obj:8 Id:6 Ind:9 Comp:8 Eff:7 Nov:8 Rep:9)

  • Strength: 跨 6 架构 × 3 seeds 的 18 runs 一致发现 Korean chart 音乐 era offset 在 1960s–1980s 约 −4.7 至 −4.1 年、1990s 减半至 −2.4 年并持续至 2000s,框架简洁且代码数据完全开源。
  • Weakness: mel-spectrogram CNN 无法分离制作技术 lag 与风格 lag,缺少最简声学特征 baseline 对照,2000s 双峰分布使单一中位数失效,1960s Melon 仅 14 测试曲统计效力不足。
  • Full review: Claude Code 全文七公理审稿

Popular music circulates globally while being locally reinterpreted, yet this process of cross-cultural style diffusion has rarely been quantified. We propose an era-classification framework for measuring temporal alignment between chart cultures. CNN classifiers trained from scratch on Billboard Hot 100 audio are applied to Korean Melon chart songs. Korean chart songs from the 1960s through the 1980s are consistently inferred as belonging to earlier Billboard eras, by a median of about four to five years, while the same models remain unbiased on held-out Billboard audio. The offset then halves at the 1990s, to roughly two to three years, and holds there through the 2000s. Reverse inference shows a complementary narrowing, and the pattern holds across architectures and seeds. We interpret these results as reflecting how globally circulating pop styles were locally adopted and progressively synchronized. The framework can be applied to other pairs of chart cultures beyond the US-Korea case examined here.


3. A Dataset and Benchmark for Optical Music Recognition of String Quartet Scores

Authors: Dongmin Kim, Brian Liu, Jose J. Valero-Mas, Dasaem Jeong

Categories: cs.CV, cs.SD | 8 pages, 2 figures, 5 tables. Accepted at the ISMIR 2026 Score: 7.05/10 (Obj:8 Id:5 Ind:8 Comp:6 Eff:8 Nov:8 Rep:9)

  • Strength: 首个多声部 OMR 数据集,24,544 系统图像 + 98,172 谱表图像,手动视觉对齐可审计,三格式编码使表示选择可比较,最佳基线达 3.6%/5.9% OMR-NED 确认任务可行
  • Weakness: 34 配置缺乏系统消融隔离各因素独立贡献,扫描退化原因未识别,缺少最简 heuristic baseline 作为下界参照
  • Full review: Claude Code 全文七公理审稿

Optical music recognition (OMR) transcribes music scores into digital formats. While the field has advanced significantly on monophonic and piano-form scores, multi-part score transcription remains underexplored, largely due to the absence of a suitable dataset. We introduce OpenScore String Quartet for Optical Music Recognition (OSSQ-OMR), the first dataset dedicated to multi-part OMR. Built on the OpenScore String Quartet corpus, OSSQ-OMR pairs digitally encoded scores with their original scanned editions from IMSLP, with all images visually aligned to their transcriptions. The dataset is released with score images at system and staff levels, and paired transcriptions in three encoding formats: Extended Linearized MusicXML (LMXE), **kern, and ABC. In total, OSSQ-OMR contains 24,544 system images and 98,172 staff images drawn from 116 string quartet scores. We accompany the dataset with a benchmark protocol and baseline results from two representative OMR models, evaluated across four random score-level splits with mutually exclusive test sets. Baselines reach OMR-NED as low as 3.6% on synthetic and 5.9% on scanned inputs; results reveal substantial effects of encoding and segmentation choices, with the LSTM-based baseline degrading on scanned inputs roughly 2.6 times less than the Transformer-based baseline.


4. Beyond Dry References: Learning Relative Audio Effects Representations via Contrastive Distance Learning

Authors: Xinlu Liu, Huibin Lin, Weixing Wei, Zhenhai Yan

Categories: cs.SD | 8 pages (6 pages of main text), 3 figures, 3 tables. Accepted at the 27th International Society for Music Information Retrieval Conference (ISMIR 2026). Project page: https://relative-fx.github.io Score: 6.95/10 (Obj:9 Id:5 Ind:7 Comp:5 Eff:9 Nov:6 Rep:8)

  • Strength: 提出 relative Fx representation 重构,去除 dry reference 依赖,在标准 MUSDB18 协议下四种乐器全面超越 Fx-Encoder++(Avg Ld 1.388 vs 1.601, 13.3% 提升),仅用 MoisesDB 训练仍有 9.5% 提升。
  • Weakness: 多组件同时变更导致归因不够干净(cross-segment sampling 贡献大于 relative encoding 本身),cross-attention 在核心指标上无贡献但仍作为架构贡献,antisymmetric variant 的实际应用价值未验证。
  • Full review: Claude Code 全文七公理审稿

Audio effects (Fx) representation learning plays a key role in intelligent music production, including automatic mixing and Fx style transfer. Existing methods typically rely on dry or nearly dry references for effect modeling, yet truly unprocessed audio is rarely available in practice, as real recordings inevitably reflect the microphone, room acoustics, and preceding signal processing. Instead of pursuing absolute effect encodings, we argue that the relative effect distance between audio signals is more meaningful for real-world music production. Motivated by this, we propose RelFx, a contrastive learning framework that learns relative effect transformations from general audio collections without requiring dry references during representation training. Our approach uses a dual-branch Siamese encoder equipped with cross-attention and differential gating fusion to infer the shared effect transformation from a reference clip and an effect-processed, content-related clip. We further propose an antisymmetric fusion variant for bidirectional effect encoding, such that swapping the input order directly produces a nearly sign-reversed embedding, a property not explored in earlier work. Moreover, our dry-reference-free formulation eliminates the reliance on dry multitrack datasets and enables training on effect-bearing audio. Experiments on Fx style transfer demonstrate state-of-the-art performance under the standard Fx-Encoder++ MUSDB18 evaluation protocol, consistently outperforming existing approaches across all four instrument categories.


5. Whisper-Aware LLM: Self-Supervised Uncertainty Learning for Robust Whispered Speech Recognition

Authors: Gaopeng Xu, Zhenyu Wang, Zheng Xue, Yinfeng Xia, Haitao Yao

Categories: cs.SD, cs.AI Score: 6.50/10 (Obj:9 Id:5 Ind:6 Comp:5 Eff:8 Nov:5 Rep:5)

  • Strength: 在 AISHELL6-Whisper 上实现 SOTA(CER 1.31%,相对 Seed-ASR 降 17%),同时将噪声幻觉率从 25%+ 降至 4.5%,且通用 ASR 性能未退化。
  • Weakness: 各组件(F0 预测、掩码频谱重建、attention bias)均为已有技术拼装,缺少与简单 confidence 阈值方法的直接对比、UPM 两个自监督任务的独立消融、以及代码/Noise Hallucination Set 的开源。
  • Full review: Claude Code 全文七公理审稿

The signal ambiguity of whispered speech drives ASR systems toward two opposing failure modes: failing to capture whispered speech or hallucinatory transcription of noise. This paper introduces the Whisper-Aware LLM, a framework that teaches an Audio-LLM to perceive and react to this uncertainty. Our model develops an intrinsic self-awareness by learning to quantify the physical deficiencies of acoustic signals through targeted self-supervised tasks. This learned uncertainty is then operationalized via a novel Confidence-Fused Decoding mechanism, which provides both high-level instructions and frame-level attention modulation to the LLM decoder. Our experiments confirm the effectiveness of this approach. The model sets a new state-of-the-art on whispered speech with a 17% relative CER reduction on AISHELL6-Whisper. At the same time, it directly addresses the reliability trade-off, with hallucination rates dropping from over 25% to 4.5%.


6. X2-Turn: Frame-Synchronous Dual-Head Modeling for Joint Streaming ASR and Turn State Prediction

Authors: Kaiqi Fu, Rime Wen, Altman Lin, Shawn Qin, Roy Gan et al.

Categories: cs.CL, eess.AS Score: 6.37/10 (Obj:8 Id:5 Ind:5 Comp:8 Eff:6 Nov:6 Rep:5)

  • Strength: 双头帧同步设计简洁且工程合理,但贡献主要是将现有的骨干模型与新的头相结合,而不是一种新的机制见解。
  • Weakness: 论文将流式轮次预测的改进归因于帧同步双头建模和 ASR 锚定监督。然而,在识别真正有效的原因方面存在关键缺失
  • Full review: Claude Code 全文七公理审稿

Accurate and responsive turn-taking is essential for spoken dialogue systems, which must distinguish in real time between user interruptions, backchannels that should be ignored, and the completion of an utterance. Prior modular approaches typically optimize turn state prediction at the utterance or fixed-chunk level, creating a mismatch with the continuous turn state estimate, and often depend on an auxiliary ASR model, which limits responsiveness and increases overall system complexity. Therefore, we present X2-Turn, a frame-synchronous turn state prediction method via delayed-stream modeling. Specifically, building on the pretrained Voxtral Realtime model, we introduce a frame-synchronous turn state head that operates in parallel with the ASR head on shared streaming representations, jointly predicting ASR tokens and fine-grained turn states at the frame level. We evaluate our method on the bilingual Chinese-English Easy-Turn test sets, and the results demonstrate its effectiveness in achieving accurate turn-taking detection while maintaining low latency.


7. DuplexWorld: Can voice agents help you get through the day?

Authors: Aryan Vijay Bhosale, Harshit Rajgarhia, Akhil Pothanapalli, Asif Shaik, Abhishek Mukherji et al.

Categories: cs.SD, cs.AI, cs.CL Score: 6.32/10 (Obj:8 Id:6 Ind:5 Comp:8 Eff:6 Nov:8 Rep:5)

  • Strength: 首个面向全双工语音智能体的导航世界(Pathfinding)加五个企业世界统一在单一 harness 和 12 指标三支柱体系下评测 5 个商用系统,发现“沉默可赢对话但不可赢导航”“探索越多到达越少”“声学质量不预测能力”等可靠、可迁移、部分反直觉的经验规律。
  • Weakness: 全部被评系统和全部 harness 侧模型均为闭源商用 API 且无 dated snapshot,复现性严重受限;Pathfinding 同时引入三个新变量且无单独识别;LLM judge 与被评系统存在同供应商家族的潜在偏好风险,未做 judge 中立性 ablation。
  • Full review: Claude Code 全文七公理审稿

Speech-to-speech (S2S) voice agents are increasingly being incorporated into enterprise for customer care and as daily companions for consumers owing to the ease of the conversational modality over text. However, existing benchmarks fail to holistically evaluate voice agents along axes that really matter and are shaped as tests of agentic tool calling against a database. We believe they fail to adequately account for the diversity of conversational dialogue that mundane activities introduce and further, never test how faithfully an agent can assist on tasks that move beyond database manipulation. To tackle this DuplexWorld introduces six worlds where voice agents are especially useful: banking, insurance, travel, healthcare and logistics, and Pathfinding. Agents are evaluated on eleven different types of conversations across 156 scenarios (350+ hours of conversation), each testing conversational and analytical capability to varying degrees. Through extensive evaluation comprising agentic, conversational and speech-naturalness metrics, we show that even the best voice agents leave substantial room for improvement on all 3 axes (Pass@1: 0.490, turn-taking: 0.653, DNSMOS: 3.378). We perform extensive analysis on agentic v conversational performance, world- and conversation type-wise performance, failure modes exploring the explore v exploit lens for Pathfinding conversations and voice agent reliability over all six worlds.


8. Pitch Contour Tokenization using VQ-VAE and Its Application on Korean Traditional Music Analysis

Authors: Seonguk Ju, Seola Cho, Sooin Chung, Danbinaerin Han, Dasaem Jeong

Categories: cs.SD | 8 pages, 3 figures, 2 tables. Accepted at ISMIR 2026 Score: 6.30/10 (Obj:8 Id:6 Ind:8 Comp:6 Eff:5 Nov:8 Rep:8)

  • Strength: 提出了一种新的基于 VQ-VAE 的无监督音高轮廓 tokenization 方法,配以 transformation-minimized reconstruction loss,在分割一致性上显著优于 baseline(KLD 0.531 vs 1.747),且学习到的 token 在无监督条件下与韩国传统音乐的 sigimsae 类别和 pansori mode 对齐。
  • Weakness: 在 sigimsae 分类探针任务上,最佳模型与 Autoencoder baseline 的提升极小(F1 macro All: 0.285 vs 0.284),与监督上界差距大(0.285 vs 0.463),且作者选择性引用最有利的指标比例;baseline 覆盖不足,缺少其他无监督离散化方法的比较。
  • Full review: Claude Code 全文七公理审稿

Computational analysis of music often relies on discrete representations, yet many musical traditions are organized around continuous pitch movement that resists segmentation into note-like units. For such traditions, the discrete units that analysis would build on are not given in advance. We address this gap by learning a vocabulary of local pitch-contour patterns directly from unlabeled audio, using a VQ-VAE that quantizes fixed-length contour segments into a finite codebook. To make the learned tokens stable across segmentation positions and small variations in timing and pitch range, we train the model with a reconstruction objective evaluated under the best alignment among a set of candidate temporal and pitch-domain transformations. Applied to Korean traditional music, the learned tokens recover information about expert-defined sigimsae categories without supervision, and in pansori individual tokens align with the two principal modes, Gyemyeonjo and Ujo, supporting their use as units for corpus-level analysis of contour-centric traditions.


9. ASR-Roundtrip Evaluation Can Mask Context- and Convention-Dependent Reading Errors in Chinese News TTS

Authors: Shijun Luo, Lizhi Wan

Categories: cs.CL | 5 pages, 4 tables. Conference-format manuscript. Supporting materials are available at https://github.com/Jayden-X-L/cn-newstts-asr-roundtrip-masking and archived at https://doi.org/10.5281/zenodo.21454402 Score: 6.20/10 (Obj:8 Id:6 Ind:5 Comp:8 Eff:5 Nov:6 Rep:8)

  • Strength: 完整分母的针对性审计(46 masked + 9 exposed + 55 no-error = 110)配合音频优先人工标注和 span-isolation 诊断(re-exposes 18/46),以极低任意性压缩了 ASR-roundtrip 在中文新闻 CDRD span 上的评估盲区这一真实且先前未被精确表征的机制。
  • Weakness: 主要证据链(MiMo-TTS + MiMo-ASR 同厂商闭环)存在结构性独立性问题,且 110 例靶向审计规模有限、无随机抽样生产基线,跨 ASR 机制差异(Paraformer 2/97 vs Qwen 40/97)未充分隔离根因。
  • Full review: Claude Code 全文七公理审稿

ASR-roundtrip evaluation is widely used as a scalable proxy for text-to-speech (TTS) intelligibility, but it can produce false negatives for reading errors perceived by listeners. We study Chinese news TTS spans whose correct reading depends on context or domain conventions, such as sports scores, aircraft models, technical units, and membership names. In these cases, Raw TTS can choose a plausible but wrong reading while ASR transcribes the audio as the intended or surface-correct text. A targeted audit over 110 high-risk MiMo TTS cases, reported with a complete denominator, confirms 46 masked false negatives, 9 exposed TTS errors, and 55 cases with no Raw TTS error. A span-isolation diagnostic re-exposes 18/46 previously masked errors. A Raw-only CosyVoice audit on the same targeted pool confirms 51 masked cases. Across the 97 TTS-specific audio files labeled confirmed masked across the two audits, Qwen3-ASR surface-recovers 40 cases, whereas Paraformer does so in only 2. The results suggest that ASR-roundtrip is useful for screening but insufficient as standalone ground truth for Chinese news reading-risk evaluation.


10. The GENEA Challenge 2026: A Large-Scale Disentangled Evaluation of Speech-Driven Gesture Generation on the Seamless Interaction Dataset

Authors: Rajmund Nagy, Silvia Arellano García, Hendric Voss, Mihail Tsakov, Taras Kucherenko et al.

Categories: cs.CV, cs.GR, cs.HC, cs.SD | 15 pages, 14 figures. Preprint Score: 5.60/10 (Obj:8 Id:5 Ind:9 Comp:6 Eff:6 Nov:5 Rep:9)

  • Strength: 23,000+ 票、869 名评测者的大规模人类评测揭示了当前手势生成系统在 Seamless Interaction 数据集上的能力边界——motion realism 和 speech alignment 有梯度差异,但 dyadic alignment 和 semantic alignment 全部接近随机水平(mocap 上限分别为 65% 和 79%,最佳系统仅 8% 语义得分)。
  • Weakness: 仅评估 5 个参赛系统而未包含 GENEA Leaderboard 上已达 mocap 平价的 7 个已发表系统作为 baseline,无法定位参赛系统相对于领域 SOTA 的位置;T3 仅 46 个 segment 限制了统计功效。
  • Full review: Claude Code 全文七公理审稿

This preprint presents the results of the fourth GENEA Challenge, a large-scale human evaluation of five speech-driven gesture-generation systems trained by participating teams on the Seamless Interaction dataset of dyadic conversations. As in the 2023 GENEA Challenge, we used a disentangled evaluation methodology to assess motion quality and speech alignment without confounding between the two, and performed a dyadic mismatching study to isolate the effect of listening and reacting to the interlocutor. We additionally introduce a new semantic gesture-generation task and a text-mismatching evaluation methodology using the Grounded Gestures subset of the data. In total, we ran four large-scale user studies, collecting over 23,000 votes from 869 test-takers. In the motion-realism study, the dataset’s filtered segments had substantially higher motion quality than all challenge submissions (68-95% pairwise winrate). In the speech-alignment study, the motion-capture segments provided a conceptual ceiling at 62% alignment score, with the top submission significantly behind at 32% and the rest only slightly above the 0% expected of an input-independent system. In the dyadic study, motion capture again set the ceiling at 65% appropriateness score, but no submission scored substantially above chance, indicating that the systems could not yet respond to the interlocutor. Finally, the semantic mismatching evaluation found highly expressive gestures in the dataset (test-takers identified the matching transcript 79% of the time), yet almost all submissions failed to generate semantically expressive motion, with the best achieving only an 8% appropriateness score. The collected votes and outputs will be made publicly available at https://genea-workshop.github.io/2026/challenge/ to facilitate reproducibility and further research.


11. MazzikaAI: A knowledge-based performance-to-prompt compiler for real-time Arabic maqam accompaniment with a streaming text-to-music model

Authors: Jiaxin Du, Boulbaba Abdeljaouad, Yong Zhuang, Haoyu Li

Categories: cs.HC, cs.AI, eess.AS Score: 5.60/10 (Obj:8 Id:5 Ind:5 Comp:8 Eff:5 Nov:6 Rep:5)

  • Strength: 完整的实时 maqam 伴奏系统,knowledge-based compiler 层成本 <1ms(vs generator 三数量级),263ms median end-to-end latency,179 re-prompts/min 零失败流稳定性,maqam grounding 使 quarter-tone voiced frames 从 14.9% 显著增至 22.4%。
  • Weakness: 两大 prompt-level 控制机制之一(instrument suppression)完全失效;beat-level entrainment 缺失(专家反馈系统保持固定 tempo);核心 quarter-tone rendering 仅实现 inflection-without-anchoring(无稳定 150-cent peak);依赖闭源 Lyria RealTime 导致结果不可独立复现;缺乏与 ReaLJam 和 au。
  • Full review: Claude Code 全文七公理审稿

Arabic maqam music microtonal, modal, and built on ornamented call and response is among the traditions most underserved by generative music models, whose training frameworks remain predominantly Western and equaltempered. Real time accompaniment sharpens this gap: an AI partner must listen, adapt dynamically, and respect idiomatic microtonal structures. Streaming text to music models provide strong generative capabilities but lack precise control interfaces. We present MazzikaAI, a knowledge based system that uses natural language as the actuator of a realtime control loop. By compiling live MIDI, gesture, and inferred harmony into continuously updated text prompts, MazzikaAI steers an unmodified streaming generator, Google Lyria RealTime, without requiring model finetuning. The system embeds expert knowledge of six core maqamat, characteristic ornaments, and ensemble dynamics, maintaining realtime responsiveness with subsecond keytoaudibleupdate latency. Empirical evaluations demonstrate that dynamic prompt compilation reliably grounds generation in microtonal scales, significantly increasing offgrid quartertone content over baseline generation. Beyond its core implementation, MazzikaAI illustrates how deterministic knowledgebased rules can effectively bridge expert, nonWestern musical traditions and unfinetuned foundation models. This architecture establishes a scalable paradigm for realtime humanAI cocreation, offering a generalizable blueprint for interactive accompaniment, adaptive music education, and culturally inclusive generative audio across diverse global idioms.


12. MAJEPPA: Morphing and Assessing in a Unified Piano Performance Space

Authors: Jinwen Zhou, Huan Zhang, Weixi Zhai, Jinhua Liang, Aidan O. T. Hogg et al.

Categories: eess.AS, cs.MM Score: 5.60/10 (Obj:8 Id:5 Ind:5 Comp:8 Eff:5 Nov:6 Rep:5)

  • Strength: 贡献了首个覆盖 beginner→virtuoso 全技能谱的 ~4,000 录音标注数据集,并将 LLM-JEPA 非平凡地迁移到 score→performance two-view 表示学习,在 Chopin ranking 和 PISA quality regression 上显著优于 Aria/Moonbeam 等 baseline。
  • Weakness: 生成能力完全未量化评估(仅 demo 网站),ablation 不隔离 data vs objective 贡献,LoRA rank=512 缺乏全量微调对照,Pianist8 上输给 baseline,UMP 绝对 F1 仅 0.16–0.34。
  • Full review: Claude Code 全文七公理审稿

We present MAJEPPA, a self-supervised framework to learn piano performance representations that span the full skill spectrum, from beginner practice sessions to virtuoso concert recordings. We curate the MAJEPPA dataset, comprising ~4,000 annotated recordings across six expertise levels and six recording contexts. We adapt a single pre-trained MIDI autoregressive model with a joint objective: next-token prediction learns score-conditioned performance generation at various skill levels, while InfoNCE and supervised contrastive losses align abstract score and performance representations in a joint embedding space. The proposed model both generates and understands performances in a unified framework. By introducing the EVPMR benchmark, a suite of downstream tasks spanning quality assessment, competition ranking, mistake and technique classification, we evaluate the learnt representations, demonstrating progress towards a real-world model for the piano performance space.


13. Never Stop Speaking: a Denial-of-Service Attack on End-to-End Speech Language Models

Authors: Shuozhe Cheng, Kunlan Xiang, Mingxuan Li, Ji Zhang, Dongxiao Liu et al.

Categories: cs.SD, cs.AI Score: 5.00/10 (Obj:7 Id:5 Ind:4 Comp:4 Eff:5 Nov:5 Rep:5)

  • Strength: 首次针对 E2E 音频大语言模型提出 DoS 攻击,在 3 个开源模型上白盒 ASR 达 83-87%,输出长度从 ~200 增至 ~950 tokens,消融实验清晰表明 EOS 抑制(Leos)是核心机制(移除后 ASR 0.84→0.12)。
  • Weakness: Simple Loss baseline 已达 0.77-0.80 ASR,完整多损失方法的增量提升仅 4-10 个百分点;黑盒迁移性仅 7-13%;缺失与最近语音 DoS 工作 SlothSpeech 的直接比较;关键超参数(ε、PGD 步数)未报告且无代码发布。
  • Full review: Claude Code 全文七公理审稿

Many studies have shown that specially crafted inputs can induce large language models (LLMs) to generate excessively long outputs, resulting in significant computational overhead and resource consumption. While most existing denial-of-service (DoS) attacks target text-only LLMs, end-to-end (E2E) speech LLMs are rapidly emerging. Existing text-based DoS attacks primarily rely on prompt engineering, such as adversarial suffixes or semantic inducement, which exploit the discrete nature of text inputs and therefore cannot be directly transferred to continuous speech inputs. Moreover, prior studies on speech model security mainly focus on ASR or TTS systems, leaving the DoS vulnerability of E2E speech LLMs largely unexplored. To address this gap, we propose the perturbation-based DoS attack targeting E2E speech models. Instead of inducing long outputs through prompt manipulation, our method optimizes imperceptible acoustic perturbations to directly influence the model’s autoregressive generation process while preserving the original input length. Specifically, we formulate the attack as a composite optimization objective that jointly suppresses EOS generation, encourages prolonged decoding, and largely preserves semantic consistency by integrating weighted EOS loss, top-k logit loss, length loss, and semantic alignment loss. To further improve stealthiness, we employ voice activity detection (VAD) to inject perturbations only into voiced regions. Extensive experiments on three open-source E2E speech LLMs demonstrate that our method achieves stable attack success rate while significantly increasing generation length and GPU resource consumption, revealing security risks in modern ALLMs.


14. VoxSumm: A Multilingual Corpus of Long-Form Spoken News for Joint Summarization and Translation

Authors: Yejin Jeon, Marie Maltais, Virginia Ceccatelli, Min Ma, David Ifeoluwa Adelani

Categories: cs.SD, cs.CL Score: 4.95/10 (Obj:6 Id:5 Ind:3 Comp:5 Eff:5 Nov:5 Rep:6)

  • Strength: 构建了首个覆盖 24 语言、703 小时的多语言语音摘要+翻译 benchmark,规模和语言覆盖显著超越已有工作(如 arXiv 2408.06484 仅西-英),并提供了 3 模型 × 3 prompting 策略的系统比较,发现 translate-first 策略加剧 instruction-following 失败。
  • Weakness: 全部语音由 TTS 合成而非真实录制,benchmark 可能测的是 ASR 能力而非语音摘要能力;缺少标准 cascade baseline(ASR+MT+摘要)对比;G-Eval 使用同家族 Gemini 模型做 judge 存在评测循环风险。
  • Full review: Claude Code 全文七公理审稿

As information increasingly traverses linguistic boundaries, users require concise cross-lingual representations of long-form content. Nevertheless, long-document summarization research remains text-centric, whereas multilingual speech research has largely prioritized translation, preserving source content rather than compressing it. We address this methodological gap by formalizing joint speech summarization and translation (JSumT): the generation of a succinct, faithful target-language summary directly from a long spoken document in a source language. We additionally introduce VoxSumm, the first multilingual and cross-lingual benchmark for this task, comprising 10,045 BBC article-summary pairs across 24 languages and encompassing approximately 703 hours of speech data. Our evaluation of representative speech-language models reveals pronounced variation across models and generation settings: Gemini3.1-Pro demonstrates the greatest consistency, summarization into English generally surpasses generation into non-English target languages, and translating an entire document before summarization compounds instruction-following failures. Through the release of VoxSumm, we establish a foundation for developing and evaluating multilingual systems capable of jointly interpreting, compressing, and translating long-form speech.


15. DINO-A: Adapting Self-Distillation Vision Transformers to General Audio Representation Learning

Authors: Tomasz Radzikowski, Mateusz Modrzejewski, Przemysław Rokita

Categories: cs.SD Score: 4.26/10 (Obj:6 Id:5 Ind:8 Comp:3 Eff:3 Nov:2 Rep:6)

  • Strength: 在严格 matched conditions 下确认 canonical DINO 迁移到音频后落后 BYOL-A v2 11.96pp,并系统比较了 patch 分辨率和 CNN/ViT 架构在不同音频任务上的 trade-off。
  • Weakness: 方法为零创新组合(DINO + BYOL-A aug),核心归因假说无消融验证,仅与一个 baseline 比较,遗漏 ATST/EAT 等关键工作,无代码发布声明。
  • Full review: Claude Code 全文七公理审稿

We present DINO-A, an adaptation of self-distillation from vision to general audio representation learning. While DINO has become a canonical method in self-supervised vision and prior audio work has explored latent prediction (BYOL-A) and masked modeling (Audio-MAE, BEATs), no prior work has brought canonical DINO to general audio classification in the way BYOL-A brought BYOL. DINO-A retains DINO’s multi-crop, EMA teacher, and high-dimensional projection, replacing only the input modality and augmentations with log-mel spectrograms and the BYOL-A v2 augmentation block. We pretrain three backbones, two Vision Transformers with 8x8 and 16x16 patches and a convolutional encoder, on FSD50K and evaluate them with linear probing on ESC-50, Speech Commands v2, UrbanSound8K, and GTZAN. Three findings characterize the resulting representations. Patch resolution within the Vision Transformer family has consistent effect on representation quality, with smaller patches winning across all four tasks. The choice between Vision Transformer and convolutional backbone interacts with task type: convolutional networks lead on speech while Vision Transformers lead on environmental sounds and music. Under identical pretraining and evaluation conditions, DINO-A and BYOL-A v2 differ by 11.96 percentage points on average, and we trace this difference to two mechanisms: the interaction between DINO’s high-dimensional projection space and FSD50K’s limited scale, and the additional cost of multi-crop augmentation, which DINO uses but BYOL-A v2 does not. The high-dimensional projection space, central to DINO’s success in vision, becomes a liability at FSD50K scale.