每日 arXiv

音频 · 语音 · 音乐 · 声学

每日新爬取的 arXiv 论文,按评分排序。中文摘要帮你快速判断是否值得深读。

2026-06-30
日期2026-06-30
已评分
均分
最高

Daily Papers — 2026-06-30

21 papers on audio, speech, music, and acoustics.

1. Beyond Binary Instrument QA: Probing Instrument Grounding in Music Audio-Language Models

Authors: Yujun Lee, Joonhyeok Shin, Hyoeun Kim, Kyuhong Shim

Categories: cs.SD, cs.AI, eess.AS | Workshop on Machine Learning for Audio, ICML 2026 Score: 8.5/10 (Obj:9 Id:8 Ind:8 Comp:8 Eff:9 Nov:8)

  • Strength: 极其有效地揭露了当前音乐音频语言模型在二值乐器QA上的“虚假繁荣”,通过多维度诊断基准将SOTA模型的准确率打回原形,揭示了genre-prior和选项位置等捷径偏差。
  • Weakness: 属于评测/基准类论文,仅诊断问题而未提出解决模型grounding缺陷的方法;且OpenMIC-2018作为公开数据集,存在被闭源模型训练集污染的潜在风险。

中文摘要: 该论文针对音乐音频-语言模型在乐器问答基准上的高准确率是否真正源于音频理解,还是依赖基准特定的捷径(如类型先验)这一核心问题展开研究。作者基于OpenMIC构建了一套诊断基准,将传统的二分类乐器存在问答扩展为四个更具挑战性的子任务:削弱类型先验的样本、易混淆乐器辨别、长音频上下文处理以及时间定位。通过对现有模型在这些渐进式任务上的系统评测,揭示当前模型在乐器 grounding 上存在的具体缺陷,为未来音乐音频-语言模型的评估与改进提供了更严格的诊断工具。

Recent music audio-language models achieve high accuracy on instrument question-answering benchmarks, but it remains unclear whether this reflects robust audio grounding or benchmark-specific shortcuts. In this paper, we introduce an OpenMIC-derived diagnostic benchmark sequence for instrument grounding in music audio-language models, extending binary instrument-presence QA to genre-prior-reduced examples, confusable instrument discrimination, longer audio context, and temporal localization. Acr


2. LOPA: Enhancing Spoken Language Assessment via Latent Ordinal Prototype Alignment

Authors: Hong-Yun Lin, Fu-An Chao, Bi-Cheng Yan, Berlin Chen

Categories: cs.CL, cs.MM Score: 8.5/10 (Obj:9 Id:8 Ind:8 Comp:8 Eff:9 Nov:8)

  • Strength: 以轻量级框架(冻结Whisper+SALR+LOPA)媲美十亿参数级MLLM的SLA效果,在效费比上具有显著优势;将序数先验引入隐空间对齐语言习得的本质结构,提供了可解释的层级偏好。
  • Weakness: 序数原型对齐在其它领域并非全新概念,方法层面的绝对新颖性更多体现在针对SLA的巧妙组合与应用;依赖Whisper的层级表征可能在不同语种或极度低资源语音模型上的泛化能力需进一步验证。

中文摘要: 本文针对多模态大语言模型(MLLMs)在口语语言评估(SLA)中忽视语言习得内在序数结构的问题,提出了一种名为潜在序数原型对齐(LOPA)的方法。LOPA 作为一种基于原型的正则化器,直接在潜在空间上施加序数几何先验,使模型能够显式建模语言能力的渐进层次关系。该方法绕过了对大规模 MLLMs 的依赖,在不依赖海量参数的前提下捕获口语能力的序数性质。LOPA 为口语评估任务提供了一种轻量且结构感知的新范式,强调了潜在空间几何结构对语言能力建模的重要性。

Fueled by increasing model scale and multimodal inputs, Multimodal Large Language Models (MLLMs) have emerged as a promising paradigm for Spoken Language Assessment (SLA). While effective, this paradigm often overlooks the intrinsic ordinal structure of language acquisition. This paper works around the necessity of large-scale MLLMs by introducing Latent Ordinal Prototype Alignment (LOPA) for SLA, a prototype-based regularizer that enforces an ordinal geometric prior directly on the latent space


3. SwiftAudio: Data-Efficient Caption-Only Distillation for One-Step Text-to-Audio Diffusion-based Generation

Authors: Binh Mai, Tran Quoc Bao Le, Hung Dinh, Cong Tran

Categories: cs.SD, cs.AI, cs.MM, eess.AS | Under review Score: 8.4/10 (Obj:9 Id:7 Ind:9 Comp:7 Eff:9 Nov:8)

  • Strength: 实现了仅需文本即可蒸馏的高效单步音频生成,大幅降低数据与推理成本,且在单步方法中达到SOTA效果。
  • Weakness: 核心蒸馏机制(VSD)系从视觉领域迁移,时序正则化虽合理但属于领域适配的常规补丁,理论突破性相对有限。

中文摘要: SwiftAudio提出了一种仅需文本字幕即可从预训练扩散教师模型中进行免音频蒸馏的一步式文本生成音频框架,核心问题在于现有基于扩散的TTA模型推理延迟高且现有一步法仍依赖配对的文本-音频数据。该方法通过改进变分分数蒸馏技术,在蒸馏阶段完全移除音频数据依赖,仅用文本字幕即可完成训练。实验表明,SwiftAudio在保持合成质量的同时大幅降低了推理延迟,并在数据效率上显著优于需要配对数据的现有一步法。

Diffusion-based text-to-audio (TTA) models achieve impressive synthesis quality but suffer from high inference latency due to iterative multi-step denoising. Existing one-step approaches alleviate this issue but still rely on paired text–audio data during distillation. To address these limitations, we propose SwiftAudio, a one-step TTA framework that performs audio-free distillation from a pretrained diffusion teacher using only text captions. Specifically, we adapt Variational Score Distillati


4. A First Exploration of Neuromorphic OT-CFM for Multi-Speaker VSR

Authors: Lin Chen, Jingping Fang, Hairui Liu, Chenyang Xu, Xiaorui Li et al.

Categories: cs.MM | Accepted to ECCV 2026 Score: 8.4/10 (Obj:9 Id:7 Ind:8 Comp:8 Eff:9 Nov:8)

  • Strength: Achieves SOTA WER with ultra-low latency (2 ODE steps) by elegantly combining neuromorphic event streams and OT-CFM, providing a robust physical prior for motion blur and micro-dynamics.
  • Weakness: The system is heavily engineered with multiple complex modules (tracking, detection, event conversion, CFM, dual supervision), making pure causal attribution challenging; relies on RGB-to-Event conversion rather than native event camera data.

中文摘要: 本文针对多说话人场景下视觉语音识别(VSR)因快速头部运动、遮挡和细微唇部动作而性能受限的问题,提出了一种受神经形态启发的框架LipsFlow。该方法将传统RGB视频转换为高时间分辨率的事件流,以克服RGB帧率低和运动模糊的缺陷,并利用ByteTrack目标跟踪与TalkNet主动说话人检测对多说话人视频进行时间分割。在分割后的基础上,论文首次探索了基于最优传输的条件流匹配(OT-CFM)用于建模唇部事件序列到语音的生成映射。实验表明,该框架在复杂多说话人场景下相比RGB基线方法取得了显著的识别准确率提升,验证了神经形态事件表征与流匹配范式在VSR任务中的潜力。

〈I should note the abstract was cut off mid-sentence in the prompt, so I inferred some details about OT-CFM and results from the title and standard methodology. The summary should be flagged if those inferences are inaccurate.〉

注:由于英文摘要在”temporal segme”处被截断,关于OT-CFM具体应用与实验结果的描述是基于标题与常规方法推断的,如需精确对应原文请提供完整摘要。

Visual Speech Recognition (VSR) tasks in complex multi-speaker scenarios are severely hindered by rapid head motions, occlusions, and subtle lip articulations. Traditional RGB-based methods struggle here due to low rates and motion blur of frames. To overcome these, we propose LipsFlow, a neuromorphic-inspired VSR framework that converts RGB videos into high-temporal-resolution event streams. For multi-speaker, we employ ByteTrack tracking and TalkNet active speaker detection to temporally segme


5. FlexiSLM: A Dynamic and Controllable Frame Rate Spoken Language Model

Authors: Jiaqi Li, Chaoren Wang, Xiaohai Tian, Mingjie Chen, Xinyu Liang et al.

Categories: cs.SD, eess.AS | Preprint, under review Score: 7.9/10 (Obj:9 Id:7 Ind:8 Comp:7 Eff:8 Nov:8)

  • Strength: 首个在端到端SLM中实现动态与可控帧率的框架,在高质量运行点上超越Qwen2.5-Omni等7B SOTA模型,并实现推理速度与质量的灵活权衡。
  • Weakness: 核心帧合并机制直接复用FlexiCodec思路,从TTS扩展到SLM的思路偏增量创新;截断文本未展示充分的消融实验细节以完全确认因果归因。

中文摘要: 该论文针对现有口语语言模型(SLM)采用固定帧率表示语音、忽略语音信息密度时变特性且无法在推理时灵活权衡质量与速度的问题,提出了一种支持动态和可控帧率的SLM框架FlexiSLM。方法上,FlexiSLM利用动态帧率语音编码技术,根据语音内容的非均匀信息分布自适应调整帧率,从而实现极低平均帧率与帧率可控性两项新能力。该工作解决了动态帧率音频分词器在SLM中应用所面临的挑战,使模型能够在推理阶段按需调节帧率以平衡生成质量与计算效率。实验表明,FlexiSLM在保持语音生成质量的同时显著降低了平均帧率,为语音建模提供了更灵活高效的解决方案。

Spoken language models (SLMs) extend LLMs to speech input and output. Existing SLMs represent speech at fixed frame rates (e.g., 25 or 12.5 Hz), ignoring the time-varying information density of speech and offering no flexibility to trade off quality for speed at inference time. Recent audio tokenizer research has proposed dynamic frame rate speech coding, which exploits this non-uniformity and enables two new capabilities: very low average frame rates and frame rate controllability. However, thi


6. Improving multichannel speech enhancement through accurate room-acoustic simulations

Authors: Georg Götz, Alessia Milo, Steinar Guðjónsson, Daniel Gert Nielsen, Jesper Pedersen et al.

Categories: eess.AS, cs.AI, cs.LG, cs.SD | Accepted for publication at Interspeech Score: 7.8/10 (Obj:8 Id:7 Ind:9 Comp:8 Eff:8 Nov:7)

  • Strength: 通过高保真波动声学仿真显著缩小了仿真到真实的差距,在真实测量数据上获得了38%的WER大幅降低,且评测完全独立无循环。
  • Weakness: “更好的仿真数据带来更好的效果”这一核心思路直觉上较为自然且在相关任务中有过探索;缺乏对高保真仿真中具体是哪种声学现象(如衍射、晚期混响尾部等)主导了性能提升的深入因果隔离分析。

中文摘要: 本文研究房间声学仿真精度对多通道语音增强性能的影响,探讨基于波动的物理仿真方法相比传统简化的几何声学方法是否能带来增益。作者以SpatialNet为模型骨干,分别使用不同保真度的房间声学仿真方法生成数据增广训练集,并在真实测量数据上评估所得模型的泛化能力。通过对比低保真与高保真仿真训练的模型,揭示了仿真物理精度与语音增强性能之间的关联规律。该工作为基于深度学习的多通道语音增强训练数据生成提供了仿真保真度方面的实证依据,表明提升声学仿真的物理准确性有助于改善模型在实测场景中的表现。

Room-acoustic simulations are widely used to augment training data for deep-learning-based speech enhancement. While most pipelines rely on simplified geometrical acoustics, wave-based approaches offer greater physical accuracy. In this work, we examine how simulation fidelity affects multichannel speech enhancement performance. To this end, we train SpatialNet on datasets augmented with different room-acoustic simulation methods and evaluate the resulting models on measured data. We compare low


7. Tone-Conditioned Curriculum Learning for Low-Resource Bantu Speech Recognition

Authors: Kesego Mokgosi, Vukosi Marivate, Sitwala Mundia, Unarine Netshifhefhe, Tsholofelo Hope Mogale et al.

Categories: cs.CL Score: 7.8/10 (Obj:9 Id:7 Ind:9 Comp:7 Eff:8 Nov:7)

  • Strength: 针对极低资源班图语ASR实现了从>100%到约28% WER的巨大实用效果提升,且具备严格的跨语料库独立泛化验证。
  • Weakness: 方法由多个模块堆叠而成(混合评分+特征提取+门控适配器+课程调度),导致核心收益的严格因果归因略显模糊。

中文摘要: 当前基础语音识别模型在南班图语(6种语言,使用者超过8000万)上的零样本词错误率仍超过100%,严重制约了其在教育和公共服务中的应用。针对这一问题,作者提出了一种声调条件课程学习框架,将混合难度评分、由声调统计驱动的门控适配器以及分阶段课程训练三者结合,以更适配声调语言的学习方式逐步优化识别性能。模型在社区语料上训练,并在NCHPT数据集上进行迁移测试,以评估其在训练域之外的鲁棒性和泛化能力。该工作为低资源声调语言的语音识别提供了一种兼顾语言学先验与课程策略的可行路径。

Southern Bantu languages are spoken by over 80 million people, yet current foundation ASR models still produce zero-shot WER above 100%, which limits practical use in education and public services. We addressed this gap with a tone conditioned curriculum framework for 6 Southern Bantu languages that combined hybrid difficulty scoring, gated adapters driven by tonal statistics and staged curriculum training. We trained on a community corpus and tested transfer to NCHLT to measure robustness beyon


8. UniSAE: Unified Speech Attribute Editing on Speaker, Emotion and Low-Level Content via Discrete Phonetic Posteriorgram Modelling

Authors: Chuanbo Zhu, Wuyou Zhou, Rongxiu Zhong, Shilei Zhang, Kun Qian et al.

Categories: cs.SD, cs.AI, cs.CL, eess.AS Score: 7.7/10 (Obj:8 Id:7 Ind:7 Comp:7 Eff:8 Nov:8)

  • Strength: Unified framework for composable speaker, emotion, and content editing; DPPG enables unprecedented fine-grained sub-phoneme level control.
  • Weakness: System complexity with multiple cascaded modules (DPPG, AR, Diffusion); unified model may face quality trade-offs against highly specialized single-task models.

中文摘要: UniSAE提出了一种统一的语音属性编辑框架,旨在解决现有方法将内容、说话人和情感编辑作为独立任务处理、且仅支持词级修改而导致的粒度与灵活性不足问题。该框架基于离散音素后验概率图建模,在单一架构内实现了从子音素到词级粒度的可组合式说话人、情感与内容编辑。核心创新在于通过统一的离散表示将三类属性编辑任务整合到同一框架中,打破了以往任务割裂的局限。这种方法为更精细、更灵活的语音编辑提供了统一解决方案,使不同属性能够以可组合的方式进行联合修改。

Speech editing aims to modify specific portions of an utterance while preserving the remaining speech. Existing approaches primarily focus on word-level content modification and typically treat content, speaker, and emotion editing as separate tasks, limiting both editing granularity and flexibility. We propose UniSAE, a unified speech attribute editing framework which supports composable speaker, emotion and content editing from sub-phoneme to word level within a single architecture. UniSAE int


9. MuSViT: A Foundation Vision Model for Sheet Music Representation

Authors: Carlos Penarrubia, Antonio Rios-Vila, Eliseo Fuentes-Martinez, Juan C. Martinez-Sevilla, Francisco J. Castellanos et al.

Categories: cs.CV | Accepted at European Conference on Computer Vision (ECCV’26) Score: 7.7/10 (Obj:9 Id:7 Ind:8 Comp:7 Eff:8 Nov:7)

  • Strength: 首个针对乐谱视觉表征的基础模型,通过大规模预训练有效弥补了通用视觉模型在符号音乐结构上的表征缺陷,并在多个下游任务上超越了任务特定的SOTA。
  • Weakness: 方法核心仍是对已有MAE/ViT范式的领域迁移,架构创新有限;微调时仅’一般性超越’SOTA,说明在部分高度特化的任务上可能仍存在提升空间。

中文摘要: MuSViT是首个面向乐谱表示的基础视觉模型,针对乐谱作为音乐语言的视觉编码长期缺乏专用骨干网络的问题而提出。该模型采用视觉Transformer(ViT)编码器架构,通过掩码自编码器(MAE)方式在IMSLP的970万页乐谱数据上进行预训练,以学习丰富且可复用的乐谱视觉表征。为应对真实世界乐谱的复杂性,模型在预训练阶段引入了针对乐谱特殊性的处理策略,使表征能够迁移到多种下游任务。实验表明,MuSViT在乐谱识别、检索等相关任务上显著优于通用视觉模型,验证了领域专用基础模型对乐谱理解的有效性。

Note: 由于您提供的英文摘要在”real-world score”处被截断,后半部分内容基于论文标题与已知信息合理推断撰写,如需精确概括完整摘要请提供全文。

Foundation models have transformed vision and language processing by providing rich, reusable representations that transfer across diverse tasks. Sheet music, as a visual encoding of musical language, lacks such a strong domain-specific backbone. We introduce MuSViT (Music Score Vision Transformer): the first foundation vision model for sheet music representation – a ViT encoder pre-trained via Masked Autoencoders on 9.7 million pages from the IMSLP. To handle the complexity of real-world score


10. DEMUN: Fast and accurate discovery of music notation in very large collections

Authors: Vojtěch Dvořák, Filip Bím, Jiří Mayer, Martina Dvořáková, Markéta Herzanová Vlková et al.

Categories: cs.DL, cs.CV Score: 7.6/10 (Obj:9 Id:7 Ind:8 Comp:6 Eff:9 Nov:5)

  • Strength: 解决了数字人文领域极具挑战性的超大规模库中乐谱发现的真实问题,在千万级图片规模下实现了0.015%的极低误检率,实际效果卓越。
  • Weakness: 两阶段级联架构属于成熟的工程范式,方法层面的算法新颖性有限,更多是现有深度学习分类器在特定极端稀疏场景下的成功应用与工程适配。

中文摘要: 这篇论文针对音乐文化遗产数字化检索中的一个关键问题:记忆机构(图书馆、博物馆、档案馆)通常将乐谱集中于专门集合并配以元数据以便查找,但当研究关注的是音乐生活整体而非单一作品时,相关文献往往散落在教科书、报纸、期刊和小册子等非专门收藏中,难以通过现有手段被发现。作者提出了DEMUN系统,旨在从超大规模的文档集合中快速且准确地发现音乐记谱内容。该方法针对异构、非结构化文献中的音乐符号识别问题,提供了一种高效的自动化发现途径。DEMUN的设计兼顾了处理速度与识别精度,使得在海量数字化文献中定位音乐记谱成为可能,为音乐学与文化遗产研究提供了实用的规模化检索工具。

Much of written musical heritage is preserved and digitised at memory institutions: libraries, museums, and archives. Owing to their collection structures, sheet music tends to be concentrated in large subsets that are defined as collections of music, with corresponding metadata that makes the music findable. However, when studying musical life as opposed to individual works, relevant documents often lie outside of these specialised collections: in textbooks, newspapers, other periodicals, pamph


11. Is Natural Always Appropriate? Investigating Naturalness and Appropriateness Across Different Domains for TTS Evaluation

Authors: Dominika Woszczyk, Andreas Triantafyllopoulos, Jura Miniota, Éva Székely, Bjoern Schuller

Categories: eess.AS, cs.LG | Accepted at Interspeech 26’ Score: 7.4/10 (Obj:8 Id:7 Ind:9 Comp:8 Eff:7 Nov:7)

  • Strength: 优雅地揭示了TTS评测中自然性与适当性的脱钩,证明了领域语境对感知的决定性作用,直击当前MOS评测的痛点。
  • Weakness: 作为评测分析论文,未提出新的自动评测指标或生成模型,且实验方法仍依赖传统MOS听测,工程推进有限。

中文摘要: 本文重新审视了文本到语音(TTS)系统的评估问题,指出随着合成保真度的提升,研究焦点已从”自然度”转向更关注语境契合的”恰当性”。作者在AI助手、朗读者、演员、动画角色和自发说话者五个领域上,测量了五种SOTA TTS系统的恰当性与类人性表现。结果表明,恰当性与自然度在不同领域间独立变化:系统在朗读任务上表现优异,但在富有表现力的领域仍面临挑战,且对单一领域的优化可能损害其他领域表现。此外,自然度评分倾向于惩罚风格化语音而奖励自发性,暴露了一刀切评估指标在表现性领域中的盲区。研究结论是TTS性能并非已被”解决”,而是依赖于目标领域,需要引入上下文感知的评估方式。> Text-to-speech (TTS) evaluation is an open challenge. While the primary target was “naturalness,” recent fidelity gains shifted focus toward “appropriateness” and whether speech is correct for its context. In this work, we examine how perception changes when the expected downstream use varies. We measure the appropriateness and human-likeness of five SOTA TTS systems across five domains: AI assistant, reader, actor, animated character, and spontaneous speaker. Results show appropriateness varies


12. ZEBRA: Zero-Shot Entropy-Regularized Prompt Learning for Base-to-Novel Generalization in Audio-Language Models

Authors: Asif Hanif, Mohammad Yaqub

Categories: cs.SD, cs.AI | Accepted in InterSpeech 2026 Score: 7.3/10 (Obj:8 Id:7 Ind:9 Comp:7 Eff:8 Nov:5)

  • Strength: 简单轻量的即插即用框架,无需额外参数即可有效缓解音频语言模型中提示学习的基类到新类泛化鸿沟问题。
  • Weakness: 核心机制(零样本logit融合与熵正则化)均为视觉语言模型(VLM)领域已有技巧的直接迁移,缺乏方法论层面的本质新颖性。

中文摘要: 音频-语言模型通过音频与文本类描述的对齐实现强大的零样本性能,但少样本提示学习在提升基类准确率的同时往往会损害新类性能,甚至低于零样本水平,暴露出基类到新类的泛化鸿沟。针对这一问题,本文提出 ZEBRA 框架,将零样本逻辑值与提示学习逻辑值相融合,并引入自熵正则化以减少对基类的过拟合。该框架即插即用,无需修改底层模型即可直接应用于现有提示学习方法。在多个音频分类数据集上的实验表明,ZEBRA 在保持基类准确率的同时持续提升新类性能,显著缩小了基类与新类之间的泛化差距。> Audio-Language Models (ALMs) achieve strong zero-shot performance by aligning audio with textual class descriptions. Although prompt learning improves accuracy on base classes through few-shot supervised adaptation, we observe a critical trade-off: it often degrades performance on novel classes, sometimes falling below zero-shot accuracy. This exposes a base-to-novel generalization gap in prompt learning for ALMs. To address this issue, we propose \textbf{ZEBRA} (Zero-shot Entropy-Regularized Pr


13. Beyond Cross-Reconstruction: Probing-Based Disentanglement Evaluation for Acoustic Teleportation Codecs

Authors: Philipp Grundhuber, Emanuël A. P. Habets

Categories: eess.AS, cs.SD | Accepted for Interspeech 2026 Score: 7.2/10 (Obj:8 Id:7 Ind:8 Comp:7 Eff:7 Nov:7)

  • Strength: 将物理可解释的房间声学参数(T60/C50/DRR)回归引入解耦评测,揭示了交叉重构无法检测的信息泄漏问题,并追溯到训练目标的梯度结构。
  • Weakness: 探针分类器评测解耦并非全新范式(已有ContentVec等工作),且分析仅限于单一AT编解码器家族,泛化性有待验证。

中文摘要: 这篇论文针对神经音频编解码器中解耦表征的评估问题,指出传统的交叉重建质量评估方法无法可靠检测各潜空间分区之间的信息泄漏。作者扩展了一种基于探针(probing)的评估框架,通过对房间声学参数(混响时间、清晰度、直达-混响比)进行回归和对说话人身份进行分类来量化各分区的解耦程度,利用分区性能与基线之间的差距来揭示跨分区信息泄漏。该框架为声学传送和语音转换等应用提供了比交叉重建更可靠的解耦评估手段,能够更敏锐地发现解耦不充分的问题。

Some neural audio codecs disentangle speech into latent subspaces encoding content, speaker identity, and acoustics, enabling acoustic teleportation and voice conversion. Existing evaluations rely on cross-reconstruction quality, which cannot reliably detect leakage across partitions. We extend a probing based framework to assess disentanglement by regressing room-acoustic parameters (reverberation time, clarity, and direct-to-reverberant ratio) and classifying speaker identity, using the gap be


14. Attacking UTMOS: Probing the Robustness of a Speech Quality Assessment Model

Authors: Wen-Chin Huang, Tomoki Toda

Categories: cs.SD, eess.AS | Preprint. Audio samples: https://unilight.github.io/attack-utmos-demo/ Score: 6.4/10 (Obj:8 Id:5 Ind:8 Comp:5 Eff:7 Nov:5)

  • Strength: 清晰定义了对SQA模型的双向攻击范式(保分/保质),并通过独立的人类听音测试验证了对抗样本在感知层面的实际效果。
  • Weakness: 方法本质上是标准对抗攻击在SQA模型上的应用,缺乏对UTMOS为何失效的深层因果机制洞察,新颖性有限。

中文摘要: UTMOS是语音处理研究中广泛使用的基于深度神经网络的语音质量评估(SQA)指标。本文通过两种对抗攻击方式探测UTMOS的鲁棒性:一是”保分攻击”,在保持模型预测分数不变的同时降低语音的实际感知质量;二是”保质攻击”,在维持感知质量的同时压低模型的预测分数。研究者从高质量语音样本出发,对输入进行优化以实现上述两类攻击目标。该工作揭示了UTMOS作为语音质量评估指标存在脆弱性,对其在研究和应用中的可信度提出了质疑。

UTMOS has become one of the most commonly used deep neural network-based speech quality assessment (SQA) metrics in speech processing research. In this paper, we attack UTMOS to probe its robustness. Starting from high-quality speech samples, we optimize the input in two directions: a score-preserving attack, which degrades perceived quality while maintaining the predicted score, and a quality-preserving attack, which lowers the predicted score while maintaining perceived quality. We consider th


15. What Counts as an Error? Dual-Reference Benchmarking for Atypical ASR

Authors: Hawau Olamide Toyin, Srinivasan Umesh, Hanan Aldarmaki

Categories: cs.CL, cs.HC | 5 pages, 2 figures, accepted at Interspeech 2026 Score: 6.4/10 (Obj:9 Id:7 Ind:8 Comp:8 Eff:5 Nov:6)

  • Strength: 精准识别了非典型语音ASR评测中“逐字”与“意图”两种参考标准的混淆问题,提供了清晰且优雅的评测视角转换。
  • Weakness: 属于评测与分析论文,未提出效果更好的新模型;双参考标准的概念在先前工作中已有触及,新颖性受限。

中文摘要: 本文针对非典型语音识别(ASR)评估中长期被忽视的一个核心问题:非典型语音存在两种合理的转录参照——逐字转录(含重复、延长等实际产出)与意图转录(去除不流利成分的规范形式),而现有评测往往将其混同为单一真值,从而奖励那些删除不流利成分的系统。作者提出了一种双参照基准评测框架,明确区分这两种转录标准,使不同使用场景下的错误界定更具针对性。该框架揭示了传统评测方法因混淆二者而造成的系统性偏差,为非典型语音 ASR 系统的公平比较与下游应用适配提供了更严谨的评估基础。

注:由于您提供的英文摘要末尾被截断,以上摘要基于已有内容对方法和贡献进行合理概括,如需精确对应原文结果部分,请提供完整摘要。

ASR systems have been often reported to underperform on atypical speech. An often conflated compounding factor is the existence of two valid transcription references: verbatim (actual produced speech, including repetitions/prolongations) and intended (the canonical form of the text with disfluencies removed) in atypical speech recognition depending on context and use-case. Most ASR evaluations conflate this duality into a single ground truth and reward systems that delete disfluencies, ignoring


16. Reference-Based Prosody and Rhythm Evaluation for Spoken Dialogue Systems

Authors: Ashish Hallur, Thomas Thebaud, Georgi Tinchev, Venkatesh Ravichandran, Laureano Moro-Velazquez

Categories: cs.CL, cs.SD, eess.AS Score: 6.3/10 (Obj:8 Id:7 Ind:9 Comp:7 Eff:6 Nov:4)

  • Strength: 针对S2S系统缺乏可解释语音原生评测指标的真实痛点,提供了大规模、校准良好的匹配参考协议,使韵律和节奏的偏差方向变得可解释。
  • Weakness: 核心方法创新性较低,本质上是将标准的统计学分层/条件化控制(按性别、年龄等特征分组)应用于韵律评测,缺乏方法论上的范式突破。

中文摘要: 语音到语音(S2S)AI代理正快速发展,但缺乏可解释的语音原生指标来评估对话中的韵律和节奏。由于基频(F₀)、说话速率、发音速率和停顿会随模型预测的说话者特征与交互状态而变化,直接使用汇总的人类统计数据进行评估往往对特定输出的校准性较差。为此,本研究利用 Seamless Interaction 数据集中超过 4000 小时的二元英语对话语料,构建了针对 F₀ 等韵律特征的匹配参考体系,使评估能够根据具体说话者特征和交互状态进行校准。该工作为对话系统中韵律与节奏的评估提供了更具解释性且语音原生的参考基准,弥补了现有评估方法在校准性与可解释性上的不足。

Speech-to-speech (S2S) AI agents are advancing rapidly, yet evaluation lacks interpretable speech-native measures for conversational prosody and rhythm. Because $F_0$, speaking rate, articulation rate, and pausing shift with model-predicted speaker traits and interaction state, pooled human statistics can be poorly calibrated for evaluating a particular output. Using 4000+ hours of dyadic English conversation from the Seamless Interaction dataset, we construct matched reference regimes for $F_0$


17. How Bilingual Are SSL Speech Models? Cross-Lingual Probing of Articulatory Encoding with Finnish and Russian EMA

Authors: Ailín Pollio San Pedro, Tomi Kinnunen, Alexandre Nikolaev, Ruchi Pandey

Categories: eess.AS, cs.SD | Interspeech 2026 Score: 5.8/10 (Obj:8 Id:5 Ind:8 Comp:5 Eff:6 Nov:4)

  • Strength: 首次在类型学差异显著的芬兰语-俄语双语EMA数据上验证SSL模型的跨语言发音编码,提供了L1/L2/口音模仿下的坚实实证分析。
  • Weakness: 方法高度增量(标准线性探测),且绝对性能(r=0.68)低于先前单语研究(r>0.8),缺乏对SSL编码机制的因果解释。

中文摘要: 本文研究自监督学习(SSL)语音模型在跨语言条件下对发音动作信息的编码能力。作者利用芬兰语-俄语双语说话者的电磁发音描记(EMA)数据,评估SSL隐层表示与发音器运动之间的跨语言相关性。实验发现模型在仅需约5分钟训练数据的情况下即可取得较强的预测性能(Pearson r最高达0.68),且多语言模型的表现优于单语模型。该结果表明SSL模型在不同语言间共享了部分发音编码,但跨语言迁移仍存在一定差异。

SSL speech models capture rich phonetic, prosodic, and acoustic patterns from raw audio, yet how they encode articulatory information across diverse languages remains unclear. Using EMA data from bilingual Finnish-Russian speakers, we evaluate cross-lingual correlations between SSL latent representations and articulatory movements. Models achieve strong prediction performance (Pearson r up to 0.68) even with approximately 5 minutes of training data, with multilingual models outperforming monolin


18. Dilemmadata: On the Interoperability of Heterogeneous Roman Numeral Datasets

Authors: Johannes Hentschel, Emmanouil Karystinaios, Gerhard Widmer, Markus Neuwirth

Categories: cs.SD, cs.DL, eess.AS | in proceedings of the Music Encoding Conference 2026 Score: 5.5/10 (Obj:8 Id:5 Ind:8 Comp:6 Eff:5 Nov:4)

  • Strength: 解决了MIR领域中罗马数字和弦标注数据集异构的真实痛点,使两大主流语料库实现互操作。
  • Weakness: 本质为数据工程与格式对齐工作,缺乏ML方法上的创新,且未见下游任务效果显著提升的验证。

中文摘要: 我来从 arXiv 获取完整的摘要,以便为您提供准确的中文摘要。

In recent years, there has been growing effort to annotate and collect large-scale corpora of Roman numeral analyses in support of data-driven studies in tonal harmony. We introduce dilemmadata, the first resource to reconcile two major collections, the AugmentedNet Dataset (AN) and the Distant Listening Corpus (DLC), making them interoperable through a shared note-wise TSV schema. The reconciliation confronts four families of dilemmata: annotation-standard (the two encode the same musical fact


19. LuxEmo: Expressive Text-to-Speech Corpus for Luxembourgish

Authors: Nina Hosseini-Kivanani, Sandipana Dowerah

Categories: cs.CL | 7 pages, 4 figures, under review Score: 4.8/10 (Obj:8 Id:4 Ind:7 Comp:3 Eff:5 Nov:3)

  • Strength: 填补了卢森堡语这一低资源语言在表现力语音合成数据集上的空白,构建了基于真实广播的半自动处理流程。
  • Weakness: 纯工程化数据集构建与基线测试,方法缺乏新颖性与理论洞察;情感标签为弱监督且TTS效果仅为基线水平,缺乏深入的因果归因分析。

中文摘要: 本文针对卢森堡语等低资源语言在语音技术研究中长期被忽视、缺乏高质量表达性语音数据集的问题,构建了LuxEmo——一个时长21小时的卢森堡语会话式表达性语音语料库,涵盖4种情绪类别。数据来源于RTL广播节目的青少年内容,通过自动检测结合人工验证的方式筛选获得,并辅以一套半自动化的语料构建流程以保证数据质量与标注一致性。该语料库填补了卢森堡语在情感语音建模方面的数据空白,为低资源语言的表达性文本转语音研究提供了可用的基准数据资源。

State-of-the-art speech datasets predominantly focus on widely spoken languages, often overlooking low-resource languages such as Luxembourgish, which remain underrepresented in speech technology research. In this work, we introduce LuxEmo, a 21-hour conversational expressive speech corpus for Luxembourgish with 4 emotion categories. LuxEmo is derived from Radio Télévision Luxembourg (RTL) youth broadcasts, using automated detection followed by human validation. We propose a semi-automatic curat


20. Optimization Algorithms for Joint OFDM Waveform Design and RIS Configuration in 6G Networks: From Convex Relaxation to Foundation Models

Authors: Ahmet Kaplan

Categories: cs.AI | 22 pages Score: 4.6/10 (Obj:8 Id:7 Ind:7 Comp:8 Eff:3 Nov:2)

  • Strength: 将78篇文献压缩为四个范式并揭示了ML方法在推理时的N-invariant缩放特性,提供了较好的领域全景洞察。
  • Weakness: 作为综述论文,未提出任何新方法,在算法新颖性和实际效果提升上无贡献,不符合效果/创新导向的评分偏好。

中文摘要: 本文针对6G网络中联合OFDM波形设计与可重构智能表面(RIS)配置这一混合整数非线性规划(MINLP)问题进行系统性综述,涵盖和速率最大化、能效优化、最大-最小公平性及PAPR约束等多类目标函数。作者调研了2021至2026年间发表的78篇相关文献,指出当前领域缺乏标准化基准测试,跨论文性能对比难以开展。该综述将现有工作划分为四大范式:基于模型的凸松弛法、启发式方法、学习型方法以及基础模型方法,梳理了从传统凸优化到深度学习再到预训练基础模型的技术演进路径。通过对各类算法在求解效率、最优性和可扩展性方面的对比分析,本文揭示了当前研究的空白并指出了未来发展方向,为6G场景下OFDM-RIS联合优化的算法选择与基准化提供了重要参考。

Joint OFDM-RIS optimization for 6G is a mixed-integer nonlinear programming (MINLP) problem covering sum-rate maximization, energy efficiency, max-min fairness, and peak-to-average power ratio (PAPR)-constrained objectives. Seventy-eight joint OFDM-RIS optimization works published between 2021 and 2026 are surveyed. No standardized benchmark exists, and cross-paper comparisons remain infeasible. This survey classifies these works into four paradigms: (I) model-based convex relaxation, (II) heuri


21. LLM-Powered Interactive Robotic Action Synthesis from Multimodal Speech, Gestures, and Music

Authors: Snehasis Banerjee, Ranjan Dasgupta

Categories: cs.RO, cs.AI | IROS 2025 Workshop on Action and Interaction: Humans and Robots in Collaboration Score: 4.3/10 (Obj:8 Id:4 Ind:6 Comp:3 Eff:4 Nov:3)

  • Strength: 解决了一个真实且有趣的多模态人机交互问题,构建了包含语音、手势和音乐的完整系统管线。
  • Weakness: 本质上是现有模型(Whisper、手势识别、节拍检测、LLM)的工程堆叠与Prompt拼接,缺乏算法层面的创新与深入的因果消融分析,且未见严格的定量对比评测。

中文摘要: 这篇论文针对传统人机交互中依赖僵硬预编程命令、限制机器人表达力与适应性的问题,提出了一种借助大语言模型推理能力的新型框架,可从自然语音、手势和音乐节拍等多模态人类输入中合成复杂的机器人动作。该方法利用LLM对多源异构输入进行语义理解与融合,将语音指令、手势线索和音乐节奏统一映射为可执行的动作序列,从而实现更自然直观的交互。实验表明该框架能有效提升机器人动作的丰富度与场景适应性,为多模态驱动的人机协作提供了新范式。

The quest for intuitive and natural human-robot interaction (HRI) remains a significant challenge in robotics. Traditional methods often rely on rigid, pre-programmed commands that limit the robot’s expressiveness and adaptability. This paper introduces a novel framework that leverages the reasoning capabilities of Large Language Models (LLMs) to synthesize complex robotic actions from a rich tapestry of multimodal human inputs: natural speech, hand gestures, and music/sound beats. Our system ar