每日 arXiv

音频 · 语音 · 音乐 · 声学

每日新爬取的 arXiv 论文,按评分排序。中文摘要帮你快速判断是否值得深读。

2026-08-16
日期2026-08-16
已评分
均分
最高

Daily Papers — 2026-08-16

9 papers on audio, speech, music, and acoustics.

1. Adding Voice Cloning to Text-to-Audio-Video Models with a Single Zero-Initialised Layer

Authors: Ivan Mikheev, Viacheslav Vasilev, Anna Dmitrienko, Alexey Letunovskiy, Ivan Kirillov et al.

Categories: cs.SD, cs.AI, cs.LG, cs.MM Score: 7.70/10 (Obj:9 Id:8 Ind:9 Comp:8 Eff:9 Nov:7 Rep:3)

  • Strength: 零初始化单层配方把 5B T2AV 模型变成语音克隆器,674 样本 VCTK 基准上六列 SECS 全最高(WavLM vs-ref 0.944)、对最强外部基线三网平均领先 0.041(Wilcoxon p<10^-89),音频独走推理约 30× 加速。
  • Weakness: k6a5b 的 WER 5.76% 对 Qwen3-TTS 的 0%(WER0 86.6% 对 95.7%),智能度明显落后专用 TTS,且代码、微调权重与基准划分均未释放,缺 GPU 小时数等训练开销证据。
  • Full review: Claude Code 全文七公理审稿

Text-to-audio-video (T2AV) generation models produce a video and its soundtrack from a textual description, but offer no control over whose voice speaks in the output. We show that a base T2AV model can be turned into a voice-cloning model by adding a single zero-initialized linear layer on top of its audio backbone, fine-tuning for a comparatively short training schedule, and conditioning on a short reference recording at inference time. The reference is injected through two complementary signals: its diffusion latents are prepended to the audio stream, and a global speaker embedding modulates token of the target audio. On a benchmark of 674 speaker-text pairs spanning 30 speakers we compare against five strong voice-cloning text-to-speech baselines: our enhanced 5B model attains the highest speaker-encoder cosine similarity (SECS) across three independent verification networks (ECAPA-TDNN, WavLM-SV, Resemblyzer), statistically significantly outperforming every baseline. A side product of the architecture is that the audio path can be evaluated without the video path at inference time, yielding a ~30x speed-up over the full audio-video diffusion loop while preserving the voice-cloning behaviour.


2. ARENA: Automated Red-Teaming for Large Audio Language Models

Authors: Jiaming He, Zhicong Huang, Tian Jin, Zhen Sun, Cheng Hong et al.

Categories: cs.SD, cs.AI Score: 6.95/10 (Obj:8 Id:6 Ind:8 Comp:5 Eff:8 Nov:7 Rep:6)

  • Strength: ARENA 在 520 条 held-out AdvBench 目标上对 AF3/Qwen2-Audio/MiMo-Audio/GPT-Audio 实现 FDR 87.9/71.5/68.1/75.4%,把静态基线 AJailBench(最高 31.2%)与 JALMBench(最高 12.6%)提升约 3-6 倍,PSR 达 96.3-100%。
  • Weakness: 缺少文本-only 无音频的自适应红队对照来量化音频通道边际贡献,成功判定全部依赖单一 Llama Guard 3 评估器且无人工核验,训练控制器权重与 2,000 例种子池数据也未随代码仓库发布。
  • Full review: Claude Code 全文七公理审稿

Large audio-language models (LALMs) make it possible to interact with language models through speech, music, and environmental sound, but they also introduce a safety surface that is difficult to expose with text-only red-teaming. We study automated audio-grounded red-teaming, where a text query must remain safe in isolation while the joint text-audio input induces harmful target behavior. We propose ARENA, a closed-loop framework that trains a controller on an independent 2,000case text-audio dataset. MD-Judge supplies training rewards and adaptive search feedback, while a separate, non-adaptive Llama Guard 3 evaluator alone labels final outcomes. On 520 held-out AdvBench objectives, ARENA achieves FDR/PSR of 87.9/100.0%, 71.5/96.3%, 68.1/100.0%, and 75.4/98.5% on Audio Flamingo 3, Qwen2-Audio, MiMo-Audio, and GPTAudio, respectively. Ablations show that feedback-based refinement and audio-variant search substantially improve attack discovery.


3. VARM-Bench: Benchmarking Verifiable Structured Reasoning in Chinese Abusive Speech Moderation

Authors: Mingyu Yuan, Shengtao Wen, Lingbing Guo, Zhen Bi, Xiang Chen

Categories: cs.AI Score: 6.84/10 (Obj:8 Id:6 Ind:6 Comp:7 Eff:7 Nov:7 Rep:7)

  • Strength: 在 8,000 条中文有害言论审核基准上证明标签级准确率可掩盖完整记录错误,GPT-5.5 Label 97.6% 时 JREM 仅 55.4%、38.1% 的类别正确输出仍含记录错误,且 Shapley 分解将 76.6–80.6% 的差距归因于 referent 识别。
  • Weakness: 完整复现缺失——代码库排除全部模型预测文件与审计工件、闭源模型名(GPT-5.5、Qwen3.7-Max、DeepSeek-V4-Pro)未披露版本,且缺少词典最小基线与既有中文基准(COLD、STATE-ToxiCN)的迁移对照。
  • Full review: Claude Code 全文七公理审稿

The widespread circulation of abusive online content has increased the need for reliable moderation of Chinese social-media text. Existing Chinese benchmarks support label classification, fine-grained toxicity categorization, and target-aware extraction, but do not provide a unified representation for deterministically verifying the stated basis of a moderation decision. We introduce VARM-Bench, a benchmark for field-anchored chain-of-thought rationales in Chinese abusive-speech moderation. Each instance contains a concise natural-language rationale with explicit anchors for six decisions: target, target type, target explicitness, author stance, harmfulness label, and fine-grained category. Our deterministic protocol evaluates field correctness, target alignment, output validity, complete-record agreement, and hidden record errors conditioned on correct final decisions, without relying on an LLM judge. Under a common structured-output protocol, we evaluate language models across multiple model families using zero-shot prompting, taxonomy guidance, and structured CoT supervision, and analyze lexical-cue sensitivity and field-level errors. Results show that strong label-level performance can conceal substantial errors in complete moderation records. VARM-Bench provides an auditable and reproducible benchmark for evaluating verifiable moderation rationales in Chinese abusive-speech moderation.


4. CineDub: Scaling End-to-End Video Dubbing to Multi-Speaker Dialogues with Coherent Sound Effects

Authors: Yusheng Dai, Kangdi Wang, Baolong Gao, Yuxuan Jiang, Weiqiang Wang et al.

Categories: eess.AS, cs.MM, cs.SD | Accepted to ACM MM 2026 Score: 6.79/10 (Obj:8 Id:7 Ind:6 Comp:7 Eff:7 Nov:7 Rep:5)

  • Strength: 在自建的 CineDub-Multi 多说话人基准上,CineDub 把 cpWER 从最强基线 FunCineForge 的 43.47% 降到 13.93%(相对错误率降低 68.0%),并在 V2SA 联合生成上以 FDVGG 1.24、IS 4.00 全面超越级联管线(2.13–2.95 / 3.01–3.29)。
  • Weakness: 模型代码与权重、Gemini 2.5 Pro 标注管线均未发布,评测 Desync 指标与视觉条件同源于 SynchFormer、cpWER 依赖 Gemini 转写,且 CineDub-Multi 仅 139 条样本、全程无人类听感评测。
  • Full review: Claude Code 全文七公理审稿

Automatic video dubbing in the wild remains fundamentally limited by two competing constraints: hierarchical methods depend on brittle, multi-stage preprocessing pipelines that severely restrict data scalability and practical deployment, while holistic approaches operating on uncropped video suffer from weak temporal alignment and speaker-utterance ambiguity in multi-speaker settings. To overcome these limitations, we propose CineDub, a unified diffusion-based model that achieves precise multi-speaker dialogue dubbing directly from uncropped videos, without face cropping or speaker diarization. Central to our approach is the Implicitly-Coupled Holistic Conditioning (ICHC) paradigm, where holistic visual representations and a semantic-bundled transcription format are encoded independently, yet implicitly coupled through cross-modal training to resolve speaker ambiguity and enable precise multi-speaker multi-turn dialogue dubbing. Building on the unified temporal cues captured by holistic visual features, we further extend CineDub to joint speech and audio generation. We introduce an Ambient-to-Linguistic Curriculum Learning (ALC) to mitigate sub-task degradation, and a decoupled textual branch control mechanism to resolve cross-prompt interference during simultaneous generation. We also release two in-the-wild benchmarks, CineDub-Multi for multi-speaker dialogue dubbing and CineDub-SA for video-to-speech-and-audio (V2SA) generation, to enable evaluation under realistic conditions. Experiments show that CineDub achieves state-of-the-art results on established single-speaker dubbing and video-to-audio benchmarks while excelling in multi-speaker dialogue dubbing and acoustically coherent joint generation.


5. The Null Token Knows: Reducing Message-Free Hallucination in ASR and NMT

Authors: Kirill Borodin, Vasiliy Kudryavtsev, Ivan Viakhirev, Grach Mkrtchian

Categories: cs.CL, cs.LG, cs.SD | Submitted to the Thirty-Ninth AAAI Conference on Artificial Intelligence (AAAI-27) Score: 6.68/10 (Obj:8 Id:9 Ind:8 Comp:7 Eff:6 Nov:5 Rep:5)

  • Strength: 在预冻结的源分离 300/300 锁定集上,trained EOT row 把 LOR 从 91.7% 降到 3.3%,且 Canary 在 β=8 处 LOR 79.3→0.0%、WER 7.02→6.94 近乎无损,同时提出 Cκ= F+κD 双代价评测框架。
  • Weakness: 代码与权重未发布(无 artifact locator)、锁定评测仅覆盖 Whisper-small 且未在锁定集上解码外部门控基线,干预方法在 Cκ 网格上从不优于 stock 或标量 bias。
  • Full review: Claude Code 全文七公理审稿

Modern encoder-decoder systems can produce fluent text even when their input contains no recoverable message. We study this failure in ASR and NMT through the models’ reserved null tokens, asking whether the score for ending generation already carries a usable abstention signal. Across speech recognizers and translation models, we audit native null-token scores and scalar logit shifts. In Whisper, we additionally probe decoder states and compare supervised row edits with conventional external gates. The evaluated models often expose a useful abstention signal, but stock decoding does not reliably act on it. Raising the null-token score can sharply suppress fabrication, but aggressive intervention also deletes valid speech or shortens legitimate translations. These findings turn the null token into a diagnostic lens on hallucination and motivate evaluating abstention methods by both suppression and deletion costs, rather than by hallucination reduction alone.


6. Iterative Self-Learning for Expressive Text-to-Speech Synthesis

Authors: Nicholas Sanders, Gustav Eje Henter, Simon King, Korin Richmond

Categories: eess.AS, cs.CL, cs.SD Score: 6.58/10 (Obj:9 Id:5 Ind:6 Comp:8 Eff:8 Nov:6 Rep:4)

  • Strength: ISL 在 5% ESD 切分把 Emo2vec 情绪遵循 F1 从单轮控制的 39.38 提升至 49.58(反超 100% GT 的 45.95),Prominence 1% 主观听测偏好 win prob 0.678±0.078,伪标签 F1 在 1% Emotion 切分相对提升超 20%。
  • Weakness: 没有发布任何代码或权重(GitHub 检索 invert-classify/IPL-TTS 均为 0 结果),且未与任何外部 classifier-based 标注系统或现有半监督 TTS 做受控对比。
  • Full review: Claude Code 全文七公理审稿

Expressive text-to-speech (TTS) systems that use explicit conditioning labels provide direct and interpretable control over expressive attributes, in contrast to reference-based or prompting-based approaches, but require labeled data. Obtaining these labels at scale is costly and time-consuming, yet no prior semi-supervised framework addresses this specific bottleneck. Existing semi-supervised TTS methods instead target scarcity of paired speech-text data or transcriptions. To address the scarcity of expressive labels, we propose an Iterative Self-Learning (ISL) framework for expressive TTS, built on Invert-Classify, a classifier-free method that recovers discrete expressive labels by inverting a frozen generative model. The framework iteratively pseudo-labels unlabeled speech using the current model, retrains on the combined labeled and pseudo-labeled data, and repeats, progressively refining label quality and synthesis. We validate on two expressive tasks, word-level prominence and utterance-level emotion, across multiple low-resource data splits. We find that iterative refinement can improve pseudo-label accuracy over single-pass baselines. Furthermore, we observe that these improvements in pseudo-labeling of expressivity translate to gains in expressive label adherence and synthesis quality, confirmed by objective metrics and human listening tests. In the most data-scarce conditions, ISL-trained models outperform single-pass pseudo-labeling and further approach fully supervised performance, demonstrating that gradient-based ISL is an effective solution to expressive label scarcity in low-resource TTS.


7. Efficient Audio-Visual Generation via Synchrony-Aware Cross-Modal Sparse Attention

Authors: Shengchuan Gao, Teng Hu, Bohao Feng, Luchen Li, Wenqiang Wang et al.

Categories: cs.CV Score: 6.58/10 (Obj:8 Id:7 Ind:6 Comp:6 Eff:8 Nov:6 Rep:4)

  • Strength: 在 Ovi 骨干上以 1.992× 加速(448.75s→225.30s)同时取得所有加速方法中最佳的视频质量(PSNR 25.69)、音频质量(VISQOL 3.6902)与音视频同步(Sync-C 4.606)。
  • Weakness: 未发布代码与评测样本,且 LTX-2.3 骨干上加速仅 1.281×,跨骨干通用性与可复现证据缺失。
  • Full review: Claude Code 全文七公理审稿

Recent audio-visual generation models can synthesize synchronized video and sound in a unified diffusion process, but their inference cost remains high because long video token sequences require repeated attention computation across denoising steps.A variety of acceleration techniques have been developed for video generation models, including low-bit quantization, attention sparsification, and feature caching.However, since these methods are originally designed for video generation, directly applying them to audio-visual models overlooks the interactions between the audio and video branches and may therefore disrupt audio-video synchronization.We present a synchronization-aware acceleration framework for efficient audio-visual generation.Our key observation is that bidirectional audio-video cross-attention reveals structured interactions between the two branches, with high responses often concentrated on a few sound-related visual and temporal regions.Guided by this interaction pattern, we introduce a protected sparse attention strategy that preserves high-fidelity computation for synchronization-critical tokens while sparsifying redundant attention interactions.By explicitly accounting for cross-modal dependence during acceleration, our method improves inference efficiency while keeping video quality, audio quality, and audio-video synchronization.


8. FlowDance: Music-Driven Dance Video Generation with Parallel Pose and RGB Streams

Authors: Genying Li, Boda Lin, Jiachen Li, Zijian Jia, Haojie Zheng et al.

Categories: cs.CV Score: 6.16/10 (Obj:8 Id:6 Ind:7 Comp:5 Eff:7 Nov:6 Rep:5)

  • Strength: FlowDance 以 243.37 FVD 领先最强端到端基线 Wan-S2V(328.02)25.8%,BAS 0.2450 与 FIDg 7.501 均最优,40 人用户研究中四项偏好全部领先(M-A 69.6%)。
  • Weakness: 代码、权重与约 165K 条 FlowDanceSet 数据集全部标注 Coming Soon 且无任何发布链接,且评测仅在自建同源数据集上闭环、缺少 AIST++ 等第三方基准交叉验证。
  • Full review: Claude Code 全文七公理审稿

Music-driven dance video synthesis aims to animate a reference person according to a given music clip. The task is challenging because it requires a model to jointly learn music-to-motion correspondence, identity-preserving human animation, temporal coherence, and visually realistic video generation. We present FlowDance, a music-driven dance video generation framework that integrates explicit motion modeling with reference-preserving visual synthesis through parallel pose and RGB streams. We further introduce timestep-aware pose injection to adapt structural guidance across denoising steps and persistent identity injection to preserve the reference appearance over long video. To support this task, we further build a popularity-curated, high-resolution in-the-wild dance video dataset with synchronized music, RGB videos, 3D body motion, camera parameters, and projected 2D pose annotations. Extensive experiments show that FlowDance achieves strong performance in both dance motion generation and music-driven dance video synthesis.


9. Using the Mimi codec for metalinguistic representations

Authors: Artem Saloev, Erin Pacquetet, Nicolas Ballier

Categories: cs.CL | 11 pages, accepted for the Proceedings of the Third Workshop on the Bridges and Gaps between Formal and Computational Linguistics (BriGap-3), Paris 2026 Score: 5.84/10 (Obj:8 Id:5 Ind:6 Comp:6 Eff:6 Nov:6 Rep:4)

  • Strength: 将 Mimi codebook0(2048 token、80 ms/帧)与 TIMIT 音素对齐重映射后测得 PNMI 0.66(高于 HuBERT 0.43、EnCodec 0.28)且词级重转写一致性达 96.43%,量化支持语义 token 编码次音位/同位异音单位。
  • Weakness: 核心批评与解释缺少受控证据——未实证演示 ABX 评测失效、无随机/最小基线对照、PNMI 对比数值借用他人论文,且代码与重对齐数据仅承诺发布而尚未公开。
  • Full review: Claude Code 全文七公理审稿

In this paper, we focus on the dictionary of 2048 tokens used in Mimi semantic token codebook, the neural codec of the Moshi language model. We show that the ABX experiment carried out with Mimi fails to capture the mapping of the semantic tokens to phone realisations. By realigning Mimi representations to the TIMIT corpus transcriptions, we show that the 2048 tokens IDs of the semantic codebook map to quadphone, triphone, biphone, phone and subphone realisations.