每日 arXiv

音频 · 语音 · 音乐 · 声学

每日新爬取的 arXiv 论文,按评分排序。中文摘要帮你快速判断是否值得深读。

2026-09-03
日期2026-09-03
已评分
均分
最高

Daily Papers — 2026-09-03

19 papers on audio, speech, music, and acoustics.

1. Two-Stage Reinforcement Learning for Sound and Adversarial Test Generation in Code LLMs

Authors: Jiacheng Xu, Wentao Zhang, Zhiyi Lyu, Fuxiang Zhang, Chaojie Wang et al.

Categories: cs.CL, cs.LG | 21 pages, 7 figures. Accepted to Findings of the Association for Computational Linguistics: EMNLP 2026 Score: 7.60/10 (Obj:9 Id:8 Ind:7 Comp:7 Eff:8 Nov:7 Rep:7)

  • Strength: 两阶段 RL(声音性→对抗性)让测试生成同时可校验、可击杀:TACO pass@1 1.5B 5.63→12.31、7B 14.36→24.09,自生成测试选择后 7B LCB 达 48.79,平均超过 InternLM2-7B-reward 排序(44.57 vs 41.30),且 TCS 测试对更强外部模型的选择增益大于其自生成测试。
  • Weakness: 最邻近先验(Wang et al. 2025 共演化 coder/tester RL、CodeT)距离偏近使新颖性属增量级,TACO 训练/验证同源去重未报告、δ>α 假设无经验测量,且官方仓库截至 2026-08-31 仅含 README、训练评测代码未上传,端到端复现暂不可行。
  • Full review: Claude Code 全文七公理审稿

Reinforcement learning (RL) has substantially advanced code generation with large language models (LLMs) through executable feedback. The feedback for coding problems mainly comes from specific test cases, where high-quality test cases are often scarce since they should be both sound and discriminative. We thus turn to study the auto-generation of test cases using the learned model. We find this is naturally an adversarial RL problem: the model is expected to generate effective test cases as counterexamples, depending on the solver’s current failure modes. We propose Test Cases Scaling (TCS), a two-stage RL framework for effective test generation. Both stages train a test generator from a rolling policy-aligned buffer: Stage 1 generates tests consistent with the reference solution, and Stage 2 restricts the buffer to current failure modes and learns counterexample tests. Across TACO and LiveCodeBench, TCS improves both pass@1 and inference-time answer selection according to generated tests. We find the learned test generator also enables effective selection among other LLM outputs.


2. Building and Evaluating Fixed-Voice Thai TTS from Synthetic Speech

Authors: Kunat Pipatanakul, Potsawee Manakul, Warit Sirichotedumrong, Sittipong Sripaisarnmongkol, Pakorn Nathong et al.

Categories: cs.CL, cs.AI | 18 pages, technical report Score: 7.37/10 (Obj:9 Id:8 Ind:5 Comp:7 Eff:8 Nov:7 Rep:7)

  • Strength: 82M 端侧泰语固定音色学生达 Gemini 3.1 Flash Keyword 的 85.5%(68.2%)与停顿精确率的 94.8%(91.4%,超过 600M 教师的 89.9%),消融证明严格过滤损害难例覆盖(Keyword 69.6%→67.2%)而重采样被拒候选可同时恢复覆盖与韵律(PPER 减半至 6.2%),模型+评测框架+基准已开源。
  • Weakness: Challenge Set 与训练语料共用 WangchanThaiInstruct 来源且未报告去污染检查,停顿参考由同族 Gemini 模型标注,全文无任何主观听感测试(MOS/ABX 缺失),训练流水线代码未开源。
  • Full review: Claude Code 全文七公理审稿

In low-resource settings, deploying TTS typically requires choosing between a large voice-cloning model with costly inference or a compact fixed-voice system that requires a speaker-specific corpus. We study a third route: using a large voice-cloning model as a programmable data source to turn a short voice reference (e.g., 15 seconds) into a compact fixed-voice student trained entirely on synthetic speech. This setting makes pipeline design consequential: teacher errors become training targets, while filtering failed generations can reduce coverage of difficult texts. Thai further introduces challenges from ambiguous word boundaries, lexical tone, names and loanwords, numeric verbalization, and Thai-English code-switching. We study how text preparation, synthetic generation, quality filtering, rejection sampling, and frontend choices affect the resulting student, and where teacher limitations remain. We evaluate CER, Challenge-Set Keyword Accuracy, Prosody Pause Accuracy, speaker similarity, and speaking rate. The resulting 82M-parameter model, Wayu-Paxa-TTS-Edge, enables on-device Thai TTS without reference audio. It achieves 68.2% Challenge-Set Keyword Accuracy (85.5% of Gemini 3.1) and 91.4% pause precision, outperforming its OmniVoice teacher (89.9%) and reaching 94.8% of Gemini 3.1. It also achieves the lowest pause-placement error and intra-word pause rates among the three systems, and 3.7% and 1.1% CER on Thai and English, respectively. We open-source the model and evaluation framework for Thai TTS development.


3. Summary of the ChinaVoices Challenge 2026: Data, Tasks, Baseline, and Methods

Authors: Yujie Liao, Bingshen Mu, Shuiyuan Wang, Liumeng Xue, Hexin Liu et al.

Categories: eess.AS | 15 pages,1 figure Score: 7.00/10 (Obj:8 Id:7 Ind:8 Comp:6 Eff:8 Nov:6 Rep:6)

  • Strength: 统一 16 方言双任务评测平台动员 28 支队伍参赛,基线 53.62% ACC/18.10% CER 被推至 83.19% ACC/11.08% CER(ASR 相对错误减少 38.80%),并量化双任务负相关(r = −0.76)与 kejia/chaoshan 两大难点方言。
  • Weakness: 系统归因完全依赖 15 支队伍自述报告而组织方未做受控消融,r = −0.76 仅基于 16 个数据点且无显著性检验,基线权重仅走百度网盘(仓库无 license、12 stars)、参赛系统报告未统一发布,外部无法复现 Hidden Eval 结果。
  • Full review: Claude Code 全文七公理审稿

This paper summarizes the ChinaVoices Challenge 2026, which aims to establish unified task definitions and evaluation conditions for Chinese dialect speech processing and to advance multi-dialect identification and automatic speech recognition. The challenge covers 16 dialect categories and defines two tasks: Chinese Multi-Dialect Identification and Chinese Multi-Dialect Automatic Speech Recognition (ASR). It uses approximately 320 hours of speech across the Reference Set, Open Evaluation Set, and Hidden Evaluation Set. The two tasks use the same evaluation audio, and each includes restricted-data and open-data tracks. We describe the task settings, data, evaluation metrics, and Qwen3-ASR-1.7B baseline, and analyze the leaderboard results and submitted systems. In total, 28 teams submit results, 17 provide system reports, and systems from 15 teams pass the compliance review and are included in the analysis. Most eligible systems outperform the baseline, and the official top-three order remains unchanged on the Hidden Evaluation Set for both tasks. Dialect-level results show that categories with higher identification accuracy generally have lower ASR error rates, although the tasks assess related but distinct capabilities. Leading identification systems commonly exploit dialect-discriminative acoustic representations, whereas leading ASR systems emphasize data normalization, augmentation, and auxiliary CTC objectives. These results provide practical guidance for developing and evaluating Chinese multi-dialect speech processing systems.


4. Geometric Ceilings on Time-Frequency Masking for Single-Channel Separation

Authors: Maxime Baelde

Categories: eess.SP, cs.SD, eess.AS | Submitted to Signal Processing (Elsevier), manuscript SIGPRO-D-26-03635. 33 pages, 6 figures. Companion paper: “What Selects, What Reconstructs” (submitted to IEEE/ACM TASLP) Score: 7.00/10 (Obj:9 Id:9 Ind:8 Comp:9 Eff:5 Nov:7 Rep:3)

  • Strength: 给出实增益时频掩蔽类的精确闭式天花板(MUSDB18 held-out 16.62 dB,比 IRM 高 2.94 dB),并证明平方误差准则把任何条件均值估计器拉回直线:后验均值估计器比天花板低 11.44 dB,4 倍分量与 7.5 倍数据合计移动不足 0.5 dB,闭式门归因约 70%(dB 计)于后验方差。
  • Weakness: 主结果仅 9 个 held-out excerpt、单一两源数据集、无置信区间,未与任何神经基线做数值对比,估计器落后 IRM 8.50 dB,且无代码、无数据可用性声明,独立复现不可行。
  • Full review: Claude Code 全文七公理审稿

Most single-channel separators estimate a source by applying a real gain to the mixture in each time-frequency bin. The optimum of that format, which the oracle masks used as bounds do not attain, is the orthogonal projection of the source onto the line spanned by the mixture, its residual set by the angle between them. Locating an estimator reduces to the block structure of a real-linear operator on stacked spectra, giving a chain of four nested classes whose three larger terms match three assumptions on the prior: zero means, circularity and absence of inter-frequency coupling. Held fixed the chain is a cascade of four orthogonal projections; refitted per frame it collapses onto its first term, attributing the whole residual to one missing real parameter per bin, the phase. When the phase posterior is symmetric about the mixture direction, the minimum mean-square estimate falls back onto the line, with gain the posterior mean of the oracle gain and excess error its variance. On MUSDB18 a posterior mean under a non-circular Gaussian-mixture prior leaves the class yet stays 11.44 dB under the per-frame ceiling, which four times as many components and 7.5x the data do not close; a closed-form gate attributes some 70% of it, in decibels, to the predicted variance. The widest fixed class stays 6.70 dB under the same ceiling. Leaving the class and minimising squared error are conflicting requests: the barrier lies in the criterion rather than in the prior.


5. Deep Neural Compression for RIR-Characterized Acoustic Environments with Structure-Aware Constraints

Authors: Chen-Yuan Ning, Yang Ai, Hui-Peng Du, Xiao-Hang Jiang, Zhen-Hua Ling

Categories: eess.AS Score: 6.90/10 (Obj:8 Id:8 Ind:6 Comp:9 Eff:7 Nov:6 Rep:5)

  • Strength: 375 bps 下 T60 误差降至 1.14 s(基线 2.49/5.04 s)、DRR 误差 5.11 dB(基线 29.11 dB)、MOS 4.28±0.08,并给出 46.875–93.75 bps 的压缩实用下限。
  • Weakness: 全部实验限于 Motus 单一房间的单声道 RIR,跨房间泛化未验证,且未引用/对比 2024 ISCAS 同类神经 RIR 压缩工作,无开源代码。
  • Full review: Claude Code 全文七公理审稿

Room impulse responses (RIRs) characterize the acoustic environment of a room by capturing how sound propagates and decays within an enclosed space. In applications such as immersive audio rendering, accurate acoustic reconstruction often relies on spatially densely sampled RIRs. This consequently gives rise to a large volume of RIR data, imposing a substantial burden on storage. Although recent neural audio codecs provide an effective framework for low-bitrate compression, their training objectives are mainly tailored to speech and general audio, and are therefore not well aligned with the acoustic characteristics of RIRs. Therefore, we propose an EnCodec-based neural RIR compression method, which incorporates RIR structure-aware constraints at two levels. Specifically, at the RIR level, structure-aware constraints are imposed on the global decay behavior and local energy distribution of RIRs through energy decay curve (EDC) regularization and a short-time window energy constraint, while at the reverberant-speech level, reverberant-speech supervision is further introduced to constrain the consistency of the reverberant speech generated by the reconstructed RIRs. Experimental results show that, at a low bitrate of 375 bps, the proposed method achieves lower RIR reconstruction error and better reverberant-speech perceptual consistency than audio-oriented codecs.


6. Neural Music Enhancement with Dual Time-Frequency Spectral Representations for Prediction and Discrimination

Authors: Fei Liu, Yang Ai, Zhen-Hua Ling

Categories: cs.SD | Accepted by ISCSLP 2026 Score: 6.89/10 (Obj:8 Id:7 Ind:8 Comp:6 Eff:7 Nov:7 Rep:5)

  • Strength: 双谱分工(STFT 生成 + 八度分段 CQT 判别 + 色度损失)在去噪任务上 fwSSNR 14.80、FAD 0.30 全面超越重训的 MP-SENet(14.22),混合任务主观 ABX 偏好率 63–71%(p<0.01)。
  • Weakness: 全部实验仅基于合成失真、未验证 in-the-wild 场景,混合任务 FAD(1.39)落后 MP-SENet(1.27),无音乐专用基线对比且未开源代码(GitHub 检索 0 仓库)。
  • Full review: Claude Code 全文七公理审稿

Non-professional music recordings shared online often suffer from background noise and reverberation, degrading perceived quality and limiting reuse. This paper proposes DSME, a music enhancement model based on dual time-frequency spectral representations. Within a generative adversarial framework, DSME uses short-time Fourier transform (STFT) spectra for generation and constant-Q transform (CQT) spectra for discrimination. Leveraging STFT’s fixed window, invertibility, and predictability, the generator estimates clean amplitude-phase spectra from degraded inputs and reconstructs waveforms via inverse STFT. Exploiting CQT’s log-frequency, variable-window structure aligned with musical octaves, we design an octave-segmented CQT discriminator. We also introduce a chroma-spectrum loss to emphasize pitch and harmonic consistency. Experiments show DSME outperforms baselines in objective and subjective tests, validating the effectiveness of the dual-spectrum approach.


7. ToolDF: Tool-Integrated Reasoning for Mixed-Authenticity Audio Deepfake Detection

Authors: Taewoo Kim, Young Han Lee, Nam In Park, Chanwoo Kim

Categories: eess.AS, cs.AI, cs.SD | To appear in Findings of the Association for Computational Linguistics: EMNLP 2026 Score: 6.53/10 (Obj:8 Id:6 Ind:6 Comp:6 Eff:7 Nov:7 Rep:5)

  • Strength: 提出混合真实性音频深度伪造检测的形式化任务与 ALLM 工具集成框架 ToolDF,C-Avg macro-F1 达 81.89,超最强联合训练单体 3.72 点、超固定分离管线 14.39 点,且提供成分级定位(事件级 F1 87.62–94.24)与可解释证据链,消融证明按需分离决策是关键(去掉 Planning 后 C2 strict-F1 从 77.66 崩至 20.87)。
  • Weakness: 全部评测局限于自建合成基准且无外部交叉验证,领先幅度仅 3.72 点并在 C2 上被固定管线反超;SFT 轨迹观测填充真值标签导致训练-部署分布裂口未刻画;摘要声称代码数据公开但 GitHub 仓库(rlataewoo/tooldf)经核实仅有 LICENSE 与 README、无任何代码或数据。
  • Full review: Claude Code 全文七公理审稿

Audio deepfake detection is commonly formulated as clip-level binary classification of single-domain audio. However, real-world manipulated audio can exhibit mixed authenticity, where genuine and manipulated cues coexist across temporal transitions, overlapping sources, or both. This setting requires not only detecting manipulated audio but also localizing the components that provide evidence for the decision. We propose ToolDF, a tool-integrated reasoning framework for mixed-authenticity audio deepfake detection. ToolDF employs an audio large language model as an orchestrator trained with supervised tool-use trajectories. It adaptively analyzes the audio scene, selectively performs source separation, routes components to domain-specific experts, and aggregates their evidence into an interpretable verdict. We further introduce a mixed-authenticity ADD benchmark covering temporal transitions, acoustic overlaps, and hybrid mixtures. Experimental results show that ToolDF achieves the best overall performance on composite-type detection, achieving macro-F1 gains of 3.72 and 14.39 points over the strongest monolithic baseline and a fixed pipeline, respectively, while providing interpretable evidence localized to temporal regions and acoustic sources. Our source code and dataset are publicly available online.


8. DuplexSpeechBench-IFEval: Evaluating Implicit Instruction Following in Full-Duplex Voice Agents

Authors: Puneet Mathur, Dinesh Manocha

Categories: cs.AI | Under Submission Score: 6.50/10 (Obj:8 Id:6 Ind:6 Comp:6 Eff:6 Nov:7 Rep:7)

  • Strength: 首个全双工语音代理隐式指令跟随基准(1,038 用例、8 角色、5 条件协议),同音频跨条件复用实现单变量隔离,揭示架构依赖 trade-off:F-Actor/PersonaPlex 在仅人设条件下 adherence 掉 9.7%/4.5%,安全冲突下除 MiniCPM-o(60.0%)外全部系统 SafetyOverride 不足。
  • Weakness: 全系统 L0 绝对 IAS 仅 6.5%-48.4% 且无人类基线校准,IAS 条件间差异 1-4pp 未报告显著性检验,动作级结论基于 21-42 例小样本,机制归因与干预实验缺位,PAS 评委同源性未披露且 HuggingFace 数据集公开性未获确认(401)。
  • Full review: Claude Code 全文七公理审稿

Full-duplex voice agents must continuously decide when to listen, backchannel, interrupt, handle speech overlaps, take the floor, and yield. Existing benchmarks largely test these behaviors through explicit turn-management instructions, while deployed agents are often configured through roles or personas from which the appropriate conversational behavior must be inferred. We introduce DuplexSpeechBench-IFEval (DSB-IFEval) for evaluating implicit instruction-following in real-time spoken interaction. (DSB-IFEval) comprises 1,038 test cases spanning eight diverse assistant roles and evaluates five conditioning protocols for instruction-following: default behavior, explicit behavioral instructions, persona-implied behavior, combined persona–rule conditioning, and instruction conflict. We measure real-time floor management using a deterministic Instruction Adherence Score (IAS) and persona-consistent content using LLM-judged Persona Adherence Score (PAS). Across six real-time speech systems, we find architecture-dependent trade-offs. Full duplex models like F-Actor and PersonaPlex are more sensitive to whether conversational behavior is stated explicitly or must be inferred from a persona, with adherence dropping by 9.7% and 4.5%, respectively, under persona-only conditioning. In contrast, GPT-Realtime, MiniCPM-o, and Fun-Audio-Chat strongly adhere to persona-consistent content, but their floor behavior does not adapt across explicit and persona-only instructions and remains constrained on several proactive actions. We further find that even if systems reliably follow conflicting directives to their prescribed persona, they still struggle to override them under safety conflict. These results show that inferring the behavior implied by a role, executing it at the appropriate conversational moment, and resolving competing instructions remain distinct challenges for full-duplex voice agents.


9. Local Chord Corruption Is Not Recognizer Replay: Chord-Condition Propagation in MIDI-SAG

Authors: Weiwen Huang

Categories: cs.SD, cs.MM Score: 6.47/10 (Obj:8 Id:7 Ind:6 Comp:7 Eff:6 Nov:5 Rep:8)

  • Strength: 受控配对实验(30 tracks × 3 seeds)证明局部和弦扰动与完整 ACR 回放测到的条件传播不等价——CENS target 差值 29/30 tracks 为正(均值 0.462,p=2.61×10⁻⁸),匹配时间支持与关系构成后回放距离从 0.482 压至 0.098,并在第二条识别器路径上复现。
  • Weakness: 全部结论条件于单一 MIDI-SAG 生成器、30 个 12s MUSDB18-HQ 窗口,无跨生成器验证、无人类听感评估,代理余弦指标与可感知音乐差异的映射未建立,核心命题的泛化半径与感知效度均缺关键证据。
  • Full review: Claude Code 全文七公理审稿

Synthetic chord corruption provides a controlled stress test for singing accompaniment generation (SAG), whereas complete automatic chord recognition (ACR) replay measures the condition delivered to a deployed system. We compare them by replaying CNN-CRF and DeepChroma+CRF predictions through one fixed MIDI-SAG generator, holding track, seed, context, and scoring window constant. Across 30 paired tracks and three seeds, a central four-second tritone produced larger changed-target and inside-output effects than CNN-CRF replay in STFT, CQT, and CENS; the CENS target gap was positive on 29/30 tracks (mean 0.462). Matching replay support and relation composition reduced this mismatch, with joint matching giving the lowest replay distance in the full-30 analysis. Relative-root substitutions produced 2.88-fold larger CENS full-window output change than same-root quality flips at near-equal dose. The matched surrogate was closer to replay for both recognizer paths. We conclude that localized corruption tests mechanisms, whereas complete replay evaluates deployed chord-condition propagation.


10. StreamWSR: Streamable and Lightweight Waveform-Domain Neural Speech Super-Resolution

Authors: Yuan Tian, Yang Ai, Hui-Peng Du, Zhen-Hua Ling

Categories: eess.AS Score: 6.40/10 (Obj:8 Id:6 Ind:5 Comp:8 Eff:7 Nov:6 Rep:5)

  • Strength: 首个零前瞻全因果波形域语音 SR 模型,仅 9.03M 参数与 2.12G FLOPs,在 2k→16k 设定上取得全场最优 ViSQOL(3.81)与 STOI(87.14%),WER(39.33%)大幅优于 UDM+(86.64%)与 FLowHigh(60.80%)。
  • Weakness: 缺失关键验证证据:仅 VCTK 单一英语语料无跨数据集泛化测试,核心卖点流式推理未报告延迟毫秒数/RTF 实测,全程客观指标无 MOS 主观评测,且无开源代码(GitHub 检索 0 条仓库)。
  • Full review: Claude Code 全文七公理审稿

This paper proposes StreamWSR, a Streamable neural Waveform-domain model for speech Super-Resolution (SR). By adopting a fully causal architecture with compact frame-level waveform representation, the proposed StreamWSR supports zero-look-ahead streaming inference while avoiding vocoder-based reconstruction and explicit phase prediction. Specifically, StreamWSR downsamples the input waveform into a compact frame-level representation using strided causal convolutions. Then, a lightweight causal long-short-term modeling backbone is employed to capture both local waveform structures and long-range historical dependencies under causal constraints. Finally, the modeled output is converted back to the waveform domain through a causal transposed-convolution and combined with the input waveform via a residual connection to generate the final high-resolution speech. Experimental results on 16 kHz speech SR show that StreamWSR achieves competitive or superior speech quality and intelligibility compared with representative waveform- and spectrum-based baselines, while maintaining a zero-look-ahead streaming advantage with only 9M parameters and 2G FLOPs.


11. Masked Autoregressive Speech Enhancement with Continuous Neural Audio Codec Representations

Authors: Yoto Fujita, Simon Leglaive, Laurent Girin

Categories: cs.SD, cs.AI | Presented at the 19th International Workshop on Acoustic Signal Enhancement (IWAENC), Sep 2026, Cremona, Italy Score: 6.30/10 (Obj:8 Id:6 Ind:5 Comp:6 Eff:5 Nov:7 Rep:8)

  • Strength: MARSE 首次在连续 DAC 潜向量上实现掩码自回归语音增强,N=10 迭代以 1912 GFLOPs(约为全自回归 C-AR 的一半)取得介于 C-NAR 与 C-AR 之间的质量(OVRL 3.34)且 dWER 大幅优于 C-AR(12.68% 对 20.89%),统一架构/训练的解码策略对比与 MIT 许可完整开源代码构成可靠证据链。
  • Weakness: 全部质量指标依赖无参考代理(DNSMOS/wav2vec dWER)且无人工主观听测,域内结果仅 300 样本、域外仅一套数据;可部署配置算力为 DPTNet 的约 160 倍而域内 dWER 仍劣于判别式基线(12.68% 对 9.79%),最优 oracle 解码序依赖真实干净语音不可部署,误差累积机理未识别。
  • Full review: Claude Code 全文七公理审稿

Most previous work on speech enhancement (SE) based on masked generative modeling relied on discrete token representations of audio signals, obtained using neural audio codecs (NACs). However, a recent study has shown that continuous latent representations of NACs can be advantageous for SE in terms of speech quality and intelligibility. In this work, we propose masked autoregressive SE (MARSE), a method for SE based on iterative decoding of masked clean speech frames using continuous NAC representations of speech. In particular, we investigate a set of different decoding policies, ceteris paribus, that is, using the same DNN (a Conformer model), the same NAC (the DAC codec) and the same training setup. The results show that MARSE enables a flexible trade-off between SE performance and computational cost. Audio examples and code are available online.


12. PACodec: A Low-bitrate Neural Speech Codec with Parallel Additive Vector Quantization

Authors: Fei Liu, Yang Ai, Xiao-Hang Jiang, Zhen-Hua Ling

Categories: cs.SD | Accepted by APSIPA 2026 Score: 6.00/10 (Obj:8 Id:6 Ind:7 Comp:6 Eff:6 Nov:5 Rep:5)

  • Strength: 并行加性 VQ(4×128 码本、求和聚合)使 6.7M 参数的 PACodec 在 VCTK 上以 4.2 kbps 全面超过 4.5 kbps 的 MDCTCodec(LSD 0.78 vs 0.83),1.4 kbps 下感知质量与 2 kbps 基线相当(ABX p>0.01),码率节省 30%。
  • Weakness: 缺少同骨干 PAVQ↔RVQ 受控消融(跨模型对比混合骨干与量化两变量),LibriTTS 上 1.4 kbps 略逊 1.5 kbps 基线(UTMOS 3.80 vs 3.87),解耦分工无机制解释且零下游验证,且无代码/损失权重披露。
  • Full review: Claude Code 全文七公理审稿

This paper proposes PACodec, a novel low-bitrate neural speech codec based on parallel additive vector quantization (PAVQ). Unlike the mainstream residual vector quantization (RVQ) used in most neural speech codecs, where vector quantizers (VQs) are sequentially dependent, the PAVQ strategy adopted in PACodec aggregates parallel quantization results to optimize bitrate usage. Specifically, the PAVQ adopts a “global-local-global” (GLG) design: the global encoded features are quantized in parallel by multiple independent VQs, each attending to a local component of the representation, and their outputs are aggregated through addition to yield the final global quantization result for decoding. Experimental results show that PACodec, as each VQ focuses only on local information, supports smaller codebooks and reduces bitrate by 30% compared with baselines at the same decoding quality, with only minor model complexity. Further analysis shows that, owing to the GLG framework of PAVQ, the proposed PACodec is disentanglement-friendly, and each independent VQ captures different aspects of speech, e.g., content, timbre, and acoustic details, suggesting potential for application to downstream tasks such as voice conversion.


13. StrixAE: An Intelligent Agent for Audio Enhancement under Complex Distortion Coupling in Real-World Scenarios

Authors: Chenglin Wu, Junjie Wu, Jinhang Chen, Mingyang Chen, Zixu Lin et al.

Categories: cs.SD, cs.AI Score: 5.80/10 (Obj:7 Id:6 Ind:6 Comp:5 Eff:7 Nov:6 Rep:2)

  • Strength: 音频域首个 MLLM 编排 agent,URGENT 2025 track 1 实测第一(DNSMOS 2.88/ESTOI 0.82/SDR 12.66 vs baseline 2.85/0.76/10.24),SFT+RL 消融证明 APRL 结构化奖励有效(3.18 vs 3.12)。
  • Weakness: 无代码、无数据发布,全部指标为无参考感知指标且无人工 MOS,闭源 Adobe Acrobat 在 blind set 上仍领先(DNSMOS 2.92 vs 2.83、UTMOS 2.35 vs 1.98),SOTA 声称的证据链不闭合。
  • Full review: Claude Code 全文七公理审稿

Audio enhancement in real-world scenarios involves complex distortion couplings and requires personalized enhancement. Existing solutions struggle to address both simultaneously. To improve robustness and enable autonomous operation in such scenarios, we propose StrixAE, an agent based on a multimodal large language model (MLLM). StrixAE leverages the MLLM as a controller to coordinate multiple audio enhancement and personalization models. To further enhance system robustness, reduce artifacts, and improve generalization across diverse real-world scenarios, StrixAE is trained through a two-stage process: first, CoT supervised fine-tuning on AcoustBench to ground basic reasoning and tool invocation; second, Audio Perception Reinforcement Learning (APRL), a reward design specifically tailored for audio restoration pipelines that jointly optimizes format validity, structural coherence, and perceptual quality. Unlike generic RL fine-tuning, APRL introduces structured rewards that enforce executable pipelines and logical section ordering, enabling the agent to produce reliable, interpretable enhancement plans without hallucinated tools. Based on real-world test datasets, our proposed method outperforms most existing open-source and proprietary solutions, achieving state-of-the-art performance across multiple perceptual metrics and demonstrating strong generalization robustness.


14. The Attention Triangle in Audio-Video Models

Authors: Sagi Polaczek, Noa Kraicer, Gal Metzer, Zhuo Ning, Ali Mahdavi-Amiri et al.

Categories: cs.AI Score: 5.79/10 (Obj:8 Id:6 Ind:6 Comp:6 Eff:5 Nov:6 Rep:4)

  • Strength: 首个对联合音视频扩散模型做注意力级泄漏归因的工作:二阶有效注意力 P{VA}·P{AV} 锁定音视频通路为主泄漏源,训练无关引导在 2,400 视频的三重评测(Qwen 判官 + VA-Judger + 30 人用户研究 p<10⁻²⁸)下归因分 0.1216→0.1349 且 VBench 质量不掉。
  • Weakness: 主指标绝对增益仅 0.013 且需 2.5× 计算 + 人工文本锚点;机制解释(音频仅时间编码致空间欠约束)无直接检验、结论绑定 LTX-2 单模型;官方仓库截至审稿日仅有 README 空壳,挑战集未确认公开。
  • Full review: Claude Code 全文七公理审稿

Audio-video diffusion models rely on cross-modal attention to coordinate text, sound, and visual content, yet this same mechanism can introduce subtle and systematic semantic leakage. We study these models by probing and analyzing the ``attention triangle,’’ comprising the three cross-attention edges connecting the text, audio, and video streams, and examine how semantic information is routed across modalities during generation. Our analysis reveals that routing along the audio-video edge is bidirectional: audio can influence video generation, while video can influence audio generation. This edge is shaped by biases encoded in the model’s parameters and emerges as a major contributor to leakage: when prompts are in tension with learned priors, cross-modal interactions may override the intended conditioning and reroute semantics toward visually canonical but incorrect outcomes. These effects suggest that semantic artifacts arise not merely from attention spreading beyond its intended target, but from structured, bias-driven interactions along specific pathways. Building on this perspective, we extract attention-derived signals that expose how semantics are distributed and grounded across modalities, and use them as a diagnostic tool to both analyze and deliberately incur leakage under controlled conditions. This enables us to probe the internal dynamics of cross-modal routing and isolate the role of individual interactions. We further leverage these signals to guide inference-time interventions that encourage more consistent cross-modal alignment. Extensive experiments support our analysis and demonstrate improved semantic grounding while preserving generation quality.


15. Is Semantics Enough for Speech Mean Opinion Score Prediction?

Authors: Tianyu Lan, Yufei Shi, Yang Ai, Honghao Sun, Huipeng Du et al.

Categories: cs.SD | 5 pages, 2 figures, accepted by ISCSLP 2026 Score: 5.68/10 (Obj:8 Id:6 Ind:8 Comp:5 Eff:5 Nov:5 Rep:4)

  • Strength: 首次将统一神经编解码器表征(XCodec/SpeechTokenizer)纳入 MOS 预测系统比较,固定下游头 + 冻结/微调双设置 + 域内(BVCC SRCC 最高 0.882)与双 OOD(SOMOS/BC2019)的矩阵设计较完整,并以梯度分析佐证语义主导机制。
  • Weakness: 核心结论”协同表征上限更高”仅靠 0.004-0.007 的 SRCC 差距支撑且未报告种子间方差与显著性,全文零对比已发表 MOS 系统(UTMOS 0.878/DAMOS 0.885)、零代码发布,跨语言退化机制只有定性归因。
  • Full review: Claude Code 全文七公理审稿

Mean Opinion Score (MOS) is the gold standard for evaluating synthesized speech naturalness. However, current automatic MOS predictors are dominated by self-supervised learning (SSL) models that prioritize high-level semantics, potentially compromising their ability to capture critical acoustic details. In this paper, we systematically investigate representations from three paradigms: SSLs, acoustic-only neural audio codecs (NACs), and unified NACs that integrate semantics into reconstruction-based architectures. Extensive benchmarking on the standard BVCC and multiple out-of-domain (OOD) datasets demonstrates that features synergizing semantic understanding with fine-grained acoustic modeling achieve a higher performance upper bound in speech quality assessment. Ultimately, our findings highlight that semantics alone are not enough; a dual focus on semantic content and acoustic fidelity is essential for robust MOS prediction.


16. Broadband Acoustic Intensity Direction Estimation with Tight-Frame Cardioid Arrays

Authors: Akira Omoto

Categories: eess.AS Score: 5.60/10 (Obj:8 Id:6 Ind:5 Comp:6 Eff:5 Nov:5 Rep:5)

  • Strength: 紧框架心形阵列实现 50 Hz–20 kHz 宽带声强方向估计,TF24 + 20 mm 间距在 LJ/S = −20/−30 dB 时全频段中位跟踪误差 ≤2.5°,单源 200 Hz 以上误差约 −1°,经构型×间距×干扰三因素消融(每条件 1000 次蒙特卡洛)验证。
  • Weakness: 26 通道 IR 由单传声器顺序旋转合成、通道失配与真实混响均未验证,多径基准与被评系统共用同一套 IR 缺乏独立交叉验证,三个关键误差现象(低频仰角偏差、高频 −2° 至 −3° 偏移、d=100 mm 的 3–4 kHz 误差峰)成因未明,且无代码与公开数据。
  • Full review: Claude Code 全文七公理审稿

This study experimentally investigates source-direction estimation from acoustic intensity using paired cardioid microphones, with particular emphasis on well-balanced arrangements known as tight frames. Impulse responses were measured at 6 or 24 microphone positions, and two opposed-pair spacings were examined. The resulting arrays were evaluated for a single source, coherent interference between waves arriving from two orthogonal directions, and directional tracking in the presence of interfering waves from multiple directions. The results demonstrate broadband acoustic-intensity direction estimation from 50 Hz to 20 kHz, with the 24-microphone arrangement and the shorter spacing generally providing smaller errors.


17. Compressing Streaming Neural Audio Encoders via Latent-Space Distillation

Authors: Prasanth Yadla, Mohammad Samragh, Dongseong Hwang, Mingbin Xu, Yuanyuan Zhang et al.

Categories: cs.SD | 15 pages Score: 5.50/10 (Obj:8 Id:6 Ind:5 Comp:6 Eff:5 Nov:6 Rep:2)

  • Strength: 在 Apple 系统听写 tokenizer 上以 pre-quantizer latent 为单一蒸馏目标实现 2.8× 压缩,6 对师生对中 5 对相对 WER 偏差 ≤1.9% 且零微调,J2 配置反超教师 5.3%,并较同容量独立训练改善 3.9% 相对 WER,一个配方同时覆盖离散与连续两种 token 接口。
  • Weakness: 全文零开源、无绝对参数量、无训练数据与算力规模、无功耗/延迟实测,核心主张”latent 目标优于替代目标”被作者自陈未经消融验证,全部 WER 经教师组件中转测量且 J3 学生 7.7% 退化仅有一句归因,外部无法复核任何结论。
  • Full review: Claude Code 全文七公理审稿

System-wide Dictation on Apple devices runs entirely on-device, and the speech it transcribes reaches the foundation model through a tokenizer: an encoder that maps short windows of waveform onto the representation the language model reads. Because that model is sparsely activated under Instruction-Following Pruning, only a small subset of its experts occupies DRAM at any time, so the always-on tokenizer competes for the same memory, and its parameter count bears directly on power and latency. In this work we study how to compress such a tokenizer by distillation, taking as the supervision target neither the discrete token nor the output distribution but the pre-quantizer latent the model actually consumes - the last representation the two token interfaces share. We train only the student encoder to regress the teacher’s per-frame latent under a squared-error objective, with a single affine layer absorbing the teacher-student width mismatch. Because the target precedes both the quantizer and the language-model bridge, one recipe covers both token interfaces we support, and applies both to a tokenizer pretrained alone and to one jointly trained with a language model. At 2.8x compression the distilled student stays within 1.9% relative WER of its teacher on five of six teacher-student pairs without any fine-tuning, and improves on an independently trained tokenizer of identical capacity by 3.9% relative.


18. Test-time adaptation for speech enhancement with an autoregressive speech prior

Authors: Sofiene Kammoun, Simon Leglaive, Xavier Alameda-Pineda, Timo Gerkmann

Categories: cs.SD, cs.AI | Submitted to IWAENC 2026 Score: 5.42/10 (Obj:8 Id:5 Ind:7 Comp:5 Eff:4 Nov:6 Rep:4)

  • Strength: 提出基于 DAC 潜空间自回归干净语音先验的单句 KL 散度 TTA,4 数据集按噪声不匹配程度分级验证(oracle 早停下 DNS OVRL +0.18 / BAK +0.23),并以四分位分析揭示先验对数密度可预测最优自适应步数(k̃ 预测器在全部数据集保持正增益)。
  • Weakness: 可部署增益仅 0.00-0.07 OVRL(与 DNSMOS 噪声同量级、无显著性检验)且头条数字依赖 oracle 步数 k,全文零 TTA 基线与零随机先验消融导致归因不可判定,TTA 方法代码未发布(仓库仅音频演示)。
  • Full review: Claude Code 全文七公理审稿

Test-time adaptation (TTA) offers a promising direction for improving speech enhancement models under mismatched acoustic conditions, without requiring access to labeled target data. In this work, we propose a single-utterance TTA method that regularizes a pretrained speech enhancement model using an autoregressive prior trained on clean speech latent representations extracted from a neural audio codec. Adaptation is performed by minimizing the Kullback-Leibler divergence between the enhanced speech distribution and the clean speech prior. Experiments across multiple noisy speech datasets show consistent improvements in speech quality, particularly under training-testing noise mismatch conditions. Code and audio examples are available online.


19. Alignment-Free Text-Audiobox for Voice Dubbing and Full-Duplex Dialogue Synthesis

Authors: Sanyuan Chen, Min-Jae Hwang, Sho Inoue, Anna Sun, Bokai Yu et al.

Categories: cs.CL, eess.AS Score: 5.20/10 (Obj:8 Id:6 Ind:4 Comp:6 Eff:4 Nov:6 Rep:2)

  • Strength: 3B flow-matching DiT 在 1920× 压缩 DAC-VAE 潜空间上统一跨语配音与全双工对话合成,配音人工评测对内部系统自然度 +0.39/+0.45、可分享性 +0.40/+0.38,短对话人类相似度 4.53 距真实录音 4.62 仅 0.09,重排序使 WER 4.05%→2.20%。
  • Weakness: 配音与长对话主结果仅与不公开内部系统在内部基准(100 条/方向)上比较,公开基线只在情感对话单一维度出现;零代码零权重(GitHub 检索 0 命中)叠加 480k 小时/256×A100 训练成本使外部复现不可行,跨语 SFT 致 SpkSim 0.73→0.60 等关键现象无机制解释。
  • Full review: Claude Code 全文七公理审稿

We present Alignment-Free Text-Audiobox (Text-AB), a unified framework for high-quality voice dubbing and full-duplex dialogue synthesis. Building on a Diffusion Transformer trained with a flow-matching objective, Text-AB departs from the Audiobox system along three dimensions. First, it operates in a latent diffusion framework using DAC-VAE features that encode 48 kHz waveforms into a 25 Hz latent sequence, giving over 10x higher compression than previous EnCodec representations while improving resynthesis quality. Second, Text-AB is alignment-free: it consumes raw text via an off-the-shelf text encoder and learns text-speech alignment through cross-attention, removing the need for forced alignment and explicit duration prediction. Third, we scale model and data substantially, pretraining a 3B-parameter model on 480k hours of monolingual speech, followed by supervised fine-tuning on three downstream tasks: cross-lingual voice dubbing, full-duplex dialogue synthesis, and emotional full-duplex dialogue synthesis. At inference, Text-AB supports one-shot generation for up to ~1 min of speech and arbitrarily long-form generation via a multi-diffusion scheme, plus a multi-stage reranking strategy that enhances quality based on automated metrics. On a real-world dubbing benchmark, Text-AB delivers a step-change improvement over the latest internal dubbing system, with large gains in prosody similarity, voice similarity, naturalness, and shareability. For full-duplex dialogue synthesis, it approaches human recordings on short-form conversations and substantially outperforms the latest internal model on long-form human-likeness and expressivity, while natively modeling turn-taking, back-channeling, and emotional dynamics. For emotional dialogue synthesis, emotion conditioning significantly improves emotion alignment and emotional interaction quality over the unconditioned baseline.