Daily Papers — 2026-08-06
15 papers on audio, speech, music, and acoustics.
1. LILAC: An Idempotent Neural Speech Codec
Authors: June Young Yi, Dongwook Lee, Jiheum Yeom, Sungroh Yoon
Categories: cs.SD, cs.LG | 22 pages, 4 figures Score: 6.84/10 (Obj:9 Id:8 Ind:9 Comp:6 Eff:6 Nov:6 Rep:9)
- Strength: LILAC 通过 invertible analysis transform + FSQ 代数保证 codec idempotence (E(D(c))=c) — 在 7,457 个测试样本上经验证达到 100% token agreement,且在 100 次循环中无 UTMOS/dWER 退化,同时在 sub-1kb/s 范围内实现了具有竞争力的质量(LibriTTS-R 上 UTMOS 为 4.24,所有数据集上 STOI/SI-。
- Weakness: idempotence 的实际下游效用尚未得到证实(作者承认“仅为理论上的”),dWER 落后于同行,且 construction-based idempotence 原理本身已在图像压缩(Li et al. NeurIPS 2023)中确立,使得核心创新成为渐进式而非变革式的。
- Full review: Claude Code 全文七公理审稿
Neural Audio Codecs are widely adopted in speech generation and editing. However, existing neural audio codecs are not idempotent: across the paper’s twelve baseline systems, every configuration tested rewrites, on average, at least 15% of its tokens in a single decode-re-encode pass. This poses a problem for utilizing Neural Audio Codecs as token interfaces in pipelines where re-encoding decoded outputs can occur. We present LILAC, a fully convolutional 24 kHz speech codec at 9.375 Hz and 0.75 kbit/s that is codec idempotent by construction; re-encoding the decoded audio of any valid token stream returns the identical stream. LILAC achieves idempotency while maintaining competitive quality, reaching UTMOS 4.14 and 4.24 on LibriSpeech and LibriTTS-R test sets, comparable to SOTA sub-1 kbit/s Neural Audio Codecs.
2. KVAE: Family of Tokenizers for Multimodal Generative Models
Authors: Andrey Shutkin, Denis Parkhomenko, Ivan Kirillov, Kirill Chernyshev, Kirill Malakhov et al.
Categories: cs.CV, cs.LG, cs.SD Score: 6.58/10 (Obj:8 Id:5 Ind:5 Comp:6 Eff:8 Nov:5 Rep:8)
- Strength: KVAE 分词器在所有三种模态上均达到或超越了前沿开源基线——视频 PSNR 比 HunyuanVideo-1.5/Wan-2.2 高 +1 dB,音频重构在参数少 5 倍的情况下匹敌 SAME-L,并在所有基线上取得了 0.54-0.74 的人工偏好胜率——且代码和权重均在 MIT 许可下发布。
- Weakness: 核心新颖性在于已有方法的工程组合(CogVideoX 式因果 Conv3D,DAC 到连续 VAE 的转换,iREPA 的 CDS 指标),且关键音频对齐组件(目标模型 F 和损失形式 Lalign)未公开,削弱了识别和可复现性。
- Full review: Claude Code 全文七公理审稿
Latent diffusion modeling (LDM), a prominent paradigm, utilizes tokenizers to map input signal to compressed representation. This dependency positions tokenizer as an integral part of generation process itself, since it affects learning speed, quality of synthesized samples and lay foundation for later applications. This report presents series of KVAE tokenizers for audio, image and video, all designed for subsequent text-conditioned generation: KVAE-Audio, a continuous full-band 48 kHz tokenizer with a 50 Hz latent of 64 channels; KVAE-3D – two causal video tokenizers for 4x16x16 and 4x8x8 compression; KVAE-2D, an image model, compressing input by factor of 8 with 32 channels. We demonstrate that reconstruction (PSNR, LPIPS, PESQ, etc.) and generation results on objective (Frechet Distance, CLIP score, CLAP score, etc.) and subjective (side-by-side evaluation) metrics matches or surpasses frontier opensource tokenizers, such as VAEs from Wan-2.2, HunyuanVideo-1.5, FLUX.2, MovieGen, StableAudio and MMAudio. Considering difficulty of development, we share with community training details, model selection method and ablation on design choices. The code is publicly available at https://github.com/kandinskylab/kvae and https://github.com/kandinskylab/kvae-audio.
3. Audio-to-Score Transcription using Pre-trained Features, Data Augmentation, and the New SheetSage-A2S Dataset
Authors: Eoin Cummins, Zhongyi Huang, Alexandre D’Hooge, Zhuoro Mo, Yaolong Ju
Categories: cs.SD, cs.AI, cs.MM | Accepted at the 34th ACM International Conference on Multimedia (MM ‘26) Score: 6.58/10 (Obj:8 Id:5 Ind:6 Comp:5 Eff:8 Nov:6 Rep:9)
- Strength: 提出首个流行音乐 kern A2S 数据集(61h, 9468 clips, 真实录音),并在 Quartets 上将 SOTA SER 从 15.3% 降至 4.98%(67.5% 相对降低),数据集与代码完全公开。
- Weakness: 方法为三项已有技术的直接组合(Pre-Norm/FFN 扩展 + MuQ 冻结编码器 + 标准增强),无机制创新;MuQ 在 MSD 上预训练与 SheetSage-A2S 流行音乐存在风格重叠风险,未量化分离;训练超参数与架构修改混同消融不充分。
- Full review: Claude Code 全文七公理审稿
Existing audio-to-score (A2S) systems primarily focus on classical music, and the application to popular music remains underexplored. This paper first presents the new SheetSage-A2S Dataset, which includes 61 hours of audio with \texttt{**kern} score encodings for 9,468 clips originating from 6,066 unique songs, the first of its kind to facilitate A2S research for popular music. Additionally, we improve on existing A2S approaches by using data augmentation and MuQ, a pretrained feature-extraction model for music audio, to enhance generalisation abilities and extract meaningful audio features. Results show that the proposed A2S model achieves 4.98\% symbol error rate (SER) on the Quartets collection for classical music, which significantly outperforms the 15.3\% SER from the existing state-of-the-art \cite{alfaro-contrerasTransformer2024}. Additionally, our model achieves 20.92\% SER on the SheetSage-A2S dataset for popular music, serving as a strong benchmark for future research. The dataset, model, and code are made publicly available at: https://github.com/Multimodal-Music-Research-Lab/SheetSage2Kern_model.
4. Explicit and Stable Pseudospectral Time-Domain Method for the Föppl-von Kármán Equations
Authors: Victor Zheleznov, Stefan Bilbao
Categories: cs.SD, physics.comp-ph | To be presented at Forum Acusticum 2026, Graz, Austria, September 2026 Score: 6.50/10 (Obj:8 Id:5 Ind:9 Comp:6 Eff:5 Nov:6 Rep:9)
- Strength: 将 FvK 板模态合成的非线性耦合计算从 O(Mₓ⁴) 降至 O(Nₓ² log Nₓ),并证明了非线性势能非负性,实现显式且能量守恒到机器精度 (10⁻¹³) 的稳定时间积分,漂移调控降低两个数量级。
- Weakness: 核心 claim “降低计算代价” 无运行时间实证,无与模态方法 [2] / FDTD [1] / Kirby-Yosibash [8] 的精度或速度对比,无消融实验隔离各组件贡献,声音样本无正式评估。
- Full review: Claude Code 全文七公理审稿
Modal synthesis is a widely-used technique for simulation of musical instrument dynamics. In the linear case, a modal decomposition leads to an uncoupled system of damped and forced harmonic oscillators which can be efficiently solved by standard time-stepping methods. However, extensions to nonlinear problems are challenging due to the presence of products of modal expansions in the governing equations. In the case of the Föppl-von Kármán plate, the nonlinear coupling between the modes is described by a fourth-order tensor and is prohibitively expensive to evaluate in the modal domain. In this work, we propose a pseudospectral method in which the products are evaluated on a grid in the spatial domain while spatial derivatives are computed exactly in the modal domain. Discrete sine and cosine transforms between the modal and spatial domains are used to impose simply supported boundary conditions for the plate. Finally, we prove non-negativity of the nonlinear potential energy of the system and employ a scalar auxiliary variable technique for explicit and stable time integration in the modal domain. As a result, we reduce the computational cost of modal synthesis while preserving its advantages like a precise control over the simulated frequency range. Sound examples are presented.
5. Vorch-Streamer: Extending Human Audio-Visual Generation to Real-Time Long-Form Streaming
Authors: Menglin Han, Yang Ding, Yulei Lu, Haoran Yu, Xin Ma et al.
Categories: cs.CV, cs.SD | Project page: https://vorch-project.github.io/Vorch-Streamer-project/ Score: 6.30/10 (Obj:8 Id:6 Ind:5 Comp:5 Eff:8 Nov:5 Rep:6)
- Strength: 在 native T2AV 设定下首次实现 27.12 FPS 实时长序列(~2 分钟)流式生成,WER 7.92% 接近离线 LTX2.3 的 7.59%,远优于 OmniForcing (98.79%) 和 Hallo-Live (94.13%),且长序列端点指标几乎无退化(Sync-C 仅降 0.49%,Dynamic Degree 增 10%)。
- Weakness: 核心组件均为已有方法组合(Self Forcing + DMD + bounded KV-cache + Fun-CosyVoice),训练数据/teacher/初始化全部来自 LTX2.3 形成数据闭环,Stage 3 中 self forcing 与 DMD 的各自贡献未分离消融。
- Full review: Claude Code 全文七公理审稿
Real-time long-form avatar audio–video generation requires causal, continuous synthesis while maintaining audiovisual synchronization and visual consistency. Adapting a pretrained bidirectional model to this setting presents two key dilemmas. First, autoregressively reusing generated blocks as context creates exposure bias, causing errors and visual drift to accumulate over long rollouts. Second, a global speech utterance does not indicates a causal generator which portion should be spoken next when only limited local audio–video context is available. We present \textbf{Vorch-Streamer}, a post-training framework that addresses these challenges and enables real-time long-form Text-to-Audio-Video (T2AV) streaming. We construct a synthetic corpus of 80K avatar clips spanning 12–21 seconds and first train a causal generator with mixed Teacher Forcing and Diffusion Forcing. We then apply long-horizon Self Forcing with DMD distillation, exposing the model to its own rollout distribution while preserving the quality of the pretrained bidirectional teacher. To explicitly control speech progression, an external language model predicts discrete 25-Hz speech-planning tokens, whose continuous features condition the audio diffusion branch and align each causal block with the content it should speak. With bounded causal context and four-step denoising, Vorch-Streamer jointly generates audio and video from text at 27.12 FPS, exceeding the 24-FPS real-time playback rate while maintaining competitive audio–lip synchronization and strong identity preservation over long-form generation.
6. EG-VAE: A Unified Framework for Electric Guitar Tone Transfer and Removal
Authors: Yen-Tung Yeh, Yun-Ning, Hung, Yi-Hsuan Yang
Categories: eess.AS Score: 6.20/10 (Obj:8 Id:6 Ind:5 Comp:6 Eff:8 Nov:6 Rep:6)
- Strength: 首次统一 EGTT 和 EGTR,通过 tone masking 单一机制同时实现解耦训练和 removal 推理,EGTT Mel distance 在 seen tones 上较最强 baseline 降低 44%(0.86 vs 1.53),主观 tone similarity 接近 Oracle 上界(4.30 vs 4.48)。
- Weakness: 方法由 5 个解耦机制 + 两阶段训练组成,组件较多且多数直接来自语音转换已有工作(DSAE/CLN/pvpGD/β-VAE);EGTT baseline 的 “w/ EGTR” 配置依赖 EG-VAE 自身输出形成循环比较,且数据生成与评测同源于闭源商业插件,独立复现困难。
- Full review: Claude Code 全文七公理审稿
Electric guitar tone transfer (EGTT) and tone removal (EGTR) are two fundamental tasks in guitar tone modeling: EGTT replaces a recording’s tone with that of a reference, while EGTR recovers the dry direct-input (DI) signal from a wet, processed recording. Despite their highly related nature, prior work has addressed them independently, and both works have yet to achieve satisfactory results. In this paper, we propose EG-VAE, a unified framework that jointly models EGTT and EGTR by disentangling frame-level content and global tone representations from wet recordings with a variational autoencoder. EGTT is achieved by recombining a source’s content with a reference’s tone, while EGTR is attained by a novel tone masking objective that enforces content-tone disentanglement during training and realizes the removal procedure at inference. To improve transfer to tones unseen in training, a second training stage shapes a smooth tone space through variational sampling and audio-effects augmentation. Experimental results from both objective and subjective evaluations demonstrate that EG-VAE outperforms task-specific baselines on transfer and removal. Demos are available at https://guitar-tone-demo.vercel.app/.
7. Rethinking Automatic Music Mixing as Sequential Stem Blending
Authors: Yen-Tung Yeh, Chung-Jui Chan, Yun-Ning, Hung, Yi-Hsuan Yang
Categories: eess.AS Score: 6.16/10 (Obj:8 Id:5 Ind:6 Comp:8 Eff:5 Nov:8 Rep:6)
- Strength: 提出 sequential stem blending 作为 AMM 的新范式重构,在 stem blending benchmark 上大幅超越 baseline(KAD-CLAP -0.01 vs DMC 46.45),perceptual MUSHRA 测试 6 例中 5 例最高分,且自然支持任意 stem 数量和 interactive workflow。
- Weakness: 缺少同 backbone/同数据/同 compute 的 parallelized vs sequential 公平比较(识别公理缺口),AMM 全任务在 tonal balance 和 mixing style similarity 上未达 SOTA(论文自认),且 stem blending benchmark 与训练数据合成方式相同存在循环论证风险。
- Full review: Claude Code 全文七公理审稿
Automatic music mixing, the task of automatically combining individual audio tracks into a cohesive mixture, is typically addressed by parallelized architectures that process all input tracks in a single pass. In this work, inspired by how human mix engineers process stems one at a time, we propose a paradigm shift and ask whether automatic music mixing can be reformulated as a sequential stem blending task, where each stem is blended into a growing submix. Specifically, we train a latent flow matching model conditioned on the submix context, enabling sequential processing of an arbitrary number of input tracks. To train the model, we introduce a degradation-based data synthesis strategy that simulates realistic stem blending scenarios from existing multitrack and source separation datasets. Experimental results on both stem blending and automatic music mixing benchmarks demonstrate the effectiveness of the proposed approach. We provide audio examples on the accompanying demo page\footnote{https://sequential-mixing-demo.vercel.app/}.
8. AffectDF: The Most Comprehensive Benchmark for Speech Deepfake Detection against Emotionally Expressive Attacks
Authors: Aurosweta Mahapatra, Xiutian Zhao, Shreeram Suresh Chandra, Zihan Zhang, Zongyang Du et al.
Categories: eess.AS Score: 6.10/10 (Obj:9 Id:6 Ind:5 Comp:8 Eff:6 Nov:5 Rep:8)
- Strength: 21 种攻击 × 5 种情感 × acted/spontaneous 的 ~260h 基准数据集,系统揭示 SDD 系统在情感 spoofing 下 EER 高达 59.71%(接近 random),且大规模情感训练不持续提升跨域鲁棒性
- Weakness: 训练集仅含 LALM-EVC+TTS 攻击(无 VC)、仅 4 speakers,且 Qwen 同时作为攻击源和检测器存在潜在循环;因果识别深度不足,核心矛盾现象(如 RawNet2 最简模型在 AffectDF 训练后最优)未获解释
- Full review: Claude Code 全文七公理审稿
Speech deepfake detection (SDD) systems achieve strong performance on conventional benchmarks; however, existing datasets provide limited coverage of emotionally expressive and recent large audio-language model (LALM)-based attacks. Existing emotional spoofing datasets are also limited in scale and attack diversity, typically covering only voice conversion (VC) or text-to-speech (TTS) attacks. We introduce AffectDF, the most comprehensive benchmark for emotionally expressive speech deepfakes, spanning TTS, VC, emotional VC, and LALM-based spoofing attacks across both acted and spontaneous emotional speech. AffectDF contains approximately 260 hours of speech generated using 21 spoofing attacks across five emotional states. We benchmark state-of-the-art SDD systems under conventional and emotional spoofing conditions, including LALM-based detectors evaluated with both inference-only prompting and supervised fine-tuning. Our experiments reveal severe robustness degradation when models trained on conventional benchmarks are evaluated on AffectDF, with several systems approaching near-random performance. Surprisingly, even large-scale emotional training does not consistently improve cross-domain robustness, indicating that current SDD systems fail to learn generalized spoof representations under emotional and prosodic variability. Robustness further varies substantially across emotional states, attack families, and acted vs spontaneous emotional speech conditions. These findings expose fundamental limitations of current SDD systems and establish AffectDF as a benchmark for developing more robust spoof detection models.
9. Enhancing Social Intelligence in LLMs with Hierarchical Reasoning and Utterance-Level Goal Rewarding
Authors: Xiaofeng Wang, Kakam Chong, Shuai Xiao, DeXin Kong, Qingyuan Tian et al.
Categories: cs.CL Score: 6.10/10 (Obj:8 Id:6 Ind:5 Comp:5 Eff:8 Nov:5 Rep:8)
- Strength: 在 SOTOPIA benchmark 上用 Qwen2.5-7B 达到 SOTA(AVG 6.02),超越 GPT-4o (5.60) 和 AMPO (5.92),variance-gated reward routing 机制无已有工作,代码和数据开源。
- Weakness: 方法是已有组件(HRL + GRPO + boundary-aware scaling)的组合加一个启发式路由规则,TPB 理论包装缺乏实验验证,训练-评测存在 GPT-4o 同源闭环,人类评测仅 50 样本。
- Full review: Claude Code 全文七公理审稿
Large language models (LLMs) excel in structured tasks but struggle with dynamic social interactions, where success requires long-term goal coordination and rapid adaptation. Current methods often apply uniform goal-based rewards to every utterance, overlooking the specificity of objectives at each dialogue turn and failing to account for the rationale of potential strategies. Inspired by the Theory of Planned Behavior, we propose the Think-Strategy-Response (TSR) framework, which decomposes social dialogue into two hierarchical stages: high-level strategic planning and low-level linguistic execution. To optimize TSR, we introduce Linearized Hierarchical Reinforcement Learning with Variance-Gated Rewards (LHRL-VGR), a novel algorithm that dynamically routes rewards - balancing goal completion and strategy adherence - based on the variance of goal achievement scores. Experiments on the SOTOPIA benchmark show that our approach fine-tunes a Qwen2.5-7B agent to surpass the GPT-4o baseline by 7.32% in goal completion success, demonstrating state-of-the-art performance in multi-agent social negotiation tasks.
10. How to Recognize New Words: A Comparison Between Context Biasing Methods and Speech LLMs
Authors: Christian Huber, Alexander Waibel
Categories: cs.CL Score: 5.90/10 (Obj:8 Id:5 Ind:8 Comp:6 Eff:6 Nov:5 Rep:5)
- Strength: 系统性 head-to-head 对比两种 context biasing 与三种最新 speech LLM,发现 context biasing method B 在 N=0 时 BWER 最多降低 88% relative,且 speech LLM 加 250 distractors 后 BWER 退化 39%-570%。
- Weakness: 跨方法对比存在未控制变量(模型规模 1.5B-30B、训练数据量、架构),baseline 覆盖仅 2 类 context biasing 方法,训练/prompt 细节不完整影响复现。
- Full review: Claude Code 全文七公理审稿
Recognizing new and rare words - named entities, acronyms, domain specific special words, and other items scarce in training data - remains a key challenge for automatic speech recognition (ASR). We compare two strategies for this: context biasing methods, where an ASR model is extended such that during inference a word list can be supplied, and speech large language models (LLMs) prompted with context directly. We evaluate two context biasing methods based on Whisper against three speech LLMs across read and non-read speech, reporting biased, unbiased, and overall word error rate (WER). The context biasing methods cut biased WER by up to 88% relative while leaving other words largely unaffected. Speech LLMs excel on read speech but generalize less well to non-read speech, and prove sensitive to distractor count and prompt word order. We characterize the resulting trade-offs to guide method selection.
11. Whence the Voice? Self-supervised Dual-source Audio-Visual Localisation via Selective Convergence
Authors: Han Hu, Dongheng Lin, Yuqi Hou, Haotian Li, Hyung Jin Chang et al.
Categories: cs.MM Score: 5.42/10 (Obj:8 Id:5 Ind:6 Comp:6 Eff:5 Nov:5 Rep:4)
- Strength: 发现对比学习在多源 audio-visual localisation 中的 selective convergence 现象并加以利用,两阶段框架在 self-supervised 方法中取得最优 CI ratio/cIoU,同时指出 heatmap-vs-bbox 评测的系统偏差并引入 pixel-level mask 评测。
- Weakness: 选择性收敛的成因未隔离,两阶段框架缺乏充分的消融实验证明各阶段独立必要性,绝对指标仍远离实用水平,且仅验证双源设置泛化性受限,代码与新增标注的开源状态未明确。
- Full review: Claude Code 全文七公理审稿
Localising multiple sound sources in visual scenes remains a fundamental challenge in multimodal perception due to an inherent circular dependency: separating mixed audio requires knowing source locations, while identifying sound-producing regions requires separated audio signals. In this paper, we focus on the dual-source setting and discover a selective convergence in self-supervised audio-visual learning: when presented with multiple sound sources, contrastive models naturally converge to the most salient audio-visual correspondence rather than attempting to represent all sources equally. This emergent phenomenon, analogous to human selective auditory attention, enables us to break the above circular dependency through a progressive two-stage framework: first, leveraging selective convergence to identify dominant sources, and then exploiting these learned priors to uncover remaining sources. Our self-supervised approach achieves the best performance among self-supervised methods on dual-source benchmarks without requiring any manual annotations, and even surpasses some weakly-supervised approaches \red{on certain metrics. Furthermore, we identify a fundamental evaluation inconsistency in existing benchmarks: comparing continuous localisation heatmaps against bounding-box annotations creates systematic biases, particularly for non-axis-aligned objects where the bounding box includes substantial background regions. To address this, we introduce pixel-level segmentation masks to the existing benchmark, enabling spatially-aligned evaluation. Together, these results suggest that embracing rather than suppressing selectivity offers a scalable, annotation-free route to multi-source localisation.
12. Decolonizing Linguistic Policies in Automated Speech Recognition: A Framework for Cross-Culturally Competent Speech AI
Authors: Jay L. Cunningham, Mark Atta Mensah, Richard Martinez, Joao Vieira da Silva Neto, Efi Dawodu
Categories: cs.CL, cs.CY, cs.HC | 10 Pages, 2 Figures, 2 Tables, Interspeech 2026 - Sydney, Australia Score: 5.30/10 (Obj:6 Id:5 Ind:8 Comp:5 Eff:5 Nov:5 Rep:4)
- Strength: 论文将四条社会学理论传统综合应用于 ASR 领域,提出 3M taxonomy(Misrecognition/Misalignment/Mistrust)将 ASR 伤害从单一 WER 扩展为三个正交维度,填补了 ASR fairness 研究中语用和关系层面伤害的分类空白
- Weakness: 框架无任何实证验证(无 pilot study、无 case study、无与 ASR-FAIRBENCH 等已有工具的比较),audit protocol 缺少操作化参数(评估者数量、测试集规模、评分量表),其他团队无法直接复现
- Full review: Claude Code 全文七公理审稿
This paper focuses on automatic speech recognition (ASR) and ASR-mediated voice interfaces that shape access to public services, healthcare, and education. We argue that persistent failures for low-resource, Indigenous, and non-standard language varieties are not only technical errors, but also implicit linguistic policies that reproduce colonial language hierarchies. Drawing on linguistic capital, raciolinguistic ideology, language policy research, and decolonial computing, we show how data, metrics, and model priors determine whose voices become machine-legible. We introduce the Three Harms (3M) taxonomy—Misrecognition, Misalignment, and Mistrust—and a seven-layer situatedness model for linguistic diversity in ASR and ASR-mediated voice interfaces. We then propose a participatory framework and minimum audit protocol for culturally competent ASR, positioning affected communities as co-designers, evaluators, and governance partners.
13. Numerical Model of a Multiple-Input-Multiple-Output Distributed Acoustic Sensing System with Joint Phase and Birefringence Estimation
Authors: Diane Prato, Mehran Mokhtari Sheramin, Renaud Gabet, Elie Awwad
Categories: eess.SP, physics.optics Score: 4.90/10 (Obj:7 Id:3 Ind:8 Comp:5 Eff:5 Nov:5 Rep:4)
- Strength: 通过仿真和实验验证了 MIMO-DAS 系统中相位与双折射的联合估计可实现纵向/横向应变判别(1.3m 空间分辨率,3km 仿真光纤 + 2km 实验光纤 spool),物理模型基于第一性原理且实验验证独立。
- Weakness: 缺少与任何 baseline(∆ϕ-OTDR、P-OTDR)的定量对比和消融实验,所有结果为定性图示,核心判别逻辑直接来自光弹理论已知推论,且代码/数据均未开源。
- Full review: Claude Code 全文七公理审稿
In this work, we introduce and experimentally validate a numerical model for a Multiple-Input-Multiple-Output Distributed Acoustic Sensing (MIMO-DAS) system that accounts for dynamic perturbations of fiber birefringence and of the common optical phase of the backscattered signal (or polarization-averaged phase, shared by both polarization tributaries). The MIMO-DAS system probes the fiber using polarization-multiplexed constant-power coded sequences that are suited for coexistence of DAS with WDM data transmission over the same fiber. We study the effect of both axisymmetric and anisotropic events on the two quantities. We demonstrate, through numerical simulations and lab experiments, the estimation of effective birefringence magnitude in static conditions, and the joint estimation of common phase and effective birefringence magnitude in the case of dynamic longitudinal strain and anisotropic transverse strain. This allows for event discrimination and increased sensitivity to disturbances that act transversely on the fiber, since polarization will be responsive to perturbations that break cylindrical symmetry, while the phase will strongly respond to longitudinal strain.
14. Beyond Residual Connections: Manifold-Constrained Hyper-Connections for Robust Speaker Representation Learning
Authors: Zezhong Jin, Xiaoyu Wang, Zhe Li, Chong-Xin Gan, Zilong Huang et al.
Categories: cs.SD, cs.MM | Accepted to INTERSPEECH 2026 Score: 3.95/10 (Obj:5 Id:5 Ind:9 Comp:2 Eff:4 Nov:1 Rep:6)
- Strength: mHC 作为 drop-in replacement 在 4 种 speaker recognition backbone 上实现一致 EER 降低,以近乎零参数/FLOP 开销在 ECAPA-TDNN-L 上将 VoxCeleb1-O EER 从 0.87% 降至 0.77%,且 HC vs mHC 消融验证了双随机约束的有效性。
- Weakness: 核心方法 mHC 完全来自 Xie et al. (arXiv:2512.24880) 的已有工作,本文仅为将其迁移到 speaker recognition 的直接应用,无领域适配创新、无机制分析、无与同期 SOTA 方法的对比,且未提及代码开源。
- Full review: Claude Code 全文七公理审稿
Residual connections are fundamental to deep speaker recogni- tion models, such as ECAPA-TDNN and ResNet. However, standard identity mapping limits information flow to a sin- gle path, constraining representation capacity. We introduce Manifold-Constrained Hyper-Connections (mHC), reformulat- ing residual paths as a multi-stream evolution where informa- tion is mixed through a doubly stochastic matrix. By employing Sinkhorn-Knopp iterations, mHC ensures energy conservation by preserving signal intensity and feature mean, which stabi- lizes gradients and mitigates signal degradation in complex net- works. We evaluate mHC by replacing standard residual con- nections in backbones including ECAPA-TDNN, ResNet-34, Res2Net, and E-Res2Net. Extensive experiments on VoxCeleb1 demonstrate that mHC connections consistently enhance per- formance across all architectures, highlighting its effectiveness for robust speaker representation learning.
15. Vorch-Omni: Multi-Task Orchestration of Sight and Sound
Authors: Vorch Team, Xiaoyu Chen, Yang Ding, Cong Han, Menglin Han et al.
Categories: cs.CV | Project Page: https://vorch-project.github.io/Vorch-Omni-project/ Score: 3.70/10 (Obj:6 Id:2 Ind:4 Comp:5 Eff:5 Nov:3 Rep:2)
- Strength: 在单一 flow-matching DiT 上集成 VLM+VAE 双条件通路与 token-level mask/task-id/position-type 机制,覆盖 10+ 音视频任务并在 VBench/FAD 上展示定量结果。
- Weakness: 未与 UniForm/APOLLO/UniVerse-1 等直接竞争者定量对比,未引用同团队前序工作 TIDE (8 位重叠作者),无消融隔离各组件贡献,无代码/数据/checkpoint 开源。
- Full review: Claude Code 全文七公理审稿
Recent advances in generative video modeling have enabled diverse generation, reference-based synthesis, extension, and editing, but existing approaches often rely on fragmented task-specific models. A general model must distinguish heterogeneous target, source, and reference signals to determine what to generate, preserve, or use as guidance, while reducing interference among tasks. Joint audio-visual generation further increases this challenge by introducing diverse conditioning and output configurations across modalities. We present Vorch-Omni, a unified multi-task framework for audio-visual synthesis based on an arbitrary-condition-to-arbitrary-output formulation. It flexibly treats video and audio signals as either conditioning inputs or generation targets. Token-level conditioning masks and task identifiers distinguish targets, source content, and references, while position types separate temporal context from independent conditions. To capture semantic and structural information, Vorch-Omni employs complementary visual conditioning pathways: a vision-language model interprets sampled frames with text instructions, and a video VAE encodes conditions into latent tokens for direct guidance. We further build a distributed data pipeline to curate diverse temporally aligned audio-visual clips, generate structured captions and metadata, and balance heterogeneous task distributions. Built on a single flow-matching diffusion transformer without task-specific architectural changes, Vorch-Omni supports over 10 tasks, including text-to-video, text-to-audio-video, image- and reference-conditioned generation, temporal extension, audio-driven generation, video transformation, and audio-visual editing. This unified framework provides a scalable foundation for general-purpose audio-visual generation and manipulation.