Daily Papers — 2026-07-25
8 papers on audio, speech, music, and acoustics.
1. Infinite Canons: Maximally Self-Similar Melodic Lines and Canons with Infinite Solutions
Authors: Clifton Callender
Categories: cs.SD, math.CO, math.HO | 19 pages, 15 figures. Versions of this material were presented at the Joint Mathematics Meetings (2011), Nancarrow in the 21st Century (2012), and the joint AMS/SEM/SMT meeting (2012). A more developed treatment, with musical applications, an interactive program, and greater mathematical detail, is in preparation Score: 6.26/10 (Obj:8 Id:5 Ind:8 Comp:8 Eff:5 Nov:7 Rep:4)
- Strength: 通过素因子分解同态 φ:(Q₊,×)→(R,+) 构造在所有有理速度比下和声一致的极大自相似旋律线,并发现逆行倒置在极大自相似下坍缩为移位(公式10),使任意声部数/速度比/方向的卡农成为可能。
- Weakness: 多项关键证明明确标注 deferred/forthcoming(§2.5 非周期性扩展、§6 更完整数学处理),连续构造的量化误差未分析,无代码或交互程序发布,当前版本是不完整的研究报告。
- Full review: Claude Code 全文七公理审稿
Infinite Canons is an ongoing series of canons with infinite solutions. More specifically, each canon is based on a melodic line that can be combined in any number of voices, at any tempo ratios (rational or irrational), and with each voice moving either forward or in retrograde inversion, while maintaining harmonic consistency. This paper describes the structure of these maximally self-similar melodic lines based on two different constructions: 1) a discrete prime-factorization approach yielding self-similarity under all rational ratios, and 2) a continuous logarithmic approach extending this to irrational ratios. In both cases the vertical interval between voices in a tempo ratio of $λ_i / λ_j$ is given by a homomorphism $φ(λ_i / λ_j)$, which is a constant independent of time. Furthermore, under these constructions retrograde inversion collapses to transposition, allowing for all manner of table canons. These structures are demonstrated with suggestive realizations of several different infinite canons. Future work includes a more complete mathematical treatment, musical applications, and an interactive program that allows users to explore an unlimited number of realizations of these pieces.
2. Speech Signals Complement LLMs for Predicting Interpersonal Attraction in Speed Dating
Authors: Yuriko Kikuchi, Takato Hayashi, Ryusei Kimura, Naoya Inoue, Ryo Ishii et al.
Categories: cs.CL | Accepted at the 28th ACM International Conference on Multimodal Interaction (ICMI 2026). 16 pages total, including a 7-page supplementary appendix Score: 6.20/10 (Obj:8 Id:8 Ind:9 Comp:5 Eff:5 Nov:5 Rep:8)
- Strength: 在单一控制研究中,采用经过仔细隔离的分支(Claude 零样本文本 vs. 冻结 HuBERT 语音)和诚实的条件发现,首次经验性测试了独立训练的语音预测器是否在基于文本的 LLM 预测之上,为快速约会吸引力预测提供增量价值——在所有四个条件下 PW 显著提升(+.025 到 +.054, Holm p<.05),而 Δr 未达到显著性。
- Weakness: 所有增益都很小且未校准(最佳 r = .323,远低于共识参考 r ≈ .42),结果仅限于 147 名日本参与者和一种模型代际,Δr 是主要指标但从未达到显著性,且对于哪些声学信号驱动互补性缺乏机制洞察。
- Full review: Claude Code 全文七公理审稿
Large language models (LLMs) can predict interpersonal attraction from conversation transcripts, but it remains unclear what a speech predictor can add beyond transcript-only LLM prediction. Using Japanese speed-dating conversations, we combine predictions from a transcript-only LLM and a supervised speech predictor to estimate participants’ reported liking of their partners. We show that speech can complement transcript-only LLM prediction, but that this complementarity is conditional rather than universal. Combining the two predictions significantly improves pairwise ranking accuracy over the transcript-only LLM alone in all evaluated conditions. By contrast, gains in per-participant Pearson $r$ vary across conversation rounds and rating directions, with none significant after correction. Retrospectively, these $r$ gains are concentrated among participants for whom the speech predictor is more accurate. Speech can therefore retain predictive value even when an LLM predicts attraction from transcripts. The relevant question is not simply whether speech helps, but where its complementarity emerges.
3. OmniMate: Open-Ended Real-Time Streaming Audio-Visual Generation for Interactive Avatars
Authors: Quanyue Song, Yishan He, Yanbo Ding, Zhixiang He, Yongxiang Li et al.
Categories: cs.CV Score: 5.84/10 (Obj:8 Id:5 Ind:4 Comp:5 Eff:8 Nov:5 Rep:4)
- Strength: 在 LTX2.3 backbone 上以 GPC+MRCM 实现最高 FPS(27.64)、最低 TTFF(3.49s)、最佳 Speech↓(0.097)与 V-ID(0.957),240s 长时段无显著 identity 衰减。
- Weakness: GPC 核心编码机制继承自 OmniAvatar,代码/checkpoint/自建数据未开源,无最简 baseline 对照与组件 synergy 论证,评测 benchmark 为作者群自建 VerseBench 扩展,独立性存疑。
- Full review: Claude Code 全文七公理审稿
Recent advances in diffusion-based generative models have enabled real-time audio-driven avatar generation and unified audio-visual synthesis, providing a promising foundation for interactive avatar systems. However, extending these models to real-time interactive streaming remains challenging, as the generation horizon is unknown in advance and cross-modal identity consistency gradually degrades during long-term generation. To address these challenges, we propose OmniMate, a unified framework for open-ended real-time interactive audio-visual avatar generation. OmniMate jointly synthesizes visual content, speech, and audio effects in real time, enabling natural and immersive multi-turn interactions. To achieve adaptive response progression, we introduce a Generation Progress Controller (GPC) that explicitly models the generation progress of each streaming chunk, allowing the model to complete responses according to the desired progress and achieve seamless transitions between execution and listening states. To preserve long-term cross-modal identity consistency, we propose a Multi-Reference Conditioning Module (MRCM), which leverages multiple reference images and a reference speech segment to provide persistent visual and speaker identity cues throughout long-duration streaming interactions. Extensive experiments on an interaction-oriented adaptation of VerseBench demonstrate that OmniMate achieves high-quality, low-latency streaming generation while maintaining strong long-term audio-visual consistency. The results further show that OmniMate supports realistic, coherent, and responsive interactive avatar experiences over extended multi-turn conversations.
4. PathRIR: Physics-Guided Acoustic Path Selection and Late-Tail Compensation for Fast Room Impulse Response Simulation
Authors: Shaoheng Xu, Chunyi Sun, Jihui Zhang, Amy Bastine, Prasanga N. Samarasinghe et al.
Categories: eess.AS, cs.LG, cs.SD | Accepted for publication in the Proceedings of the 19th International Workshop on Acoustic Signal Enhancement (IWAENC 2026). Project page: https://github.com/ShaoHenry/PathRIR Score: 5.42/10 (Obj:8 Id:5 Ind:4 Comp:5 Eff:5 Nov:6 Rep:5)
- Strength: 子树级 ISM 路径修剪是一种新颖的应用,在 Omax=10 时实现了 90.5% 的节点减少和 279 倍的速度提升(相对于 Pyroom),且 Compensation-MLP 显著改善了衰减指标(EDC-Err 从 18.60 到 4.69 dB)。
- Weakness: 评估处于仿真器与仿真器的闭环中(Full-ISM 生成标签,Pyroom 作为参考),没有真实世界的 RIR 验证,并且缺失关键基准(FRA-RIR, gpuRIR),同时绝对精度中等(RT60 误差为 36.84 ms,采样率仅为 8 kHz)。
- Full review: Claude Code 全文七公理审稿
Image-source-method (ISM)-based room impulse response (RIR) simulation is a useful and physically interpretable tool for acoustic scene modeling, but full-order ISM becomes computationally expensive as the reflection order and room complexity increase. We propose a physics-guided framework for fast RIR simulation that preserves the geometric structure of ISM while learning to retain only acoustically important image-source paths during online traversal. To recover energy removed by pruning, the proposed PathRIR uses a lightweight compensation multilayer perceptron to predict the missing late-tail energy envelope and generate a compensation tail whose energy follows that envelope. Experiments on irregular 3D rooms show that PathRIR reduces image-source computation and improves runtime efficiency over a full-order ISM simulator, while achieving low waveform- and decay-related errors. Ablation results show that adding the compensation tail improves waveform fidelity and reduces energy-decay-curve error, reverberation-time error, and direct-to-reverberant-ratio error, with modest runtime overhead.
5. Music-Source-Separation-Training (MSST): A Unified Framework for Training and Evaluating Music Demixing Models
Authors: Roman Solovyev, Ilya Kiselev, Alexander Stempkovskiy, Tatiana Gabruseva
Categories: cs.SD, cs.LG, eess.AS Score: 5.26/10 (Obj:8 Id:3 Ind:5 Comp:5 Eff:5 Nov:3 Rep:9)
- Strength: 开源框架代码完整、MIT 许可、社区生态活跃(3 个衍生 GUI、7+ 类社区模型),填补了 Demucs/nussl archived 后多模型族 MSS 训练工具的维护空白。
- Weakness: 唯二 ablation(TTA ΔSDR +0.00~0.06 dB、ensembling +0.01~0.2 dB)处于噪声级且无统计检验,LoRA 零定量结果,未与 Asteroid/Open-Unmix 做同 compute head-to-head 对比,所有组件均有已有来源且被论文自引用,无新方法或新认知。
- Full review: Claude Code 全文七公理审稿
Music Source Separation (MSS), the task of recovering individual sound components (stems) from a polyphonic mixture, is central to applications ranging from karaoke and remixing to audio restoration and content production. The separation quality depends on engineering decisions across the entire pipeline: model choice, training data preparation and augmentation, loss function and metrics choice, training configuration, validation, and post-processing. This paper presents MSST (Music-Source-Separation-Training) - a universal open-source framework for MSS tasks, which unifies training, validation, and inference for a broad range of modern demixing model families under a single, configuration-driven interface. The framework supports various model architectures, data preprocessing and augmentations, multiple loss functions and evaluation metrics, which helps with fast iterations and ablation studies. Additionally, the framework supports a range of practical techniques that improve separation quality, such as sliding-window inference with cross-fading, test-time augmentation, model ensembling, and fine-tuning via Low-Rank Adaptation (LORA). Our ablation studies demonstrate improvements of MSS using the above techniques. By consolidating these components into a reproducible, YAML-configurable framework, MSST lowers the barrier to systematic experimentation and enables rapid iteration from idea to verifiable result.
6. Low-Latency Turn-Taking via Context-Aware Preface Generation in a Real-World Dialogue Robot
Authors: Yuki Okafuji, Koji Inoue, Yoshiki Ohira
Categories: cs.RO, cs.CL, cs.SD | 5 pages, 4 figures. Accepted at ICMI LBR 2026 Score: 4.74/10 (Obj:6 Id:5 Ind:6 Comp:5 Eff:5 Nov:4 Rep:4)
- Strength: 在真实购物商场部署两阶段增量响应框架并报告延迟权衡,contextual-preface 相对 fixed-filler 将 initial-to-main gap 从 0.74s 显著缩短至 0.43s(p<.001),且诚实报告 initial latency 反而显著更长(1.15s vs 0.94s, p=.027)及 8.0% 的 preface breakdown 率。
- Weakness: 识别公理存在缺口——contextual-preface 与 fixed-filler 同时在 LLM 调用、生成成本、VAP 触发比例上不可比,initial-to-main gap 的缩短可能源自 gpt-4o vs gpt-4o-mini 生成速度差而非框架设计;主观效用全部 null(4 项 Likert all p>.05),代码/数据/模型均未开源且核心依赖闭源 API,新颖性主要来自已有组件(VAP + intent cl。
- Full review: Claude Code 全文七公理审稿
Large language model (LLM)-based dialogue systems suffer response delays because generation begins only after final speech recognition. While fixed fillers are a workaround, they become unnatural over time. We propose a two-stage incremental framework that decouples prefatory-response preparation from speech onset. Once user intent becomes predictable, an intent readiness detector triggers LLM-based generation of a short prefatory response. Concurrently, a voice activity projection (VAP) model determines when to deliver it. Through a field experiment with a route-guidance robot in a shopping mall, we evaluated three conditions: no-filler, fixed-filler, and contextual-preface. Both fixed-filler and contextual-preface significantly reduced initial response latency relative to no-filler. Relative to fixed-filler, contextual-preface had significantly longer initial response latency but a significantly shorter initial-to-main gap. Exploratory ratings showed no significant differences. These results indicate a timing trade-off.
7. Singlish, Can or Not? Fine-Tuning and Evaluating Zero-Shot TTS for Singapore English
Authors: Ivan Kukanov, Zheng Xin Chai
Categories: eess.AS | 7 pages, 5 figures, 6 tables. Submitted to SLT 2026 IEEE Score: 4.47/10 (Obj:6 Id:5 Ind:5 Comp:3 Eff:4 Nov:4 Rep:5)
- Strength: 首次使用两个现代 ZS-TTS 模型对 Singlish 口音 TTS 进行系统基准测试,表明微调使 ACC-SIM 提高多达 +0.126,并泛化至未见的说话人(域外 0.5617 > 现成域内 0.5114),证明了口音学习而非记忆。
- Weakness: 无人类评估,无与替代口音适配方法的比较(AccentBox, adapters),同一 ASR 用于数据过滤和评估导致循环偏差,微调后绝对 ACC-SIM 仍然很低(0.56-0.65),且“首次”的新颖性声明忽略了先前的 NTU 学位论文和 Singlish TTS 的开源项目。
- Full review: Claude Code 全文七公理审稿
Zero-shot text-to-speech (ZS-TTS) achieves near-human quality for standard English, but it copies regional accents poorly. Prompted with a short Singlish utterance, state-of-the-art systems reproduce a speaker’s timbre while flattening the accent toward generic English. We investigate whether targeted fine-tuning off-the-shelf ZS-TTS can close the gap for Singapore English (Singlish). We fine-tune two cutting-edge ZS-TTS models, Chatterbox and CosyVoice 3, on 50 Singlish speakers from the IMDA National Speech Corpus. Three speech distributions are evaluated: real recordings against off-the-shelf and fine-tuned generation driven by the same Singlish audio prompts. The evaluation covers four dimensions: naturalness, intelligibility, speaker similarity, and accent similarity. We separate adaptation (in-domain speakers seen during fine-tuning) from consistency (held-out speakers) to test whether accent transfer generalises beyond the training data. Fine-tuning raises accent similarity on in-domain and out-of-domain speakers for both Chatterbox and CosyVoice 3. It moves the generated distribution measurably toward real Singlish, with the gain persisting on held-out speakers. To our knowledge, this is the first systematic study of Singlish-accented TTS.
8. Explainable AI through the Lens of Material Agency: Enabling Musical Interface Design with Neural Audio Models
Authors: Shuoyang Jasper Zheng, Anna Xambó Sedó, Nick Bryan-Kinns
Categories: cs.HC, cs.SD | Under review for “Explainable AI for the Arts” (N. Bryan-Kinns, Ed.), Springer Score: 3.05/10 (Obj:5 Id:2 Ind:4 Comp:2 Eff:2 Nov:2 Rep:8)
- Strength: 开源 Latent Terrain 工具箱(GitHub + 文档 + 教程 + 预训练模型)为 NIME 社区提供了可用的神经音频模型探索资源,4个艺术家项目展示了设计空间多样性。
- Weakness: 核心概念”material explainability”来自 XAIxArts Manifesto(论文自承),工具技术来自作者已发表的 Frontiers 2025/NIME 2026 论文,三个建议来自 Wekinator/FluCoMa 已有经验——无增量知识贡献,无任何因果识别或效用评估。
- Full review: Claude Code 全文七公理审稿
Recent work in Human-Computer Interaction (HCI) increasingly treats AI models as design materials that have distinctive computational properties to shape design artifacts. Artists learn to work with the model “at play” to explore their emerging properties. The aim of explainability, in this view, is to make visible a crafting and hacking space to enable sustained creative practices with AI. In this chapter, we propose material explainability as a range of activities and artifacts that transform AI models into accessible and inclusive design materials in the workspace of artists, designers, and makers. We present a case study of building a repository of resources to enable artistic explorations of neural audio models in New Interfaces for Musical Expression (NIME) design. Reflecting on our community-building journey and the making of a collection of musical interface designs with a group of artists, we raise three recommendations on enabling the exploration of AI as materials in artistic practices to inspire future XAI design for artists.