Sitemap
A list of all the posts and pages found on the site. For you robots out there, there is an XML version available for digesting as well.
Pages
每日论文
每日 arXiv 音频/语音/音乐/声学论文 — 按评分排序
Posts
paper_reviews
A Text-Steerable Instrument for Sketching Procedural Soundscapes via Language Models
Published:
Paper Review: A Text-Steerable Instrument for Sketching Procedural Soundscapes via Language Models
Enhancing Flow Matching with A Unified Guidance Framework for Efficient and Robust Speech Synthesis
Published:
Paper Review: Enhancing Flow Matching with A Unified Guidance Framework for Efficient and Robust Speech Synthesis
From Objectives to Applications: Aligning Architectural Biases in Audio Self-Supervised Learning
Published:
Paper Review: From Objectives to Applications: Aligning Architectural Biases in Audio Self-Supervised Learning
Speech Playground: An Interactive Tool for Speech Analysis and Comparison
Published:
Paper Review: Speech Playground: An Interactive Tool for Speech Analysis and Comparison
AmbiDrop: Ambisonics-Based Array-Agnostic Neural Speech Enhancement
Published:
我已经掌握了完整的情况。让我整理一下关键发现并撰写评审。
What’s a Credit Worth? A Market Framework for Attribution-Aware Compensation in Generative Music
Published:
我现在已经掌握了充分的事实依据。我找到了关键的先前研究 —— Deng 等人 2023 年的 “Computational Copyright” (arXiv:2312.06646),这是由同一团队(Deng、Jiang、Donahue、Ma)完成的最直接的先前研究,此外还有 Choi 等人 2025 年关于通过 unlearning 进行音乐归因的研究,以及 Barnett 等人 2024 年的研究。我还发现了 “ARIA: A Diagnostic Framework for Music Training Data Attribution” 和 “Attribution-by-design” 的相关研究。OMC wiki 没有返回任何相关结果,paper-wiki 中也没有归因与补偿相关的条目。我现在已经有足够的信息来撰写评审了。
AV-SyncBench: Decoupled Benchmarking of Temporal and Semantic Audio-Visual Synchronization
Published:
Paper Review: AV-SyncBench: Decoupled Benchmarking of Temporal and Semantic Audio-Visual Synchronization
Evaluating Pretrained Music Embeddings for Cross-Performance Jazz Standard Recognition
Published:
Paper Review: Evaluating Pretrained Music Embeddings for Cross-Performance Jazz Standard Recognition
Positive-Incentive Noise Predictor for Adversarial Purification in Speaker Verification
Published:
现在我已经掌握了所有研究资料。让我来整理最终的结构化审稿报告。以下是我掌握的关键事实:
A Geometric Perspective on Composable Emotion Steering in Text-to-Speech Models
Published:
现在我已经具备了所有必要的事实依据。让我来撰写最终的结构化评审。
NPUsper: Eliminating Redundant Computation for Real-Time Whisper on Mobile NPUs
Published:
我已掌握足够的事实依据。关键发现如下:
Disentangling Speaker and Language Effects in Cross-Lingual Speaker Verification for Iberian Languages
Published:
我已经有了足够的锚点。现在开始撰写评论。
CNN Models for Microphone Array Covariance Matrix Upsampling and Acoustic Imaging
Published:
Paper Review: CNN Models for Microphone Array Covariance Matrix Upsampling and Acoustic Imaging
Few-Shot Open-Set Audio Classification Using Attention Information-Fused Prototypes
Published:
我已阅读了这篇 14 页论文的全文并收集了所有必要的证据。现在我将输出结构化的 Markdown 格式评审。
TurnNat: Automatic Evaluation of Turn-Taking Naturalness in Dyadic Spoken Dialogue
Published:
我已读完论文全文(PDF 8页)并建立了事实锚点,下面输出最终 review。
From Monolingual to Multilingual: Evaluating Mamba for ASR in South African Languages
Published:
Paper Review: From Monolingual to Multilingual: Evaluating Mamba for ASR in South African Languages
Quantifying the Uncertainty of Blindly Estimated Room Embeddings Using a Dispersion-Calibrated Score
Published:
我已经掌握了足够的证据。关键事实如下:
H-SAGE: Holistic Speaker-Aware Guided Experts for MoE-based Multi-Talker ASR
Published:
Paper Review: H-SAGE: Holistic Speaker-Aware Guided Experts for MoE-based Multi-Talker ASR
Enhancing Acoustic-to-Articulatory Inversion with Multi-Target Pretraining for Low-Resource Settings
Published:
Paper Review: Enhancing Acoustic-to-Articulatory Inversion with Multi-Target Pretraining for Low-Resource Settings
UT-AISTimprt submission for ICME 2026 Grand Challenge on Academic Text-to-Music Generation
Published:
我已阅读了完整的 HTML 论文,并执行了 wiki/paper-wiki/free-search 查询。现在开始撰写评论。
Pmeta-TLA: Backdoor Attacks for Speech Classification Models via Meta-Learning with Timbre Leakage Attack
Published:
我已经确立了关键事实锚点。从论文 PDF 阅读和搜索中获得的关键发现:
DRL-CLBA: A Clean Label Backdoor Attack for Speech Classification via DDPG Reinforcement Learning
Published:
已获取全文及所有外部查询结果。现基于七条公理输出完整评审。
Rethinking Speech-LLM Integration for ASR: Effective Joint Speech-Text Training by Interleaving
Published:
Paper Review: Rethinking Speech-LLM Integration for ASR: Effective Joint Speech-Text Training by Interleaving
Self-Supervised Test-Time Tuning for Packet Loss Concealment
Published:
Paper Review: Self-Supervised Test-Time Tuning for Packet Loss Concealment
RT-Tango: Real-Time Distributed Binaural Speech Enhancement for Low-Power Hearing Aid Devices
Published:
Paper Review: RT-Tango: Real-Time Distributed Binaural Speech Enhancement for Low-Power Hearing Aid Devices
Decomposer: Learning to Decompile Symbolic Music to Programs
Published:
论文评审:Decomposer: Learning to Decompile Symbolic Music to Programs
Neural Audio Codec with Adjustable Token Temporal Resolution Using Sampling-Frequency-Independent Convolutional Layers
Published:
Paper Review: Neural Audio Codec with Adjustable Token Temporal Resolution Using Sampling-Frequency-Independent Convolutional Layers
Towards a Phonology-Informed Evaluation of Multilingual TTS
Published:
太好了。我现在已经掌握了全面的事实锚点。关键发现:
A Multi-Branch Hierarchy-Aware Framework for Heterogeneous Audio Classification
Published:
已建立充分的事实锚点。层级分类头(GC、LCL)是标准技术(HiClass 库,Ding 等人 2023 — 论文中已引用)。KNN 蒸馏(Kong 等人 2020,Yang 等人 2022 — 均已引用)。CLAP 是基础。BST/BSD10k 是官方数据集。DCASE 2026 挑战赛排行榜确实存在。该论文是一篇挑战赛技术报告(5 页,DCASE 格式)。现在我已经有足够的信息来撰写审稿意见了。
Using embeddings to predict spoken word duration and pitch in Mandarin monosyllabic words
Published:
我已经收集了足够的证据。研究阶段已完成。正在输出最终的结构化评审报告。
LMPAN: A Lightweight Multi-Path Alignment Network for Joint Full-Duplex Acoustic Echo Cancellation and Noise Suppression
Published:
我已经掌握了所需的事实锚点。研究阶段已完成(已使用 15 轮中的 13 轮)。现在直接输出结构化的 Markdown 评审。
An Efficient vLLM-Based Inference Pipeline for Unified Audio Understanding and Generation
Published:
现在我已经掌握了确凿的事实依据。我已获取所需的所有信息。现在基于全文和外部证据撰写评审。
Speaker head orientation estimation with a single microphone array using phase spectrogram features
Published:
Paper Review: Speaker head orientation estimation with a single microphone array using phase spectrogram features
Unlocking Speech-Text Compositional Powers: Instruction-Following Speech Language Models without Instruction Tuning
Published:
Paper Review: Unlocking Speech–Text Compositional Powers: Instruction-Following Speech Language Models without Instruction Tuning
Cross Domain Few-Shot Class-Incremental Audio Classification Via Adversarial Contrastive Learning
Published:
论文评审:基于对抗对比学习的跨域少样本类增量音频分类
Spatial Speech Perception Systems: A Survey of Sound Source Localization, Directional Enhancement, and Speech Recognition
Published:
我已经阅读了完整的 9 页 HTML 全文(171K 字符,包含 8 个章节,106 条参考文献,7 个表格),并完成了所有必要的研究查询。以下是结构化评审。
SelectTSL: Prompt-Guided Selective Target Sound Localization in Complex Scenarios
Published:
我已经阅读了全文(15页)并收集了相关的前沿背景信息。以下是结构化的审查报告。
Audio-Based Understanding of Audiobook Narration Appeal
Published:
已完成研究。已阅读全文 HTML 及摘要,并已运行 paper-wiki(空结果)和 free-search 以进行先前工作锚定。现在输出最终审查报告。
Reasoning LLM Improves Speaker Recognition in Long-form TV Dramas
Published:
Paper Review: Reasoning LLM Improves Speaker Recognition in Long-form TV Dramas
QuaSR: Quality-Aware Sample Reweighting for Pacific Indigenous Speech Recognition
Published:
我已经获取了全文并完成了研究。现在让我来撰写最终的评审。关键发现如下:
TRACE-EVC: Text-Guided Relative Affective Control for Zero-Shot Emotional Voice Conversion
Published:
我已经获取了全文并建立了所有事实锚点。现在我将输出最终的评审报告。
Probing Low-Level Acoustic Attribute Encoding in CLAP Audio Embeddings
Published:
论文中未提及代码发布。我已经阅读了完整的论文全文(HTML 提取 + PDF 可视化),并完成了所有维基/论文库/自由搜索查询。现在我将撰写完整的审稿报告。
Q-TriM: Question-Guided Tri-Modal Attention for Audio-Visual Question Answering
Published:
我已经掌握了所需的所有事实。现在输出最终的结构化审查意见。
EEG-Based Imagined Speech Decoding Using a Hybrid CNN-SNN Architecture
Published:
Paper Review: EEG-Based Imagined Speech Decoding Using a Hybrid CNN-SNN Architecture
TokAN: Accent Normalization Using Self-Supervised Speech Tokens
Published:
Paper Review: TokAN: Accent Normalization Using Self-Supervised Speech Tokens
NouveauVoice: Generating Novel Pseudo Speakers for Voice Anonymization
Published:
Paper Review: NouveauVoice: Generating Novel Pseudo Speakers for Voice Anonymization
Speaker-Disentangled Chunk-Wise Regression for Syllabic Tokenization
Published:
我已有足够材料撰写评论。已确认的关键事实锚点:
DELTA-TTS: Adapting Autoregressive Model into Diffusion Language Model for Text-to-Speech
Published:
关键事实:OmniVoice (2026年4月1日,早于 DELTA-TTS 2026年7月5日约3个月) 使用了“diffusion language model-style discrete NAR”架构,并采用“从预训练 LLM 初始化”——这很接近,但并非明确意义上的 AR-to-dLLM 转换。这是并发相关工作,并非严格的先验工作,但在时间线上非常接近(刚好在2个月的宽限期之外,因此可以视为相关工作,而非并发工作)。DELTA-TTS 的核心创新点——“AR-to-dLLM 转换”、基于置信度排序的解码以及用于语音的 Conformer 卷积——依然具有其独特性。
Information-Geometric Superposed Vowel Evaluation: Part 1. Moraic Syllabary (Japanese)
Published:
研究阶段已完成。PDF 已完整读取(1199 行,8 页,4 张图表)。paper-wiki CLI 报错(traceback,记录为失败)。free-search 找到大量 SOTA 语音 deepfake 检测工作(SONAR、SLIM NeurIPS 2024、BiCrossMamba-ST、WaveSP-Net、CtrSVDD Challenge 2024 等)。下面直接输出最终 review。
Noisy Environment Adaptation of Neural Speech Codec via Focal Mask and Noise Feature Separation
Published:
我已经获取了所有需要的研究数据。关键发现如下:
MOSAIC: Interpretable Multi-Token Cross-Attention of Biophonetic and Self-Supervised Representations for Unified Voice Anti-Spoofing
Published:
好的 —— 代码已公开,采用 MIT 许可证,包含 Python 源码。现在我已经掌握了进行完整评审所需的所有事实。让我来整理最终的结构化评审报告。
Doppelganger: Sound Effects and Their Synthetic Twins
Published:
Paper Review: Doppelganger: Sound Effects and Their Synthetic Twins
Auto-AEG: Scalable Data Construction for Open-Vocabulary Audio Event Grounding
Published:
Paper Review: Auto-AEG: Scalable Data Construction for Open-Vocabulary Audio Event Grounding
Sampling Bias Compensation for Robust Evaluation of Audio Classification Systems with Partially Labeled Evaluation Datasets
Published:
论文审查:针对具有部分标注评估数据集的音频分类系统鲁棒评估的采样偏差补偿
Weakly Guided and Autoregressive Beamformer Parameterization for Generalizable Moving Speaker Extraction in Higher-Order Ambisonics
Published:
已建立所有事实锚点。现在输出最终的结构化 Markdown 评审。
EVAS: Efficient Multimodal Temporal Forgery Localization via Audio-Visual Synergy and Steered Boundary Calibration
Published:
未找到代码或 GitHub 仓库。现在我有足够的信息来撰写审稿了。让我整理一下我从论文全文中收集到的所有关键发现:
UniSkip-Mamba: A Frequency-Aware State Space Model for Audio-Visual Temporal Forgery Localization
Published:
事实锚点已足够。关键发现:
- UniCaCLF (arXiv 2506.08493, 2025年6月) — 最近的 SOTA 基线,已确认
- AV-Deepfake1M++ (arXiv 2507.20579, 2025年7月) — 一个更新的基准数据集,可能未被本文采用
- DiMoDif (2411.10193), UMMAFormer (2308.14395) — 均为真实且被引用的基线
- paper-wiki 中没有关于 AV-TFL 的记录;不存在与 skip-scan 等同的先前研究
- UniSkip-Mamba 是通过 “skip scanning mamba deepfake frequency-aware” 搜索得到的唯一匹配结果(确认了其新颖性声明)
Towards Digital Preservation of Efik: TTS for a Low-Resource African Language
Published:
Paper Review: Towards Digital Preservation of Efik: TTS for a Low-Resource African Language
Training-Free Model Selection and Domain-Aware Score Calibration for First-Shot Anomalous Sound Detection
Published:
我已经有了足够的事实锚点。论文全文已阅读,且之前的工作已经确认(GenRep 排名第 2,LDN 由 Wilkinghoff 等人提出,Mondrian/group-conditional conformal 是已知的,MetaOD 是已有的)。现在开始撰写评审。
CARD: Cross-component Audio Representation Distillation for Encoder-Free Audio Captioning
Published:
根据搜索结果,我已掌握了关键事实:AudioCaps SOTA CIDEr 约为 91.1(来自 LAMB),SLAM-AAC 是近期强劲的基准模型。CARD 仅达到 55.4,而基于编码器的基准模型为 66.4 —— 这虽然存在显著差距,但消融实验清晰地支持了核心论点。AuRA (arXiv 2606.11033, 2026年6月9日) 是被明确指出为同期工作的相关工作。我已掌握了足够的信息来撰写评审。
Evaluating the Effect of Linguistic Relatedness on Cross-Lingual Transfer in Large Multilingual Automatic Speech Recognition
Published:
我已经有了足够多的锚点。论文全文已阅读。正在撰写最终评审。
Ranking the Impact of Contextual Specialization in Neural Speech Enhancement
Published:
Paper Review: Ranking the Impact of Contextual Specialization in Neural Speech Enhancement
Semantic Homogenization in Italian Popular Music: A Diachronic Analysis
Published:
我已经获取了论文全文并进行了基础研究。OMC Wiki 中没有关于这些主题的记录。Paper-wiki 中也没有匹配项。Free-search 找到了关键的相关工作:Parada-Cabaleiro et al. (2024) 在《Scientific Reports》上发表的论文是主要的英语语言先行研究,而 Kim & Akama (2024) 的论文则涉及歌词相似性感知。现在我将撰写完整的评审。
SynSFX: Multi-Model Sound Effects Synthesis Dataset for Deepfake Detection and Evaluation
Published:
Paper Review: SynSFX: Multi-Model Sound Effects Synthesis Dataset for Deepfake Detection and Evaluation
Towards Robust Uncertainty-Aware Speaker Modeling
Published:
我已有足够的事实依据。关键锚点:
- U3-xi (arXiv 2601.15719, Li & Lee, 2026 年 1 月) 是其直接前置工作/强基线 —— UAAM-Softmax 仅使用说话人间的可分性。
- Xi+ (arXiv 2509.05993) 是更早的不确定性监督研究(ICASSP 2026)。
- Wang 等人 (ICASSP 2023, 2302.11763) 和基于不确定性的余弦评分 (2403.06404) 是关键先驱。
- ECAPA-TDNN、WeSpeaker、CAM++、SphereFace2、AM-Softmax 均为既定标准基线 —— 均在论文中作为基线使用。
- 针对 SRE 的基于不确定性的领域适配此前尚未有人探索;特征级 DA (CORAL/MMD/DANN) 是标准做法。
- 该论文的 ECAPA512 基线 EER 在 Vox1-O 上约为 1.07%,在 CNCeleb 上为 15.3% —— 这些数字是合理的。
DuplexChat: Constructing Speaker-Separated Full-Duplex Dialogue Speech at Scale for Spoken Dialogue Language Modeling
Published:
Paper Review: DuplexChat: Constructing Speaker-Separated Full-Duplex Dialogue Speech at Scale for Spoken Dialogue Language Modeling
Quantum-Inspired Harmonic Decision Models: A Computational Framework for Music Generation
Published:
Paper Review: Quantum-Inspired Harmonic Decision Models: A Computational Framework for Music Generation
Listen, Think, Transcribe: Continuous Latent Test-Time Scaling for ASR
Published:
Paper Review: Listen, Think, Transcribe: Continuous Latent Test-Time Scaling for ASR
Context-Aware ASR for Mandarin Technical Lectures
Published:
已有足够锚点。现在输出最终评审。
Towards Language-Agnostic Speech Inversion
Published:
Paper Review: Towards Language-Agnostic Speech Inversion
RABBiT: Rapidly adaptive BOLD foundation model via brain-tuning for accurate zero-shot and few-shot prediction of speech-elicited responses in the brain
Published:
Paper Review: RABBiT: Rapidly adaptive BOLD foundation model via brain-tuning for accurate zero-shot and few-shot prediction of speech-elicited responses in the brain
Unified Audio Intelligence Without Regressing on Text Intelligence
Published:
好的,我已经整理好事实锚点。关键发现如下:
- UALM (Tian et al., 2026) 是由同一团队(NVIDIA)发表的先前工作,也是 Audex 明确承认并在此基础上构建的基础。
- Qwen3-Omni 声称实现了“无退化”,但 Audex 的论文对其提出了反驳,指出在推理基准测试中存在明显的退化。
- Nemotron-Cascade-2 (Yang et al., 2026) 是基础 LLM 主干,这是一篇真实的 NVIDIA 论文 (arXiv:2603.19220)。
- 模型检查点已在 HuggingFace 上发布。
Progressive Refinement: An Iterative Pseudo-Labeling Approach for Mandarin-English Code-Switching ASR
Published:
我已经获取了完整的论文文本并进行了充分的研究。现在我来整理最终的评审报告。
Streaming Neural Speech Codecs through Time-Invariant Representations
Published:
Paper Review: Streaming Neural Speech Codecs through Time-Invariant Representations
ProPS: Prompted Profile Synthesis for Natural Language-Conditioned Speaker Embedding Distributions
Published:
我已经阅读了完整的论文,并收集了关于现有工作的证据。现在开始输出结构化的综述。
REDDIT: Correcting Model-Generated Timestamp Drift in ASR without Forgetting via Replay-Based Distribution Editing
Published:
我已掌握了充分的信息。论文全文已阅读(8页),相关的过往工作已通过 free-search 识别(包括 Whisper 幻觉、In-Sync 并行工作、WhisperX、CrisperWhisper、LwF、持续学习 ASR)。paper-wiki 索引了 ASR/Whisper 相关的论文,但没有直接匹配“时间戳漂移”(timestamp drift)的内容。现在我将直接输出最终的结构化评审报告。
SPEARBench: A Benchmark for Naturalness Evaluation in Streaming Speech-to-Speech Language Models
Published:
Paper Review: SPEARBench: A Benchmark for Naturalness Evaluation in Streaming Speech-to-Speech Language Models
LEGATO 2: Toward Multimodal Sheet Music Recognition and Understanding
Published:
我已经确立了强有力的事实锚点。研究阶段的关键发现如下:
From Textural Counterpoint to Feature Encoding: A Multi-Dimensional Machine Representation Study of Haydn’s “The Lark” Integrating Electroacoustic Analysis
Published:
论文评审:从文本对位到特征编码:整合电声分析的海顿《云雀》多维机器表示研究
Few-Shot Class-Incremental Audio Classification Using Pseudo-Incrementally Trained Embedding Learner and Continually Updated Stochastic Classifier
Published:
现在我对整体情况有了全面的了解。让我阅读论文 PDF 的实际内容以核实具体主张,然后撰写审稿意见。我已经阅读了全部 15 页。让我整理一下从 PDF 和研究中获得的关键发现:
Multimodal Video-to-Music Recommendation via Semantic Retrieval and Temporal Reranking
Published:
以下是最终的结构化审查报告。
Escaping the Procrustean Bed: Groupwise Orthogonal Connectors for Audio-Language Models
Published:
以下是结构化评审:
Music I Care About: Automated Multimodal Benchmarking of LLM Music Perception Skills on (Almost) Any Music
Published:
Paper Review: Music I Care About: Automated Multimodal Benchmarking of LLM Music Perception Skills on (Almost) Any Music
Fréchet Distance Loss on Speech Representations for Text-to-Speech Synthesis
Published:
Paper Review: Fréchet Distance Loss on Speech Representations for Text-to-Speech Synthesis
BlueMagpie-TTS: A Token-Efficient Tokenizer, Language Model, and TTS for Taiwanese-Accent Code-Switching Speech
Published:
Paper Review: BlueMagpie-TTS: A Token-Efficient Tokenizer, Language Model, and TTS for Taiwanese-Accent Code-Switching Speech
Determinantal point process sampling for bioacoustic active learning
Published:
Paper Review: Determinantal Point Process Sampling for Bioacoustic Active Learning
Flow Matching-Based Speech Source Separation with Best-of-N Biometric Sampling
Published:
我已经掌握了所需的全部事实锚点。让我来汇总评审意见。我的研究主要发现如下:
TriA Pipeline: A Large-Scale Automatic Audio Annotation Pipeline For Audio Classification In Specific Scenarios
Published:
我已阅读论文全文并收集了研究背景信息。以下是最终的结构化评审。
Goodbye Equal Error Rate, Hello Local Information Disclosure: Evaluating Voice Anonymisation against 1-to-N Linkage Threats
Published:
目前还没有针对该论文的代码或 GitHub 仓库。我现在已有足够的事实依据。请让我再确认一件事——关于作者是否是 SRD 的共同开发者(其中一位作者是 Bäckström,他是 SRD 论文的合著者)。
Learning-based Physics-Constrained Neural Kernel for Sound Field Estimation With Source-Position-Dependent Directional Weighting
Published:
Paper Review: Learning-based Physics-Constrained Neural Kernel for Sound Field Estimation With Source-Position-Dependent Directional Weighting
From Sinhala to Dhivehi: Cross-Lingual Transfer Learning for Low-Resource Speech Recognition
Published:
我已经获取了完整的论文文本和所有搜索结果。现在我将整理结构化的评审意见。
Designing Maintainable Hybrid Generative Systems: A Quantum-Inspired Approach to Automated Music Harmony Generation
Published:
我已经阅读了完整论文(通过 pdftotext 提取的 8 页 PDF,共 741 行文本),并完成了所有必要的研究查询。以下是结构化的评审意见。
ForestIR: Physics-Informed Forest Sound Simulation for Array-Based Bioacoustic Remote Sensing
Published:
Paper Review: ForestIR: Physics-Informed Forest Sound Simulation for Array-Based Bioacoustic Remote Sensing
InsideSSL: Understanding Self-Supervised Speech Representations using a Model-Centric Perspective
Published:
我现在已经掌握了完整的信息。让我整理一下最终的审稿意见。我已经:
- 阅读了完整的论文文本(1431 行,从摘要到参考文献)
- 通过
wiki_query查询了 OMC Wiki(3 次查询,全部返回空) - 查询了
paper-wiki(3 次查询,全部返回空) - 使用
free-search查找了前期工作(4 次查询,获得了有效结果)
Precise Video-to-Audio Generation with Cross-Modal Alignment in Latent Space
Published:
Paper Review: Precise Video-to-Audio Generation with Cross-Modal Alignment in Latent Space
Andha-Dhun: A First Look at Audio Descriptions in Hindi
Published:
Now I have all the research I need. Let me compose the final review.
WordVoice: Explicit and Decoupled Multi-Dimensional Word-Level Control for LLM-Based TTS
Published:
Paper Review: WordVoice: Explicit and Decoupled Multi-Dimensional Word-Level Control for LLM-Based TTS
Hierarchical Acoustic-Semantic Modeling: Modality Separation and Semantic Coherence for Full-Duplex SLMs
Published:
我现在已经从论文全文、paper-wiki 以及 free-search 结果中获取了足够的信息,可以产出一份详尽的审稿意见。请允许我整理这份结构化的 Markdown 审稿报告。
Audio Sentiment Analysis via Distillation and Cross-Modal Integration of Generated Multilingual Transcripts
Published:
我已获取足够信息来完成评审。现在我来整理研究发现并撰写结构化的评审意见。
Compress the Cache, Not the Speech Embedding: KV Compression for Efficient Speech LLMs
Published:
我已经掌握了所需的所有信息。现在我来撰写最终的结构化评审。
Gradient-Based Speech-to-Text Alignment for Any ASR Model: From CTC to Speech LLMs
Published:
我现在已经有了全面的研究基础。让我整理一下关键发现并生成最终评审。
UBG-Net: An Uncertainty-aware Bayesian Gating Network for Robust Audio-Visual Speech Recognition
Published:
Paper Review: UBG-Net: An Uncertainty-aware Bayesian Gating Network for Robust Audio-Visual Speech Recognition
MADB: A Large-Scale Music Aesthetics Dataset with Professional and Multi-Dimensional Annotations
Published:
Paper Review: MADB: A Large-Scale Music Aesthetics Dataset with Professional and Multi-Dimensional Annotations
MMGenre: Benchmarking Singing Voice Synthesis across Multiple Musical Genres
Published:
我已经掌握了所需的所有信息。现在我将撰写完整的结构化评审。
EscFOA: Enhancing Spatial Learning for Visually Impaired Learners via Generative Spatial Audio in 360-Degree Educational Environments
Published:
Paper Review: EscFOA: Enhancing Spatial Learning for Visually Impaired Learners via Generative Spatial Audio in 360-Degree Educational Environments
Decoupling Conversational Dynamics in Full-Duplex Spoken Models through Reinforcement Learning
Published:
Paper Review: Decoupling Conversational Dynamics in Full-Duplex Spoken Models through Reinforcement Learning
Rag Classification of Tagore Songs using Symbolic Music Notation and Novel Weighted Distance Measures
Published:
我已经完成了全文阅读并建立了所有事实锚点。现在输出完整的评审意见。
Multimodal Voice Activity Projection for Turn-Taking in Social Robots with Voice-Activity-Related Pretrained Encoders
Published:
没有针对语义一致性损失的消融实验,也没有针对说话人间注意力(inter-speaker attention)的消融实验。现在我已经掌握了所有证据来撰写这篇评审。
Text-Independent Speaker Verification Using Discrete Audio Tokens
Published:
我已经掌握了所有关键事实。研究的关键发现如下:
Multimodal Unlearning Across Vision, Language, Video, and Audio: Survey of Methods, Datasets, and Benchmarks
Published:
我已经掌握了所有模态的充分证据。以下是撰写评审所需的关键事实:
A Reliability Assessment of LALM Audio Judges for Full-Duplex Voice Agents
Published:
Paper Review: A Reliability Assessment of LALM Audio Judges for Full-Duplex Voice Agents
PS4: Proxy-Supervised Joint Training for Real Target Speaker Extraction
Published:
我已经有足够的研究数据来撰写评论了。让我整理一下调查结果:
VSRo-200: A Romanian Visual Speech Recognition Dataset for Studying Supervision and Multimodal Robustness
Published:
我已经掌握了完整的信息。让我来整理这篇综述。我已确认的关键事实如下:
COALA: Robust Contextualized Speech-augmented Language Modeling for ASR via Contrastive Regularizer and Biasing Score Estimation
Published:
论文评审:COALA:基于对比正则化器和偏置分数估计的鲁棒上下文语音增强语言建模
MuScriptor: An Open Model for Multi-Instrument Music Transcription
Published:
Paper Review: MuScriptor: An Open Model for Multi-Instrument Music Transcription
Diarization-Guided Qwen-ASR Adaptation for Multilingual Two-Speaker Conversational Speech
Published:
我已经掌握了所需的所有证据。让我来撰写最终的结构化审查报告。
Best-of-$N$ TTS Evaluation is Confounded by ASR Family Alignment
Published:
Paper Review: Best-of-$N$ TTS Evaluation is Confounded by ASR Family Alignment
On the Role of Conversational Timing in Synthetic Training Data for ASR
Published:
我已经有足够的信息来撰写评审了。让我整理一下我收集到的关键事实:
When Synthetic Speech Is All You Have: Better Call GRPO
Published:
Paper Review: When Synthetic Speech Is All You Have: Better Call GRPO
A Quantized Native Runtime for On-Device Semantic Audio Generation
Published:
我已经阅读了完整的论文,并完成了所有 Wiki/paper-wiki 的查询。现在我将输出最终的结构化审阅报告。
Structural Bottlenecks on Frequency Representation in End-to-End Audio Models
Published:
Paper Review: Structural Bottlenecks on Frequency Representation in End-to-End Audio Models
Why Do You Say It Like That? A Phoneme-Level Framework for Explainable Speech Deepfake Detection
Published:
Paper Review: Why Do You Say It Like That? A Phoneme-Level Framework for Explainable Speech Deepfake Detection
It Takes Few to TANGO: A Quantized Distributed Model for Binaural Speech Enhancement
Published:
我已经掌握了所需的所有研究锚点。现在让我来整理这篇综述。我已获取完整论文文本并确认了以下事实锚点:
MulTTiPop: A Multitrack Transcription Dataset for Pop Music
Published:
我已阅读全文并收集了相关研究资料。现在开始输出最终评审。
Optimal Transport-based Semantic Alignment for LLM-based Audio-Visual Speech Recognition
Published:
Paper Review: Optimal Transport-based Semantic Alignment for LLM-based Audio-Visual Speech Recognition
Phone Segmentation and Recognition through Phonological Activation Mapping
Published:
Paper Review: Phone Segmentation and Recognition through Phonological Activation Mapping
Technical Report for MERL’s Real-TSE Challenge Submission
Published:
我现在已经拥有了所需的全部事实锚点。研究阶段已完成。我来输出最终的结构化审稿报告。
Beyond Time Shifts: Adapting Omni-LLM as a Reference-Free Evaluator for Generative Audio-Visual Models
Published:
我已经掌握了完整论文(共28页)以及事实锚点。以下是综合评审意见。
Event-Based Token Sequences for Audio-Conditioned Music-Game Level Modeling
Published:
Paper Review: Event-Based Token Sequences for Audio-Conditioned Music-Game Level Modeling
ReGen: Hierarchical Multi-Prompt Representation Generation for Efficient Waveform Diffusion Models
Published:
太好了!我现在已经掌握了所需的所有关键研究背景。以下是我的关键发现:
Wan-Dancer: A Hierarchical Framework for Minute-scale Coherent Music-to-Dance Generation
Published:
我现在已经有了全面的事实锚点。让我整理最终审查报告。关键发现总结:
Tokenizer Transplantation: Mitigating Autoregressive Collapse in Edge-Efficient Bengali ASR
Published:
我已获取全文和足够的事实锚点。现在输出完整 review。
WaveNet-Style Guitar Amplifier Model Pruning for Real-Time iOS Deployment
Published:
Paper Review: WaveNet-Style Guitar Amplifier Model Pruning for Real-Time iOS Deployment
Efficiently Adapting Spoken Language Models for the Singaporean Context
Published:
Paper Review: Efficiently Adapting Spoken Language Models for the Singaporean Context
CoFi-Lite: Pushing the Limits of Ultra-Lightweight Speech Enhancement
Published:
我现在已经掌握了所有必要的信息。让我整理一下研究中的关键事实:
Evaluating SSL and ViViT Architectures for Cross-Corpus Audio MOS Prediction via LODO Validation
Published:
我已从论文全文(通过 arXiv HTML 获取的摘要、引言、相关工作、方法、结果、讨论、结论)以及研究中获取了所有必要信息。现在我来撰写最终评审。
Hearing Like Humans? Sound Symbolism and Perceptual Alignment in Speech Language Models
Published:
Paper Review: Hearing Like Humans? Sound Symbolism and Perceptual Alignment in Speech Language Models
Breaking the Quality–Intelligibility Trade-off in Streaming Target Speaker Extraction via Deep-Feature-Anchored Preference Optimization
Published:
我现在已经掌握了充足的事实依据。关键发现:StarTSE (arXiv 2604.19635) 被确认为基础模型,而 MeanFlow-TSE (Interspeech 2026)、FlowTSE (Interspeech 2025)、TargetVoice (Interspeech 2025) 以及 Discriminative-Generative TSE (arXiv 2601.06006) 是相关的最新 TSE 工作。关于语音的 DPO:已存在针对 TTS 的 MPO、DiffRO、Preference Alignment 等方法。WavLM 用于说话人验证是标准做法。我已经获取了全文及所有附录表格。现在我将输出最终评审。
Graph Representation of RaagBase: A Unique Dataset for Hindustani Music
Published:
Paper Review: Graph Representation of RaagBase: A Unique Dataset for Hindustani Music
MeloBottleneck: Self-Supervised Melody Skeleton Extraction with a Latent Subsequence Bottleneck
Published:
我已掌握所有所需信息。现在输出完整的评审。
Which Languages Transfer Best to Warlpiri? A Similarity-Based Study for Low-Resource ASR
Published:
Paper Review: Which Languages Transfer Best to Warlpiri? A Similarity-Based Study for Low-Resource ASR
Partial Contracts Suffice: Sound, LLM-Inferred Regression Verification
Published:
所有 Wiki 查询均未返回结果。OMC Wiki 不包含有关软件验证主题的知识。paper-wiki 数据库侧重于音频/语音领域,因此没有匹配项。我已经阅读了完整的论文。现在我将输出审阅结果。
Empowering Long-form Omni-modal Understanding with Robust Audio Perception
Published:
我已经完成了研究阶段,现在输出最终的综述。
Learn2Chat: Rethinking Dyadic Talking Heads via Interaction-Modulated Monologic Priors
Published:
INFP 使用的是 DyConv 数据集,而不是 DualTalk 数据集。这很重要——INFP 是一个相关并行工作,但使用了不同的数据集。Learn2Chat 在 DualTalk 上对其进行比较可能不公平,但 INFP 在二元对话头部生成领域是一个知名的 CVPR 2025 方法,应该被引用并在相关工作中讨论。
PC-Mix: Partial-Component Audio Spoofing Detection under Mixed Speech and Environmental Sound Conditions
Published:
Paper Review: PC-Mix: Partial-Component Audio Spoofing Detection under Mixed Speech and Environmental Sound Conditions
Perceived Annoyance in Multi-source Electric Vehicle AVAS Environments
Published:
我已获取全文和所有研究锚点。以下是完整的评审意见。
GigaAM Multilingual: Foundation Model for Underrepresented Languages
Published:
太好了。我现在已经掌握了充分的事实依据。我已有足够的信息来撰写评审。让我汇总一下关键发现:
GigaChat Audio: Time-aware Large Audio Language Model
Published:
我现在已经获取了全文和补充研究资料。我已掌握足够的信息来撰写评审。关键研究要点如下:
FdAudio: MeanFlow-Anchored Fréchet-Distance Post-Training for One-Step Text-to-Audio Generation
Published:
Paper Review: FdAudio: MeanFlow-Anchored Fréchet-Distance Post-Training for One-Step Text-to-Audio Generation
Dance to Music Generation leveraging Pre-training with Unpaired data and Contrastive Alignment
Published:
我现在已经掌握了足够的事实依据。让我来整理一下主要发现:
ECHOv2: Two-Level Band-Splitting Representation Learning for Anomalous Sound Detection
Published:
Paper Review: ECHOv2: Two-Level Band-Splitting Representation Learning for Anomalous Sound Detection
An Objective Intelligibility Metric Evaluation on Spanish Speech
Published:
很好——SpInt 已在 Zenodo 上公开确认。未发现先前的西班牙语语音清晰度数据集。现在我已经掌握了所有必要的事实锚点。让我来撰写审稿意见。
Data Augmentation for L2 English Speaking Assessment using TTS
Published:
Paper Review: Data Augmentation for L2 English Speaking Assessment using TTS
Tight-Frame Reconstruction for Acoustic Intensity Estimation Using Cardioid Microphone Pairs
Published:
我已获取完整的论文文本,并完成了所有必要的研究查询。以下是结构化的评审意见。
The SonicAGI System for the REAL-TSE Challenge
Published:
论文审稿:The SonicAGI System for the REAL-TSE Challenge
Difference-Driven Gating: Adaptive Feature Fusion for U-Net Decoder
Published:
我已获取完整的论文文本,并从所有三个维度的搜索中获得了充足的背景信息。现在我将撰写最终的结构化评审。
CHARM: Charge Calibration and Acoustic Rescue for LLM-based Multimodal Sarcasm Detection
Published:
我已经完成了全文审阅和所有研究查询。现在输出最终的结构化评审。
MusicMark: A Robust Generative Watermarking Framework for Music Generation
Published:
Paper Review: MusicMark: A Robust Generative Watermarking Framework for Music Generation
Simple Features and Honest Calibration for Ambivalence and Hesitancy Recognition in Video
Published:
Paper Review: Simple Features and Honest Calibration for Ambivalence and Hesitancy Recognition in Video
BeatEdit: Symbolic Music Generation as Explicit Editing
Published:
Paper Review: BeatEdit: Symbolic Music Generation as Explicit Editing
Anysynth:Zero-Shot Instrument Cloning via In-Context Learning and Asymmetric Hierarchical Guidance
Published:
我已经阅读了完整的论文(HTML 全文)并收集了所有证据。以下是结构化评审意见。
Where Speech Enhancement Hurts Recognition: An Inference Time Polar Projection Diagnosis
Published:
我现在已经获取了所有必要的事实锚点。以下是汇总的研究发现:
Unified Gradient Projection: Language-Balanced Continual Learning for Multilingual Low-Resource ASR
Published:
论文评审:Unified Gradient Projection: Language-Balanced Continual Learning for Multilingual Low-Resource ASR
Semantic Sampling via Learnable Observation Front Ends
Published:
我已从全文、维基查询(全部无结果)、paper-wiki 查询(全部无结果)以及 free-search(已获取关于基线、可学习前端和压缩感知的相关工作)中收集了所有信息。以下是完整的结构化评审。
Verifier-Guided Twelve-Tone Composition: A Generate-Verify-Repair Harness for Symbolic Music Generation
Published:
我已经掌握了所需的所有信息。现在我将输出完整的评审意见。
Evidence Subspace Projection: Measuring How Much Evidence Explains Deepfake Detection in Self-Supervised Speech Models
Published:
我已经掌握了所有研究数据。现在开始撰写审稿意见。
Beyond Benchmarks: Exposing the Hidden Crisis in Bangla Hate Speech Detection
Published:
Paper Review: Beyond Benchmarks: Exposing the Hidden Crisis in Bangla Hate Speech Detection
Teaching Speech Enhancement Models to Sing: Domain Adaptation from Speech Enhancement to Singing Voice Separation
Published:
Paper Review: Teaching Speech Enhancement Models to Sing: Domain Adaptation from Speech Enhancement to Singing Voice Separation
Qwen-Music Technical Report
Published:
我现在已经从全文阅读和外部搜索中获得了足够的事实依据。让我来整理最终的评审意见。
VoxENES 2026: Benchmarking Generalization of Speech Spoofing Detectors Against LLM-Era TTS and Voice Conversion
Published:
Paper Review: VoxENES 2026: Benchmarking Generalization of Speech Spoofing Detectors Against LLM-Era TTS and Voice Conversion
Qwen-Audio-VAE Technical Report
Published:
我已完整阅读了这篇 15 页的 PDF 文件,并执行了所有必要的 wiki/paper-wiki/free-search 查询。现在开始输出结构化的评审意见。
Synchronized Three-Dimensional Vocal-Tract Motion for Speech Synchronization via Joint-Embedding Predictive Architecture Alignment
Published:
太棒了——我找到了一个非常相关的前置工作:“ArtNet: A JEPA-Like Articulatory Predictive Framework for Robust Zero-Shot Phoneme Recognition” (arXiv:2606.16595),这是一个将 JEPA 应用于发音语音学的同期工作。我现在已经从全文和外部搜索中获得了足够的事实锚点。让我来撰写最终的评审。
Casting Everything to Online API Services? A Survey of Integrating Localized Speech Recognition Models in Robotic Systems
Published:
Paper Review: Casting Everything to Online API Services? A Survey of Integrating Localized Speech Recognition Models in Robotic Systems
StoryTeller: Training-Free Narrative Grounding for Long-Form Audio Description
Published:
我已经获取了所需的所有信息。以下是完整审查:
Encoder-Side Neuron Identification and Amplification for Acoustic Perception in Large Audio-Language Models
Published:
Paper Review: Encoder-Side Neuron Identification and Amplification for Acoustic Perception in Large Audio-Language Models
The Sound of Absence: Audio-Language Embedding Models Struggle with Negation
Published:
已确认:Vasilakis et al. (arXiv 2601.13931, 2026年1月) 这篇论文未被引用——该研究探讨了音乐 CLAP 模型中的否定建模,提出了基于不相似性的对比损失(contrastive loss)和文本增强(text augmentation)方法,并设计了检索和分类评估协议。这是一项直接的前序工作,比当前论文早了 6 个月,文中提到“扩展到其他领域,如语音内容和音乐属性,是未来工作的重要方向”——然而 Vasilakis 已经完成了音乐领域的研究。这是一个重大的遗漏。
Real-time Generation of Listener Nodding via Prediction of Kinematic Parameters for Avatar Dialogue Systems
Published:
Paper Review: Real-time Generation of Listener Nodding via Prediction of Kinematic Parameters for Avatar Dialogue Systems
PolarBM: Complex-valued Boltzmann Machine for Modeling Audio Signals in Polar and Log-polar Coordinates
Published:
Paper Review: PolarBM: Complex-valued Boltzmann Machine for Modeling Audio Signals in Polar and Log-polar Coordinates
An Omnilingual-ASR-Based Speech-LLM System for the 2nd MLC-SLM Challenge
Published:
我已获取全文,并通过 wiki、paper-wiki 和 free-search 建立了事实锚点。以下是结构化的评审意见。
ZipL-Dialog: Memory-Efficient Long-Form Spoken Dialog Synthesis via Latent Flow Matching
Published:
我已经获得了完整的论文文本并建立了事实锚点。这是最终的结构化审查。
Open-Source Intelligence and Music Information Retrieval for Geographic Attribution of Musical Affect and the Ecological Limits of Population Inference
Published:
Paper Review: Open-Source Intelligence and Music Information Retrieval for Geographic Attribution of Musical Affect and the Ecological Limits of Population Inference
Listen first: Output-based multi-microphone speech enhancement
Published:
我现在已经从全文和搜索中获得了足够的证据。让我来撰写这篇评论。
Traceback Translators Against Forgetting in Continual Fake Speech Detection
Published:
Paper Review: Traceback Translators Against Forgetting in Continual Fake Speech Detection
UD-ASD: A Unified Diffusion Model for Anomalous Sound Detection
Published:
Paper Review: UD-ASD: A Unified Diffusion Model for Anomalous Sound Detection
Explainable-by-Design Audio Deepfake Detection via Wiener-Hopf Linear Prediction
Published:
Paper Review: Explainable-by-Design Audio Deepfake Detection via Wiener-Hopf Linear Prediction
What is a Musical Scale? Regularity and Convention in the Organization of Pitch
Published:
我已完成对该论文及相关文献的全面阅读。现在我将撰写结构化的评论。
Investigating the Integration of Spatial Information in Foundation-Model-Based Speaker Diarization
Published:
我已经掌握了所需的所有事实锚点。现在输出完整的审稿。
Contrasting statistical patterns in melodic and molecular evolution reveal distinctive constraints in a culturally evolving system
Published:
Paper Review: Contrasting statistical patterns in melodic and molecular evolution reveal distinctive constraints in a culturally evolving system
Audio Diarization: A New Paradigm for Exploring Audio Recordings with Unknown Event Classes
Published:
Paper Review: Audio Diarization: A New Paradigm for Exploring Audio Recordings with Unknown Event Classes
AutoSIFT: Automatic Style Sifting for Controllable Speech Generation with Arbitrary Style Infilling
Published:
我已经获取了所有必需的研究锚点。现在我将撰写最终的评审意见。
Neural Morphing: Sequence-Optimized Token-Level Morphing in Neural Audio Codecs
Published:
Paper Review: Neural Morphing: Sequence-Optimized Token-Level Morphing in Neural Audio Codecs
Spatial-Frequency Cued Generative Fixed-Filter Active Noise Control Based on Deep Learning in Reverberant Environments
Published:
我已经阅读了完整论文并完成了所有研究查询。以下是最终的评审。
AVSCap: Orchestrating Audio-Visual Synergy for Omni-modal Video Captioning
Published:
我已经获取了完整的论文文本(正文及所有附录),并从多个来源收集了研究背景。现在我将输出结构化的评审意见。
ChartGenEval: Corruption-Tested Multi-Dimensional Feedback for Rhythm-Game Chart Generation
Published:
Paper Review: ChartGenEval: Corruption-Tested Multi-Dimensional Feedback for Rhythm-Game Chart Generation
Low-Latency Neural Models for Real-Time Music Enhancement
Published:
我现在已经获取了全文并完成了事实锚定。以下是我在撰写评审前总结的关键发现:
Audio-Native Speech Recognition with a Frozen Discrete-Diffusion Language Model
Published:
Paper Review: Audio-Native Speech Recognition with a Frozen Discrete-Diffusion Language Model
Audio-Text Cross-Attention with Psycholinguistic Support Features for Ambivalence/Hesitancy Recognition
Published:
Paper Review: Audio-Text Cross-Attention with Psycholinguistic Support Features for Ambivalence/Hesitancy Recognition
Improving Text-to-Audio Instruction Following via Fine-Grained Feedback from Audio-Aware Large Language Models
Published:
已收集所有研究资料。现在输出最终的结构化审查报告。
Bring Music The Horizon: Music-Driven 360$^\circ$ Video Generation
Published:
Paper Review: Bring Music The Horizon: Music-Driven 360° Video Generation
Auditing Protocol-Level Shortcuts in Large Audio Language Model Judges for Speech Evaluation
Published:
论文评论:审计大型音频语言模型评委在语音评估中的协议级捷径
Greedy Volume Maximization of Gradient Embeddings for Long-Tailed Frame-Level Bioacoustic Active Learning
Published:
Paper Review: Greedy Volume Maximization of Gradient Embeddings for Long-Tailed Frame-Level Bioacoustic Active Learning
Cover First, Disagree Softly: Rethinking Mismatch-First Active Learning for Frame-Level Audio Classification
Published:
现在我已经掌握了完整信息。让我来汇总这篇综述。
From Prediction to Collaboration: Interactive Symbolic Music Analysis
Published:
我已经掌握了完整的事实依据。论文全文已阅读(通过 pdftotext 获取 8 页文本),所有 OMC wiki 查询均未返回结果,paper-wiki 中没有相关记录,已使用 free-search 补充了外部 SOTA 锚点。正在输出最终评审。
Learning Speaker Identity Beyond Language and Modality Constraints: Insights from the POLY-SIM 2026 Challenge
Published:
I have read the full paper text (563 lines via pdftotext) and gathered facts from free-search. OMC Wiki returned no relevant records. paper-wiki returned one tangentially related paper (AudioJudge, arXiv:2507.12705). Here is the final review.
Self-supervised Speech Comparison for L2 Phone, Rhythm, and Intonation Scoring
Published:
论文评审:基于自监督语音对比的 L2 音素、节奏及语调评分
From Continuous Deployment to Queryable Dataset: Terabyte-Scale AIS-Aligned Passive Acoustic Labelling
Published:
论文评审:从连续部署到可查询数据集:TB级与AIS对齐的被动声学标注
Rethinking Speech Foundation Model Fine-tuning: Better SFT or Better Match?
Published:
Paper Review: Rethinking Speech Foundation Model Fine-tuning: Better SFT or Better Match?
Genre Bias or Aesthetic Perception? Identifying and Mitigating Shortcut Learning in Music Evaluation
Published:
我已经有了足够的事实锚点。已确定的关键现有工作:
CF-Net: Conflict Fusion with Speaker Normalisation and Certainty Weighting for Ambivalence/Hesitancy Recognition
Published:
我现在已经掌握了足够的上下文。主要的锚点事实如下:
Music-to-Dance Generation via Atomic Movements
Published:
I now have sufficient information to produce the review. Key findings from the full-text reading and research:
MetaPerch: Learning from metadata for bioacoustics foundation models
Published:
评审已完成。以下是最终的结构化评审:
Can Tokens Compete? Token Representations against Supervised CNN Backbones for BirdCLEF+ 2026
Published:
Paper Review: Can Tokens Compete? Token Representations against Supervised CNN Backbones for BirdCLEF+ 2026
MIDI-RAE-JEPA: Hierarchical Representation Learning and Generation for Symbolic Music
Published:
我现在有足够的事实依据了。关键发现:
WanSong v1.0 Technical Report
Published:
我已经获取了所有必要的信息。现在开始输出结构化审稿意见。
Large Audio Language Models for Spoofing-Aware Speaker Verification
Published:
我已经获取了全文和足够的背景信息。现在我将输出最终的结构化评审。
RW-Voice-EQ Bench: A Real World Benchmark for Evaluating Voice AI Systems
Published:
Paper Review: RW-Voice-EQ Bench: A Real World Benchmark for Evaluating Voice AI Systems
SLT 2026 REAL-TSE Challenge: Real-world Target Speaker Extraction from Conversational Recordings
Published:
Paper Review: SLT 2026 REAL-TSE Challenge: Real-world Target Speaker Extraction from Conversational Recordings
What does the model actually see? Evaluation protocols and input availability in data-driven prediction of room acoustic parameters
Published:
Paper Review: What does the model actually see? Evaluation protocols and input availability in data-driven prediction of room acoustic parameters
SceneBind: Binding What and Where Across Vision, Audio and Language
Published:
我已收集好所有必要的研究数据。现在我将输出最终的评审。
StemFX: Learning Mixing Style Representations via Autoregressive FX Chain Prediction on Source-Separated Stems
Published:
Paper Review: StemFX: Learning Mixing Style Representations via Autoregressive FX Chain Prediction on Source-Separated Stems
A Geometry-Limited Identification Floor and Its Consequences for Voice-Clone Attribution in Professional Voice Actors
Published:
论文评审:专业声优语音克隆归因中的几何限制识别下限及其影响
SpeechGuard: Online Defense against Backdoor Attacks on Speech Recognition Models
Published:
关键发现已确认。该论文引用了 [7] = STRIP (Gao et al. 2019, ACSAC),但并未引用 STRIP-ViTA (arXiv 1911.10312, 同一作者,同年度 2019),该文献明确将 STRIP 应用于 Speech Commands Dataset 的音频任务,使用了 1D-CNN 和 2D-CNN,并报告了音频任务的 FAR 结果。这是一个重大的遗漏基线。
Natural Backdoor Attacks on Speech Recognition Models
Published:
我已阅读全文(10 页,LNCS 会议论文,最初发表于 2023 年,于 2026 年 7 月重新发布至 arXiv),并收集了 wiki/paper-wiki/free-search 的证据。OMC Wiki 没有返回任何结果;paper-wiki 没有关于此主题的具体匹配信息。Free-search 找到了相关的先前和同期工作。以下是评审结果。
AuEmoChat: Authentic Emotion Understanding and Rendering for Conversational Speech Synthesis
Published:
Paper Review: AuEmoChat: Authentic Emotion Understanding and Rendering for Conversational Speech Synthesis
Contextual Semantic Relevance Tracks fMRI BOLD Responses During Naturalistic Speech Comprehension
Published:
我已经阅读了完整的论文。我已经有足够的事实锚点来撰写评论。让我整理这些发现:
Controlling Implicit Shortcut Reliance in L2 Spoken English Auto-markers
Published:
我已经阅读了全文 (HTML+PDF),并完成了所有 wiki/paper-wiki/free-search 锚定。现在开始输出评审。
Audio-Visual Flamingo: Open Audio-Visual Intelligence for Long and Complex Videos
Published:
Paper Review: Audio-Visual Flamingo: Open Audio-Visual Intelligence for Long and Complex Videos
Is One Score Enough? Assessing Singing Quality of Songs with Temporal Score Curves
Published:
Paper Review: Is One Score Enough? Assessing Singing Quality of Songs with Temporal Score Curves
Just A Rather Very Intelligent Spoken Agent
Published:
Paper Review: Just A Rather Very Intelligent Spoken Agent
Falsification-Based Verification of LLM-Generated Optimization Models: Sound Test Batteries and Their Detection Limits
Published:
我现在已经掌握了所有必要的事实依据。ReLoop (Lian et al., 2026, arXiv:2602.15983) 是最接近的现有工作,发表于 2026 年 2 月,本文发表于 2026 年 7 月——时间跨度约为 5 个月,因此它不属于同期工作,构成了合理的创新性比较。我已经阅读了完整的 PDF(32 页)和完整的 HTML 全文(81KB 纯文本)。我已具备输出审稿所需的一切材料。
HARP: Harmonic-Aware Residual Partitioning for Neural Audio Codecs
Published:
论文评审:HARP:神经音频编解码器的谐波感知残差分区
Pseudo-label distillation for discriminative anomalous sound detection
Published:
研究阶段完成,输出最终 review。
NABEATs: Noise-Aware Audio Representation Learning
Published:
Paper Review: NABEATs: Noise-Aware Audio Representation Learning
RealDESED: A Real-World Domestic Sound Event Detection Benchmark
Published:
Paper Review: RealDESED: A Real-World Domestic Sound Event Detection Benchmark
Explainable Lightweight Compact Deep Models for Speech Emotion Recognition
Published:
论文审查:基于可解释轻量级紧凑深度模型的语音情感识别
Do Speech Tokens Leak Voiceprints? Speaker Inversion Attacks Against End-to-End Speech Language Models
Published:
Paper Review: Do Speech Tokens Leak Voiceprints? Speaker Inversion Attacks Against End-to-End Speech Language Models
Efficient Audio-Visual Event Recognition via Knowledge Distillation and Dynamic INT8 Quantization of a Hybrid Cross-Attention Network
Published:
Paper Review: Efficient Audio-Visual Event Recognition via Knowledge Distillation and Dynamic INT8 Quantization of a Hybrid Cross-Attention Network
SALMONN-2: Advancing General-Purpose Hearing Abilities with Self-Supervised Representations
Published:
我现在已经掌握了充分的信息。论文文本已全文检索完毕,wiki 查询结果均为空(但已记录),paper-wiki 中有 Kimi-Audio、Audio Flamingo、MOSS-Audio-Tokenizer 和语音 LLM 综述的记录。我已经有足够的信息来撰写一份严谨的评审。
Robust Assamese Speech Recognition through Controlled Fine-Tuning of Whisper Models
Published:
我已建立了强有力的事实锚点。关键发现如下:
Adaptive Momentum Enhanced Distributed Multichannel Active Noise Control for Faster Convergence under Communication Delays
Published:
研究阶段已完成。现在输出最终的结构化 Markdown 评审。
AI_LectureNote: A Retrospective Pilot Study of a Post-ASR Workflow for English-Script Rendering and Semantic Drift in Korean-English Medical Lectures
Published:
我已获取全文并进行了事实核查。论文的 Related Work 中引用了几个关键锚点(Ma et al. 2023 LLM post-ASR correction, Adedeji et al. 2024/2025 medical ASR+LLM, Sudo et al. 2024 contextual biasing, Chapman et al. 2001 negation, Zheng et al. 2026 dysarthric post-ASR)。Wiki 记录为空;free-search 确认了该领域已有先前研究。现在我将输出最终评审。
FlowSonic: Stable Zero-Shot Music Editing via High-Order Trajectory Integration
Published:
论文审查:FlowSonic:通过高阶轨迹积分实现稳定的零样本音乐编辑
X-Translator: A Real-Time Multilingual Speaker-Aware Speech-to-Speech Translation System
Published:
论文审稿:X-Translator:一种实时多语种说话人感知语音到语音翻译系统
SSTMark: Robust Training-Free Semantic-Level Speech Watermarking
Published:
在公开搜索中未找到 SSTMark 的代码/仓库。我已经掌握了足够的事实锚点。现在我来输出最终的结构化评审。
Time-Frequency Consistency Learning for Robust Speech Deepfake Detection
Published:
我已经掌握了所有必要的事实依据。让我来撰写最终的评审报告。
When to Use Extra Context: Evidence-Grounded Terminology Adaptation for Simultaneous Speech Translation
Published:
Paper Review: When to Use Extra Context: Evidence-Grounded Terminology Adaptation for Simultaneous Speech Translation
FillGauss: Fine-Grained Filling-Aware Impact Sound Generation for 3D Gaussian Splatting
Published:
我已经掌握了足够的事实依据。已确认的关键基准如下:
- SonicGauss (arXiv:2507.19835, 2025年7月26日) — 最新的三维感知撞击声 SOTA,发表于 ACM MM 2025。这是本文的基础工作。
- Physics-Driven Diffusion Models (Su et al., CVPR 2023) — 撞击声扩散领域的重要先驱工作。
- RealImpact (Clarke et al., CVPR 2023) — 真实世界撞击声数据集。
- TangoFlux — 文本到音频生成模型。
ESCUCHA: A Spanish Speech Benchmark for Heterogeneous Acoustic Conditions
Published:
论文评审:ESCUCHA:面向异构声学条件的西班牙语语音基准测试
The tttAI System for the TSA-ASR Task of the SmartGlasses Challenge 2026
Published:
我已经掌握了足够的背景信息。现在我来整理最终的审稿报告。关键事实如下:
Harness TTS: Towards Context-Aware Expressive Speech Synthesis with Harness Layer
Published:
我已经掌握了足够的资料。论文全文已阅读,维基基准也已建立(OMC wiki 中没有匹配记录;paper-wiki 中已收录 CosyVoice3、VoxCPM2、CapTalk)。现在开始撰写审稿报告。
SpEmoC: A Balanced Speaker-Segment Multimodal Emotion Benchmark
Published:
我已有足够的依据。EMOE (CVPR 2025) 是一个强有力且近期的基线模型,文中已包含。论文中引用的五个基线模型(MulT, MISA, EmotionCLIP, TCL-MAP, EMOE)均为已发表的模型。研究阶段结束,现在开始撰写审稿意见。
Dense-Sparse Dynamic Time Warping for Customizing Piano Concerto Accompaniments
Published:
我已掌握足够的事实依据。关键发现:该论文发表于 ICASSP 2025,代码开源,未发现先前关于 Dense-Sparse DTW 的工作。基准测试很小(4 个乐章,150 个场景,1 位 pianist)。现在我将输出最终评审。
Audio Cross Verification Using Dual Alignment Likelihood Ratio Test
Published:
我现在有足够的信息来撰写评论了。论文中的所有组件(MFCC、DTW/HSTW、Shazam 风格的 fingerprinting、LRT)都是现有的技术。关键的先前研究是同一作者于 2020 年发表的 arXiv:2010.12173 (Shan & Tsai),它提出了相同的“音频交叉验证”问题框架,并提出了 SNWTW + BiLSTM。当前的 DA-LRT 是一种无需学习的后续研究。
End-to-End Markov State Sequence Learning for Auditory Attention Decoding
Published:
我现在已有足够的信息来输出审稿意见。关键发现总结:
CS-ETS: Chaos-Inspired Samba-Based EMG-To-Speech Synthesis with Nonlinear Chaotic Losses
Published:
Paper Review: CS-ETS: Chaos-Inspired Samba-Based EMG-To-Speech Synthesis with Nonlinear Chaotic Losses
Towards Array-Invariant Speech Enhancement via Geometry-Aware Dynamic Convolution
Published:
论文评审:通过几何感知动态卷积实现阵列不变的语音增强
Fusion Embedding: A Unified Embedding Space for Text, Image, Video, and Audio
Published:
Paper Review: Fusion Embedding: A Unified Embedding Space for Text, Image, Video, and Audio
Cyclic-Prefix-Free OFDM With Tail-Reuse Reconstruction for Distributed Acoustic Sensing
Published:
我已经掌握了全文、所有 wiki 查询结果(均为空)、paper-wiki 搜索结果(均为空)以及涵盖多个来源的 free-search 结果。现在开始撰写评审。
What the Waveform Knows: Transparent-first Speech and Audio Intelligence with Caption Studio
Published:
Paper Review: What the Waveform Knows: Transparent-first Speech and Audio Intelligence with Caption Studio
Summary of DCASE 2026 Task 5: Audio-Dependent Question Answering
Published:
我已经获取了完整的论文文本(5页,33364字符)以及知识库查询结果。OMC Wiki 中没有关于音频语言模型、DCASE 或文本捷径基准测试的条目。Paper-wiki 中仅包含 MOSS-Audio (2606.01802);关键引用文献 AudioMCQ/2509.21060、MMAU/2410.19168 未被收录。我将继续进行评审。
Towards a reproducible cross-venue method for quantifying crowd noise in stadiums
Published:
Paper Review: Towards a reproducible cross-venue method for quantifying crowd noise in stadiums
Transcription Policy as a Latent Variable: Activating Controllable Verbatim ASR with Word-Level Timing
Published:
现在我已经掌握了所有必要的证据。让我来撰写最终的评审。
Content is What Remains: Invariant Speech Tokenization from Parallel Utterances
Published:
我已经从 HTML 全文中获取了足够的证据。所有的 wiki 查询均未返回结果;free-search 确认了先前的研究 (ContentVec, DC-Spin, NAST, StableToken)。正在撰写评估。
Teleportation Game: Quantum Teleportation in Multi-Agent Systems for Interactive Music
Published:
Paper Review: Teleportation Game: Quantum Teleportation in Multi-Agent Systems for Interactive Music
OmniReasoner: Thinking with Long Audio-Video via Native Tool Use
Published:
Paper Review: OmniReasoner: Thinking with Long Audio-Video via Native Tool Use
Multimodal Speaker Verification as a Threat to Speaker Anonymization
Published:
我已经阅读了完整论文,并从 OMC Wiki、paper-wiki 和 free-search 中收集了相关的研究背景。论文全文共 5 页(完整 PDF 已通过 pdftotext 提取)。现在我将输出结构化的审稿意见。
Black-Box Optimization for Identifying and Inverting Audio Dynamic Range Control Effects
Published:
我已收集到足够的事实锚点。关键发现:
- DeepAFx (Ramírez et al., ICASSP 2021) — 已有 black-box audio effect 优化(SPSA),是本文最直接的前置工作
- arXiv 2504.04751 — “Unsupervised Estimation of Nonlinear Audio Effects” — diffusion-based blind system identification,是同期/近期同类工作
- arXiv 2411.04337 — 作者自己之前的工作(MEE),是本文的直接前作
- CleanUMamba — 语音去噪模型,不是 audio effect inversion 的自然 baseline
- OMC Wiki 和 paper-wiki 均无该领域记录
A Diagnostic Evaluation Framework for AI-Generated Cover Songs Using Music-Theoretic and Acoustic Features
Published:
Paper Review: A Diagnostic Evaluation Framework for AI-Generated Cover Songs Using Music-Theoretic and Acoustic Features
Ultra-Compact CNN Architectures for Tropical Bird Audio Detection on Microcontrollers
Published:
很好。我已经阅读了完整论文并进行了所有研究查询。OMC Wiki 中没有匹配的页面,paper-wiki 中没有相关记录,但 free-search 返回了有用的结果。我发现了以下内容:
- TinyChirp (Huang et al. 2024) — 基准模型,真实且已确认
- Benhamadi et al. 2026 — AudioMoth 能量模型,Goertzel 基准
- SEABAD (Zabidi et al. 2026) — 同一作者的配套数据集
- BioDCASE 2025,BirdNET,BirdVoxDetect — 相关工作已确认
- GitHub 仓库存在 (mun3im/seabadnet, mun3im/drongonet)
Learning the Arabic Dialect Continuum as a Continuous Space: A Regression Approach to Speaker Origin Prediction
Published:
Paper Review: Learning the Arabic Dialect Continuum as a Continuous Space: A Regression Approach to Speaker Origin Prediction
RPPNet: Perceptually-Grouped Rhythm-Pitch Primitives for Long-Term Structure Melody Generation via Boundary-Aware Modeling
Published:
Paper Review: RPPNet: Perceptually-Grouped Rhythm-Pitch Primitives for Long-Term Structure Melody Generation via Boundary-Aware Modeling
SimulS2ST-Omni: Data-Efficient Streaming Speech-to-Speech Translation via Explicit Trajectory Supervision
Published:
我已经收集了足够的信息。现在让我来汇总这份评审。从研究和全文阅读中确定的关键事实如下:
StellarTTS: Sparse Temporal Embedding for Low-Latency and Robust Speech Synthesis
Published:
我已经阅读了完整的 PDF,并确定了论文 wiki 的基准。现在我有足够的信息来撰写评审了。
Nonlinear Bias-Compensated Adaptive Filter and Its Application for Time-Series Prediction
Published:
Paper Review: Nonlinear Bias-Compensated Adaptive Filter and Its Application for Time-Series Prediction
Scalable Keyword Spotting via Modular Network Expansion
Published:
Paper Review: Scalable Keyword Spotting via Modular Network Expansion
Efficient Chain-of-Modality Reasoning via Progressive Compression for Spoken Language Models
Published:
我已经阅读了全文(HTML + PDF,包括附录)。现在我将整理评审报告。
Layer-Wise Decision Fusion for Fake Audio Detection Using XLS-R
Published:
Paper Review: Layer-Wise Decision Fusion for Fake Audio Detection Using XLS-R
Cumsum-Composable Phase Transport for Low-Cost Streaming Keyword Spotting
Published:
我现在已经掌握了所有需要的信息。以下是最终的评审意见。
Audio-Zero: Label-Free Self-Evolution for Fine-Grained Audio Reasoning
Published:
Paper Review: Audio-Zero: Label-Free Self-Evolution for Fine-Grained Audio Reasoning
Pushing the Frontier of Full-Song Generation: Hierarchical Autoregressive Planning Meets Flow-Matching Rendering
Published:
Paper Review: Pushing the Frontier of Full-Song Generation: Hierarchical Autoregressive Planning Meets Flow-Matching Rendering
Sound Probabilistic Safety Bounds for Large Language Models
Published:
我现在有足够的锚点来进行全面审稿。论文引用的关键相关工作包括:
Improved Monitoring of Honey bee Colony Strength via Audio IoT Sensors, Modulation Tensorgrams and Recurrent Neural Networks
Published:
Paper Review: Improved Monitoring of Honey bee Colony Strength via Audio IoT Sensors, Modulation Tensorgrams and Recurrent Neural Networks
Spectrogram-Based Joint Detection, Localization, and Classification of Events in Continuously Recorded IBR Waveforms
Published:
Paper Review: Spectrogram-Based Joint Detection, Localization, and Classification of Events in Continuously Recorded IBR Waveforms
Designed Vocalizations Dataset: Sound-Designed Human and Animal Voices for Non-human Voice Conversion
Published:
Paper Review: Designed Vocalizations Dataset: Sound-Designed Human and Animal Voices for Non-human Voice Conversion
Latent Variable-Mediated Cross-Learning for Few-Shot Acoustic Impedance Imaging
Published:
现在我已经掌握了所有必要的事实依据。根据阅读 PDF(全部 9 页)和 free-search 的结果,我对论文的研究背景有了全面的了解。以下是我的结构化审稿结果:
Faster IndexTTS-2: Accelerating and Streaming Autoregressive Zero-Shot Text-to-Speech Synthesis on GPUs
Published:
自由的学术搜索未能找到针对 TTS 使用 TensorRT-LLM 的直接相关工作,这确认了该研究方向的创新性。我已经掌握了足够的证据。现在我来整理这篇评论。
VibeVoice-ASR-BitNet Technical Report
Published:
我已阅读完整的 arXiv HTML 全文(174KB,涵盖所有章节:引言、方法、实验、结论、局限性、参考文献),查询了 OMC wiki(3次查询,均无结果),paper-wiki(概念搜索 + 4篇论文搜索,基础论文无直接命中),并使用 free-search 作为补充(BitNet b1.58 = 2402.17764,VibeVoice-ASR = 2601.18184,BitNet.cpp = 2410.16144 — 均来自同一 Microsoft 团队)。现在输出评审意见。
Toward Interpretable Speech Deepfake Detection using Artifact-Specific Experts and Calibrated Detection Scores
Published:
我已掌握所有必要信息。AASIST(Jung et al., ICASSP 2022)是公认的 SOTA 基准模型,在 ASVspoof 2019 评估集上的 EER 约为 0.83%。本文报告的 EER 为 19.76% —— 存在显著差距。现在我来撰写完整的审稿意见。
TF-MossFormer: Integrating Convolution Gated Local-Global Attentions for Enhanced Time-Frequency Domain Monaural Speech Separation
Published:
我已经掌握了所有必要的信息。现在我来整理最终审查报告。
Investigating Codec-Internal Latent Audio Watermarking for Neural Codec Robustness
Published:
关键事实锚点已充分建立。现在输出完整 review。
From Read Speech to Spoken Digits: A Task-Specific Evaluation of Speech Privacy With Informed Attackers
Published:
我已经完成了全文阅读和事实锚点建立。研究阶段结束,现在输出最终评审。
An Evaluation Framework for Structured Audio Captions Validated by Controlled Perturbations
Published:
Paper Review: An Evaluation Framework for Structured Audio Captions Validated by Controlled Perturbations
DONDO: Open w2v-BERT Speech-Recognition Base Models for African Languages
Published:
Paper Review: DONDO: Open w2v-BERT Speech-Recognition Base Models for African Languages
X$^3$-OPD: Distilling Reasoning into Large Audio-Language Models via On-Policy Alignment
Published:
我已经从全文(HTML,已提取 67K 字符)和 paper-wiki 的事实锚点中获得了足够的材料。OMC Wiki 中没有相关记录(记录了所有查询的失败)。paper-wiki 确认 X-OPD (2603.24596) 是作者团队之前的直接前身工作——相同的作者重叠(Cao, Fu, Jin),这使得 X3-OPD 成为具有扩展数据覆盖范围的后继工作。现在我将输出最终的结构化审查报告。
CODA: Cascaded Online Discontinuity-Aware Alignment for Real-Time Image-Based Score Following
Published:
Paper Review: CODA: Cascaded Online Discontinuity-Aware Alignment for Real-Time Image-Based Score Following
Listen, Do Not Copy: Internalizing Audio-Grounded Scaffold Context for Robust Omni-Model Speech Understanding
Published:
Paper Review: Listen, Do Not Copy: Internalizing Audio-Grounded Scaffold Context for Robust Omni-Model Speech Understanding
Music-JEPA: Learning a World Model of Sound from Action
Published:
我已经掌握了足够的事实依据。研究阶段结束(第 0-3 步,耗时 <15 轮)。现在输出完整的审稿报告。
How Meta-Learning Shapes LoRA Adapter Geometry in Speech Deepfake Detection
Published:
Paper Review: How Meta-Learning Shapes LoRA Adapter Geometry in Speech Deepfake Detection
MemNMF: Memory-Augmented NMF on LPC Spectra for Anomalous Sound Detection
Published:
我已经完成了所有必要的研究。我已经获取了论文全文,并确认了以下信息:
- 论文类型:方法型 (AE-based ASD with LPC + NMF memory)
- 先前研究:Gong et al. ICCV 2019 (memorizing normality), NMF-based ASD (PHMAP 2017), 基于记忆的 ASD (DCASE 2024 tech report, OpenReview PeV5LKMrxb)
- 基线模型:AE [5], AE-IDNN [6], ANP-IDNN [9], PAE [10], AudDSR [11], VAE/GRLNet [23]
- 数据集:MIMII, DCASE 2020 Task 2
Transforming Keystroke Noise to Text: Self-Supervised Acoustic Eavesdropping Attacks on Keyboards
Published:
Paper Review: Transforming Keystroke Noise to Text: Self-Supervised Acoustic Eavesdropping Attacks on Keyboards
MEUSLI: a Multilingual Projector for LLM-based ASR and Beyond
Published:
Paper Review: MEUSLI: a Multilingual Projector for LLM-based ASR and Beyond
IQ-JEPA: A Joint-Embedding Predictive Architecture with a Hermitian Vision Transformer for Sound Speed and Attenuation Estimation from Ultrasound IQ Data
Published:
我已获取完整论文正文及足够的 wiki/paper-wiki 锚点。以下是结构化审查报告。
Reflector: Arrangement-Aware Harmonic Retrieval for Sample-Based Composition
Published:
Paper Review: Reflector: Arrangement-Aware Harmonic Retrieval for Sample-Based Composition
OmniMate: Open-Ended Real-Time Streaming Audio-Visual Generation for Interactive Avatars
Published:
Paper Review: OmniMate: Open-Ended Real-Time Streaming Audio-Visual Generation for Interactive Avatars
Singlish, Can or Not? Fine-Tuning and Evaluating Zero-Shot TTS for Singapore English
Published:
我现在掌握了所有关键研究发现。最关键的发现如下:
- AccentBox (ICASSP 2025) - 针对零样本(zero-shot)口音生成的现有工作,论文引用了该工作但未将其作为基线进行比较
- 两篇 NTU 学位论文探讨了新加坡英语零样本 TTS 及 Singlish 神经语音合成——这直接对“首次系统研究”的新颖性声明提出了挑战
- malaya-speech 和 sil-ai/tts-singlish 项目 - 先前存在的 Singlish TTS 实现
- 存在多种口音 TTS 方法(Accent Vector、Scalable Controllable Accented TTS、Joycent),但均未被用作基线
Speech Signals Complement LLMs for Predicting Interpersonal Attraction in Speed Dating
Published:
论文评审:《语音信号与 LLMs 在预测快速约会中人际吸引力方面的互补性》
Low-Latency Turn-Taking via Context-Aware Preface Generation in a Real-World Dialogue Robot
Published:
Paper Review: Low-Latency Turn-Taking via Context-Aware Preface Generation in a Real-World Dialogue Robot
Infinite Canons: Maximally Self-Similar Melodic Lines and Canons with Infinite Solutions
Published:
Paper Review: Infinite Canons: Maximally Self-Similar Melodic Lines and Canons with Infinite Solutions
PathRIR: Physics-Guided Acoustic Path Selection and Late-Tail Compensation for Fast Room Impulse Response Simulation
Published:
我已阅读全文,并完成了所有维基、论文维基和自由搜索查询。现在开始撰写最终评审。
Explainable AI through the Lens of Material Agency: Enabling Musical Interface Design with Neural Audio Models
Published:
Paper Review: Explainable AI through the Lens of Material Agency: Enabling Musical Interface Design with Neural Audio Models
Music-Source-Separation-Training (MSST): A Unified Framework for Training and Evaluating Music Demixing Models
Published:
Paper Review: Music-Source-Separation-Training (MSST): A Unified Framework for Training and Evaluating Music Demixing Models
Improving Zero-Shot Phonetic Classification through Language-Agnostic Articulatory Features
Published:
Paper Review: Improving Zero-Shot Phonetic Classification through Language-Agnostic Articulatory Features
Expose Your Disguise: Recovering Source Speaker Identity From Voice Conversion
Published:
我已经掌握了足够的证据。现在输出最终的结构化审查报告。
Indic DiarBench: A Multilingual Joint Diarization and ASR Benchmark for Indian Languages
Published:
我已经阅读了完整的论文。我已经掌握了所有必要的信息来撰写评审。现在开始生成结构化的 Markdown 评审。
Memory Efficient Audio Synthesis with Decoupled Temporal Depth Diffusion Transformers
Published:
我已经掌握了足够的依据。通过 free-search 发现的关键点:T-Mimi (2601.20094) 和 Qwen3-TTS 是用于端侧 TTS 的相关并行/近期工作,而 “Efficient Streaming TTS Acoustic Model with Depthwise RVQ Decoding Strategies in a Mamba Framework” (Interspeech 2025) 是另一项直接针对 RVQ 深度解码效率的相关工作。现在我已有足够的信息来撰写评估报告。
Earnings25: A Comprehensive 500-Hour Speech Benchmark for Finance
Published:
我已经掌握了足够的事实依据。正在综合撰写审稿意见。核心锚点如下:
Automatic Audio Equalization with Semantic Embeddings
Published:
Paper Review: Automatic Audio Equalization with Semantic Embeddings
OmniVAE: An Audio-Video VAE with Cross-Modal Alignment for Joint Generation
Published:
Paper Review: OmniVAE: An Audio-Video VAE with Cross-Modal Alignment for Joint Generation
Qwen-Audio-3.0-TTS: Freely Controllable and Highly Robust Speech Synthesis with Multi-Stage Training Paradigm
Published:
Paper Review: Qwen-Audio-3.0-TTS: Freely Controllable and Highly Robust Speech Synthesis with Multi-Stage Training Paradigm
Modeling Stylistic Co-evolution in Symbolic Music Heritage Collections
Published:
我已阅读了完整的 28 页 PDF,并完成了所有必要的 wiki 和免费搜索查询。现在输出结构化的评审意见。
Leveraging Gradient Reversal Loss and Multitask Learning for Datasets-Aware Audio Deepfake Detection
Published:
我已经获得了完整的论文全文,并从之前的搜索中获得了足够的背景信息。现在我来撰写最终的结构化审稿报告。
AptAvatar: Fast and Vivid Long-Form Audio-Driven Video Generation for Production-Ready Avatars
Published:
Paper Review: AptAvatar: Fast and Vivid Long-Form Audio-Driven Video Generation for Production-Ready Avatars
MoLGE: Mixture of Language Group Experts for Efficient Scaling of Massively Multilingual Speech Recognition
Published:
我已经获取了所有必要的信息。仅存在两个基线:dense 和 random grouping。没有与 Kwon & Chung (2023) 的原始 MoLE、Wang 等人 (2023) 的 language-routing MoE,或任何近期基于 MoE 的多语言 ASR 进行比较。这是一个重大的基线缺失。现在开始撰写评论。
Looking for Affect in Spontaneous Finnish Speech through Linguistic Interpretability
Published:
我现在已有充分的事实依据。论文本身引用了 Wagner 2023,其揭示了纯音频模型会隐式地学习语言结构——这直接确立了多模态文本与音频情感识别(SER)已经是一个有充分研究基础的研究方向,而本文的研究结果证实了 Wagner 的早期发现。我现在已经有了完整的事实基础。让我输出最终评审。
Revisiting Vocos: That Phasiness Business in Time-Frequency Neural Vocoding
Published:
我已经掌握了所需的全部证据。让我来整理这份评审。这项并发工作(2509.18806,发表于 2025 年 9 月)在 2 个月的并发窗口期之外(该论文发表于 2026 年 7 月),因此它是相关的前期工作。让我来撰写最终的评审。
TaoMate: Anchor-Guided Memory Bridging Evolving and Reference States for Real-Time Audio-Video Digital Human Generation
Published:
我已经建立了全面的事实锚点。我已具备撰写评论所需的所有信息。让我总结一下我所掌握的内容:
Let Me Look at You: Advanced Facial Expression Modeling for Conversational Speech Synthesis
Published:
Paper Review: Let Me Look at You: Advanced Facial Expression Modeling for Conversational Speech Synthesis
Mind the Microphone Gap: Benchmarking Array Upsampling Strategies for Latent Acoustic Mapping
Published:
我已经获取了完整的论文和所有必要的研究锚点。现在开始输出最终的结构化审稿意见。
A Cross-lingual Comparison of Human and Classification Model Entrainment Behavior in Code-switched Speech Settings
Published:
Paper Review: A Cross-lingual Comparison of Human and Classification Model Entrainment Behavior in Code-switched Speech Settings
AMRD: Adaptive Multi-Teacher Relational Distillation for Lightweight Speech Emotion Recognition
Published:
Paper Review: AMRD: Adaptive Multi-Teacher Relational Distillation for Lightweight Speech Emotion Recognition
faster-enhancer.c: A Dependency-Free int8 Runtime for Streaming Speech Enhancement on Commodity CPUs
Published:
现在我已经掌握了所有事实锚点。让我根据全文撰写结构化的评审。
Extracting Voice Styles from Frozen TTS Models via Gradient-Based Inverse Optimization
Published:
研究已完成。OMC Wiki 未查询到任何相关记录。paper-wiki 中收录了 VALL-E、Seed-TTS、CosyVoice 和 WavLM,但未收录 SupertonicTTS、VoiceLoop、kNN-VC、ECAPA-TDNN 或 Marras et al. 的相关内容。现在我已经获取了完整论文文本以及研究基准,准备输出评审意见。
From Semantics to Readout: Mechanistic Understanding of Audio Tokens after Fine-Tuning for Temporal Audio Grounding
Published:
我已经收集到了足够的研究依据。现在输出最终的综述。
Finding the noise: Zero-shot AI Music Detection
Published:
我现在已经收集好了所有证据。让我来撰写最终的评审。
Less is More: Modality-Decoupling for General AIGC Audio-Video Detection
Published:
太好了——我现在有了详尽的研究结果。我发现了以下几项关键的前置工作:
- MVAD 数据集论文 (arXiv:2512.00336,发表于 2025 年 11 月) —— 这是本文使用的基础数据集
- “Divide and Conquer” (arXiv:2602.00209) —— 针对 DDL 挑战赛的多模态 deepfake 检测
- X-AVDT (CVPR 2026) —— 用于 deepfake 检测的音视频交叉注意力机制
- HAVIC (CVPR 2026, 作者同一团队) —— 整体音视频内在一致性
- 关于 deepfake 检测中分数级融合规则的分析 (MDPI 2022)
- “A Robust Approach to Multimodal Deepfake Detection” —— 展示了模态解耦的优势
Evaluation of forced alignment of code-mixed speech: the case of Hindi-English
Published:
代码仓库已存在(HTTP 200)。我现在拥有所需的全部事实锚点。让我来撰写最终的结构化审查报告。
Joint Text-Audio Alignment for EEG-to-Text Decoding in Chinese Speech Production and Perception
Published:
Paper Review: Joint Text-Audio Alignment for EEG-to-Text Decoding in Chinese Speech Production and Perception
SpeechLLM Meets Federated Learning for End-to-End ASR: English and Italian Case Studies
Published:
Paper Review: SpeechLLM Meets Federated Learning for End-to-End ASR: English and Italian Case Studies
GraphIDyOM: A graph-native Python reimplementation of IDyOM for musical expectation modelling
Published:
Paper Review: GraphIDyOM: A graph-native Python reimplementation of IDyOM for musical expectation modelling
VAD to the Bone: Ultra-Tiny Speech Activity Detection for Edge Deployment
Published:
Paper Review: VAD to the Bone: Ultra-Tiny Speech Activity Detection for Edge Deployment
Device Invariance using Domain Adaptation on Acoustic Scene Classification
Published:
现在我已经有了坚实的事实依据。DCASE 2020 Task 1A 的排行榜显示,最佳准确率约为 76-78%。这篇论文在设备 A 上匹配的 PaSST 基线达到了 84.5%,但未匹配的准确率下降到了 61-65%。让我来汇总最终的审稿意见。
Depression Markers in Speech: An Approach based on Tract Variables Dynamics
Published:
论文评审:语音中的抑郁标志物:一种基于声道变量动力学的方法
Spacing Out: On the Reliability of Binaural Music Source Separation Metrics
Published:
我已经收集到了足够的事实依据。现在让我来汇总这份评审。关键发现如下:
LLM4OSC: Profile-Bound Natural Language Control with Deterministic Validation for Open Sound Control
Published:
我已获取全文并建立了事实锚点。现在我来撰写这篇综述。
Voice Memory for Agentic Speech Recognition
Published:
I have sufficient evidence from the full paper text and external searches. Key facts gathered:
Explicit Note-Event Tokenization and Pitch-Validity Constrained Decoding for MIDI-to-Tablature Transcription
Published:
我已经掌握了完整论文文本和研究结果。现在开始输出最终评审。
Audio-Anchored Fusion of Multi-Ratio DiT Reconstruction Residuals for Cross-Domain Audio Deepfake Detection
Published:
我已经掌握了所需的所有信息。现在开始输出最终的结构化审查结果。
Prosody-driven Jailbreaks in Audio LLMs: A Controlled Study and Mechanistic Analysis
Published:
我现在已经掌握了全文和所有研究数据。以下是最终的结构化审查报告。
ThinkOmni: A Reasoning-Driven Omni-Modal LLM Framework for Audio Forgery Detection and Localization
Published:
我已阅读了论文全文(包括正文和附录),并完成了研究查询。以下是结构化评审。
Speech2Grasp: Data-Efficient Transfer of Text-Conditioned Grasp Detection to Speech in Humanoid Robots
Published:
Paper Review: Speech2Grasp: Data-Efficient Transfer of Text-Conditioned Grasp Detection to Speech in Humanoid Robots
Unfolded Recursive Expectation-Maximization Neural Network For Speaker Tracking
Published:
论文评审:基于展开递归期望最大化神经网络的单说话人追踪
Few-Shot Open-Set Audio Classification via Transductive Prototype Refinement and Class Logit Enhancement
Published:
Paper Review: Few-Shot Open-Set Audio Classification via Transductive Prototype Refinement and Class Logit Enhancement
A Study on Online Mask-based Beamforming Using Per-channel Masking for Spatially Distributed Microphones
Published:
Paper Review: A Study on Online Mask-based Beamforming Using Per-channel Masking for Spatially Distributed Microphones
MPEcho: A Melody and Phoneme-Aware Generative Framework for Controllable Cover Song Generation
Published:
Paper Review: MPEcho: A Melody and Phoneme-Aware Generative Framework for Controllable Cover Song Generation
Zero-Shot Face-to-Speech Synthesis via Latent Space Adaptation of a Style-Diffusion TTS Model
Published:
我掌握了充分的依据:自由搜索证实了该领域已有先前的工作(Face2Speech, FVTTS, Face-TTS/Imaginary voice, Face-StyleSpeech, Progressive Facial Granularity)。我已经获得了论文全文并进行了必要的研究。现在我将撰写审稿意见。
Phoneme- vs. Character-Level Targets and Selective State-Space Models for Intracortical Brain-to-Text
Published:
Paper Review: Phoneme- vs. Character-Level Targets and Selective State-Space Models for Intracortical Brain-to-Text
Ripple: Real-Time Streaming Audio-Video Generation With Cross-Modal Recurrent Memory
Published:
我已经掌握了所需的所有研究数据。让我来整理最终的评审意见。从我的研究中所确立的关键事实如下:
Detection of AI-generated stems within hybrid human-AI music
Published:
论文评审:混合人机音乐中 AI 生成音干的检测
Latent-IM: Latent Interaction Management for Speech LLMs
Published:
未找到代码仓库。我已掌握所需的所有信息。现在输出最终审查报告。
Qwen-Audio-3.0-Gen-Preview Technical Report
Published:
Paper Review: Qwen-Audio-3.0-Gen-Preview Technical Report
MMAC: A Massive Multi-dimensional Benchmark for Audio Captioning
Published:
Paper Review: MMAC: A Massive Multi-dimensional Benchmark for Audio Captioning
Cocktail-Talker: Multi-Speaker Dialog Modeling in Noisy Social Environments with Turn Action GRPO
Published:
Paper Review: Cocktail-Talker: Multi-Speaker Dialog Modeling in Noisy Social Environments with Turn Action GRPO
VocalRender: Score-Native Singing Voice Synthesis for Real-World Composition
Published:
我已经获取了完整文本并确认了关键事实。我已经有足够的信息来撰写评审了。关键事实总结如下:
RIPPLE: Generating Multi-Channel Phase, Not Recovering It
Published:
我已经完成了足够的研究。让我整理一下调查结果并撰写评审。研究阶段总结如下:
CrowdioSet and PaRIRset: Two Datasets Towards Live Music Source Separation
Published:
我已阅读完整论文,并从 wiki/paper-wiki/free-search 中收集了证据。所有 wiki 查询均未返回结果;paper-wiki 未返回结果;我使用了 free-search 作为备用。现在我将输出最终的结构化审稿意见。
Integrating Contextual Embeddings into Evaluation of Expressive MIDI Piano Performances
Published:
Paper Review: Integrating Contextual Embeddings into Evaluation of Expressive MIDI Piano Performances
Face and Voice Cross-modal Association with Learning Convex Feature Embedding
Published:
我已经掌握了足够的研究背景信息。现在开始撰写最终的审稿意见。
AgenticASR: Refining Speech Recognition in Real-World Scenarios via an Agentic Approach
Published:
Paper Review: AgenticASR: Refining Speech Recognition in Real-World Scenarios via an Agentic Approach
Teffic-Audio: Tell Fact from Fiction
Published:
我已经阅读了完整的16页论文,并从wiki、paper-wiki和free-search中收集了研究证据。现在输出最终的结构化综述。
Correlation between prosody and pragmatics: A case study of the discourse marker hālā `now’ in Persian
Published:
我已阅读完整论文(35页)并收集了研究背景。以下是我的评审意见。
Do Music Foundation Models Embed Pitch in Helical Structure?
Published:
现在我已经掌握了所有必要的事实锚点。让我输出最终的结构化评审报告。
DoubleHelix: Structured Cross-Modal Fusion for Audio-Visual Speech Recognition with LLMs
Published:
我已收集好所有证据。现在根据完整的 PDF 阅读和外部研究输出最终评审。
Model-Agnostic Meta-Learning Initialization for Distributed Multichannel Active Noise Control
Published:
我已经有了所需的所有证据。让我来整理最终的审稿报告。
M3-DuplexBench: A Multi-Turn, Multilingual, Multidomain Benchmark for Full-Duplex Spoken Dialogue Models
Published:
现在我已经掌握了所有事实依据。以下是最终的结构化审查报告。
Exploring Efficient Waveform Diffusion Models for Foley Sound Generation
Published:
事实锚点已建立。关键发现:
ParaASR: Multi-Token Prediction for Fast and Long-Context LLM-Based Speech Recognition
Published:
论文审查:ParaASR:基于多 token 预测的快速且长上下文 LLM 语音识别
Leveraging Beam Search Information for Confidence Estimation in E2E ASR
Published:
Paper Review: Leveraging Beam Search Information for Confidence Estimation in E2E ASR
Versatile On-device Adaptation at the Edge by Unifying Few-shot, Zero-shot, Continual, and In-context Learning
Published:
现在我已经获取了所有研究锚点。OMC wiki:无结果(在所有 8 次查询中均为空)。paper-wiki:失败(缺少 papers 目录)。Free-search:已确认 Chameleon [25] 是作者之前发表的关于 FSL+CL 的工作;FSL-HDnn [18] 和 Clo-HDnn [19] 是基于硬件的 FSL/CL 竞争对手;SAPIENS [20] 和 Kim et al. [21] 是 FSL 硬件竞争对手;Liu et al. [50] 是基于硬件的 ZSL 竞争对手;RegBench [51] (Akyurek et al. 2024) 和 Fluent Speech Commands [49] 是已确立的基准。目前没有先前的工作在边缘设备上统一这四种学习场景。现在开始输出审查结果。
Stable Autoregressive Speech Generation with Low-Frame-Rate High-Dimensional Continuous Tokens
Published:
我已从已阅读的全文 PDF 中获取了所有必要信息。论文内部引用了 ELD(该论文提出了一种编解码器 + AR 框架,并命名其组件为 Locodec 和 MP-ELD)。现在我将输出最终的审稿意见。
Anomalous Sound Detection Meets Noise-Aware Self-Supervised Learning
Published:
我已经掌握了所有必要的事实。现在让我来撰写最终的评审。
Scene2Sound: Auditory-Grounded Soundscape Generation for 3D Gaussian Worlds
Published:
HTML 导航导致无法轻易提取章节内容。我已在 PDF 阅读中获取了足够的消融实验和研究结果。现在,我将根据收集到的事实输出最终的审稿意见。
Hidden-Domain Routing for All-Type Audio Deepfake Detection
Published:
我已经有了足够的事实依据。让我来整理一下从全文和外部搜索中收集到的关键信息:
Beyond Prompt Adherence: Auditing Attribute-Level Voice Control in Speech Generation
Published:
我现在已经有足够的研究依据了。让我来整理这篇综述。我已经获得了完整的论文文本,执行了所有必要的 wiki 查询(均未返回结果),搜索了 paper-wiki(找到了 CosyVoice3、Fish-Speech-S2、VoxCPM2、InstructTTSEval、SPAM、DMP-TTS、DisCo-Speech),并进行了自由搜索(找到了关键的相关研究 “Do You Hear What I Mean?” 2509.13989)。
AnyBand: Unified Multi-Bandwidth Speech Extension via Frequency-Aware In-Context Spectral Infilling
Published:
我已经完成了研究阶段。我已阅读完整的论文(通过 pdftotext 获取的 2611 行),并从 free-search 中获得了关键发现:
UOT-IR: Structured Routing of High-Polyphony Symbolic Music into Fixed-Budget Representations
Published:
Paper Review: UOT-IR: Structured Routing of High-Polyphony Symbolic Music into Fixed-Budget Representations
Simulation-Based Plate-Reverb Parameter Estimation from a Single Impulse Response
Published:
Paper Review: Simulation-Based Plate-Reverb Parameter Estimation from a Single Impulse Response
Band-Count Dense Modal Estimation with Fixed-Frequency Differentiable Resonator Refinement
Published:
所有三次 OMC wiki 查询均未返回结果。OMC wiki 中没有关于此主题的记录。free-search 工具找到了关键的相关工作(Engel 等人的 DDSP [ICLR 2020],Diaz 等人的 Differentiable Modal Resonators [ICASSP 2023],Yu 等人的 Differentiable All-pole Filters [OpenReview],以及 ESPRIT/matrix-pencil 经典方法)。paper-wiki 中没有相关记录。我已掌握足够的事实锚点来撰写审稿报告。现在开始生成结构化的 Markdown 审稿报告。
Experience-Calibrated Contrastive Decoding for Mitigating Hallucinations in LM-Based Text-to-Speech
Published:
我已获取全文并完成所有维基/论文库/自由搜索查询。OMC Wiki 对所有查询均未返回相关记录(TTS 幻觉领域无经验知识)。论文库中收录了 CosyVoice、Seed-TTS、GLM-TTS 和对比解码综述论文 (2603.09232)。Free-search 确认了该论文及两篇密切相关的先行工作:Liu 等人 2025 (EMNLP, GFlowNet 分布对齐) 和 Wang 等人 2025 (注意力引导)。现在开始输出最终评审。
SoniSpeech: A Large-Scale Open-Vocabulary Tri-Modal Dataset for Wearable Silent Speech Interfaces
Published:
Paper Review: SoniSpeech: A Large-Scale Open-Vocabulary Tri-Modal Dataset for Wearable Silent Speech Interfaces
REIMU: Efficient Heterogeneous Hierarchical Reasoning for SSL-Based Speech Deepfake Detection
Published:
论文评审:REIMU:基于 SSL 的语音深度伪造检测中的高效异构分层推理
DRONEAUDIONET: Noise Suppression for Drone Audition-based Search and Rescue
Published:
我已获取完整论文。现在输出最终评审。
Beyond One-Size-Fits-All: Personalized and Culturally Adaptive Emotional TTS via Interactive Optimization of Individual Emotion Perception Spaces
Published:
Paper Review: Beyond One-Size-Fits-All: Personalized and Culturally Adaptive Emotional TTS via Interactive Optimization of Individual Emotion Perception Spaces
Separate-and-Detect: Unified Drum Transcription and Stem Generation via Latent Diffusion
Published:
我已经掌握了足够的事实锚点。现在让我输出最终的结构化审查报告。
JoyAI-Talker: Full-Duplex Speech Interactive Large Model Built for Empathetic Voice Agents
Published:
已收集足够的研究锚点。论文 wiki 确认了以下内容:
- Moshi (2410.00037):双流全双工,7B Helium LLM,Mimi codec,Inner Monologue — 已记录
- Qwen3-Omni (2509.17765):Thinker-Talker MoE 架构,这是首个在多模态之间没有性能下降的单模型 — 已记录
- Freeze-Omni (2411.00774):未在 wiki 中收录,但论文引用其为 baseline
- JoyVoice (2512.19090):同一团队的 Talker 所采用的架构 — 已记录
- TurnGuide (2508.07375):全双工 turn-taking SOTA — 已记录
- Full-Duplex-Bench v3 (2604.04847):存在较新的 v3 版本,但论文使用的是 v1.5
InteracVid: Building a Real Interactive Audio-Visual Response Dataset from Live-Chat Videos
Published:
我已经阅读了完整的 49 页 PDF,并执行了所有必要的研究查询。现在我将输出最终的结构化评审。
Latent Softmax for Data-Efficient Phoneme-Based Multilingual ASR Across Tonal and Non-Tonal Languages
Published:
我已有足够的研究锚点。OMC wiki 没有返回任何结果;paper-wiki 有一些 ASR 论文 ID,但该概念下没有相关的数据集或具体论文;free-search 确认了 Whistle 论文(相同作者群,前置工作)、Autosegmental Neural Nets(参考文献 14,音素-声调建模)、Universal Phone Recognition(参考文献 16),并且未发现已有的“latent softmax”或针对音素 ASR 的声调边缘化方法。现在我将输出最终评审。
FATE: Frame-Level Audio-Visual Temporal Embedding
Published:
我已经获取了所需的所有信息。全文已阅读(742行),知识库已查询(OMC wiki 和 paper-wiki 中无记录,已用 free-search 补充),基线已通过搜索验证。正在输出最终评审。
Interpretable MEG Decoding of Perceived Speech: Cortical Sources and the Stimulus Features That Drive Retrieval
Published:
Paper Review: Interpretable MEG Decoding of Perceived Speech: Cortical Sources and the Stimulus Features That Drive Retrieval
Discriminative Axis, Not Data Volume: What a Contrastive Corpus Teaches an Audio Embedding
Published:
我已经掌握了所有必要的事实锚点。现在让我来综合撰写这份评审。
SGAD: A State-Guided Adaptive Decision Framework for Robust EEG-Based Auditory Attention Switch Decoding
Published:
未找到 SGAD (EEG) 的代码。现在我已经掌握了所有需要的事实。让我来整理最终评审。
SAGE: Switch-Aware EEG-Guided Soft Gating for Target Speaker Extraction with In-Trial Switching
Published:
我现在已经掌握了充分的依据。我已经阅读了完整的论文正文,并收集了相关的研究背景。让我来整理最终的审稿意见。
StreamTalk: Streaming Co-Speech Gesture Generation with Key-Pose Anchoring
Published:
太好了,我已经有足够的信息来撰写评论了。关键发现如下:
Allocation Before Ranking: Decoupled Token Compression for OmniLLMs
Published:
我已经掌握了所有必要的事实。让我来撰写评审。
Comparative Validation of GPT-4o-mini and Teacher Mean Scores for Automated Scoring of Music Analysis Responses: Single-Pass Deployment, Repeatability, and Strategy-Specific Bias
Published:
Paper Review: Comparative Validation of GPT-4o-mini and Teacher Mean Scores for Automated Scoring of Music Analysis Responses: Single-Pass Deployment, Repeatability, and Strategy-Specific Bias
Multi-Backbone Self-Supervised Ensembles for Audio Deepfake Detection and a Cross-Track Analysis of Generation-Detection Asymmetry
Published:
I now have all the fact anchors I need. Key findings from research:
Hear, Invoke, and Understand: A Skill-Calling Multimodal Agent for Large Audio Language Models
Published:
我已经获取了足够的研究锚点。现在让我汇总关键发现并撰写审稿。
P-MUSE: Prompt-MIDI-Optional Model for Unified Instrumental Music Synthesis and Editing
Published:
Paper Review: P-MUSE: Prompt-MIDI-Optional Model for Unified Instrumental Music Synthesis and Editing
Music Restoration via Latent Operator Optimization and Diffusion Model Priors
Published:
现在我已经有足够的研究数据来撰写评审了。让我整理一下我已知的信息:
SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks
Published:
Paper Review: SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks
AcoustiTrace: When Plausible Sound Violates Physics
Published:
论文评审:AcoustiTrace: When Plausible Sound Violates Physics
Uncertainty-Aware Crossmodal Fusion for Classification of Animal Behavior
Published:
Paper Review: Uncertainty-Aware Crossmodal Fusion for Classification of Animal Behavior
The Role of Disfluencies in Speech Translation
Published:
我已阅读全文(共 20 页,已转换为 3876 行文本)并完成了所有 wiki 查询。OMC Wiki 中未找到关于语音翻译不流利主题的相关记录。paper-wiki 中虽然收录了关于音频基础模型 (2602.16687) 和带有不流利现象的语音智能体 (2604.04847) 的论文,但它们是相关工作,而非直接的基准模型或重复研究。以下是结构化评审结果。
MEMS Microphones as Ultrasonic Transducers: Nonlinear Electrostatic Actuation and a Parametric Array Prototype
Published:
研究阶段完成。研究轮次统计:1 次 PDF 下载 + 3 次 wiki_query(均为空)+ 3 次 paper-wiki 搜索(均为空)+ 3 次自由搜索(共 9 个来源,33 个结果)。总计约 8 轮。输出最终评审。
Sounding Canvas: Embedding Algorithms in Networked, Sensorial Sound Art
Published:
我已经收集到了足够的资料。自由搜索(free-search)结果确认了一些重要的先例工作。论文全文已阅读完毕。现在开始撰写评论。
Domain-Specific Evaluation of Text-to-Speech Systems: A Multi-Metric Benchmarking Study
Published:
Paper Review: Domain-Specific Evaluation of Text-to-Speech Systems: A Multi-Metric Benchmarking Study
An End-to-End Workflow for Fin Whale Song Detection, Note Characterization, and Localization with Distributed Acoustic Sensing
Published:
我现在已经掌握了充分的证据基础。让我来汇总关键发现,并生成最终的评审。
Can Foundation Models Hear What Made That Sound? A Tiered Benchmark of Audio-Language Models and Traditional Classifiers for Closed-Set Sound Source Identification
Published:
Paper Review: Can Foundation Models Hear What Made That Sound? A Tiered Benchmark of Audio-Language Models and Traditional Classifiers for Closed-Set Sound Source Identification
Deep Learning-Based Active Trim Panels for Enhanced Aircraft Interior Noise Control
Published:
我现在已经掌握了所需的所有信息。以下是基于全文阅读和调研结果得出的完整评审意见。
EchoCache: Energy-Guided Cross-Modal Caching for Efficient Audio-Driven Video Generation
Published:
Paper Review: EchoCache: Energy-Guided Cross-Modal Caching for Efficient Audio-Driven Video Generation
Towards Real-world Environment-aware Zero-shot Text-to-speech Synthesis via Disentangled Audio Infilling
Published:
Paper Review: Towards Real-world Environment-aware Zero-shot Text-to-speech Synthesis via Disentangled Audio Infilling
MeloCodec: Harnessing Melodic Priors for High-Fidelity Singing Voice Representation
Published:
paper-wiki 对这些查询没有返回结果。我已经有足够的锚点:EnCodec、SpeechTokenizer、NaturalSpeech 3 (FACodec) 和 Vevo2 已被记录。现在我将输出最终的结构化评审。
DDSynth-RL: Audio Synthesizer Inversion via Discrete Diffusion with Reinforcement Learning
Published:
我已经完成了所有必要的研究。以下是结构化的评审意见。
Calliphony: A Calligraphy-Driven Interface for Real-Time Generative Music Performance
Published:
Paper Review: Calliphony: A Calligraphy-Driven Interface for Real-Time Generative Music Performance
Learning Music Style for Piano Arrangement Through Cross-Modal Bootstrapping
Published:
arXiv 摘要页面或论文正文中未提供代码链接。该论文提供了一个演示页面 (https://zhaojw1998.github.io/bossa/),但没有公开的代码存储库。现在我已经掌握了所需的所有信息。让我来撰写完整的评审。
Towards More Expressive Spoken LLMs: Fine-Grained Intent Benchmarking and Acoustic-Lexical Decoupled Policy Optimization
Published:
我已经有了足够的证据。现在让我根据收集到的所有证据,产出最终的审稿意见。
GROW: Group-Relative Advantage-Weighted On-Policy Reinforcement Learning of Autoregressive-Diffusion Text-to-Speech model
Published:
Paper Review: GROW: Group-Relative Advantage-Weighted On-Policy Reinforcement Learning of Autoregressive-Diffusion Text-to-Speech Model
CLASVS: Continuous-Latent Autoregression for Melody-Preserving Lyric Editing in Singing Voice Synthesis
Published:
Paper Review: CLASVS: Continuous-Latent Autoregression for Melody-Preserving Lyric Editing in Singing Voice Synthesis
Hear to See: Discerning Stateful Listening for Audio-Visual Instance Segmentation
Published:
我已获取完整数据。论文已全文阅读(HTML:引言、相关工作、包含 ASD/HCM/ADM 在内的方法、损失函数、包含表 1-8 和附录 A-B 在内的实验、消融实验、可视化、局限性)。已对 OMC Wiki(无记录)、paper-wiki(无记录)及免费搜索(找到 AVISM CVPR’25、ACVIS ICASSP’26、AVS-Mamba、SeaVIS、Learning What to Hear ICASSP’26)进行了查询。现已有足够信息输出评审。
Multi-Task Multi-Frame Visual Piano Transcription
Published:
Paper Review: Multi-Task Multi-Frame Visual Piano Transcription
Cross-cultural evaluation of taste-sound correspondences in AI-generated music
Published:
Paper Review: Cross-cultural evaluation of taste–sound correspondences in AI-generated music
Language-Specialized Multi-Teacher On-Policy Distillation for Multilingual LLM-Based ASR
Published:
我已经完成了全文阅读和所有 wiki/free-search 查询。现在输出最终的结构化 review。
Speaker Verification Under Real Classroom Conditions for English Speech
Published:
EDSI 是马里兰大学的一个项目,目前正计划开放获取,但目前尚未提供下载链接,且需要通过专用平台才能访问 PII。没有找到该论文的代码/GitHub 仓库。我已经掌握了足够的事实锚点。现在我将撰写最终的评审报告。
Echo-Aware Modulation for Compact-Latent Frequency-Time Modeling in Lightweight Acoustic Echo Cancellation
Published:
我已经掌握了足够的事实依据。论文全文已阅读,前人研究成果也已通过 free-search 确认(wiki/paper-wiki 中没有相关记录)。现在开始输出最终评审。
On the Geometry of Music Bandwidth Extension in Latent Spaces of Audio Codecs
Published:
研究完成。我已经阅读了完整的 HTML 全文(44KB 提取文本)、查询了 OMC Wiki(所有 6 次查询均未返回结果)、查询了 paper-wiki(无相关记录),并使用了 free-search 找到了相关工作。现在输出最终审稿。
AI-Based Sound Effect Generation: A Narrative Review of Generative Models Across Input Modalities
Published:
Paper Review: AI-Based Sound Effect Generation: A Narrative Review of Generative Models Across Input Modalities
Equivariant Music Transformer
Published:
我现在已经掌握了充分的信息。我已确认以下事实:
- Moonbeam (arXiv:2505.15559) 的作者是同一位第一作者 (Zixun Guo) —— EMT 正是建立在 Moonbeam 的基础之上
- Anticipatory Music Transformer (arXiv:2306.08620) 是公认的 SOTA 基线
- Music102 是一篇关于 D12-等变 Transformer 的相关工作,但它专注于和弦进行伴奏,而非生成式建模
- PESTO、STONE 和其他等变自监督工作侧重于判别式任务(音高/节奏/调性估计),而非自回归生成
Identity-Faithful Audio-Visual Target Speaker Extraction with QIANGDA and VOXBLINK2-AVSE
Published:
现在我已经掌握了确凿的基准数据。关键发现如下:
Should We Type or Talk to LLM Agents? A Comprehensive Study of Voice and Keyboard Input Perturbations
Published:
我已经掌握了充分的证据。现在我将输出最终的结构化审查报告。
Transfer Learning for Avian Bioacoustics under Sparse Positive Labels
Published:
我现在已经建立了所有必要的事实依据。BirdCLEF+ 2025 的获胜解决方案在 private leaderboard 上实现了约 0.930 的 ROC-AUC,而本文在 BirdCLEF+ 2026 上报告了 0.860 的 macro AUC。该论文的核心方法是一个基于源可靠性特征的 L2 正则化逻辑回归元分类器——本质上是带有特征工程的 stacking。让我来撰写最终的评审意见。
Agogic: Performance-Timed Music Tokens for LLM-Native Text-to-Symbolic-Music Generation
Published:
我已从 arXiv HTML 读取了完整论文,并对现有的基线进行了 Wiki/论文库/自由搜索的事实锚定。现在输出完整的评审意见。
Searching for Sound-Meaning Collisions: Graph-Based Affordance Retrieval and Multi-Evaluator Ranking for Pun Translation at CLEF 2026 JOKER Task 2
Published:
我已经掌握了所需的所有事实锚点。已确定的关键发现:
HyPASE: Hyperbolic Geometry for Parameter-Efficient Speech Emotion Fine-Tuning Framework for Large Audio-Language Models
Published:
我已经获取了全文和所有研究背景,现在开始输出最终评审。
Helping Music Co-Creation Agents ‘Listen’ Well: Hierarchical Self-Supervised World Models for Understanding and Generation
Published:
Paper Review: Helping Music Co-Creation Agents ‘Listen’ Well: Hierarchical Self-Supervised World Models for Understanding and Generation
AudioScape-TTA: A Structured Soundscape Benchmark for Fine-Grained Text-to-Audio Evaluation
Published:
我已经掌握了全面的证据。该论文中没有代码/数据发布声明——这是一个关于可复现性的重要发现。让我再核实一下 AnyAudio-Judge 的作者重叠情况,并检查 AQAScore 的细节以进行新颖性比较,然后撰写评论。
A Dual Evaluation for Music Transcription
Published:
Paper Review: A Dual Evaluation for Music Transcription
Towards Robust Version Identification in the Wild: A Dataset, Benchmark, and Fine-Tuning Study
Published:
OMC Wiki 没有关于音乐信息检索(MIR)或版本识别(version identification)的记录。所有事实锚点均来自 free-search。我已经拥有完整的论文文本和充分的外部证据来撰写审稿报告。现在让我来整理这篇结构化的审稿 Markdown。
Breaking the Curse ofMultilinguality inMany-to-Many Speech-to-Text Translation via a Resource-AwareMixture of Speech Encoders
Published:
我已经通过全文阅读和外部搜索获取了足够的证据。现在我来整理这份评审。我收集到的关键发现如下:
Masked diffusion enables coherent beat tracking
Published:
Paper Review: Masked Diffusion Enables Coherent Beat Tracking
A Modular Part-of-Speech Tagger for Scottish Gaelic using spaCy
Published:
Paper Review: A Modular Part-of-Speech Tagger for Scottish Gaelic using spaCy
Visual Representation Matters: Exploiting Temporal Differences in Video-to-Audio Generation
Published:
Paper Review: Visual Representation Matters: Exploiting Temporal Differences in Video-to-Audio Generation
Spoken Function Calling: A New Perspective on Spoken Language Understanding for Large Audio Language Models
Published:
论文审稿:口语功能调用:面向大型音频语言模型的口语理解新视角
Objects as Audio-Visual Modal Sound Fields
Published:
论文评审:Objects as Audio-Visual Modal Sound Fields
Rethinking Automatic Music Mixing as Sequential Stem Blending
Published:
Paper Review: Rethinking Automatic Music Mixing as Sequential Stem Blending
AffectDF: The Most Comprehensive Benchmark for Speech Deepfake Detection against Emotionally Expressive Attacks
Published:
Paper Review: AffectDF: The Most Comprehensive Benchmark for Speech Deepfake Detection against Emotionally Expressive Attacks
EG-VAE: A Unified Framework for Electric Guitar Tone Transfer and Removal
Published:
Paper Review: EG-VAE: A Unified Framework for Electric Guitar Tone Transfer and Removal
Beyond Residual Connections: Manifold-Constrained Hyper-Connections for Robust Speaker Representation Learning
Published:
我现在已经掌握了所需的所有事实。主要发现:
Vorch-Streamer: Extending Human Audio-Visual Generation to Real-Time Long-Form Streaming
Published:
我已经收集到了所需的所有研究锚点。现在我将输出最终的结构化审查报告。
LILAC: An Idempotent Neural Speech Codec
Published:
我现在已经掌握了全面的研究基础。论文的重要事实锚点如下:
How to Recognize New Words: A Comparison Between Context Biasing Methods and Speech LLMs
Published:
我已经获取了所有研究数据。现在开始撰写完整的评审意见。
KVAE: Family of Tokenizers for Multimodal Generative Models
Published:
我已阅读完整论文(1810 行,涵盖所有章节)并收集了研究锚点。现在输出结构化评审。
Vorch-Omni: Multi-Task Orchestration of Sight and Sound
Published:
我已经从论文全文(35页)和所有 wiki/搜索查询中获得了足够的事实依据。现在输出完整的评审。
Whence the Voice? Self-supervised Dual-source Audio-Visual Localisation via Selective Convergence
Published:
论文评审:Whence the Voice? Self-supervised Dual-source Audio-Visual Localisation via Selective Convergence
Enhancing Social Intelligence in LLMs with Hierarchical Reasoning and Utterance-Level Goal Rewarding
Published:
Paper Review: Enhancing Social Intelligence in LLMs with Hierarchical Reasoning and Utterance-Level Goal Rewarding
Explicit and Stable Pseudospectral Time-Domain Method for the Föppl-von Kármán Equations
Published:
Paper Review: Explicit and Stable Pseudospectral Time-Domain Method for the Föppl–von Kármán Equations
Decolonizing Linguistic Policies in Automated Speech Recognition: A Framework for Cross-Culturally Competent Speech AI
Published:
我已经获取了所有需要的信息。现在我将输出完整的评审报告。
Audio-to-Score Transcription using Pre-trained Features, Data Augmentation, and the New SheetSage-A2S Dataset
Published:
Paper Review: Audio-to-Score Transcription using Pre-trained Features, Data Augmentation, and the New SheetSage-A2S Dataset
Numerical Model of a Multiple-Input-Multiple-Output Distributed Acoustic Sensing System with Joint Phase and Birefringence Estimation
Published:
Paper Review: Numerical Model of a Multiple-Input-Multiple-Output Distributed Acoustic Sensing System with Joint Phase and Birefringence Estimation
Do Audio Language Models Use Paralinguistic Evidence? Counterfactual Audits for Response Evaluation
Published:
我现在已经建立了全面的事实基础。让我整理一下在阅读完整 32 页 PDF 后的主要发现和先前的工作:
Comparative Analysis of Spread-Spectrum Codes for Fibre-Optic Distributed Acoustic Sensing
Published:
我已经获取了完整的论文文本并完成了所有研究查询。现在开始输出最终的结构化审稿意见。
MMAG: A Multi-Control Mixed Audio Generation Benchmark
Published:
我现在已经有了全面的研究发现。用于审查的关键事实:
AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward
Published:
Paper Review: AVCap: Reinforcing Audio-Video Joint Caption with Detail-Aware Reward
How Much AI Is in This Track? Quantifying the Proportion of AI-Generated Stems in Hybrid Music Mixtures
Published:
重要发现:arXiv:2607.26874 (Rigaud et al., 2026-07-29) ——《Detection of AI-generated stems within hybrid human-AI music》——这是一篇并发工作(于 2026 年 7 月 29 日发布,而本论文发表于 2026 年 8 月 7 日——仅相隔 9 天)。根据该技能的 2 个月并发工作规则,这不应被视为新颖性扣分。
Assessing AI-generated music detection in real-world broadcast monitoring
Published:
已完成全部研究阶段(Step 0-3):读完全文731行 PDF 文本、执行了 6 次 wiki_query(均无记录)、3 次 paper-wiki 查询(均无结果)、3 次 free-search(找到前序工作 Afchar et al. ICASSP 2025、Rahman et al. ICLR 2025、López-Ayala et al. ICASSP 2026,以及同期工作 MusicDET、Improved Robustness、Finding the noise)。现在输出最终 review。
Beyond Call and Response: Modelling Reciprocal Coordination in Human-AI Vocal Ensembles
Published:
我现在已经拥有了完整论文文本和扎实的事实锚点。关键发现如下:
Cloud-Boosted Low-Compute Multi-Channel Speech Enhancement
Published:
Paper Review: Cloud-Boosted Low-Compute Multi-Channel Speech Enhancement
SemBridge: Semantic Token Anchoring for Continuous-Latent Autoregressive Speech Generation
Published:
好的,我已经掌握了充分的背景信息。现在让我输出最终评审。
SCoPE: Training-Free Audio-Visual Event Perception via Sparse Cross-Modal Prior Exchange
Published:
Paper Review: SCoPE: Training-Free Audio-Visual Event Perception via Sparse Cross-Modal Prior Exchange
VIOLET: High-Fidelity Violin Synthesis with Techniques and Dynamics
Published:
所有研究查询已完成。OMC Wiki 对所有查询均未返回结果。paper-wiki 中虽然有语音合成相关的论文,但没有关于小提琴/乐器合成的记录。free-search 找到了关键基准(ViolinDiff, ICASSP 2025)及相关工作。我已经阅读了完整的 9 页论文。现在我将输出评审意见。
The Voiceprint Fallacy: Why Voices Are Not Unique Biometric Imprints
Published:
我已经阅读了这篇论文的全文(第1-8节,包括所有小节、结论和建议),并从独立搜索中获得了相关事实依据。现在我将生成结构化的评审意见。
DialectS2S: End-to-End Speech Dialogue Modeling for Low-Resource Chinese Dialects
Published:
Paper Review: DialectS2S: End-to-End Speech Dialogue Modeling for Low-Resource Chinese Dialects
SAMOT: State-Aware Step Modulation and Optimal Transport Matching for Audio-Visual Instance Segmentation
Published:
我已经获取了全文并建立了事实锚点。现在让我整理关键发现并输出评审意见。
SraVaani 1.0: Scaling Inclusive Speech Recognition for Indic Languages
Published:
我已掌握所有必要的研究锚点。关键发现如下:
ReLMCodec: Designing Predictable Speech Tokens from Pre-Quantization Phoneme Structure
Published:
Paper Review: ReLMCodec: Designing Predictable Speech Tokens from Pre-Quantization Phoneme Structure
Dramarrator: Object-Based Audio Editing for Audio Drama Production from Books
Published:
我已经获取了所需的所有信息。论文已全文阅读(PDF 转 txt,2824 行),并完成了研究查询。现在我来撰写完整的评审。
CtrlSpeech: Coarse-to-Fine Control for Expressive Speech Synthesis
Published:
我已获取完整论文文本并建立了事实锚点。以下是最终评审。
BAMU: Bitstream-Aware Marginal-Utility Allocation for Frozen Pretrained Neural Speech Codecs
Published:
Paper Review: BAMU: Bitstream-Aware Marginal-Utility Allocation for Frozen Pretrained Neural Speech Codecs
From Speech to Interaction: Analyzing Multimodal Systems in Cocktail-Party Scenarios
Published:
Paper Review: From Speech to Interaction: Analyzing Multimodal Systems in Cocktail-Party Scenarios
VoxZip: Semantic-Anchored Temporal KV Cache Compression for Long-Context Audio Inference
Published:
我现在已经获取了完整的论文文本(1744行)。我已经掌握了所需的所有信息。让我来整理最终的评审报告。
CuteTTS: Efficient and High-Quality Speech Synthesis via Autoregressive Modeling of Continuous Latents
Published:
Paper Review: CuteTTS: Efficient and High-Quality Speech Synthesis via Autoregressive Modeling of Continuous Latents
A Unifying Perspective on Audio Generative Modeling: Latent Representations and Modeling Strategies
Published:
我已经有了完整全文和充分的事实锚定。现在让我生成最终的评审。
Steering dense music retrieval with open-vocabulary concept discovery
Published:
Paper Review: Steering Dense Music Retrieval with Open-Vocabulary Concept Discovery
Multilingual Emotion Neurons in Large Audio-Language Models
Published:
我已经阅读了完整的 19 页论文(正文 10 页 + 附录 9 页)。我从 HTML 提取和 PDF 阅读中获得了全面的内容。主要发现如下:
Beyond Reconstruction: Full-Context Generative DiT for Music Generation
Published:
Paper Review: Beyond Reconstruction: Full-Context Generative DiT for Music Generation
Deferred Audio Pruning with Local Audio-Visual Dynamics for Omni-LLMs
Published:
Paper Review: Deferred Audio Pruning with Local Audio–Visual Dynamics for Omni-LLMs
Visual Token Codec: Unleashing Spatial Redundancy for ViT Feature Coding
Published:
Paper Review: Visual Token Codec: Unleashing Spatial Redundancy for ViT Feature Coding
AeroReformer2: Spoken-Query Referring Segmentation for Aerial Images
Published:
自由搜索证实,目前没有关于遥感图像口语查询引用分割(spoken-query referring segmentation for remote-sensing imagery)的前序研究。已找到关键的 RRSIS 基线(CroBIM, LSCF, FIANet, SBANet, STDNet),它们与论文中作为对比基线使用的方法一致。现在让我来撰写最终评审。
Towards an LLM-based method for quantifying the sexual content in song lyrics
Published:
Paper Review: Towards an LLM-based method for quantifying the sexual content in song lyrics
Physics-Informed Learning for Robust Acoustic Localization with Calibrated Uncertainty
Published:
我现在已经获取了完整的论文文本。我已经掌握了进行全面审查所需的所有信息。让我来汇总一下调查结果:
Beyond Piano: Cross-Instrument MIDI Velocity Estimation via Differentiable SoundFont Proxies
Published:
Paper Review: Beyond Piano: Cross-Instrument MIDI Velocity Estimation via Differentiable SoundFont Proxies
AI-Guided Learning: Research on Knowledge and Skill Acquisition Support Methods Using Deep Learning Audio-Video Processing Techniques
Published:
我已经完成了全文阅读(共119页)及所有研究查询。以下是结构化的审查报告。
Speaker Role and Language Diarization for Analyzing Multilingual Interviews for Language Proficiency of Older Adults
Published:
Paper Review: Speaker Role and Language Diarization for Analyzing Multilingual Interviews for Language Proficiency of Older Adults
MusicLayout: Explicit Structural Planning for Controllable Text-to-Music Generation
Published:
Paper Review: MusicLayout: Explicit Structural Planning for Controllable Text-to-Music Generation
A Multi-Scale Temporal Framework with Dynamic Fusion for EEG-Based Emotion Recognition
Published:
我已经获取了全文并完成了研究。以下是最终的结构化审稿意见。
CodecArena: Codec Quality Assessment via Visual Reinforcement Learning
Published:
我已从论文全文及外部搜索中获取了充足的背景信息。以下是从全文中提取的关键事实:
From Inaudible Inputs to Model Failures: Low-Frequency Safety Risks in LALMs
Published:
我已获取论文全文(包括所有附录),并进行了所有必要的研究查询。以下是最终的结构化审查。
EmoS: A Theory-Grounded Framework for Evaluating and Aligning Emotional Intelligence in Spoken Language Models
Published:
我已从全文和外部搜索中获得了足够的事实依据。现在输出最终的评审。
DAVE: A Decoupled Audio-Visual Enhancement Framework for Real-World Speech Separation
Published:
我现在已经掌握了所需的所有事实依据。以下是基于全文文本(722 行)的完整结构化审查报告。
RAG-Audio: Retrieval-Augmented Generation for Faithful Brain-to-Audio Reconstruction
Published:
我已获取完整论文文本(约 59KB HTML,涵盖所有章节)并完成了所有研究查询。以下是我收集到的信息汇总:
Neural Array-Generic Direction-of-Arrival Estimation Exploiting Array Transfer Functions
Published:
我已经从全文、wiki 查询(均为空)、paper-wiki 查询(极少)以及 free-search 中获得了足够的依据。现在开始输出评审。
Listen, See and Track: Spatio-Temporal Audio-Visual Sound Event Reasoning for Omni-Modal Language Models
Published:
Paper Review: Listen, See and Track: Spatio-Temporal Audio-Visual Sound Event Reasoning for Omni-Modal Language Models
Dynamic Clustering for Cross-Segment Permutation Alignment in Long Speech Separation
Published:
我已获取完整论文全文并建立了全面的文献锚点。以下是完整的结构化审查。
Agreement-Based Audio-Visual Segmentation:Champion Report for the MeViS-Audio Track in the 8th LSVOS Challenge
Published:
我已经获取了完整的论文文本和研究数据。现在输出最终的评审意见。
Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework
Published:
Paper Review: Towards Expressive and Faithful Audio-to-Image Generation: A Unified Multimodal Dataset and Synthesis Framework
AudioMap: Cloze-and-Choice Reinforcement Learning for Time-Aware Dense Audio Captioning
Published:
我现在已经掌握了全面的研究背景。让我整理一下关键发现并输出最终评审。
SonicWeave: Chunk-Routed Mixture-of-Experts for Unified Audio Scene Generation
Published:
Paper Review: SonicWeave: Chunk-Routed Mixture-of-Experts for Unified Audio Scene Generation
MADBench: A Benchmark for Modality-Aware Audio Deepfake Detection
Published:
我已经获取了完整的论文内容,并收集了来自 wiki 查询和免费搜索结果的所有必要研究数据。OMC Wiki 中没有关于音频深度伪造检测的记录,并且 paper-wiki 也返回了空结果。我现在将基于完整的论文文本和搜索结果中发现的相关工作进行综合评估。
REFRAMED: Towards Realistic Audio Description Generation for Movies
Published:
Paper Review: REFRAMED: Towards Realistic Audio Description Generation for Movies
Structured Phonological Representations for Audio-Articulatory rtMRI Speech Classification
Published:
Paper Review: Structured Phonological Representations for Audio-Articulatory rtMRI Speech Classification
Comparing British and American Audio Description of Movies
Published:
我已获取全文及外部研究背景,现在开始输出最终评审。
Beyond Naturalness: Probing Automated Text-To-Speech Evaluators on Linguistically Grounded Dimensions
Published:
论文评审:《Beyond Naturalness: Probing Automated Text-To-Speech Evaluators on Linguistically Grounded Dimensions》
VoxSumm: A Multilingual Corpus of Long-Form Spoken News for Joint Summarization and Translation
Published:
现在我已经掌握了所有必要的事实锚点。让我来撰写评审。
MazzikaAI: A knowledge-based performance-to-prompt compiler for real-time Arabic maqam accompaniment with a streaming text-to-music model
Published:
Paper Review: MazzikaAI: A knowledge-based performance-to-prompt compiler for real-time Arabic maqam accompaniment with a streaming text-to-music model
Never Stop Speaking: a Denial-of-Service Attack on End-to-End Speech Language Models
Published:
Paper Review: Never Stop Speaking: a Denial-of-Service Attack on End-to-End Speech Language Models
Beyond Dry References: Learning Relative Audio Effects Representations via Contrastive Distance Learning
Published:
Paper Review: Beyond Dry References: Learning Relative Audio Effects Representations via Contrastive Distance Learning
ASR-Roundtrip Evaluation Can Mask Context- and Convention-Dependent Reading Errors in Chinese News TTS
Published:
论文审稿:ASR-Roundtrip 评估可能掩盖中文新闻 TTS 中依赖上下文和惯例的朗读错误
DINO-A: Adapting Self-Distillation Vision Transformers to General Audio Representation Learning
Published:
我已掌握所需的所有信息。以下是基于全文阅读及知识库锚点整理的完整综述。
Seeds Before Objectives: Rethinking Evaluation for Low-Resource Garhwali ASR
Published:
I now have all the research anchors. Let me produce the final review.
DuplexWorld: Can voice agents help you get through the day?
Published:
我已经掌握了所有需要的信息。现在我将输出最终的结构化审查报告。
Whisper-Aware LLM: Self-Supervised Uncertainty Learning for Robust Whispered Speech Recognition
Published:
Paper Review: Whisper-Aware LLM: Self-Supervised Uncertainty Learning for Robust Whispered Speech Recognition
The GENEA Challenge 2026: A Large-Scale Disentangled Evaluation of Speech-Driven Gesture Generation on the Seamless Interaction Dataset
Published:
我现在已经掌握了完整信息。OMC Wiki 对所有手势生成查询均未返回结果。paper-wiki 也没有返回任何结果。所有知识基准均来自通过 free-search 进行的外部搜索。我已经阅读了 arXiv HTML 的完整文本。现在我将输出最终的审稿报告。
X2-Turn: Frame-Synchronous Dual-Head Modeling for Joint Streaming ASR and Turn State Prediction
Published:
论文评审:X2-Turn:用于联合流式 ASR 和轮次状态预测的帧同步双头建模
A Dataset and Benchmark for Optical Music Recognition of String Quartet Scores
Published:
我已经掌握了足够的研究背景。我已完整阅读了这篇 8 页的论文,并收集了以下事实依据:
Pitch Contour Tokenization using VQ-VAE and Its Application on Korean Traditional Music Analysis
Published:
我现在已经掌握了全面的事实依据。以下是我收集到的信息:
Measuring Cross-Cultural Style Diffusion Through Era Classification: US and Korean Popular Music
Published:
Paper Review: Measuring Cross-Cultural Style Diffusion Through Era Classification: US and Korean Popular Music
MAJEPPA: Morphing and Assessing in a Unified Piano Performance Space
Published:
Paper Review: MAJEPPA: Morphing and Assessing in a Unified Piano Performance Space
Dialogue-Aware Video-to-Music Generation Using Public Domain Film Collections
Published:
我已阅读完整论文并收集了所有研究数据。现在我将输出最终评审。
CookVoice: Unified Framework for Style Controllable Multi-Modal Human Voice Generation
Published:
以下是审稿结果。我已经阅读了完整的 PDF(共 17 页),通过 paper-wiki 和 OMC wiki 进行了查询,并进行了自由搜索,以确定关键的先前研究(Vevo2, UniVoice, Make-A-Voice, F5-TTS, NaturalSpeech 2, CosyVoice 2)。
Luna-TTS Family Technical Report
Published:
我现在已经掌握了足够的上下文信息。我已经阅读了完整的论文(共 1897 行,21 页),并且获得了外部搜索结果。现在让我来汇总最终的审稿意见。
Deep Learning Based Relative Transfer Matrix Estimation for Multiple Sources and Multiple Microphones
Published:
我已获取论文全文并完成了所有 wiki/搜索查询。现在我将输出最终的结构化评审。
Easper: An Accessible ASR Pipeline for Language Documentation
Published:
Paper Review: Easper: An Accessible ASR Pipeline for Language Documentation
Confucius4-TTS: Transcript-Free Cross-Lingual Zero-Shot TTS with a Learnable Speaker Encoder
Published:
我现在已经掌握了所有必要的事实依据。我已阅读完整论文,并收集了关于相关工作的信息。接下来我将撰写最终评审。
Phoenix TTS: High-Fidelity Synthesis and Voice Conversion via Flow-Matching-Driven Speech Tokenization
Published:
我已经完成了充分的调研。让我汇总最终的结构化审稿。关键发现如下:
UniSwap: Streaming Audio-Visual Identity Swapping for Talking Videos
Published:
我已经阅读了完整的论文,并从wiki、paper-wiki 和 free-search 中收集了证据。现在我来整理最终的评审。
MuseCritic: Learning Multi-Aspect Song Rewards through Natural-Language Aesthetic Critiques
Published:
我已经获得了所需的所有事实锚点。现在开始撰写完整的评审意见。
MiDashengLM-Gen: Unified Audio Scene Generation via LLM-Driven Autoregressive Flow Matching
Published:
这是一个关键发现。我现在已经掌握了全面的信息。让我来综合一下评审意见。
ResPCC: A Loss-Resilient Neural Point Cloud Codec over Lossy Networks
Published:
Paper Review: ResPCC: A Loss-Resilient Neural Point Cloud Codec over Lossy Networks
On-Policy Self-Distillation for Multi-Dialect ASR: Mastering Dialects, Retaining Mandarin
Published:
我已获取所有所需的研究数据。以下是我收集到的关键事实:
Do Text-to-Music Models Really Follow Instructions? A Counterfactual Evaluation of Key and Beat Grouping
Published:
我已获取全文及研究背景。以下是我的结构化综述。
HarmoniDPO: Video-guided Audio Generation via Preference-Optimized Diffusion
Published:
现在我已经掌握了所有必要的事实依据。让我来汇总这份评审。
The SLT 2026 SmartGlasses Challenge: Benchmarking Egocentric Multi-Talker Speech Recognition and Understanding with Audio-Language Models
Published:
我已经建立了足够的证据基础。现在让我来撰写完整的评论。
Rethinking Language Model-Based Generative Speech Enhancement in the Latent Space of a Neural Audio Codec
Published:
我现在已经有了全面的事实依据。关键发现如下:
RT-SEMamba: Real-Time Speech Enhancement Mamba via Progressive Knowledge Distillation
Published:
我现在已经有了充分的研究背景。让我整理一下关键发现并撰写评审。
Alignment Drift in Single-Model Speculative Decoding for ASR: Mechanism, Correction, and Cost
Published:
我已掌握了足够的研究背景。现在让我来整理最终的结构化审稿报告。我已全文阅读了这篇 23 页的论文(包括正文和补充材料),并收集了以下基准信息:
HybridSB-MoE: Dual-Domain Schrödinger Bridges with Scene-Adaptive Expert Routing for Speech Enhancement
Published:
我已经获取了完整的 21 页论文文本、所有 wiki/paper-wiki 查询结果以及 free-search 的结果。现在我将汇总生成最终的结构化审稿意见。
FastThaiG2P: Lightning-fast Thai Grapheme-to-phoneme Conversion for Voice Agent Pipelines
Published:
我已获取完整论文、外部搜索结果以及知识库查询状态。现在我将生成结构化的审稿意见。
VoxAudio: Vocalized Audio Synthesis via Multi-Reward Autoregressive Flow Matching
Published:
我已经掌握了所有必要的事实。接下来我将整理最终的评审意见。
CASA: Content-Acoustic Speaking Assessment with Speech Encoder and Large Language Model
Published:
Paper Review: CASA: Content-Acoustic Speaking Assessment with Speech Encoder and Large Language Model
Reasoning for Social Audio-Visual Question Answering: Where Do We Stand?
Published:
Paper Review: Reasoning for Social Audio-Visual Question Answering: Where Do We Stand?
The MPB Corpus: A Dataset of Melody, Rhythm, Harmony, and Melody-Harmony Relationships in Brazilian Popular Music
Published:
Paper Review: The MPB Corpus: A Dataset of Melody, Rhythm, Harmony, and Melody-Harmony Relationships in Brazilian Popular Music
H2H Music Improv: A Communication Model and Audio-Visual Dataset for Music Improvisation
Published:
Paper Review: H2H Music Improv: A Communication Model and Audio-Visual Dataset for Music Improvisation
S2Dialog: Multimodal Dialogue Retrieval with Semantic and Acoustic-Style Modeling
Published:
Paper Review: S2Dialog: Multimodal Dialogue Retrieval with Semantic and Acoustic-Style Modeling
Ambisonics Encoding of Room Impulse Responses using a Device-Agnostic Diffusion Model
Published:
Paper Review: Ambisonics Encoding of Room Impulse Responses using a Device-Agnostic Diffusion Model
AT-ADD: All-Type Audio Deepfake Detection Challenge Summary
Published:
Paper Review: AT-ADD: All-Type Audio Deepfake Detection Challenge Summary
Acoustic UAV Detection in Battlefield Scenarios: Handling Noise, Domain Shift, and Weak Labels
Published:
Paper Review: Acoustic UAV Detection in Battlefield Scenarios: Handling Noise, Domain Shift, and Weak Labels
Singer-Informed Vocal Source Separation for Multi-Singer Music Mixtures
Published:
Paper Review: Singer-Informed Vocal Source Separation for Multi-Singer Music Mixtures
Separate First, Then Associate: A Two-Stage Approach for Real-World Audio-Visual Speech Enhancement
Published:
Paper Review: Separate First, Then Associate: A Two-Stage Approach for Real-World Audio-Visual Speech Enhancement
Beyond the pale: Assessing prevalence and contents of extremist speech in LLM training data
Published:
Paper Review: Beyond the pale: Assessing prevalence and contents of extremist speech in LLM training data
What Makes a Good Layer? Assessing the Layer-Wise Intrinsic Properties of Music Foundation Models
Published:
Paper Review: What Makes a Good Layer? Assessing the Layer-Wise Intrinsic Properties of Music Foundation Models
A Parameter-Free Few-Shot Evaluation for Elephant Vocalisation Classification
Published:
Paper Review: A Parameter-Free Few-Shot Evaluation for Elephant Vocalisation Classification
Distinguishing AI-Generated Music from Edited Audio as a Hard-Negative Robustness Task
Published:
Paper Review: Distinguishing AI-Generated Music from Edited Audio as a Hard-Negative Robustness Task
Prototype-Rectified Iterative Self-supervised Manifold Denoising under Severe Acoustic Shift
Published:
Paper Review: Prototype-Rectified Iterative Self-supervised Manifold Denoising under Severe Acoustic Shift
Zipf’s Law of Abbreviation in a Logographic Script: Coding-Theoretic Bounds on Chinese Character Stroke Counts
Published:
Paper Review: Zipf’s Law of Abbreviation in a Logographic Script: Coding-Theoretic Bounds on Chinese Character Stroke Counts
AudioTQ: A Data-Oblivious 6-Bit CPU Audio Codec via Randomized Hadamard Rotation and Lloyd-Max Quantization
Published:
Paper Review: AudioTQ: A Data-Oblivious 6-Bit CPU Audio Codec via Randomized Hadamard Rotation and Lloyd-Max Quantization
Semantic Space of Parts of Speech
Published:
Paper Review: Semantic Space of Parts of Speech
Efficient Audio-Visual Generation via Synchrony-Aware Cross-Modal Sparse Attention
Published:
Paper Review: Efficient Audio-Visual Generation via Synchrony-Aware Cross-Modal Sparse Attention
ARENA: Automated Red-Teaming for Large Audio Language Models
Published:
Paper Review: ARENA: Automated Red-Teaming for Large Audio Language Models
VARM-Bench: Benchmarking Verifiable Structured Reasoning in Chinese Abusive Speech Moderation
Published:
Paper Review: VARM-Bench: Benchmarking Verifiable Structured Reasoning in Chinese Abusive Speech Moderation
Adding Voice Cloning to Text-to-Audio-Video Models with a Single Zero-Initialised Layer
Published:
Paper Review: Adding Voice Cloning to Text-to-Audio-Video Models with a Single Zero-Initialised Layer
CineDub: Scaling End-to-End Video Dubbing to Multi-Speaker Dialogues with Coherent Sound Effects
Published:
Paper Review: CineDub: Scaling End-to-End Video Dubbing to Multi-Speaker Dialogues with Coherent Sound Effects
Using the Mimi codec for metalinguistic representations
Published:
Paper Review: Using the Mimi codec for metalinguistic representations
FlowDance: Music-Driven Dance Video Generation with Parallel Pose and RGB Streams
Published:
Paper Review: FlowDance: Music-Driven Dance Video Generation with Parallel Pose and RGB Streams
Iterative Self-Learning for Expressive Text-to-Speech Synthesis
Published:
Paper Review: Iterative Self-Learning for Expressive Text-to-Speech Synthesis
The Null Token Knows: Reducing Message-Free Hallucination in ASR and NMT
Published:
Paper Review: The Null Token Knows: Reducing Message-Free Hallucination in ASR and NMT
Cached LLM Probability Retrieval for Speech Recognition
Published:
Paper Review: Cached LLM Probability Retrieval for Speech Recognition
Feedforward Active Speech Suppression Based on Time Series Prediction of Speech Signals Using Neural Networks
Published:
Paper Review: Feedforward Active Speech Suppression Based on Time Series Prediction of Speech Signals Using Neural Networks
Navigating Speech Enhancement for Real-Time MRI: A Systematic Assessment of Signal Quality, Source Preservation, and Downstream Tasks
Published:
Paper Review: Navigating Speech Enhancement for Real-Time MRI: A Systematic Assessment of Signal Quality, Source Preservation, and Downstream Tasks
AnyTalk: Speech Animation for Arbitrary Characters Leveraging a Video Generation Model
Published:
Paper Review: AnyTalk: Speech Animation for Arbitrary Characters Leveraging a Video Generation Model
ACE-Cap: Active Evidence Acquisition via Agentic Co-Evolution for Long-Paragraph Fine-Grained Audio Captioning
Published:
Paper Review: ACE-Cap: Active Evidence Acquisition via Agentic Co-Evolution for Long-Paragraph Fine-Grained Audio Captioning
INSPIRE: A Benchmark for Instruction-Aware Speech Retrieval
Published:
Paper Review: INSPIRE: A Benchmark for Instruction-Aware Speech Retrieval
SingDance: Compositional Zero-Shot Singing-and-Dancing Video Generation with Role-Aware Audio Conditioning
Published:
Paper Review: SingDance: Compositional Zero-Shot Singing-and-Dancing Video Generation with Role-Aware Audio Conditioning
Speaker-Normalized Semantic Speech Tokens via Iterative S2U-T2U Refinement
Published:
Paper Review: Speaker-Normalized Semantic Speech Tokens via Iterative S2U-T2U Refinement
Geometry-adaptive Ambisonic encoding for sparse microphone arrays of variable topology using physics-informed diffusion
Published:
Paper Review: Geometry-adaptive Ambisonic encoding for sparse microphone arrays of variable topology using physics-informed diffusion
Audio-Visual Segmentation via Depth-Guided Collaborative Modeling
Published:
Paper Review: Audio-Visual Segmentation via Depth-Guided Collaborative Modeling
Contrastive Learning with Variational Regularization for Multi-Session EEG-to-Speech Decoding
Published:
Paper Review: Contrastive Learning with Variational Regularization for Multi-Session EEG-to-Speech Decoding
Unadapted Multilingual ASR on a Garrusi Kurdish Evaluation Set: A Common-Reference Staged Normalization Analysis
Published:
Paper Review: Unadapted Multilingual ASR on a Garrusi Kurdish Evaluation Set: A Common-Reference Staged Normalization Analysis
Sonifying I2S Transport Signals to Detect Transmission Faults
Published:
Paper Review: Sonifying I2S Transport Signals to Detect Transmission Faults
Listen, Reason, and Segment: Aligning LALMs with Editorial Judgment for Media Chapterization
Published:
Paper Review: Listen, Reason, and Segment: Aligning LALMs with Editorial Judgment for Media Chapterization
How Fragile Is Your Watermark? Training-Free Structural Removal of Neural Audio Watermarks
Published:
Paper Review: How Fragile Is Your Watermark? Training-Free Structural Removal of Neural Audio Watermarks
Closing the Affective Loop: Multimodal Speaker-Listener Emotion-Dynamics-Aware Empathetic Social Robots
Published:
Paper Review: Closing the Affective Loop: Multimodal Speaker-Listener Emotion-Dynamics-Aware Empathetic Social Robots
Numerical and perceptual validity of synthetic Head-Related Transfer Functions at scale
Published:
Paper Review: Numerical and perceptual validity of synthetic Head-Related Transfer Functions at scale
A Data-Efficient Analytical Prior Machine Learning Framework for Sound Reduction Frequency Prediction in Helmholtz Resonators
Published:
Paper Review: A Data-Efficient Analytical Prior Machine Learning Framework for Sound Reduction Frequency Prediction in Helmholtz Resonators
Emotion Across Speech and Faces: Shared Affective Mechanisms in Multimodal Foundation Models
Published:
Paper Review: Emotion Across Speech and Faces: Shared Affective Mechanisms in Multimodal Foundation Models
A Multiplication-Free Feature Extractor for Signal Classification: Keyword Spotting Case Study
Published:
Paper Review: A Multiplication-Free Feature Extractor for Signal Classification: Keyword Spotting Case Study
Automatic Transcription of Microtonal Free-Rhythm Vocal Music: A Case Study in Iranian Classical Music
Published:
Paper Review: Automatic Transcription of Microtonal Free-Rhythm Vocal Music: A Case Study in Iranian Classical Music
Fair ASR: Re-Evaluating Black-Box Jailbreaks under Shared Target-Call Budgets
Published:
Paper Review: Fair ASR: Re-Evaluating Black-Box Jailbreaks under Shared Target-Call Budgets
DNN-Based Frequency-Dependent Estimation of Speech, Music, and Noise Power in Acoustic Mixtures for Hearing-Aid Scene Analysis
Published:
Paper Review: DNN-Based Frequency-Dependent Estimation of Speech, Music, and Noise Power in Acoustic Mixtures for Hearing-Aid Scene Analysis
FireRedTTS3: Unified Speech Generation and Editing with Semantically Enriched Speech Representations
Published:
Paper Review: FireRedTTS3: Unified Speech Generation and Editing with Semantically Enriched Speech Representations
The Last Mile of Deepfake Speech Detection: An Industry-Academia Experience Report
Published:
Paper Review: The Last Mile of Deepfake Speech Detection: An Industry-Academia Experience Report
Multi-turn Conversational AI from Text to Multimodal Interaction: Data, Models, Evaluation, and Open Challenges
Published:
Paper Review: Multi-turn Conversational AI from Text to Multimodal Interaction: Data, Models, Evaluation, and Open Challenges
On computational approaches to Pop music culture
Published:
Paper Review: On computational approaches to Pop music culture
UniVerse: Benchmarking and Enhancing LALMs on Culturally Inclusive Low-Resource Music Understanding
Published:
Paper Review: UniVerse: Benchmarking and Enhancing LALMs on Culturally Inclusive Low-Resource Music Understanding
SpeechSense: A Paralinguistic-Focused Dataset for Fine-Grained Speech Sentiment Analysis
Published:
Paper Review: SpeechSense: A Paralinguistic-Focused Dataset for Fine-Grained Speech Sentiment Analysis
Target Speaker Identification: A Low-Latency Streaming Pipeline
Published:
Paper Review: Target Speaker Identification: A Low-Latency Streaming Pipeline
Why GPT-Style Models Do Not Directly Transfer to Symbolic Music: Compression in the Wrong Coordinate System
Published:
Paper Review: Why GPT-Style Models Do Not Directly Transfer to Symbolic Music: Compression in the Wrong Coordinate System
ChiroEcho: extending automated bat vocalisation classification beyond the learned taxonomy
Published:
Paper Review: ChiroEcho: extending automated bat vocalisation classification beyond the learned taxonomy
FM Synthesizer Audio-Parameter Shared Embeddings
Published:
Paper Review: FM Synthesizer Audio-Parameter Shared Embeddings
Low-Power, Neuromorphic, Acoustic Anomaly Detection for Persistent Machine Monitoring
Published:
Paper Review: Low-Power, Neuromorphic, Acoustic Anomaly Detection for Persistent Machine Monitoring
VA-Judger: Reward Modeling from Human Preference Feedback for Joint Video-Audio Generation
Published:
Paper Review: VA-Judger: Reward Modeling from Human Preference Feedback for Joint Video-Audio Generation
X2Streaming-TTS: Causal Token-Level Text-to-Speech from Streaming Text with Speech-State Inheritance
Published:
Paper Review: X2Streaming-TTS: Causal Token-Level Text-to-Speech from Streaming Text with Speech-State Inheritance
Sounds Uncertain: Exploring the Affective Aspects of Sonification for Uncertainty Visualization
Published:
Paper Review: Sounds Uncertain: Exploring the Affective Aspects of Sonification for Uncertainty Visualization
Introducing the Privacy-HSD Trade-off: Hate Speech Detection, but not at the Cost of Privacy
Published:
Paper Review: Introducing the Privacy-HSD Trade-off: Hate Speech Detection, but not at the Cost of Privacy
Generalized Audio-Driven Synthesis of Precise Drummer Motion
Published:
Paper Review: Generalized Audio-Driven Synthesis of Precise Drummer Motion
Computational Features for Symbolic Melody Analysis
Published:
Paper Review: Computational Features for Symbolic Melody Analysis
Does Mapping Non-Maximal Probabilities to GMM Components Matter for S-JEPA Encoder Representations?
Published:
Paper Review: Does Mapping Non-Maximal Probabilities to GMM Components Matter for S-JEPA Encoder Representations?
Geometric Iterative Retrieval for Neural Audio Codec Resynthesis
Published:
Paper Review: Geometric Iterative Retrieval for Neural Audio Codec Resynthesis
Finetuning Strategies for Querying Sounds by Vocal Imitation
Published:
Paper Review: Finetuning Strategies for Querying Sounds by Vocal Imitation
MultiVerse: A Creator-Centered Approach to Steering Context-Adaptive Lyrics
Published:
Paper Review: MultiVerse: A Creator-Centered Approach to Steering Context-Adaptive Lyrics
A Speech Corpus for Mizo Automatic Speech Recognition: Whisper and SraVaani 1.0 Fine-Tuning with Morphology-Aware Evaluation
Published:
Paper Review: A Speech Corpus for Mizo Automatic Speech Recognition: Whisper and SraVaani 1.0 Fine-Tuning with Morphology-Aware Evaluation
Hear2Act: Benchmarking When Prosody Should Change What an Assistant Does
Published:
Paper Review: Hear2Act: Benchmarking When Prosody Should Change What an Assistant Does
DAVSS: Distilled Audio-Visual State Space Models
Published:
Paper Review: DAVSS: Distilled Audio-Visual State Space Models
Does Listening Matter? Backchanneling and Nodding in AI Clone
Published:
Paper Review: Does Listening Matter? Backchanneling and Nodding in AI Clone
Dancing Through Soundscapes: Designing a Low-Cost, Sound-Based Device for Sensing and Interpreting Movement and Dance
Published:
Paper Review: Dancing Through Soundscapes: Designing a Low-Cost, Sound-Based Device for Sensing and Interpreting Movement and Dance
Fourier is Frontier: Frequency-Aware Autoencoding for High-Fidelity Music Reconstruction
Published:
Paper Review: Fourier is Frontier: Frequency-Aware Autoencoding for High-Fidelity Music Reconstruction
Listening Forward: Next Patch Embedding Prediction Enables Scalable Audio Learners
Published:
Paper Review: Listening Forward: Next Patch Embedding Prediction Enables Scalable Audio Learners
Unified Music Identification for Tracks and Versions
Published:
Paper Review: Unified Music Identification for Tracks and Versions
Towards Quantifying Benchmark Optimization in ASR Models
Published:
Paper Review: Towards Quantifying Benchmark Optimization in ASR Models
Tracking the Trend in How Speech Synthesizers Deceive People
Published:
Paper Review: Tracking the Trend in How Speech Synthesizers Deceive People
Explainability by Design: Structured Kolmogorov-Arnold Networks over Probabilistic Attributes for Speech Deepfake Source Tracing
Published:
Paper Review: Explainability by Design: Structured Kolmogorov-Arnold Networks over Probabilistic Attributes for Speech Deepfake Source Tracing
$TCP_α$: Margin-Controlled Confidence estimation for reliable Music Information Retrieval
Published:
Paper Review: $TCP_α$: Margin-Controlled Confidence estimation for reliable Music Information Retrieval
A Regularized Block Diagonal RLS Algorithm for Acoustic Echo Cancellation
Published:
Paper Review: A Regularized Block Diagonal RLS Algorithm for Acoustic Echo Cancellation
Do SpeechLMs Hear Their Own Opinions? Diagnosing and Mitigating Previous-Belief Contamination in Streaming Emotion Understanding
Published:
Paper Review: Do SpeechLMs Hear Their Own Opinions? Diagnosing and Mitigating Previous-Belief Contamination in Streaming Emotion Understanding
Training DeepFilterNet with Accurate Room Acoustic Simulations Improves Single-Channel Speech Enhancement
Published:
Paper Review: Training DeepFilterNet with Accurate Room Acoustic Simulations Improves Single-Channel Speech Enhancement
AudioWorldSim: Realistic Binaural Audio Datasets For World Models
Published:
Paper Review: AudioWorldSim: Realistic Binaural Audio Datasets For World Models
μNet: Ultra-Low-Memory and Low-Complexity Speech Enhancement for Embedded Digital Signal Processors
Published:
Paper Review: μNet: Ultra-Low-Memory and Low-Complexity Speech Enhancement for Embedded Digital Signal Processors
DAMOS: Learning Distortion-Aware Speech Quality Assessment through Explicit Distortion Localization
Published:
Paper Review: DAMOS: Learning Distortion-Aware Speech Quality Assessment through Explicit Distortion Localization
SlimDiffuSE: Towards Efficient Diffusion-Based Speech Enhancement using Slimmable Networks
Published:
Paper Review: SlimDiffuSE: Towards Efficient Diffusion-Based Speech Enhancement using Slimmable Networks
TurboBias 2.0: Streaming Context-Biasing for Production-Efficient ASR Systems
Published:
Paper Review: TurboBias 2.0: Streaming Context-Biasing for Production-Efficient ASR Systems
MusPyExpress: Extending MusPy with Enhanced Expression Text Support
Published:
Paper Review: MusPyExpress: Extending MusPy with Enhanced Expression Text Support
Bulbul: A Dataset for Dialectal Arabic Speech Recognition
Published:
Paper Review: Bulbul: A Dataset for Dialectal Arabic Speech Recognition
SPAR-Hate: Auditor-Guided Multi-Perspective Role Reasoning for Bilingual Hate Speech Parsing
Published:
Paper Review: SPAR-Hate: Auditor-Guided Multi-Perspective Role Reasoning for Bilingual Hate Speech Parsing
Vibrato Matching for Modulation Control and Blending in Sound Mixtures
Published:
Paper Review: Vibrato Matching for Modulation Control and Blending in Sound Mixtures
Beyond Fresh Starts: Stateful Inference for Streaming ASR in Conversational Voice Agents
Published:
Paper Review: Beyond Fresh Starts: Stateful Inference for Streaming ASR in Conversational Voice Agents
FlowSep 2: Self-Supervised Flow Matching for Language-Queried Audio Source Separation
Published:
Paper Review: FlowSep 2: Self-Supervised Flow Matching for Language-Queried Audio Source Separation
AudioNoisePrints: Model-free audio watermarking using spatial correlation in flow matching TTS
Published:
Paper Review: AudioNoisePrints: Model-free audio watermarking using spatial correlation in flow matching TTS
Mitigating Speaker Leakage in Cascaded Multi-talker ASR with Diarization-based Transcript Correction
Published:
Paper Review: Mitigating Speaker Leakage in Cascaded Multi-talker ASR with Diarization-based Transcript Correction
MRMAD: A Multi-Round Multi-Audio Benchmark for Evaluating Acoustic Degradation Perception in Large Audio-Language Models
Published:
Paper Review: MRMAD: A Multi-Round Multi-Audio Benchmark for Evaluating Acoustic Degradation Perception in Large Audio-Language Models
Motion-Aware Reasoning from Speech to Mask Tracks: Runner-up Solution for the MeViS-Audio Track of the 8th LSVOS Challenge 2026
Published:
Paper Review: Motion-Aware Reasoning from Speech to Mask Tracks: Runner-up Solution for the MeViS-Audio Track of the 8th LSVOS Challenge 2026
Pre-Decoding Acoustic Triage for Budgeted Vision-Language Captioning of Untrimmed Egocentric Video
Published:
Paper Review: Pre-Decoding Acoustic Triage for Budgeted Vision-Language Captioning of Untrimmed Egocentric Video
Dual-Scale State-Space Modeling with Speaker-Wise Dynamic CRF for Speech Emotion Recognition in Conversation
Published:
Paper Review: Dual-Scale State-Space Modeling with Speaker-Wise Dynamic CRF for Speech Emotion Recognition in Conversation
Cross-Subject Generalization in Decoding Perceived Speech from Non-Invasive Brain Recordings
Published:
Paper Review: Cross-Subject Generalization in Decoding Perceived Speech from Non-Invasive Brain Recordings
WnW: Waxing-and-Waning KV Cache for Long-Form Speech LLMs
Published:
Paper Review: WnW: Waxing-and-Waning KV Cache for Long-Form Speech LLMs
DiaScriber: A Speech LLM for Joint Diarization and Transcription in Multi-Speaker Scenarios
Published:
Paper Review: DiaScriber: A Speech LLM for Joint Diarization and Transcription in Multi-Speaker Scenarios
Toward Sub-1 kB Identity-Preserving Face Compression: A Benchmark of Codecs, a Custom Learned Codec, and Studies of Resolution, Demographic Fairness, Recompression, and Adversarial Robustness
Published:
Paper Review: Toward Sub-1 kB Identity-Preserving Face Compression: A Benchmark of Codecs, a Custom Learned Codec, and Studies of Resolution, Demographic Fairness, Recompression, and Adversarial Robustness
Better Retrieval, Worse Robustness: How Multi-hop RAG Amplifies Upstream ASR Errors
Published:
Paper Review: Better Retrieval, Worse Robustness: How Multi-hop RAG Amplifies Upstream ASR Errors
Unsupervised Speech Recognition at the Syllable Level
Published:
Paper Review: Unsupervised Speech Recognition at the Syllable Level
Do Spoken Language Models Hear Speech as They Read Text? Bridging Structural Gaps Between Speech and Text
Published:
Paper Review: Do Spoken Language Models Hear Speech as They Read Text? Bridging Structural Gaps Between Speech and Text
LipsAM: Lipschitz-continuous Neural Networks for Convergent Plug-and-Play Audio Signal Recovery
Published:
Paper Review: LipsAM: Lipschitz-continuous Neural Networks for Convergent Plug-and-Play Audio Signal Recovery
Reasoning-Oriented Post-Training and Inference-Time LoRA Rescaling for Audio-Dependent Question Answering
Published:
Paper Review: Reasoning-Oriented Post-Training and Inference-Time LoRA Rescaling for Audio-Dependent Question Answering
PolyChirp: Multi-Species Birdsong Classification Using TinyML on Low-Power Acoustic Sensors
Published:
Paper Review: PolyChirp: Multi-Species Birdsong Classification Using TinyML on Low-Power Acoustic Sensors
MLLM-Assisted Audio VOS: A 3rd Place Report for the MeViS-Audio Track, 8th LSVOS Challenge
Published:
Paper Review: MLLM-Assisted Audio VOS: A 3rd Place Report for the MeViS-Audio Track, 8th LSVOS Challenge
Long-Horizon Audio-Visual Generation for Persistent Stories and Interactive Worlds
Published:
Paper Review: Long-Horizon Audio-Visual Generation for Persistent Stories and Interactive Worlds
AT-ADD: A Benchmark and Challenge for Robust and All-Type Audio Deepfake Detection
Published:
Paper Review: AT-ADD: A Benchmark and Challenge for Robust and All-Type Audio Deepfake Detection
Multi-Modal Semantic Expansion with Constrained LLM Reranking for Conversational Music Recommendation
Published:
Paper Review: Multi-Modal Semantic Expansion with Constrained LLM Reranking for Conversational Music Recommendation
EXAM$^2$: Extending Audio Understanding in Multilingual and Multimodal Analysis
Published:
Paper Review: EXAM$^2$: Extending Audio Understanding in Multilingual and Multimodal Analysis
The ISCSLP 2026 Real-World Audio-Visual Speech Enhancement Challenge
Published:
Paper Review: The ISCSLP 2026 Real-World Audio-Visual Speech Enhancement Challenge
EmoTra-TTS: Smooth Intra-Utterance Emotion Transitions for Speech Synthesis
Published:
Paper Review: EmoTra-TTS: Smooth Intra-Utterance Emotion Transitions for Speech Synthesis
Don’t Just Listen, Try Planning: Graph-based Retrieval-Generation Agent for Long-form Audio Meeting Understanding
Published:
Paper Review: Don’t Just Listen, Try Planning: Graph-based Retrieval-Generation Agent for Long-form Audio Meeting Understanding
On the Robustness of Audio Deepfake Detection under Audio Watermarking
Published:
Paper Review: On the Robustness of Audio Deepfake Detection under Audio Watermarking
FireRedAudio: A General-Purpose Audio Language Model with Decoupled Continuous Representations for Understanding and Generation
Published:
Paper Review: FireRedAudio: A General-Purpose Audio Language Model with Decoupled Continuous Representations for Understanding and Generation
‘Ghaib in Translation’ aka Unseen Harm: Measuring Cross-Script Safety Inconsistency with ‘Missed-in-Urdu’ Scores in LLM Hate Speech Detection
Published:
Paper Review: ‘Ghaib in Translation’ aka Unseen Harm: Measuring Cross-Script Safety Inconsistency with ‘Missed-in-Urdu’ Scores in LLM Hate Speech Detection
Task-disentangled Low-Rank Adaptation for Versatile Audio-visual Multi-modal Learning Tasks within a Unified Framework
Published:
Paper Review: Task-disentangled Low-Rank Adaptation for Versatile Audio-visual Multi-modal Learning Tasks within a Unified Framework
Benchmarking LLM Judges for Voice-Agent Evaluation: Reliability, Calibration, and Human Oversight
Published:
Paper Review: Benchmarking LLM Judges for Voice-Agent Evaluation: Reliability, Calibration, and Human Oversight
SeMoCo: A Semantic-First Motion Codec for Motion Language Modeling
Published:
Paper Review: SeMoCo: A Semantic-First Motion Codec for Motion Language Modeling
CoSTALA: Compositional Spatio-Temporal Audio-Language Alignment via Multi-Grain Hierarchical Contrastive Learning
Published:
Paper Review: CoSTALA: Compositional Spatio-Temporal Audio-Language Alignment via Multi-Grain Hierarchical Contrastive Learning
Array-Agnostic Ambisonics Encoding via Diffusion Posterior Sampling
Published:
Paper Review: Array-Agnostic Ambisonics Encoding via Diffusion Posterior Sampling
Visually-Guided Spatial Audio Generation for $360^\circ$ In-the-Wild Speech Scenes
Published:
Paper Review: Visually-Guided Spatial Audio Generation for 360° In-the-Wild Speech Scenes
Investigating voiced and unvoiced regions of speech for audio deepfake detection
Published:
Paper Review: Investigating voiced and unvoiced regions of speech for audio deepfake detection
REDnet: Recursive Encoder and Decoder for Speech Separation under Unknown Number of Speakers and Variable Number of Microphones
Published:
Paper Review: REDnet: Recursive Encoder and Decoder for Speech Separation under Unknown Number of Speakers and Variable Number of Microphones
TurboT2VA: Fast Large-Scale Text-to-Video-Audio Generation via Score-Regularized Consistency Distillation
Published:
Paper Review: TurboT2VA: Fast Large-Scale Text-to-Video-Audio Generation via Score-Regularized Consistency Distillation
Lost in Speech: Trilingual Spoken Hallucination Detection Across Audio and Transcripts
Published:
Paper Review: Lost in Speech: Trilingual Spoken Hallucination Detection Across Audio and Transcripts
Can We Read the Mind of an Audio LLM? A Verbalizable, Multilingual Middle-Layer Workspace
Published:
Paper Review: Can We Read the Mind of an Audio LLM? A Verbalizable, Multilingual Middle-Layer Workspace
SPECTRA: Subspace-Preserving Embedding Calibration, Transport, and Replay for Fully Few-Shot Class-Incremental Audio Classification
Published:
Paper Review: SPECTRA: Subspace-Preserving Embedding Calibration, Transport, and Replay for Fully Few-Shot Class-Incremental Audio Classification
What Do Audio-Visual Synchronization Metrics Actually Measure?
Published:
Paper Review: What Do Audio-Visual Synchronization Metrics Actually Measure?
ROMNet: a hybrid reduced order modeling and machine learning approach to waveform inversion
Published:
Paper Review: ROMNet: a hybrid reduced order modeling and machine learning approach to waveform inversion
AudioLens: Multi-Perspective Speech Clustering with Reasoning Audio-Language Models
Published:
Paper Review: AudioLens: Multi-Perspective Speech Clustering with Reasoning Audio-Language Models
LibriBrain100: One Hundred Hours of Broad and Deep MEG Data for Neural Speech Decoding at Scale
Published:
Paper Review: LibriBrain100: One Hundred Hours of Broad and Deep MEG Data for Neural Speech Decoding at Scale
TurnBench: A Multi-Domain Benchmark for Turn-Taking Dynamics in Spoken Dialogue
Published:
Paper Review: TurnBench: A Multi-Domain Benchmark for Turn-Taking Dynamics in Spoken Dialogue
AllMusicCaps: Album Reviews as Complementary Supervision for Music CLAP
Published:
Paper Review: AllMusicCaps: Album Reviews as Complementary Supervision for Music CLAP
A Training-Free Proactive Defense Against Partial Speech Manipulation via Self-Embedding Steganography
Published:
Paper Review: A Training-Free Proactive Defense Against Partial Speech Manipulation via Self-Embedding Steganography
Combining Self-Embedding Audio Watermarking with Ultra-Low-Bitrate Neural Codecs
Published:
Paper Review: Combining Self-Embedding Audio Watermarking with Ultra-Low-Bitrate Neural Codecs
Leveraging Speech Acts for Low-Data and Cross-Domain Conversation Derailment Forecasting
Published:
Paper Review: Leveraging Speech Acts for Low-Data and Cross-Domain Conversation Derailment Forecasting
Mandarin Humorous Homophone Recognition and Disambiguation in Automatic Speech Recognition
Published:
Paper Review: Mandarin Humorous Homophone Recognition and Disambiguation in Automatic Speech Recognition
CSAVocoder: A Causal Spatial Audio Vocoder Towards Real-Time Spatial Audio Generation
Published:
Paper Review: CSAVocoder: A Causal Spatial Audio Vocoder Towards Real-Time Spatial Audio Generation
Acoustic Echo Control Based on Sound Object Identification for Suppressing Howling Caused by Complicated Acoustic Paths
Published:
Paper Review: Acoustic Echo Control Based on Sound Object Identification for Suppressing Howling Caused by Complicated Acoustic Paths
PAGS: Autofocusing Photoacoustic Tomography via Speed-of-Sound-Adaptive Gaussian Splatting
Published:
Paper Review: PAGS: Autofocusing Photoacoustic Tomography via Speed-of-Sound-Adaptive Gaussian Splatting
Generative vs. Encoder Large Language Models for ASR Evaluation: A Comparative Study
Published:
Paper Review: Generative vs. Encoder Large Language Models for ASR Evaluation: A Comparative Study
Knowledge Distillation for Efficient Acoustic Echo Control
Published:
Paper Review: Knowledge Distillation for Efficient Acoustic Echo Control
Dissonance Spectrum explicitly models perceptual frequency interactions for better music understanding
Published:
Paper Review: Dissonance Spectrum explicitly models perceptual frequency interactions for better music understanding
InteractGesture: Progressive Chunk Guidance for Continuous Streaming Co-Speech Gesture Control
Published:
Paper Review: InteractGesture: Progressive Chunk Guidance for Continuous Streaming Co-Speech Gesture Control
Cooperative Multi-Agent Reinforcement Learning for Adaptive Aggregation in Semi-Supervised Federated Learning with non-IID Data
Published:
Paper Review: Cooperative Multi-Agent Reinforcement Learning for Adaptive Aggregation in Semi-Supervised Federated Learning with non-IID Data
Formal, Executable and Explainable Runtime Monitoring of Spoken Air Traffic Control Operational Procedures
Published:
Paper Review: Formal, Executable and Explainable Runtime Monitoring of Spoken Air Traffic Control Operational Procedures
Lost but not erased: Finding traces of a forgotten language in neural speech models
Published:
Paper Review: Lost but not erased: Finding traces of a forgotten language in neural speech models
VoiceMem: Streaming Dual-Brain Memory for Real-Time Interaction
Published:
Paper Review: VoiceMem: Streaming Dual-Brain Memory for Real-Time Interaction
Fine-Tuning Whisper for Automatic Speech Recognition in Baniwa: A Preliminary Study
Published:
Paper Review: Fine-Tuning Whisper for Automatic Speech Recognition in Baniwa: A Preliminary Study
papers
Daily Papers 2026-06-23
Published:
Daily Papers — 2026-06-23
Daily Papers 2026-06-24
Published:
Daily Papers — 2026-06-24
Daily Papers 2026-06-25
Published:
Daily Papers — 2026-06-25
Daily Papers 2026-06-26
Published:
Daily Papers — 2026-06-26
Daily Papers 2026-06-27
Published:
Daily Papers — 2026-06-27
Daily Papers 2026-06-28
Published:
Daily Papers — 2026-06-28
Daily Papers 2026-06-29
Published:
Daily Papers — 2026-06-29
Daily Papers 2026-06-30
Published:
Daily Papers — 2026-06-30
Daily Papers 2026-07-01
Published:
Daily Papers — 2026-07-01
Daily Papers 2026-07-02
Published:
Daily Papers — 2026-07-02
Daily Papers 2026-07-04
Published:
Daily Papers — 2026-07-04
Daily Papers 2026-07-05
Published:
Daily Papers — 2026-07-05
Daily Papers 2026-07-06
Published:
Daily Papers — 2026-07-06
Daily Papers 2026-07-07
Published:
Daily Papers — 2026-07-07
Daily Papers 2026-07-08
Published:
Daily Papers — 2026-07-08
Daily Papers 2026-07-09
Published:
Daily Papers — 2026-07-09
Daily Papers 2026-07-10
Published:
Daily Papers — 2026-07-10
Daily Papers 2026-07-11
Published:
Daily Papers — 2026-07-11
Daily Papers 2026-07-12
Published:
Daily Papers — 2026-07-12
Daily Papers 2026-07-13
Published:
Daily Papers — 2026-07-13
Daily Papers 2026-07-14
Published:
Daily Papers — 2026-07-14
Daily Papers 2026-07-15
Published:
Daily Papers — 2026-07-15
Daily Papers 2026-07-16
Published:
Daily Papers — 2026-07-16
Daily Papers 2026-07-17
Published:
Daily Papers — 2026-07-17
Daily Papers 2026-07-18
Published:
Daily Papers — 2026-07-18
Daily Papers 2026-07-19
Published:
Daily Papers — 2026-07-19
Daily Papers 2026-07-20
Published:
Daily Papers — 2026-07-20
Daily Papers 2026-07-21
Published:
Daily Papers — 2026-07-21
Daily Papers 2026-07-22
Published:
Daily Papers — 2026-07-22
Daily Papers 2026-07-23
Published:
Daily Papers — 2026-07-23
Daily Papers 2026-07-24
Published:
Daily Papers — 2026-07-24
Daily Papers 2026-07-25
Published:
Daily Papers — 2026-07-25
Daily Papers 2026-07-26
Published:
Daily Papers — 2026-07-26
Daily Papers 2026-07-27
Published:
Daily Papers — 2026-07-27
Daily Papers 2026-07-28
Published:
Daily Papers — 2026-07-28
Daily Papers 2026-07-29
Published:
Daily Papers — 2026-07-29
Daily Papers 2026-07-30
Published:
Daily Papers — 2026-07-30
Daily Papers 2026-07-31
Published:
Daily Papers — 2026-07-31
Daily Papers 2026-08-01
Published:
Daily Papers — 2026-08-01
Daily Papers 2026-08-02
Published:
Daily Papers — 2026-08-02
Daily Papers 2026-08-03
Published:
Daily Papers — 2026-08-03
Daily Papers 2026-08-04
Published:
Daily Papers — 2026-08-04
Daily Papers 2026-08-05
Published:
Daily Papers — 2026-08-05
Daily Papers 2026-08-06
Published:
Daily Papers — 2026-08-06
Daily Papers 2026-08-07
Published:
Daily Papers — 2026-08-07
Daily Papers 2026-08-08
Published:
Daily Papers — 2026-08-08
Daily Papers 2026-08-09
Published:
Daily Papers — 2026-08-09
Daily Papers 2026-08-10
Published:
Daily Papers — 2026-08-10
Daily Papers 2026-08-11
Published:
Daily Papers — 2026-08-11
Daily Papers 2026-08-12
Published:
Daily Papers — 2026-08-12
Daily Papers 2026-08-13
Published:
Daily Papers — 2026-08-13
Daily Papers 2026-08-14
Published:
Daily Papers — 2026-08-14
Daily Papers 2026-08-15
Published:
Daily Papers — 2026-08-15
Daily Papers 2026-08-16
Published:
Daily Papers — 2026-08-16
Daily Papers 2026-08-17
Published:
Daily Papers — 2026-08-17
Daily Papers 2026-08-18
Published:
Daily Papers — 2026-08-18
Daily Papers 2026-08-19
Published:
Daily Papers — 2026-08-19
Daily Papers 2026-08-20
Published:
Daily Papers — 2026-08-20
Daily Papers 2026-08-21
Published:
Daily Papers — 2026-08-21
Daily Papers 2026-08-22
Published:
Daily Papers — 2026-08-22
Daily Papers 2026-08-23
Published:
Daily Papers — 2026-08-23
Daily Papers 2026-08-24
Published:
Daily Papers — 2026-08-24
Daily Papers 2026-08-25
Published:
Daily Papers — 2026-08-25
Daily Papers 2026-08-26
Published:
Daily Papers — 2026-08-26
portfolio
Portfolio item number 2
Short description of portfolio item number 2 
