每日 arXiv

音频 · 语音 · 音乐 · 声学

每日新爬取的 arXiv 论文,按评分排序。中文摘要帮你快速判断是否值得深读。

2026-09-18
日期2026-09-18
已评分—
均分—
最高—

Daily Papers — 2026-09-18

40 papers on audio, speech, music, and acoustics.

1. Listen Before You Speak: Response Planning from Listener Facial Reactions for Conversational Speech Generation

Authors: Yunji Chu

Categories: cs.AI, cs.CL, cs.CV, cs.SD

Conversational speech depends on dialogue context and the listener’s immediately preceding behavior. We propose ReACT-TTS, a two-stage framework that uses a one-second pre-response listener facial sequence to plan the next utterance’s emotion and prosody before speech realization. On a strict dyadic MELD protocol, Temporal conditioning yields higher mean macro-F1 and VAD concordance than Text-only across ten seeds, while accuracy remains essentially unchanged. Ablations show that temporal modeling performs best among the visual variants and that an explicit early-to-late difference is unnecessary; correct listener reactions also outperform cyclic mismatches on average. In a contextual-appropriateness study with 20 speech researchers, 76% of judgments prefer Temporal, 9% Text-only, and 15% report no preference. We further connect the predicted response style to a Grad-TTS backbone for end-to-end speech realization. Overall, the results support pre-response listener dynamics as complementary cues for conversational response planning. The source code is available at https://github.com/CYJ1/ReACT-TTS_public.


2. Multi-turn Conversational AI from Text to Multimodal Interaction: Data, Models, Evaluation, and Open Challenges

Authors: Syeda Faiza Ahmed Sara, Zien Sheikh Ali, Hunzalah Hassan Bhatti, Firoj Alam, Shammur Absar Chowdhury

Categories: cs.CL, cs.AI, cs.SD

Conversational AI is moving beyond isolated text prompts toward sustained, multimodal interaction. In real conversations, users clarify goals, revise requests, interrupt responses, switch topics, and introduce new evidence while expecting systems to preserve context across turns. This makes multi-turn dialogue a distinct challenge requiring systems to maintain and update memory, ground responses across modalities, tools, and external knowledge, and adapt across languages and cultures. This study reviews multi-turn conversational AI across text-only dialogue, AudioLLMs and speech-native systems, multimodal and omni-modal systems, and tool-augmented agents. We organize the literature around datasets and benchmarks, modeling paradigms, training strategies, evaluation setups, and cross-cutting challenges. Our analysis shows that support for multiple modalities has advanced faster than the ability to sustain coherent interaction across a session. Despite stronger capabilities to perceive, speak, and act across modalities, current systems still struggle with persistent memory, cross-turn grounding, full-duplex interaction, robust evaluation, and cultural alignment. We conclude with a research agenda for systems that can remember, revise, ground, speak, listen, act, and adapt across turns, modalities, and cultures. (https://github.com/faiza-sfa/multiturn-conversational-ai-survey)


3. Omni Demand Understanding: A Benchmark for Contextual User-Intent Inference in Multimodal Interaction

Authors: Qi Chen, Yunfei Chu, Haolin He, Yifan Yang, Zihan Liu et al.

Categories: cs.CL, cs.CV, cs.MM, cs.SD

Natural audio-visual interaction is emerging as an important interface for AI assistants, allowing users to communicate through speech and vision rather than carefully composed text prompts. However, existing benchmarks of interactive capabilities still focus primarily on response quality, leaving a more fundamental question underexplored: can a model correctly infer the user’s underlying demand from complex multimodal interaction? Real-world user demands are often underspecified in speech and must be inferred from multimodal cues and dialogue history. This inference is further complicated by ambiguous or disfluent expression and noisy acoustic environments. Conversely, request-like speech may not constitute a demand to the assistant, leading to false triggers. We establish Omni Demand Understanding (ODU) as a distinct multimodal contextual inference problem: given an interaction stream, a model must detect whether a user demand is present and infer intent from multimodal and conversational context. ODU evaluates this capability along five dimensions, covering both single-turn and multi-turn interactions. We construct ODU-Bench using a challenge-driven taxonomy, taxonomy-guided agentic video generation, and human-recorded interactions, followed by media-grounded annotation and human verification. We evaluate 14 native MLLMs. Even the strongest, Gemini 3.1 Pro, recovers only 44.7% of key information that must be inferred from visual, acoustic, or conversational context. Moreover, 11 of the 14 models exhibit false-trigger rates above 50% on non-demand scenarios. These results reveal a systematic capability gap in current MLLMs’ ability to infer contextual user demands. We hope ODU can establish the evaluation of a previously underexplored yet essential capability in multimodal interaction: correctly understanding user demands before generating an appropriate response.


4. Reusing Latent Speech Representations for Query-Conditioned Topic Localization in Transcripts

Authors: Steffen Freisinger, Philipp Seeberger, Thomas Ranzenberger, Tobias Bocklet, Korbinian Riedhammer

Categories: cs.CL, eess.AS

Long transcripts are costly inputs for downstream NLP systems and often contain irrelevant context. We study query-conditioned topic localization: predicting the sentence span in a transcript that best addresses a topic-title query. To improve span localization, we reuse ASR encoder states as sentence-level representations and fuse them with textual embeddings. This lets lightweight span locators exploit speech information without running a separate audio encoder. Experiments on two public datasets show consistent gains over text-only baselines, especially under strict boundary-matching criteria. Cross-dataset experiments further indicate that the benefits are strongest for structured or semi-structured speech, while gains on spontaneous speech are limited and mixed.


5. Towards the Vision-Sound-Language-Action Paradigm: The HEAR Framework for Sound-Centric Manipulation

Authors: Chang Nie, Tianchen Deng, Guangming Wang, Zhe Liu, Hesheng Wang

Categories: cs.RO, cs.AI, cs.CV, cs.SD

While recent Vision-Language-Action (VLA) models have begun to incorporate audio, they typically treat sound as static pre-execution prompts or focus exclusively on human speech. This leaves a significant gap in real-time, sound-centric manipulation where fleeting environmental acoustics provide critical state verification during task execution. Consequently, key sounds are easily missed due to low-frequency updates or system latency. This problem is exacerbated by action chunking with open-loop execution, which creates a Blind Execution Interval where acoustic events are lost between discrete audio observation windows. Recognizing the necessity of continuous auditory awareness, we formalize Vision-Sound-Language-Action (VSLA) as a continuous control paradigm conditioned on vision, streaming audio, language, and proprioception under delayed decision loops. As an instantiation, we introduce HEAR, a VSLA framework integrating four components: (i) a streaming Historizer to maintain a compact, causal audio context across execution gaps; (ii) an Envisioner adapted from omni foundation models to reason over multi-sensory inputs; (iii) an Advancer, formulated as an audio world model, to learn temporal dynamics by predicting near-future audio codes; and (iv) a flow-matching Realizer policy to generate smooth action chunks. To address the scarcity of pretraining data and evaluations for VSLA, we construct OpenX-Sound for pretraining, alongside HEAR-Bench, the first sound-centric manipulation benchmark with strict causal timing rules. Our results suggest that robust sound-centric manipulation necessitates causal persistence and explicit temporal learning. This framework provides a practical step toward multi-sensory foundation models for embodied agents, enabling robots to perceive and interact with dynamic environments. Code and videos are available at https://hear.irmv.top


6. GenTraceBench: A Benchmark for Tracing Audio Deepfakes Across Pre- and Post-training Stages

Authors: Li Wang, Kunyu Feng, Wan Lin, Dekun Chen, Qinke Ni et al.

Categories: cs.SD

Modern text-to-speech (TTS) systems are rarely deployed as unchanged pre-trained models. They are often adapted through supervised fine-tuning (SFT) or preference optimization such as DPO and GRPO. This raises a practical question for audio deepfake forensics: do fingerprints learned from a foundation generator remain valid after adaptation? We present GenTraceBench, a controlled benchmark spanning five TTS architectures, 16 pre-/post-training variants, and 49,728 utterances generated with fixed texts and speaker prompts. Under a train-on-foundation, test-on-adapted protocol, we evaluate binary detection, closed-set attribution, and open-set verification. DPO and GRPO generally preserve fingerprints, whereas some SFT and pre-training-data changes cause substantial drift; effect sizes vary across three forensic backbones. Repeated training runs confirm the largest W2V-BERT attribution drop, while a data-mixture control with comparable speech quality shows that composition change need not cause drift. In W2V-BERT verification, multi-shot enrollment reduces EER for the SFT condition from 44.4% to 11.0%, whereas the SingNet-only condition remains at or above 45% EER.


7. Investigating the Performance and Energy Costs of Replicating Band-Split RNN for Music Source Separation

Authors: Paul Magron, Romain Serizel, Constance Douwes

Categories: cs.SD

Band-split recurrent neural network (BSRNN) is a popular music source separation model that yields close to state-of-the-art results using reasonable computational resources and public datasets. It is therefore interesting from a reproducible research perspective, but achieving its performance is not straightforward since its full code is not available. In this paper, we conduct a replication of BSRNN via implementing the full pipeline. We extend the original paper’s analysis by experimentally studying various design choices about data preprocessing, the optimization protocol, and architectural parameters. We report and discuss this project’s energy cost, and we underline how its footprint could have been substantial lower upon availability of the full pipeline, which advocates for more reproducible research practices. To comply with this objective, we publicly release our code and pre-trained models.


8. Online Algorithms for Independent Low-Rank Matrix Analysis and Rank-Constrained Spatial Covariance Matrix Estimation Based on Maximum Weighted Likelihood Estimation

Authors: Yuto Ishikawa, Norihiro Takamune, Tomohiko Nakamura, Daichi Kitamura, Hiroshi Saruwatari et al.

Categories: cs.SD

Real-time multichannel speech extraction (MSE) under diffuse noise conditions is an important task with a wide range of applications, such as speech recognition and hearing aids. In this paper, we propose online algorithms for independent low-rank matrix analysis (ILRMA) and rank-constrained spatial covariance matrix estimation (RCSCME). Previously, we proposed a real-time extension of the RCSCME-based method: an MSE method based on ILRMA and RCSCME using the blockwise batch algorithm. However, it assumes that the spatial characteristics are stationary within a single batch, and thus, in dynamic situations where the target speaker moves, its performance may degrade. To address this problem, we derive the online algorithms for ILRMA and RCSCME in the following three steps. First, we formulate framewise cost functions for ILRMA and RCSCME on the basis of maximum weighted likelihood estimation. Second, we derive the update rules for the framewise cost functions on the basis of auxiliary-function techniques. These naive update rules are computationally costly for real-time execution on a practical machine. Thus, we finally derive the online algorithms by approximating some intermediate parameters with their estimates. Furthermore, we propose stabilization and further acceleration techniques for these online algorithms. In experiments, we simulate situations where a target speaker is stationary or moves and show that the proposed method achieves superior speech extraction performance compared with conventional methods. In addition, using real-world recorded signals, we demonstrate the effectiveness of the proposed method in practical scenarios.


9. Training Music Sample Identification Models on Real Sample Pairs

Authors: R. Oguz Araz, Joan Serrà, Xavier Lizarraga-Seijas, Emilio Molina, Xavier Serra et al.

Categories: cs.SD

Sample identification (SI) is the task of matching pairs of tracks, where one track is created by musically transforming an element of the other. In the absence of sample annotations at scale, the dominant training paradigm has depended on artificially creating sample pairs. Although a recently released dataset provides annotations of real sample pairs at scale, an effective training recipe is missing. In this work, we present SI Embeddings (SIE), an SI model that achieves state-of-the-art results on three benchmarks, including a large-scale test set. We show that the previous state of the art trained on artificial pairs generalizes only partially to real pairs, and that its training data limits its performance. We also show that real pairs do not fully account for SIE’s performance: its architecture and training recipe contribute substantially. We provide the first fully supervised training recipe for real-world SI, establishing a strong foundation for future research in the field.


10. Causal limits and optimal allocation for passive frequency-selective hearing protection

Authors: Geonhwi Hwang, Jaewoo Lee, Hyunjun Kim, Hyeonwoo Na

Categories: cs.SD, cond-mat.mtrl-sci

Passive Helmholtz-resonator earplugs aim to attenuate hazardous noise while keeping speech audible; the air volume of an ear-sized shell is finite, and we treat it as the resource a design spends. For a vented bore transparent at zero frequency and loaded by reactive side branches, causality fixes a conserved integral of transmission loss. We derive it in device variables, add viscous and flanking constraints, and allocate the budget by hazard-weighted water-filling with a continuous speech-cost price. The audio-band budget B_ad = (20/ln10) pi^2 V / 2cS is set by the ratio of total cavity volume V to bore area S; occluding plugs are exempt. The f^2 price signal favours frequencies well above the source peak. Replacing strict transparency by a speech-cost price gives an exchange rate of about 1.3 dB of protection per decibel of speech cost under the present model, with the budget setting where it saturates. Across three form factors, low-frequency-selective designs stay near a decibel while hazard-matched mid/high-frequency designs reach twelve. The binding variables are bore area and speech allowance, not resonator count.


11. A Training-Free Proactive Defense Against Partial Speech Manipulation via Self-Embedding Steganography

Authors: Yigitcan Özer, Zhe Zhang, Wanying Ge, Xin Wang, Junichi Yamagishi

Categories: cs.SD, cs.AI

Partial deepfake speech, where only limited segments of an utterance are synthesized or manipulated, poses a significant challenge to existing deepfake detection systems. As the proportion of spoofed regions decreases, passive detectors become increasingly unreliable, and accurate detection and restoration remain challenging. In this paper, we revisit audio steganography from a new perspective and propose its use as a proactive defense against partially deepfaked audio. In particular, we consider a self-embedding strategy in which a clean speech signal embeds a compressed representation of itself, enabling post-hoc extraction of reference content. We demonstrate how existing audio steganography methods can be repurposed to support detection of partial deepfakes through codec-based restoration. Experiments on a benchmark dataset show that the proposed approach complements passive defenses. Remarkably, the proposed method operates without any training, providing a robust and data-efficient alternative for partial deepfake detection.


12. CPR: Combining global composing, local performing and full-sequence refining in piano rendering with continuous autoregressive modelling

Authors: Chong Jing, Junan Zhang, Zhizheng Wu

Categories: cs.SD, cs.AI

Prompt-conditioned piano MIDI-to-Music rendering aims to faithfully render target notes while reproducing the timbre of a reference recording. Existing approaches primarily follow two paradigms: autoregressive (AR) modeling and flow matching (or diffusion). Discrete-codec AR models provide causal temporal modeling, but quantization can discard acoustic detail. Flow matching better preserves acoustic structure in the cost of full-sequence attention costs and worse semantic structure. Continuous autoregressive models operate directly on continuous representations. It not only combines the condition-following ability of AR models and distribution-modeling capacity of flow matching but also bypasses the quantization bottleneck with lower computational costs. Building on this principle, we present Composer–Performer–Refiner (CPR) framework. Composer autoregressively predicts continuous hidden states, Performer generates 24kHz acoustic latents through local flow matching and Refiner then upsamples the waveform to 48 kHz. We further introduce Bottlenecked Representation Alignment (BREPA) and Modality–Time RoPE (MT-RoPE) to strengthen musical semantic structure in Composer hidden states and temporal alignments across modalities. Codes are available at https://github.com/FEAfeatherTHER/CPR_official


13. CoReLoop: Parameter-Efficient Controlled Recurrent Refinement for Audio Deepfake Detection

Authors: Kunyu Feng, Yuxiang Wang, Li Wang, Wan Lin, Zhizheng Wu

Categories: cs.SD, cs.AI

Generalizing to unseen attacks remains challenging for audio deepfake detectors, and collecting training data covering all potential attacks is impractical. We explore recurrent refinement in an already-trained SSL-based detector without additional data or changes to its original parameters. However, directly recycling encoder outputs as inputs degrades detection in our diagnostic. We propose CoReLoop, which makes this reuse effective by adapting recurrent inputs to the frozen encoder, controlling state updates, and aligning refined outputs with the frozen classifier. By training only lightweight refinement modules and loop-specific low-rank adapters on the original data, CoReLoop enables additional refinement while preserving the detector’s original first-pass prediction. On 14 cross-domain test sets, the 24-layer model reduces pooled equal error rate (EER) from 4.85% to 3.74% with two passes, with approximately 10M trainable parameters out of 598M. To selectively apply this refinement, an optional halting head chooses the depth for each utterance, achieving 3.73% pooled EER with an average of 1.18 passes.


14. SEABED: SouthEast Asian Benchmark for Evaluating Audio Reasoning

Authors: Harshit Rajgarhia, Asif Shaik, Rachuri Lokesh, Sushanta Kumar Pani, Abhishek Mukherji

Categories: cs.SD, cs.AI

Modern audio-language models are no longer judged only on what words they can transcribe, but on whether they can reason over what they hear: recovering meaning that lives in tone and prosody, telling dialects and regional languages apart, and resolving ambiguity that the written form leaves open. This capability is now measured by a growing family of audio-reasoning benchmarks, but almost entirely in English and on general-domain audio. Southeast Asia (SEA) is served instead by benchmarks that inherit an English task taxonomy of recognition, translation, and paralinguistic classification, and therefore test whether a model hears SEA speech rather than whether it can reason from it. We introduce SEABED, an audio-first question answering dataset designed to benchmark language and audio reasoning models on SEA speech. SEABED comprises a suite of six audio-reasoning tasks built entirely from real, openly available SEA speech corpora, yielding 5,404 question-answer pairs. We evaluate six frontier and region-specific audio LLMs: even the state-of-the-art model Gemini 3.5 Flash achieves only 50.3% weighted average accuracy. SEABED evaluates not only answer accuracy, but also whether models’ stated reasoning is grounded in the audio evidence. A sample of the benchmark data is available here: https://huggingface.co/datasets/CentificAIResearch/SEABED.


15. MENASpeechBank: A Reference Voice Bank with Persona-Conditioned Multi-Turn Conversations for AudioLLMs

Authors: Zien Sheikh Ali, Hunzalah Hassan Bhatti, Rabindra Nath Nandi, Shammur Absar Chowdhury, Firoj Alam

Categories: cs.SD, cs.AI, cs.CL, eess.AS

Audio large language models (AudioLLMs) enable instruction following over speech and general audio, but progress is limited by the scarcity of diverse, conversational, and instruction-aligned speech–text data. This gap is particularly pronounced for persona-grounded and dialectal interactions, where collecting real multi-speaker recordings remains costly and slow. We introduce MENASpeechBank, a reference speech bank comprising ~18K high-quality utterances from 124 speakers spanning multiple MENA countries, covering English, Modern Standard Arabic (MSA), and regional Arabic varieties. We develop a controllable data pipeline that (i) constructs persona profiles enriched with World Values Survey (WVS) inspired attributes, (ii) defines a taxonomy driven ~5Kconversational scenarios, (iii) matches personas to scenarios via semantic similarity, (iv) generates ~417K role-play conversations with an LLM where the user speaks as the persona and the assistant behaves as a helpful agent, and (v) produces speaker-conditioned user-turn audio (synthetic) from reference recordings to preserve speaker diversity. We evaluate synthetic and human recorded conversations and provide an analysis. We will make the MENASpeechBank available for the community.(\href{https://huggingface.co/datasets/QCRI/MenaSpeechBank)


16. I’ll Keep an Ear Out: Teaching AudioLLMs Proactive Audio Assistance

Authors: Amit Kumar Singh Yadav, Ritvik Shrivastava, Xuan Zhang, Seungwhan Moon, Shashank Jain et al.

Categories: cs.SD, cs.CL

Audio large language models (AudioLLMs) operate reactively, responding only when queried. We introduce proactive audio assistance, where an AudioLLM monitors an audio stream and autonomously decides when to alert the user from a single natural-language intent, motivated by wearable applications for Deaf and Hard of Hearing users. We propose Interrupt and Silent Modeling (ISM), a model-agnostic paradigm that embeds proactive decisions into LLM decoding via two special tokens: \texttt{<interrupt>} and \texttt{<silent>}, capturing four states: onset detection, sustained-relevance triggering, irrelevance suppression, and de-duplication. Applied to Qwen2-Audio-7B, ISM achieves 99.6\% interrupt F1 and perfect de-duplication recall on ESC-50. On noisy Epic-Sounds kitchen audio, ISM achieves the highest interrupt F1 without domain-specific training, the only method maintaining strong onset detection without over-triggering or over-suppression. Streaming evaluation confirms real-time viability with 3.5-second average latency.


17. Spherical Harmonic Sliced Wasserstein Displacement Interpolation for Acoustic Source and Reflection Density Modeling

Authors: Yuancheng Luo

Categories: cs.SD, eess.AS, eess.SP, math.OC, math.PR

Spatial room impulse responses (SRIRs) capture directional distributions of acoustic sound-sources and their reflections. However, collecting SRIRs of moving sound-sources remains a challenge, requiring complex interpolations across measurements that account for multi-path spatial-temporal dynamics. This paper investigates the Wasserstein metric and displacement for evaluating interpolated SRIR echo densities in the spherical harmonic domain. We present novel sum-of-magnitude square expansions for efficiently fitting probability density functions, maximizing likelihood, inverse sampling, and computing spherical sliced Wasserstein interpolations. Experiments compare the Wasserstein displacements and metric to linear and geometric interpolations of SRIR image-source densities on a line-path, and demonstrate model-order reduction.


18. CGaLore: Curvature-Guided GaLore for Memory-Efficient Continual Adaptation of ASR Foundation Models

Authors: Steven Vander Eeckt, Hugo Van hamme

Categories: eess.AS

Automatic speech recognition models suffer from catastrophic forgetting when adapted to new domains, accents, or downstream tasks. This problem becomes increasingly important with the growing use of speech foundation models, where adaptation should be both memory-efficient and safe, preserving the broad capabilities learned during pretraining. Gradient Low-Rank Projection (GaLore) has recently been proposed as a memory-efficient fine-tuning method that keeps model parameters full-rank while reducing the optimizer memory through low-rank gradient projection. However, GaLore does not account for catastrophic forgetting. We propose Curvature-Guided GaLore (CGaLore), which incorporates old-task curvature information when selecting the low-rank projection bases. Specifically, CGaLore filters current-task gradients using Kronecker-factored approximate curvature from previous tasks before computing the gradient subspace. Our experiments show that CGaLore enables effective adaptation while alleviating forgetting, outperforming state-of-the-art continual learning baselines. Extensive ablation studies confirm the practical applicability of CGaLore.


19. Confidence-Guided Markov Weighting for Semi-Supervised Tabla Stroke Transcription

Authors: Rahul Bapusaheb Kodag, Vipul Arora

Categories: eess.AS

Tabla Stroke Transcription (TST) converts tabla audio into symbolic stroke sequences, but the scarcity of annotated recordings makes fully supervised training challenging. We propose a semi-supervised framework that uses sequence-level labelled and unlabelled tabla recordings. A teacher model generates pseudo-label sequences, while a Stroke-Level Confidence Estimation Model (S-CEM) estimates confidence for each predicted stroke. To improve training with uncertain pseudo-labels, we propose Confidence-Guided Markov Weighted Alternative Temporal Classification (CMW-ATC), which weights candidate sequences within uncertain spans using learned stroke transitions and reliable neighbouring strokes. Experiments across three TST evaluation settings show consistent improvements over conventional teacher–student training and Alternative Temporal Classification in its replacement form (ATC-R). Ablation studies further show the contributions of S-CEM, Markov transition weighting, and the use of reliable strokes on both sides of uncertain spans.


20. Consensus-Guided Shared-Specific Tri-View Learning for Speech Emotion Recognition

Authors: Bing Huang, Yujian Ma, Xikun Lu, Xianquan Jiang, Jinqiu Sang

Categories: eess.AS

Speech emotion recognition (SER) benefits from heterogeneous acoustic representations, but views derived from the same utterance contain both overlapping emotional evidence and representation-dependent cues. Direct fusion may therefore propagate redundant information or obscure complementary details. To address this issue, we propose Tri-view Consensus-Guided Fusion (TriCGF) for jointly modeling spectrogram, Mel-frequency cepstral coefficients, and HuBERT representations. TriCGF organizes each view into common and view-specific components before fusion. Cross-view Consensus Learning aggregates the common components into a global reference, while View-wise Gated Integration adaptively combines this reference with each view-specific component. A soft difference regularizer further discourages excessive information overlap. Under speaker-independent evaluation, TriCGF achieves 74.19% weighted accuracy (WA) and 75.17% unweighted accuracy (UA) on IEMOCAP, and 94.36% WA and 94.28% UA on EmoDB, outperforming representative SER methods on both datasets.


21. HAMMER: Harmonic-Aware Parallel Context Modeling and Discriminator-Free Perceptual Optimization for Speech Enhancement

Authors: Shang-Fu Chen, Szu-Wei Fu, Sung-Feng Huang, Rong Chao, Wen-Huang Cheng et al.

Categories: eess.AS

Recent speech enhancement systems combine self-attention and Mamba to capture global interactions and long-range dependencies. Yet these hybrids usually operate as sequence mixers and do not explicitly exploit harmonic periodicity, a strong cue for preserving voiced speech under noise. Perceptual optimization poses another challenge. PESQ is non-differentiable, so many methods train auxiliary metric discriminators that increase complexity and introduce adversarial instability. We propose \ours, a harmonic-aware and discriminator-free speech enhancer built around two components. (i) The Time-Frequency Harmonic-aware Attention-Mamba (TF-HAM) block runs self-attention and bidirectional Mamba in parallel along both spectrogram axes, then applies a speech-adapted autocorrelation feed-forward network to encode local periodic structure. (ii) Metric-explicit perceptual refinement (MEPR) combines differentiable PESQ and log-likelihood-ratio losses to expose perceptual metric structure without a learned surrogate. On VoiceBank+DEMAND, \ours achieves 3.69 PESQ and 4.41 COVL with only 2.39\,M parameters, outperforming or matching discriminator-based systems. Inference-time perceptual contrast stretching further raises PESQ to 3.79 without retraining. The source code will be available at https://github.com/shangfuu/HAMMER.git.


22. LLMs and Speech: Integration vs. Combination

Authors: Robin Schmitt, Albert Zeyer, Mohammad Zeineldeen, Ralf Schlüter, Hermann Ney

Categories: eess.AS

In this work, we study different approaches to utilize large language models (LLMs) for automatic speech recognition (ASR). Specifically, we compare the tight integration of an acoustic model (AM) with the LLM (“speech LLM”) to the traditional way of combining AM and LLM via shallow fusion and provide ablations on the effect of different label units and LLM sizes. For tight integration, we further examine the effect of attention interfaces, encoder downsampling, and length normalization. Furthermore, we investigate joint recognition with a CTC model to mitigate hallucinations of speech LLMs and present effective optimizations. We train and evaluate on LibriSpeech and Loquacious and additionally evaluate on the HuggingFace ASR leaderboard. Across model sizes, we find that shallow fusion consistently outperforms tight integration of AM and LLM on in-domain data, highlighting the importance of strong shallow-fusion baselines when evaluating speech LLMs for ASR. On the more heterogeneous HuggingFace ASR leaderboard, however, the integrated prefix LLM achieves lower average WER than shallow fusion, with gains concentrated on out-of-domain corpora.


23. Partial Accent-Control Editing in Frozen Speech Representations for Accent Conversion

Authors: Yangyang Qu, Michele Panariello, Massimiliano Todisco, Nicholas Evans

Categories: eess.AS

Accent conversion is the task of modifying a speech recording so that it sounds closer to a target accent while preserving linguistic content and other speaker-related characteristics. Most accent conversion systems use trained, generative models. Although they can induce target-accented speech, the strength of accent modification is not controllable at inference time, making it difficult to analyse how the strength of accent conversion affects source preservation. We propose Partial Accent-Control Editing (PACE), an accent conversion framework based upon the editing of frozen WavLM representations without the training of an accent-conditioned generator. A constrained edit is first applied to source WavLM features, which are then fused with target-accent reference features retrieved from non-parallel accent examples. Fusion weights are used to control the trade-off between accent conversion strength and the degradation of other source attributes. Using a suite of five metrics, and with the cost of accent-entangled speaker similarity, we show that PACE is substantially superior to a pair of competitive baselines in terms of both accent conversion and source preservation.


24. Rethinking Music Tokenization: A Semantic Codec toward High-Fidelity LLM Music Generation

Authors: Huakang Chen, Guobin Ma, Yuepeng Jiang, Dake Guo, Jingbin Hu et al.

Categories: eess.AS

Discrete audio tokenization has become the critical interface between raw waveforms and autoregressive modeling in recent music generation. As a result, music tokenizers must simultaneously support high-fidelity reconstruction and produce discrete sequences that remain amenable to language modeling. Existing reconstruction-oriented tokenizers often mix musical structure with fine acoustic details, producing high-entropy tokens that are hard to model. In contrast, semantics-guided alternatives are designed for speech and do not fit music well, often hurting reconstruction quality. We address these trade-offs by rethinking music tokenization around a measurable notion of music semantic content grounded in downstream Music Information Retrieval tasks. Guided by this definition, we propose MuSeC, a music semantic codec that factorizes semantic and acoustic content directly from mixed signals without source separation. MuSeC preserves information required for high-fidelity reconstruction while producing more LM-friendly discrete units. Empirically, it improves reconstruction quality and yields more predictable token sequences, providing a practical foundation toward high-fidelity LLM music generation. Demos are available at https://longwaytog0.github.io/MuSeC/.


25. Towards Zero-Shot Attribution of Synthetic Speech via Audio-Text Contrastive Retrieval

Authors: Cristian-Teodor Neamtu, Serban Mihalache, Stefan Smeu, Dan Oneata, Horia Cucu et al.

Categories: eess.AS

Audio deepfake forensics is moving beyond a simple real-or-fake verdict toward attribution: which system generated the audio clip? Most source-attribution methods cast this as closed-set classification, so they cannot name a generator that was absent from training, a gap that widens with every newly released text-to-speech (TTS) system. We instead frame attribution as cross-modal retrieval: each generator is described in natural language, and a clip is attributed by retrieving the description closest to it in a shared audio-text embedding space. Adding a new system then takes nothing more than writing its description, with no retraining and no new classifier head. Our model couples a frozen Wav2Vec2-BERT audio encoder with a frozen E5 text encoder and aligns them through small trainable projection heads, using a contrastive objective that combines a cross-modal supervised-contrastive loss with intra-modal terms. We evaluate on MLAAD v9 (140 TTS models, 51 languages) under 10-fold leave-models-out cross-validation. For generators it has never encountered before, the model reaches a model-level mean reciprocal rank (MRR) of 58.4%. Even when the correct model is not identified, the audio clip is often matched to systems that share the true generator’s vocoder, acoustic model, or architecture. Because the same embedding space also answers natural-language attribute queries, one set of descriptions covers both open-set attribution and attribute-level forensic profiling.


26. Vaani Benchmark V1.0: An Inclusive Multimodal Benchmark Dataset for Hindi

Authors: Sujith Pulikodan, Agneedh Basu, Saurabh Kumar, Pranav Bhat, Pavan Kumar J et al.

Categories: eess.AS

Benchmarking is critical for the systematic evaluation of machine learning systems. While several open-source datasets are available for Hindi, existing benchmarks remain limited in terms of modality, geographic diversity, demographic representation, and transcription robustness. We introduce an inclusive, multimodal Hindi benchmark dataset collected from 102 districts across India. The dataset consists of spontaneous speech elicited using image prompts and recorded under real-world acoustic conditions across diverse demographic groups. Each audio segment is associated with the image prompt that elicited it and is annotated with three independent transcriptions, enabling multi-reference evaluation that accounts for permissible orthographic and lexical variations. We show that single-reference evaluation overstates errors, while multi-reference evaluation provides a more robust, inclusive, and realistic assessment of automatic speech recognition (ASR) systems. The pairing of each utterance with its eliciting image further enables image retrieval evaluation using both speech and text queries. We evaluate multiple multimodal embedding models for image retrieval and report their performance on the combined speech–image and text–image retrieval tasks. The results reveal a performance gap between text-based and audio-based retrieval, with text-based retrieval consistently achieving higher performance.


27. Samsone: A Family of Open Small Audio Language Models for On-Device Inference

Authors: Piotr Masztalski, Michał K. Grzeszczyk, Olaf Sikorski

Categories: eess.AS, cs.AI

The success of Large Audio Language Models has driven the development of massive multimodal networks exceeding billions of parameters. However, the demand for privacy-preserving, low-latency processing has shifted focus toward Small Audio Language Models (SALMs) capable of on-device execution. In this paper, we introduce Samsone, a family of SALMs designed for edge computing. Our core model, Samsone-134M, establishes a new state-of-the-art for its size class across multiple benchmarks. We further explore the scaling laws of SALMs by introducing Samsone-99M and Samsone-356M. Despite their compact footprint, the Samsone family delivers performance competitive with models orders of magnitude larger. To foster open research and reproducibility, we train Samsone on publicly available data. We release the training code, model weights, mobile-optimized checkpoints and provide an open-source Android application to demonstrate real-time on-device inference of Samsone.


28. Cover First, Disagree Softly: Rethinking Mismatch-First Active Learning for Frame-Level Audio Classification

Authors: Shiqi Zhang, Tuomas Virtanen

Categories: eess.AS, cs.AI, cs.SD

Sound event detection relies on frame-level strong labels whose annotation is expensive. Active learning addresses this problem by selecting the audio segments whose labels help the classifier most. One of the prevailing acquisition strategies for this task, mismatch-first farthest-traversal (MFFT), combines the disagreement between two classifiers and the diversity of the selected segments through hard sequential decisions. It selects whole groups of high-disagreement segments first and spreads only the remaining budget by farthest traversal. On two multi-label datasets we show that this design is blind to the similarity among the selected segments and fails under low budgets, with every mismatch-first variant ending below the plain geometric strategy it builds on. We propose mismatch-weighted facility location (MW-FL), which spends the entire budget through a disagreement-weighted coverage objective that penalizes similarity among the selected segments. The disagreement signal from MFFT is used to obtain the nonnegative weights of this facility-location objective, using fixed smoothing without dataset-specific tuning. Experiments across two geometric mechanisms with three ways of using disagreement show that coverage of the selected segments is the dominant factor, hard disagreement gating of selection is harmful on both mechanisms, and soft disagreement weighting helps on top of coverage. MW-FL attains the best area under the learning curve on both datasets.


29. OmniVChat: Synthesizing, Benchmarking, and Training for Native Audio-Visual Dialogue

Authors: Haolin He, Yunfei Chu, Qi Chen, Wen Huang, Yuan Feng et al.

Categories: eess.AS, cs.AI, eess.IV

We define OmniVChat (Omni Video Chat) as the task of native audio-visual dialogue between a user and an omni model. In OmniVChat, omni models directly and simultaneously receive audio and video from a user and return text. The user’s query is embedded in the audio and video, without a separate text question, external captioning, or speech recognition. Direct audio-visual input reduces external latency and computation while preserving perceptual cues. However, research on OmniVChat faces two constraints: data availability and evaluation. Recordings of people using their own devices are scarce. Furthermore, a good reply often needs to account for the user’s surroundings, facial expressions, and nearby objects, and such responses can be expressed in many different ways, making keyword matching unreliable for evaluating reply quality. Recent progress in agent systems and video generation makes generation for comprehension viable, which means using synthesized dialogues for training and evaluation. Therefore, we present OmniVChat-Studio, a multi-agent data engine for synthesizing single- and multi-turn audio-visual dialogues. We use synthesized dialogues to build OmniVChat-Bench, an evaluation benchmark that evaluates omni models’ basic dialogue abilities across five ability categories. We also present OmniVChat-RL, a reinforcement learning reward design that jointly targets reply correctness, efficiency, and style in OmniVChat. Training Qwen3-Omni-Instruct with OmniVChat-RL on synthesized dialogues improves its performance on both OmniVChat-Bench and the human-recorded OmniVChat-Bench-Human. These gains validate the reward design and show transfer to real-world dialogues in training and evaluation.


30. The Spoken Wikipedia Presentation Corpus

Authors: Thomas Ranzenberger, Steffen Freisinger, Tobias Bocklet, Korbinian Riedhammer

Categories: eess.AS, cs.CL, cs.MM, cs.SD

We present the Spoken Wikipedia Presentation Corpus, an extension of the Spoken Wikipedia Corpora featuring LLM-generated slide decks for multimodal ASR. Slides are created from LLM-segmented sections using a hybrid pipeline that combines LLM-based content planning with rule-based design decisions. For each section, an LLM generates a slide title, bullet points, a takeaway message, and a visual description that is used to create an illustration. Rule-based matching then selects layouts, themes, and styles to produce the final slides. A vision LLM extracts slide text as Markdown. We evaluate multiple ASR and spoken language models (SLMs). The best model achieves an average micro-WER of 10.23% and an average micro-CER of 6.48% on audio-only inputs. English yields the lowest error rates, followed by German and Dutch, while performance declines across lower-resource languages. Although audio-only baselines are strong, multimodal zero-shot prompting of omni models remains challenging. The aligned slide, text, and audio data show a strong potential to improve recognition through cross-modal context.


31. SG-Mamba: Sparse Graph-Guided Mamba for Audio-Visual Speech Enhancement

Authors: Guo-Ruei Tseng, Hung-Shin Lee, Hsin-Min Wang, Berlin Chen

Categories: eess.AS, cs.CL, cs.SD

Lightweight audio-visual speech enhancement (AVSE) models face a critical trade-off between computational efficiency and cross-modal alignment accuracy. While simple concatenation lacks relational expressiveness, dense cross-attention incurs computational overhead and is prone to unreliable cross-modal correspondence under strong acoustic interference. We propose Sparse Graph-Guided Mamba (SG-Mamba), a lightweight AVSE framework that integrates a sparse heterogeneous graph with a linear-complexity Mamba backbone. The graph explicitly models modality-specific relations through content-adaptive attention and cross-frame audio-visual connections, while Mamba captures long-range temporal context. We further introduce an audio skip connection to preserve spectral detail without sacrificing noise suppression. Evaluated on LRS3, SG-Mamba achieves competitive or superior performance against strong lightweight baselines and reaches 13.091 dB SI-SDR under noise-only condition. It also remains robust in cluttered multi-speaker conditions with a competitive cost of 3.45 G MACs (or 6.90 G FLOPs). Results on VoxCeleb2 further suggest that explicit structural priors improve robustness, generalizability, and computational efficiency in lightweight AVSE.


32. BLINC: Blind Calibration For Training-Free Speech Enhancement Adaptation

Authors: Tobias Raichle, Ekaterina Gavrilko, Bin Yang

Categories: eess.AS, cs.SD

Speech enhancement (SE) models degrade under domain shifts and have to adapt to unseen target domains during deployment. Most existing test-time adaptation (TTA) methods for SE do so by adapting a subset of the model weights using a self-supervised loss, which requires backpropagation at test-time and permanently alters the model. We instead recalibrate the prediction itself and propose BLINC, a training-free TTA method that remaps the predicted time-frequency mask onto a bimodal target distribution by histogram matching. At test-time, the target distribution is parameterized from blind features of the noisy recording, so neither a reference distribution from a classical algorithm nor online metric optimization is involved. BLINC improves the overall quality of both evaluated SE models on almost every target condition and matches or exceeds the loss-based TTA baselines at minimal overhead.


33. Towards Physics-Informed Neural Networks for Stiff Guitar String Vibrations

Authors: Xinmeng Luan, Kensuke Okada, Ryoya Tabata, Gary Scavone

Categories: eess.AS, physics.flu-dyn

Modeling stiff string vibrations is challenging due to their dispersive and high-frequency characteristics. This study investigates the effectiveness of Physics-Informed Neural Networks (PINNs) in simulating the transverse vibration of a one-dimensional linear stiff string with sharp initial conditions induced by plucking. The governing Partial Differential Equation (PDE), along with the associated initial and boundary conditions, is incorporated directly into the loss function of the neural network. For the reference measurements, a wire-breaking experiment was performed to excite the string, and its vibration response was captured using a laser profiler. The comparison between experimental measurements, finite-difference time-domain (FDTD) simulations, and PINN-based simulations shows good overall agreement, highlighting the potential of PINNs for modeling stiff-string vibrations.


34. A frontend-backend architecture for tool calls in full-duplex speech models

Authors: Ke Hu, Slyne Deng, Chen Chen, Elena Rastorgueva, Edresson Casanova et al.

Categories: cs.CL

Full-duplex speech-to-speech (S2S) models provide natural, low-latency conversational interaction and would benefit from the ability to use external tools and complete voice-agent tasks. We propose a frontend-backend architecture where a duplex speech-to-text frontend learns to emit a delegation token and forwards streaming ASR transcripts to a text-based backend LLM for tool calls. Tool-call results from the backend are injected back into the frontend through a lightweight prefill-and-repeat mechanism and then synthesized using streaming TTS to the user. Our approach largely preserves regular duplex turn-taking, interruption handling, and low-latency interaction as it requires minimal modifications to the frontend model. In a single-turn tool-call evaluation, our system achieves 92-97% tool-call recall, competitive tool-call prediction performance, and 81.2% accuracy in rejecting irrelevant calls. When equipped with a larger backend (e.g., Qwen3-235B-A22B), our system achieves competitive results on Full-Duplex-Bench-V3 compared to open and closed source models, and significantly outperforms GPT-realtime-mini and Qwen3-Omni-30B-A3B-Instruct on EVA-Bench. These results demonstrate that backend delegation is an effective and modular approach for combining natural duplex speech interaction with strong agentic tool-call capabilities.


35. NemotronLabs VoiceChat: An Open Full-duplex Speech-to-Speech Model with Tool Calling Capabilities

Authors: Jagadeesh Balam, Travis Bartley, Edresson Casanova, Sanjay Chauhan, Chen Chen et al.

Categories: cs.CL, cs.AI

We introduce NemotronLabs VoiceChat, an open full-duplex speech-to-speech model with native tool-calling capabilities. NemotronLabs VoiceChat combines a streaming speech encoder and decoder-only language model with parallel specialized output streams for agent text and structured function calls, an auxiliary RNN-T branch for incremental user transcription, and a streaming TTS decoder. This design enables the model to listen, transcribe, reason, invoke tools, and speak within a unified streaming architecture while preserving the temporal behavior required for natural conversation. On Full-Duplex-Bench 1.0, NemotronLabs VoiceChat achieves the lowest pause-handling takeover rates among evaluated open-weight systems, 100\% takeover following user interruptions, and a 4.33/5 post-interruption response-quality score. On Full-Duplex-Bench 1.5, it resumes its response after user backchannels in 93\% of cases. NemotronLabs VoiceChat obtains a 55.1 normalized average on VoiceBench and, on Full-Duplex-Bench 3.0 (FDB 3.0), achieves 82.5\% tool-selection F1, while argument accuracy and end-to-end tool execution remain areas for improvement. These results demonstrate that full-duplex interaction, speech recognition and generation, general language capabilities, and external tool use can be integrated in a single open speech-to-speech model without sacrificing real-time conversational behavior.


36. Vox-Infinity: Benchmarking the Limits of Long-Context Spoken Language Models

Authors: Xize Cheng, Wenxu Jia, Chenyuhao Wen, Dongjie Fu, Zehan Wang et al.

Categories: cs.CL, cs.HC

Long-context understanding remains a fundamental challenge for large language models, as excessively long inputs often lead models to forget salient information. This issue is even more pronounced in the speech domain, where audio, as a low-compression modality, requires substantially more embeddings than text to preserve both semantic content and acoustic cues. To address this challenge, we introduce \textbf{Vox-Infinity}, the first benchmark specifically designed to evaluate long-context understanding in spoken language models. Vox-Infinity systematically extends audio history along two dimensions: turn count and turn duration. It covers a diverse range of representative scenarios with varying interaction structures and semantic complexity. Crucially, Vox-Infinity provides explicit answer-provenance annotations and organizes samples according to the amount of historical context required to resolve each query, enabling precise and length-aware evaluation. Extensive evaluations of seven representative spoken language models reveal a clear overall recency effect: models generally achieve higher accuracy when answer-supporting evidence is closer to the query, but struggle to retrieve and use evidence located farther back in the dialogue history. Cases and datasets are available at https://vox-infinity.github.io.


37. GestureFAR: Streaming Co-Speech Gesture Generation with Flow Autoregression

Authors: Pinxin Liu, Haiyang Liu, Jiahao Luo, Junhua Huang, Chunhao Zou et al.

Categories: cs.CV, cs.GR, cs.HC

Generating natural co-speech gestures from streaming speech is essential for embodied conversational agents, where motion must be produced while a user is still speaking. Recent streaming gesture systems make online generation possible by autoregressing over discrete motion tokens, but this design compresses high-dimensional continuous motion into finite codebooks and can limit the realism and diversity of generated gestures. To preserve both causality and continuous expressiveness, we propose \textbf{GestureFAR}, a flow-autoregressive framework for streaming co-speech gesture generation. First, GestureFAR autoregresses over causal continuous motion latents, using a transformer to model streaming audio-motion context and a per-token flow-matching head to sample the next latent from a continuous distribution. Second, we introduce a head-only flow distillation strategy that freezes the causal backbone and distills the multi-step per-token flow head into a single network evaluation using consistency and distribution-matching objectives. This keeps the model token-causal while removing the main latency bottleneck for live interaction. Experiments on BEAT2 show that GestureFAR significantly improves the quality–latency trade-off among streaming-capable methods, preserving strong gesture quality while enabling real-time token-causal generation. Project Page: https://andypinxinliu.github.io/GestureFAR


38. Multi-Subject Pretraining Enables Short-Calibration Personalization for Closed-Corpus Surface EMG Speech Decoding

Authors: Chenqian Le, Beatrice Fumagalli, Yasamin Esmaeili, Xupeng Chen, Tianyu He et al.

Categories: cs.LG

Surface electromyography (sEMG)-based silent speech interfaces are limited by cross-user variability and calibration burden. We study a limited-data setting in which each of 27 speech-typical participants contributed less than 0.5 h of data (21.3 min on average) across Aloud and Mimed speech. Within a closed 50-sentence corpus, we used leave-one-subject-out evaluation, initializing from a released single-subject checkpoint, pretraining on non-held-out participants, and fine-tuning on the target participant. This pipeline achieved 21.7% character error rate (CER) and 31.9% word error rate (WER), compared with 49.3% CER without target-subject calibration and 68.0% CER for direct checkpoint fine-tuning. Multi-subject pretraining from random initialization followed by fine-tuning reached 44.9% CER and did not converge under the fixed schedule in 5 of 27 folds, indicating substantial optimization and accuracy benefits from checkpoint initialization. Macro-averaged CER declined from 74.4% with one pretraining participant to 21.7% with 26. Three minutes of target-subject calibration achieved 20.5% CER and 31.7% WER, with no statistically significant difference from the full approximately 13-min pool (21.7% CER and 31.9% WER). A subject-specific adapter provided no detectable benefit. Excluding the five evaluation sentences from all sEMG model-training data increased CER and WER to 78.6% and 99.9%. These results support short-calibration personalization in a standardized-montage, closed-corpus setting.


Authors: Yichen Liu, Quanwei Zhang, Haozhe Wang, Donghao Zhou, Jiankun Zhang et al.

Categories: cs.MM, cs.CV

Joint audio-video generation models have made substantial progress in visual quality and audio-visual synchronization. However, they still provide limited control over when shot transitions occur and dialogue is spoken. This limitation constrains their application in script-driven content creation, where timing errors can undermine narrative coherence and the viewing experience. Current joint generators align video and audio representations on a shared temporal axis, yet the precise timing of shots and dialogue specified in a structured prompt is encoded only in the prompt’s text representation and remains unaligned with the temporal coordinates of either modality. Consequently, video and audio may remain synchronized with each other while both fail to follow the script timeline. This mismatch motivates us to extend temporal alignment beyond video and audio to include the structured script. We therefore introduce Temporal Context Routing (TCR), which maps the script timing onto the shared temporal axis of video and audio generation and routes each prompt’s guidance to the corresponding positions in both modalities. Compared with the baseline on 200 test scripts, TCR reduces Shot Boundary MAE by 96%, from 1.11 s to 0.042 s, and raises Dialogue Acc@0.5 s from 28.3% to 84.1%. TCR achieves these improvements while maintaining visual quality and audio-visual synchronization comparable to those of the baselines. A user study further shows that participants prefer TCR on all five evaluated dimensions.


40. AcousticDiffusion: Semantically Conditioned Audio-Guided Diffusion Policy for Search-and-Rescue Assistance

Authors: Iana Zhura, Didar Seyidov, Dmitrii Plotnikov, Hajira Amjad, Miguel Altamirano Cabrera et al.

Categories: cs.RO

Navigating toward human callers is an important capability for rescue robots operating where visual contact is degraded or occluded. We present AcousticDiffusion, a semantically conditioned, audio-guided diffusion policy for human-directed navigation. A frozen pretrained audio recognizer processes 10.24 s windows, with speech gating and distress-aware prioritization converting recognition outputs into source-level navigation roles. Microphone-array direction-of-arrival measurements are recursively integrated into a robot-centric Bayesian bird’s-eye-view belief field. Ego-motion compensation aligns successive observations, progressively constraining source position while preserving bearing-induced range uncertainty. The semantic belief, recent acoustic observations, audio features, and robot state condition a diffusion model that generates waypoint trajectories. On a synthetic-navigation validation set using recorded audio, AcousticDiffusion achieves a mean end-point bearing error of 11.20 degrees, with 91.78% of trajectories aligned within 30 degrees of the caller. Distractor rejection ranges from 89.20% to 98.99%, and the policy favors a HELP-designated caller over a competing speaker in 91.07% of windows. Deployed online on a ZSL-1 quadruped without additional retraining, it achieves a mean bearing error of 64.9 degrees, compared with 98.2 degrees for A* and 90.4 degrees for RRT, with a mean planner compute time of 6.07 ms. Despite imperfect acoustic localization, the reported mean final source distance is reduced from 3.96 m for the classical planners using ODAS-derived (Open embedded Audition System) guidance to 2.48 m, a 37.4% improvement. These results demonstrate the framework’s ability to translate uncertain acoustic observations into closer approaches to human callers.