每日 arXiv

音频 · 语音 · 音乐 · 声学

每日新爬取的 arXiv 论文,按评分排序。中文摘要帮你快速判断是否值得深读。

2026-09-14
日期2026-09-14
已评分
均分
最高

Daily Papers — 2026-09-14

28 papers on audio, speech, music, and acoustics.

1. Typhoon ASR Streaming: Steerable Low-Latency Thai Speech Recognition with Real-Time Shallow Fusion

Authors: Warit Sirichotedumrong, Tanawin Samutsin, Shah Faisal Wani, Sittipong Sripaisarnmongkol, Kunat Pipatanakul

Categories: cs.CL, cs.SD, eess.AS | 4 pages, 2 figures, 3 tables. Accepted to IEEE SLT 2026 (Demo Track)

Open Thai automatic speech recognition (ASR) is dominated by offline, Whisper-based models that read the whole utterance before transcribing, ruling out low-latency uses such as live captioning and voice agents. We present a deployable system for streaming Thai ASR that lets a user steer its vocabulary at decode time, without retraining. A widely used open Thai model, trained with full context, collapses when run as a true stream; we restore streaming with a cache-aware encoder, by converting it or adapting a natively streaming one, and add a shallow-fusion layer that re-ranks candidates inside the streaming decoder with a GPU n-gram language model and phrase boosting. Across two Thai benchmarks and two model sizes, the streaming models stay usable where the full-context model fails, cutting character error rate 4.3-4.5x at a one-second look-ahead while running faster than real time. Decode-time steering then lifts keyword recall from 16.6% to 20.7% at no accuracy cost and negligible overhead; most of the gain comes from an n-gram over ordinary training transcripts, which resolves the written form of code-switched words the model hears but spells inconsistently, with phrase boosting adding targeted control over rare domain terms.


2. Building a Dataset for Music Sample Identification

Authors: R. Oguz Araz, Xavier Lizarraga, Xavier Serra, Dmitry Bogdanov

Categories: cs.SD | Extended Abstracts for the Late-Breaking Demo Session of the 27th Int. Society for Music Information Retrieval Conf., 2026

Sample identification (SI) is the task of matching an element of a musical work to its musically transformed versions used to create new works. The task has received little attention and lacks large-scale publicly available data. In this work, we mine sampling annotations from a music database and split them for training and evaluation. The resulting dataset is nearly three orders of magnitude larger than the existing SI benchmarks, with training, validation, and test sets of 114 k, 6 k, and 10 k tracks. We find that naively splitting the annotations places the same tracks in different sets. To avoid this, we construct a graph from the annotations and split it over connected components. We further find that a single mega-component contains half of the annotations, making component-wise splitting incompatible with balanced splits; we trim it, yielding a leakage-aware pipeline. We share the dataset for non-commercial scientific research purposes only and make the data-analysis and splitting code publicly available. We hope that our work fosters research on SI.


3. Cross-Lingual F5-TTS 2: A Simplified Framework for Language-Agnostic Voice Cloning

Authors: Qingyu Liu, Rixi Xu, Yushen Chen, Zhikang Niu, Haitao Li et al.

Categories: cs.SD

Zero-shot text-to-speech (TTS) can clone a speaker’s voice from a short audio prompt, yet most TTS systems still require the audio prompt transcript during inference. This dependency prevents cross-lingual voice cloning when the audio prompt transcript is unavailable, particularly for unseen languages. Cross-Lingual F5-TTS removes this dependency and enables transcript-free cross-lingual voice cloning, but it prepares its training data with forced alignment. Forced alignment is sensitive to boundary errors, and its cost grows as more languages are covered. Its speaking rate predictor is also unreliable at estimating duration when the audio prompt begins or ends with silence. In this paper, we present Cross-Lingual F5-TTS 2, a simplified framework for transcript-free cross-lingual voice cloning without forced alignment. Instead of using forced alignment to segment real utterances, we build same-speaker prompt and target pairs using a pretrained F5-TTS model and fine-tune the same model on these constructed pairs. This simplifies data preparation and preserves the acoustic modeling capability of the pretrained model, enabling adaptation with only a short fine-tuning stage. We further make the syllable-level speaking rate predictor robust to leading and trailing silence through silence-aware augmentation. Experiments show that Cross-Lingual F5-TTS 2 reaches higher speaker similarity than F5-TTS and Cross-Lingual F5-TTS while maintaining intelligibility. All related resources are publicly available.


4. Generating the Unheard: Phylogeny-Guided Latent Generation for Ancestral Sound Reconstruction

Authors: Tianyi Xu, Shrinaath Narasimhan, Evan Gorstein, Santiago Perea, Yunyi Shen et al.

Categories: cs.SD | 29 pages, 11 figures

What did an ancestral bird species sound like? Existing ancestral state reconstruction methods can infer low-dimensional traits such as morphological characters at internal nodes of a phylogenetic tree, but no one has tried to produce rich perceptual signals such as audio. Some of the challenges include inferred representations that are either too low-dimensional to decode or lie in non-generative feature spaces, so no method to date can produce ancestral audio. We introduce the first framework that generates plausible ancestral vocalizations. Our pipeline encodes bird recordings into a VAE latent space, learns a low-dimensional trait projection aligned with phylogenetic distances, performs ancestral inference in this trait space, and recovers decodable latents through an anchored inverse lift before emitting novel waveforms for each ancestral node. Because the entire pipeline stays within a decodable latent space, every internal node receives a genuinely new audio output representing plausible intermediate ancestral sounds unavailable to retrieval-based alternatives. Experiments on two phylogenetically distant bird clades, 21-species Tyrannidae and 19-species Paridae, show that our method is the only approach that simultaneously achieves genuine generation, phylogenetic consistency, and naturalistic audio quality across both datasets.


5. Graph Attention Design Choices Matter: A Controlled Study of LoRA-Adapted Audio Anti-Spoofing

Authors: Haoyu Wang, Jing Yang, Chenyu Liu, Yushan Du, Yifan Liao et al.

Categories: cs.SD | Accepted by IEEE SLT 2026

Audio anti-spoofing systems increasingly combine self-supervised learning, parameter-efficient fine-tuning, and graph-attention-based backends. However, performance gains in such systems are often entangled with concurrent changes in the backbone, fine-tuning strategy, and training protocol, making the independent contribution of graph attention design difficult to isolate. To address this issue, we conduct a systematic controlled study of the graph attention layer under a unified experimental setting. We decompose the layer into three independently testable design dimensions: scoring symmetry, temperature learnability, and routing granularity. These are instantiated as a concat-based scoring branch, a LearnT branch with learnable temperature, and a multi-temperature routing branch, respectively. Each dimension is implemented as an independently gated residual branch, enabling the evaluation of both individual variants and their combinations under the same experimental setting. Experiments on five evaluation sets with five random seeds show that the LearnT branch achieves the best average equal error rate (EER), yielding a 16.1% relative improvement over the baseline. In contrast, the multi-temperature routing branch does not improve average performance on its own, but substantially reduces cross-seed standard deviation when combined with the concat-based scoring branch. Moreover, two individually effective branches degrade performance when used together, resulting in a 25.6% relative deterioration compared with the baseline. This finding reveals strong non-additive interactions among graph attention design dimensions. Overall, the results suggest that, under parameter-constrained fine-tuning, improvements in graph attention layers depend more on capacity allocation and branch interaction than on simply adding more learnable parameters.


6. Sectional Structure and Emotional Dynamics in Chinese Pop Songs: An Empirical Analysis of Valence-Arousal Trajectories across 100 Songs

Authors: Jingyi Lyu

Categories: cs.SD | 18 pages, 4 figures, 5 tables

Music Emotion Recognition (MER) aims to identify and represent emotional information in music through computational methods and is an important research area within Music Information Retrieval (MIR). To address the limited consideration of song sectional structure in existing dynamic MER research, this study examines 100 Chinese pop songs by aligning 1,046 manually annotated sections with continuous Valence-Arousal (VA) trajectories and analyzing them from the perspectives of section type, adjacent section transitions, repeated sections, and whole-song trajectories. The results show that emotional differences across sections are reflected more strongly in Arousal. Verse typically forms a relatively low-activation baseline, Pre-chorus exhibits a progressive buildup, and Chorus produces a more pronounced high-arousal arrival, while Interlude, Bridge, and Outro tend to show transitional, divergent, and closing functions, respectively. Although whole-song sectional configurations are diverse, high-frequency local transitions are relatively concentrated. High-arousal positions occur more often in the later part of a song, but are not fixed to the final Chorus. Based on these findings, this study summarizes the emotional organization of the selected pop songs as an empirical framework of “local cycles-global accumulation”, in which local sectional cycles are accompanied by emotional pullbacks, while the overall trajectory exhibits a later-stage rise in VA and a tendency for high-Arousal positions to occur later in the song.


7. Augmenting Large Audio-Language Models with Frame-Level Grounding for Fine-Grained Temporal Perception

Authors: Yanfeng Shi, Yan Song, Junhui Li, Tinggan Huang, Wu Guo et al.

Categories: cs.SD, cs.AI | Submitted to ICASSP 2027

Large Audio-Language Models (LALMs) have substantially advanced general audio understanding, yet they remain limited in fine-grained temporal perception, particularly in precise event localization. Existing approaches primarily post-train LALMs to predict event boundaries as timestamp tokens. However, this generative formulation lacks explicit correspondence between the timestamp predictions and fine-grained acoustic evidence, limiting the precision and reliability of temporal localization. To address this issue, we augment the LALM with a dedicated frame-level grounding model while leveraging its semantic modeling capability to represent the event query. Specifically, the frozen LALM encodes the event query with audio as context, and the grounding model combines these query representations with fine-grained audio features to localize the target event at the frame level. Extensive experiments across diverse temporal grounding benchmarks demonstrate strong and consistent improvements over existing methods. Further evaluation shows that the grounding model can provide temporal evidence to support downstream reasoning.


8. MUUNRiver-Bench: Diagnosing Relation-Dependent Music Retrieval with Multimodal Instructions

Authors: Zhancheng Guo, Congren Dai, Shangda Wu, Jianhuai Hu, Danni Zhao et al.

Categories: cs.SD, cs.AI

Music retrieval is relation-dependent: given a reference track, a listener may seek its style with a new theme, a cover, or a comparable voice, and these intents demand contradictory rankings. We present MUUNRiver-Bench, a diagnostic benchmark whose reference-audio queries use natural-language instructions to define relevance. A pipeline combining expert genre priors, LLM-generated prompts and lyrics, synthesis, and expert review yields 3,440 tracks spanning 13 genres and 116 sub-genres, and seven tasks: similar-music, style-preserving lyric-rewriting, lyric-preserving style-rewriting, cover, vocal-timbre, isolated-vocal, and segment retrieval. Across six models in eight configurations, task-wise rank reversals reveal complementary biases: acoustic encoders favour local identity, whereas text-aligned encoders favour semantic relations. Frozen encoders diagnose default similarity preferences; instruction-aware and audio-text fusion systems provide exploratory tests of textual conditioning, with neither simple fusion scheme consistently improving its backbone


9. Rethinking Procedural Audio Pre-training: Source Scaling and Objective Adaptation

Authors: Jiajun Peng, Fengrui Liu, Xinyu Liu, Feng Liu

Categories: cs.SD, cs.AI | Submitted to ICASSP2027

Procedural audio has emerged as a viable source for transferable audio representation learning, but its design principles remain unclear.We revisit two questions: how a procedural source should be scaled, and whether training choices developed on natural audio should transfer unchanged to procedural data.Using a controlled source, we separate scale into formula-class coverage C and within-class rendering diversity I.Experiments with FDSL and AudioMAE show that these two forms of scale provide different benefits and depend on the learning formulation and downstream task. A matched AudioMAE study further shows that procedural audio favors low mask ratios (10%–25%), whereas AudioSet-28K favors 50%–75%. Shared-codebook analysis reveals lower patch diversity and stronger temporal predictability in procedural audio. These results motivate source-aware procedural pre-training, where source scaling and learning configuration are considered jointly.Code is available at https://github.com/Cross-Innovation-Lab/Formula-Bank.


10. CAL-MOS: Bridging Layers with Adapters for Robust MOS Prediction Across Speech Foundation Models

Authors: Alef Iury Siqueira Ferreira, Pedro Lustosa Rege Botelho, Fernanda Silva, Daniel Casanova, Rafael Faustino et al.

Categories: cs.SD, cs.AI, eess.AS

Speech Quality Assessment (SQA) is essential for modern speech technologies, and recent non-intrusive SQA predictors increasingly rely on Speech Foundation Models (SFMs). However, because SFMs expose representations from many layers, it remains unclear which depths are most informative for MOS prediction and how multi-layer information should be combined reliably across backbones and datasets. We benchmark ten SFMs on four MOS datasets under three regimes: full fine-tuning, last-layer probing with a frozen encoder, and naive cross-layer weighted aggregation. We find that the best layer is strongly backbone- and dataset-dependent, and that naive weighted fusion can be unstable across settings. We further evaluate a layer-calibrated aggregation variant that applies per-layer adapters before pooling, which improves the robustness of multi-layer fusion and narrows the gap to full fine-tuning while keeping the backbone frozen.


11. MAST: Label-Efficient, Robust, and Generalizable Sound Detection for Biodiversity Monitoring via Masked Audio Pretraining and Self-Training

Authors: Tianyi Xu, Daniel Pimentel-Alarcón, Zuzana Buřivalová, Claudia Solís-Lemus

Categories: cs.SD, cs.LG | 29 pages, 8 figures

Passive acoustic monitoring can measure biodiversity at larger scales, but time–frequency annotation of animal vocalizations is expensive, site-specific, and difficult to sustain at scale. We present a label-efficient sound detection framework that combines masked audio pretraining with a lightweight detector on mel spectrograms, then further improves robustness through iterative self-training on unlabeled audio. We first pretrain a ViT-based encoder on unlabeled recordings via masked reconstruction and transfer the encoder to a detection backbone. To better separate animal sounds from confounding background, we add a box-level contrastive loss that pulls matched event regions together while pushing noisy negatives apart. We then apply a two-stage pseudo-labeling curriculum to exploit large unlabeled pools without additional annotation. We evaluate the performance on two ecologically distinct domains: tropical rainforest soundscapes (Indonesia) and bird vocalizations in Mediterranean habitats (Spain). On both domains, masked audio pretraining and contrastive learning consistently improve time–frequency detection under temporal and cross-site distribution shift, and self-training yields further gains in out-of-distribution performance. On the rainforest domain, MAST with self-training achieves +0.22 mAP and +0.24 F1 over the strongest baseline under cross-site shift. On the bird domain, self-training achieves +0.12 mAP and +0.10 F1 over the strongest baseline under cross-site shift. Overall, our results show that MAST can effectively extend self-supervised audio representations from clip-level tasks to robust box-level localization across diverse bioacoustic settings, providing a practical path for biodiversity monitoring with limited labels.


12. Timbre Analysis of the Hulusi, a Southwestern Chinese Free-Reed Instrument, using Machine Learning

Authors: Yang Xia, Rolf Bader

Categories: cs.SD, cs.LG, nlin.AO, physics.data-an

The hulusi is a wind instrument that was invented in Yunnan Province, China, and has become tremendously popular in recent years. It consists of a mouthpiece, a gourd, and three bamboo tubes, all with free reeds made of copper. The main bamboo tube in the middle has seven finger holes. In this instrument, the pipe length, not the free reed’s eigenfrequency, determines the instrument’s pitch, unlike, for example, with the Western accordion or the blues harp. In this study, a machine learning model implemented in the COMSAR framework (https://github.com/ifsm) was used to investigate the timbre characteristics of the \emph{hulusi} to cluster different instruments and pitches. The measured \emph{hulusi} pitches C, B, A, G, and F were analyzed according to seven psychoacoustic features, among which only the spectral centroid, sharpness, and fractal correlation dimension are shown to form pitch clusters. These timbre features were used to train Kohonen self-organizing maps (SOMs) for clustering. Brightness and sharpness analysis revealed that the highest pitches were less bright and less sharp than mid- and low-range pitches were. Furthermore, the fractal correlation dimension, which mainly determines the chaoticity of the initial transients, was the best-clustering timbre feature for the hulusi, with the highest pitches showing the least chaoticity. This result is supported by defining a cluster quality index for the SOMs.


13. Tracing the Origins: Legacy Codec Identification in Neural Audio Transcoding

Authors: Wonje Heo, Shinee Youn, Yooshin Kim, Chuck Chae, Donghoon Shin

Categories: cs.SD, cs.MM, eess.AS | 5 pages, 2 figures, to appear Interspeech

Residual Vector Quantization (RVQ)-based neural audio codecs (NACs) enable high-fidelity audio distribution at unprecedentedly low bitrates through discrete token-based representations. However, this shift disrupts traditional forensics, as non-linear neural transcoding obscures the underlying traces of legacy compression. This study defines the forensic gap and proposes a Transformer-based framework designed to leverage the hierarchical and temporal dependencies inherent in RVQ sequences. By modeling inter-layer causal relationships and dynamic forensic significance, our model effectively disentangles superimposed artifacts from legacy-to-neural transcoding. Experimental results achieve 97%+ accuracy for codec identification and robust joint identification performance across 32-128 kbps. These results demonstrate that traditional codec traces persist even after neural transcoding, supporting the feasibility and necessity of neural-codec-aware audio forensics.


14. SongCraft: Unified Song Generation and Editing with Reconstructive Learning

Authors: Haohe Liu, Varun Nagaraja, Gael Le Lan, Xinhao Mei, Zhaoheng Ni et al.

Categories: cs.SD, eess.AS, eess.SP

Song generation and editing have mostly been treated as separate tasks. Existing editing methods often require noise injection and regeneration or curated paired training data. We propose a unified approach for song generation and editing based on reconstructive pretraining, in which a model is trained to reconstruct audio from varying numbers of interpretable conditions. With conditions such as text and lyrics, the model learns to generate diverse songs. With dense conditions specifying fine-grained music attributes, the model learns to reconstruct the target and enables editing by modifying any single attribute while keeping others fixed. This leads to SongCraft, a latent flow matching based model trained for both generation and fine-grained editing. To improve song generation quality, we further introduce word-level phoneme alignment that improves pronunciation learning and accelerates convergence, beat conditioning that improves general musicality, and representation alignment on VAE latent space that produces semantically meaningful latents for improved generation quality. Experiments show that SongCraft achieves the lowest word error rate among evaluated song generation baselines while maintaining competitive audio quality. We further show that a single model can support editing of lyrics, vocal melody, beats, and singer identity, and we also study the trade-off between reconstruction quality and editability.


15. Directivity-Conditioned Low-Latency Neural Filtering for Speech Enhancement in Hearing Aids

Authors: Lennart Uphaus, André Merboldt, Markus Hofbauer, Timo Gerkmann

Categories: eess.AS | Accepted at the International Workshop on Acoustic Signal Enhancement (IWAENC) 2026

Latest advances in neural directional filtering show exceptional results in adapting the direction and shape of directivity patterns during the inference phase. However, in the existing methods for adapting directivity patterns during inference, important real-world constraints have been disregarded. Particularly for hearing devices, scenarios are often much more dynamic, microphone positions vary with head diameter and hearing aid placement, head-shadow effects occur, and strict latency constraints apply. In this work, we propose a novel low-latency (10 ms) deep neural network (DNN) taking the above requirements of hearing devices into account. As in recent work, we use feature-wise linear modulation (FiLM) to steer the directivity patterns during testing. To preserve the desired directivity pattern, a loss function is proposed that maintains the spectral cross-channel relationships. Interestingly, we are able to achieve similar results to methods with relaxed latency constraints.


16. OLAC: An Overlapped Lossless Audio Codec in the Time-Domain with MDCT Compatibility

Authors: Jean-Marc Valin

Categories: eess.AS | 5 pages

Lossless audio coding is a highly mature field of research, with limited potential for significant improvements in pure compression performance. However, emerging real-time wireless applications increasingly require dynamic transitions between lossy and lossless coding to adapt to fluctuating network capacities. Existing standalone lossless codecs cannot achieve this seamless switching without discontinuity. In this paper, we propose a lossless codec based on time-domain aliasing cancellation (TDAC) that can match the overlap in the CELT mode of the Opus codec. This allows the transition between lossy and lossless coding to be achieved without discontinuity or the transmission of redundant information. We show that the proposed codec still achieves state-of-the-art lossless compression without being penalized by its use of TDAC.


17. Word Timestamps and Speaker Attribution with a Non-Autoregressive LLM

Authors: Zvi Kons, Avihu Dekel, Hagai Aronowitz, Vishal Sunder, Ron Hoory

Categories: eess.AS | Submitted to ICASSP 2027

Timestamps and speaker attribution are useful additions to speech recognition, creating a rich text transcript. This information can either be extracted during transcription or aligned to a given transcript. In this paper we present models that add timestamps and speaker information to a given transcript using a non-autoregressive LLM-based architecture. Compared to an autoregressive model built from similar components, the models are more accurate and annotate a given transcript one to two orders of magnitude faster. Compared to other models, our models achieve state-of-the-art timestamp accuracy and the best cpWER for speaker attribution.


18. Interpreting hierarchical organisation of speaker embeddings

Authors: Yanze Xu, Wenwu Wang, Mark D. Plumbley

Categories: eess.AS, cs.AI | Submit to ICASSP 2027

Speaker recognition neural networks learn latent representations (i.e. speaker embeddings) from input utterances to recognise speaker identities. However, the internal mechanisms of these networks remain largely opaque, motivating research in explainable artificial intelligence (XAI) to understand them. Nevertheless, existing studies have analysed how speaker embeddings are organised, but rarely frame these analyses within XAI. Hence, this work proposes to explain and interpret the organisation of speaker embeddings from an XAI perspective. To this end, we apply a hierarchical clustering algorithm, Single-Linkage Clustering (SLINK), to analyse whether some speaker embeddings naturally form clusters with hierarchical relationships. The resulting hierarchical organisation (i.e. hierarchical clusters) is evaluated using the Cluster-Class Matching (CCM) method. Moreover, we propose a new method, termed Hierarchical Cluster-Class Matching (HCCM), to identify which hierarchical clusters best match individual semantic classes (e.g. male) and conjunctive semantic classes (e.g. UK & male), thereby interpreting the clusters using their matched classes. The matching degree is quantified using a new metric called the L-score, which makes imperfect matches diagnosable. HCCM’s results show that hierarchical clusters analysed by SLINK are interpreted using different classes related to speaker identity, gender, and nationality, providing insight into semantics within the hierarchical organisation of our examined speaker embeddings.


19. Reducing the Output-Mode Gap in Speech Language Models via Joint-Output On-Policy Distillation

Authors: Daxin Tan, Dehua Tao, Chengxi Deng, Hanlin Zhang, Xiao Chen

Categories: eess.AS, cs.CL, cs.SD

Autoregressive generation of interleaved text and acoustic tokens is a common approach to spoken-response generation in speech large language models. Although this design enables streaming generation with explicit textual guidance, generated acoustic tokens become part of the context for subsequent text predictions. Given identical speech inputs, we observe markedly lower answer accuracy for the internal text generated in speech-to-text-and-speech (S2TS) mode than for speech-to-text (S2T) responses. We term this discrepancy the \emph{output-mode gap} (OMG). To reduce OMG, we propose \emph{Joint-Output On-Policy Distillation} (JO-OPD), which distills the model’s stronger S2T policy into joint generation using student-generated S2TS trajectories. At each text position, the S2T teacher provides soft targets from a text-only projection of the student’s preceding outputs, while the student predicts from the corresponding full interleaved history. A preservation objective further regularizes native non-text predictions. Experiments on Step-Audio-2-mini and Baichuan-Audio-Instruct reveal OMG across two interleaved generation architectures. On Step-Audio-2-mini, JO-OPD reduces OMG from 42.87 to 16.26 percentage points on Spoken-MQA and from 29.72 to 13.04 points on speech-rendered GSM8K, with little change in S2T accuracy and substantially larger reductions than matched SFT baselines. ASR-based evaluation further shows a 7.49-point improvement in spoken-answer accuracy on Spoken-MQA.


20. Learned Bow Control on a Measured Bowed-String Model: a Revised Minimum-Bow-Force Law, a Recurrent Controller, and the Domain of a Supervision Ceiling

Authors: Homayoon Beigi, Grace Conneely

Categories: eess.AS, cs.LG, cs.SD, eess.SY | 35 pages, 21 tables, 23 figures, 45 references

A finite-difference bowed-string model with implicitly resolved Stribeck friction is presented, with a regime diagnostic, the Schelleng bow-force limits on four strings, and a comparison of learned bow controllers. Implicit resolution is necessary, and quantitatively so: a lagged contact force cannot capture the string on a discrete grid, so no stick phase forms at any bow force. With friction, impedance and quality factor taken from published measurement rather than fitted, all four strings return a stick fraction of 89.1% against an ideal 90%. Schelleng’s maximum bow force is recovered on every string. The minimum is not: it follows $Z v_b β^{-1}$ rather than the predicted $Z^2 v_b β^{-2}$, reducing both squared dependences to first powers. Six controllers at matched capacity, over four strings and twenty seeds each, place a gated recurrent network ahead of a feedforward one, by most under a mid-stroke disturbance. The feedforward network completes more strokes only from a start the model’s own playability map places outside the Helmholtz region. A minimal gated variant fails because gates computed from the input alone cannot clear a latched state. Training loss selects neither the capacity nor the context length, and no learned controller improves on the lookup rule that generated its labels. That bound has a domain. Regressing the controller’s score on the rule’s gives a slope of 0.32, more than ten standard errors below unity, so the controller overtakes the rule where the rule fails and is bounded by it where it holds. Under a rigid finger stop the plant is provably invariant, so transfer loss between pitches belongs to the controller alone and is traced to one feature. A regime classifier without a stick test labels small-amplitude periodic slipping as Helmholtz motion, and a harmonicity measure rates a string the bow never grips above Helmholtz motion.


21. OpenEnded: An Open-Response Speech Corpus for Speaking Proficiency Assessment with Human Annotations and ALM Supervision

Authors: Yu-Wen Chen, Eric Zhou, Evelyn Ding, Tianyi Shen, Zhou Yu et al.

Categories: eess.AS, cs.SD | IEEE SLT 2026

The development of automated speaking assessment (ASA) is limited by the scarcity of public datasets, with most existing work relying on read-aloud speech, which limits applicability to real-world communication scenarios. In this work, we introduce OpenEnded, a corpus of English practice speech from Mandarin speakers in open-response tasks. Unlike prior open-response datasets that provide only holistic proficiency scores, OpenEnded offers utterance-level assessments of accuracy, fluency, and prosody. Approximately 10,000 utterances are collected and annotated using a hybrid framework: 1,000 are manually labeled via multi-rater scoring with discrepancy resolution to form a high-quality test set, while the remaining are pseudo-labeled by an audio language model (ALM) for training and development sets. We evaluate ALMs and existing ASA models on the OpenEnded test set and introduce VoxPA as an additional baseline. Results show that ALM-generated pseudo-labels improve training over original ALM scoring, while VoxPA achieves the best performance among all baselines.


22. Acoustic Image Source Interpolation with Optimal Transport Barycenter

Authors: Yuyang Liu, Rumeshika Pallewela, Jesper Brunnström, Isabel Haasler, Filip Elvander

Categories: eess.AS, eess.SP | Submitted to ICAASP 2027

Room impulse responses can be estimated via the image source model (ISM) using the image source point cloud (ISPC) of a physical source. However, because the source movement changes the ISPC, estimating the ISPC at a new source position typically requires repeated acoustic measurements. We propose an optimal transport (OT) barycenter framework to interpolate the ISPC of a new source location from ISPCs of known sources. The method jointly estimates image-source associations and the ISPC at the new location. The OT ground cost exploits the property that the image sources undergo the same displacement as their physical sources. This approach is realized for both grid-based and support-free configurations. The support-free method addresses the resulting nonconvex joint estimation problem by alternating between identifying image-source associations across the ISPCs and refining the target image-source locations. This enables the interpolation of ISPCs without repeated measurements, facilitating efficient and flexible room-acoustic modeling.


23. ER-EDF: A Psychology-Grounded Emotion Regulation Framework for Speech Empathetic Dialogue Generation in Large Audio-Language Models

Authors: Hongyu Jin, Wenda Zhang, Runqiu Fei, Gongping Huang, Mike Conway et al.

Categories: cs.AI

Empathetic response generation in spoken dialogue systems requires both accurate emotion perception and appropriate emotion regulation. Grounded in psychological theories such as the Perception-Action Model and emotion regulation theory, effective empathy depends not only on inferring a user’s affective state but also on regulating how it is expressed in responses. However, recent large audio-language models (LALMs) largely treat emotion as a direct conditioning signal, lacking explicit regulatory mechanisms, which often leads to affect mirroring rather than calibrated support. We propose ER-EDF, a psychology-grounded framework that explicitly decouples emotion perception and emotion regulation in LALMs. Perception tracks the user’s emotional state, while regulation determines how this state should guide empathetic response generation. The framework is model-agnostic and integrates seamlessly into existing LALMs. We further construct a spoken empathetic dialogue dataset and introduce empathy-aware evaluation metrics beyond lexical matching. Experiments across five LALMs and two datasets show that ER-EDF consistently improves empathetic response quality in both automatic and human evaluations, highlighting the importance of jointly modeling emotion perception and regulation in spoken empathetic dialogue systems, paving a new direction for psychologically grounded empathetic AI.


24. Enabling Streaming User Transcription in Full-Duplex Speech-to-Speech Models

Authors: Ke Hu, Nourchene Ferchichi, Edresson Casanova, Ankita Pasad, Elena Rastorgueva et al.

Categories: cs.CL

Full-duplex speech-to-speech (S2S) models enable natural conversational AI by allowing simultaneous listening and speaking. However, these models typically lack inherent user speech transcription, which is essential for applications such as conversation logging, accessibility features, and quality monitoring. In this work, we propose an efficient method to add streaming ASR capabilities to an existing duplex S2S model by introducing a lightweight ASR head in parallel to the agent text head. Our approach requires minimal additional parameters and no significant architectural changes to the base S2S model, enabling real-time user transcription while preserving full-duplex conversational capabilities including turn-taking and barge-in handling. Experimental results demonstrate that our method achieves streaming average WER of 10.21% on the HuggingFace Open ASR Leaderboard within the duplex S2S framework. Additionally, we show that the same architecture trained as a standalone streaming ASR model achieves competitive results (7.73% WER) compared to current SOTA models. We will open-source our training and inference code to facilitate further research in joint streaming ASR and S2S modeling.


25. Merging the Knowledge of LLMs for Automatic Speech Recognition

Authors: Hayato Futami, Tatsuya Kawahara

Categories: cs.CL | Accepted to Interspeech2026

Automatic speech recognition (ASR) systems, trained on paired speech-text data, have been improved by leveraging language models (LMs) trained on text-only data. LM fusion methods such as shallow fusion and density ratio are well-established methods that incorporate external LMs during ASR decoding. However, they incur additional computational costs due to LM inference, which is particularly problematic for recent larger LMs. In this study, we propose incorporating external LMs via model merging. This method integrates the LMs directly into the parameters of an LLM-based ASR model, requiring no additional computational cost at inference. We formulate domain extension and transfer via arithmetic operations on LoRA parameters. Experimental evaluations were conducted for the domain adaptation of LLM-based ASR trained on CSJ and LibriSpeech. We show that our LM merging consistently improved the ASR performance in the target domains, without degrading inference speed or memory footprint.


26. Sequential Adapter Stacking for Cross-Lingual Low-Resource ASR

Authors: Thai Thi Thanh Thao Dang, Mengjie Qian, Kate Knill

Categories: cs.CL

Extending large-scale multilingual automatic speech recognition (ASR) models to low-resource languages remains challenging. Model performance is skewed toward high-resource languages and degrades sharply for languages with limited labeled data and pre-training exposure. To address this, we investigate parameter-efficient approaches for transferring knowledge from resource-rich source languages to low-resource target languages on Whisper. Alongside warm initialization and attention-based fusion, we propose Sequential Adapter Stacking, which places a trainable target-language adapter on top of a frozen source-language adapter. Under controlled experiments, these approaches are evaluated on three target languages unsupported by Whisper – Asturian, Assamese, and Xhosa – using source languages with varying degrees of relatedness. Sequential Adapter Stacking with the closest related source consistently and significantly outperforms full fine-tuning across the three targets, with 5–8\% relative WER reductions. These gains largely persist with only one hour of target training data.


27. PIVOT: Physics-Grounded Verification for AI-Generated Audio-Video Detection

Authors: Bo Zheng, Kangran Zhao, Xiaoyu Zhang, Weinan Guan, Zhiheng Li et al.

Categories: cs.CV, cs.AI, cs.MM | 19 pages, 4 figures, including appendix

As generative models continue to advance, AI-generated content (AIGC) is becoming increasingly realistic, weakening the artifact cues commonly exploited by existing detectors. Nevertheless, faithfully reproducing the physical behavior of real-world events remains challenging for current generators. We therefore explore detecting AIGC by assessing whether the depicted event satisfies measurable constraints derived from physical laws. We introduce PIVOT, a physics-grounded AIGC detector, instantiated here for audio-video clips, that estimates physical quantities from video and audio, selects physical laws relevant to each clip, and verifies their measurable constraints. Beyond a real/fake decision, PIVOT returns supporting evidence that records the verification outcome, relevant time window, and supporting quantities for each applicable law. Although instantiated and evaluated here on audio-video data, the framework can, in principle, extend to other AIGC modalities whenever the physical quantities required for verification can be estimated reliably. We also introduce PhysForensics-Bench, comprising paired real and generated audio-video clips from nine event-centric scene families and two recent audio-video generators. On PhysForensics-Bench, PIVOT achieves 70.30% accuracy and 64.29% F1 score on Real+Seedance, and 72.16% accuracy and 65.82% F1 on Real+VEO. In comparison, direct inspection with Gemini 3.1 Pro obtains 53.96% accuracy and 60.09% F1 on Real+Seedance, and 57.22% accuracy and 63.44% F1 on Real+Veo. These results demonstrate the practical promise of physical-consistency verification as a structured and inspectable source of evidence that complements artifact-based AIGC detection.


28. Contact-Limited Throughput of a Buoyless Acoustic-to-LEO Gateway With Anticipatory Preparation

Authors: Muhammad Khalil

Categories: eess.SP | This work has been submitted to the IEEE for possible publication. Copyright may be transferred without notice, after which this version may no longer be accessible

A buoyless acoustic-to-LEO gateway with predictable contacts is analyzed. Advanceable cross-medium preparation competes with acoustic collection, while residual RF acquisition consumes contact time. Exact fluid and packet service, a fluid-optimal fixed preparation lead, tandem-queue stability, and one-contact reliability are derived. An explicit service-bias identity shows why treating acquisition as advanceable can select an insufficient lead. For lognormal preparation, conditions are characterized under which variability improves mean service while degrading reliability. Monte Carlo and packet-queue simulations corroborate the analysis with synthetic parameters; in a diagnostic case, accounting for residual RF acquisition raises the actual sustainable rate from 8.776 to 9.079 kbit/s.