CORTEXA
← Browse
arxiveess.ASeess.SP2026-07-04

TRACE-EVC: Text-Guided Relative Affective Control for Zero-Shot Emotional Voice Conversion

Zihan Zhang, Shreeram Suresh Chandra, Zongyang Du, Xiutian Zhao, Aurosweta Mahapatra, Hao Zhang, Philipp Koehn, Berrak Sisman

Traditional emotional voice conversion (EVC) conditions generation on explicit target emotions like labels or references, defining the target affective state but omitting the direction or nature of the transition. We introduce instruction-guided relative emotional voice conversion, a task where natural-language instructions specify source-conditioned affective transformations (e.g., "make the speech slightly calmer" or "sound noticeably more confident") instead of fixed targets. To support this task, we construct TRACE-Instruct, a dataset of relative emotion instructions covering categorical transitions, intensity modifications, and open-ended affective changes. We propose TRACE-EVC, a zero-shot framework built around Emo-Compass, a module that models each conversion as a source-anchored rectified flow. Rather than conditioning on an explicit target, it predicts the direction and degree of the affective change. Experiments demonstrate that TRACE-EVC accurately follows relative emotion instructions while preserving speaker identity, linguistic content, and speech quality, and remains competitive with conventional EVC systems on standard categorical emotion conversion.

View free PDFSource page

Related papers

arxivcs.SDcs.LGeess.ASeess.SPmath.NA2026-07-20

FlowSonic: Stable Zero-Shot Music Editing via High-Order Trajectory Integration

Ali Boudaghi, Hadi Zare

Zero-shot text-guided editing of real-world music recordings requires balancing semantic modification with faithful preservation of the original musical structure. Although recent diffusion transformers trained with rectified flow have achieved remarkable success in text-to-music…

View free PDFSource page
arxivcs.SDeess.ASeess.SP2026-06-27

Underwater Source Detection and Classification for Signal-based Surveillance: Audio Dataset Curation and Cross-Domain Evaluation

Quoc Thinh Vo, David K. Han

Machine learning for underwater acoustics is constrained by the scarcity of publicly available labeled datasets. In contrast to air-acoustic domains, where large benchmarks enable rapid model development, underwater datasets are typically small and limited in acoustic diversity,…

View free PDFSource page
arxiveess.ASeess.SP2026-07-14

Spatial-Frequency Cued Generative Fixed-Filter Active Noise Control Based on Deep Learning in Reverberant Environments

Boxiang Wang, Haowen Li, Dongyuan Shi, Junwei Ji, Ziyi Yang, Zhengding Luo, et al.

Generative fixed-filter active noise control (GFANC) effectively attenuates noise with diverse frequency characteristics through the combination of sub control filters. However, it does not incorporate the spatial information of the noise source, which limits its performance, par…

View free PDFSource page
arxivcs.SDcs.AIeess.ASeess.SPeess.SY2026-07-10

A Production-Oriented Framework for Evaluation of SFX Generation

Mélodie Desbos, Yara Bahram, Eric Granger, Mohammadhadi Shateri

Industrial sound design requires audio generation systems that not only produce realistic audio, but also preserve the perceptual identity of a reference, support controllable variation, and remain efficient for practical workflows. Existing evaluations are usually tied to text-t…

View free PDFSource page
arxiveess.ASeess.SP2026-07-04

CHILDES-Aligned: A Curated Children's Speech Dataset via Multi-Model Timestamp Ensembling

Haolong Zheng, Yuanzhuo Hu, Xinyu Liang, Vishal Sunder, Dancheng Liu, Jinjun Xiong, et al.

CHILDES is a large-scale child speech corpus containing long-form recordings of naturalistic child-adult interactions, making it a valuable resource for studying child speech and language development. However, utterance-level timestamps provided in this corpus are often noisy, in…

View free PDFSource page
arxiveess.AScs.AIeess.SP2026-07-04

Probing Low-Level Acoustic Attribute Encoding in CLAP Audio Embeddings

Héctor Martel, Joe Hennessy-Priest, Taemin Cho

Audio foundation models are widely adopted as general-purpose feature extractors, yet the internal structure of their learned representations remains insufficiently understood. In this work, we analyze CLAP audio embeddings through a probing framework, studying the encoding of th…

View free PDFSource page