CORTEXA
← Browse
arxivcs.SDcs.AIcs.CR2026-07-18

Do Speech Tokens Leak Voiceprints? Speaker Inversion Attacks Against End-to-End Speech Language Models

Ye Lu, Yihan Yan, Zhaoyang Zhang, Zhitao Ou, Runze Liu, Li Liu, Shen Wang

End-to-end speech language models increasingly represent user speech with speech tokens rather than relying exclusively on cascaded ASR--LLM--TTS pipelines. Although these tokens support expressive and low-latency spoken interaction, they may also preserve sensitive speaker characteristics. We investigate whether exposed speech tokens leak voiceprints and formulate this risk as a speaker inversion attack. We introduce Audio BERT (AuB), a trainable model that constructs token embeddings from discrete codebooks and aggregates them into speaker-sensitive representations, and propose SpInv, a two-stage inversion method built on AuB to recover embeddings in the space of an attacker-specified speaker encoder. We evaluate Moshi, Higgs3, Kimi-Audio, and Qwen3-Omni using speaker-disjoint protocols on the VoxCeleb dataset. Extensive experiments show that, with only three seconds of frontend output, SpInv achieves cosine similarities above 0.70 in the attacker-specified speaker-encoder space.

View free PDFSource page

Related papers

arxivcs.CRcs.AIcs.SD2026-07-02

Pmeta-TLA: Backdoor Attacks for Speech Classification Models via Meta-Learning with Timbre Leakage Attack

Yueming Huang, Wenhan Yao, Fen Xiao, Xiarun Chen, Weiping Wen

Recently, speech classification methods have gained widespread adoption in intelligent gadgets. Current study indicates that backdoor attacks provide a substantial security concern to these models, underscoring the pressing necessity to investigate additional potential attack tec…

View free PDFSource page
arxivcs.CVcs.AIcs.CRcs.MMcs.SD2026-07-14

Traceback Translators Against Forgetting in Continual Fake Speech Detection

Enrico Gottardis, Mattia Tamiazzo, Simone Milani

Fake speech detectors are increasingly challenged by the development of new and more accurate generative models. To cope with this problem, continual learning techniques are nowadays widely considered feasible strategies for updating models to new datasets, but they also lead to…

View free PDFSource page
arxivcs.LGcs.AIcs.CRcs.SD2026-06-26

What Was That Again? Certified Robustness for Automatic Speech Recognition

Andrew C. Cullen, Neil G. Marchant, Jiani Xie, Paul Montague, Benjamin I. P. Rubinstein

Automatic Speech Recognition systems are notoriously both sensitive to adversarial and benign perturbations. While this has been repeatedly demonstrated using reference datasets, detecting such behaviors in deployed systems is incredibly challenging, due to the absence of oracle…

View free PDFSource page
arxivcs.SDcs.AIcs.CRcs.LG2026-06-26

Room for Error: Large-Scale Simulation of Over-the-Air Acoustic Attacks

Andrew C. Cullen, Neil G. Marchant, Jiani Xie, Paul Montague, Sean Lamont, Maxwell Standen, et al.

While voice control is rapidly becoming a ubiquitous vector of human-AI communication, the risks facing these systems remain poorly understood. This is, in part, a product of the difficulties in scaling strictly digital adversarial workflows to the physical world. These scale bar…

View free PDFSource page
arxivcs.SDcs.AIcs.CRcs.MM2026-07-14

Explainable-by-Design Audio Deepfake Detection via Wiener-Hopf Linear Prediction

Mattia Tamiazzo, Simone Milani, Massimo Iuliani, Marco Fontani

The rapid advancement of synthetic speech generation methods has made audio deepfake detection a critical challenge in multimedia forensics. While recent approaches achieve high detection accuracy, they typically rely on black-box architectures that offer limited interpretability…

View free PDFSource page