Paper Detail

Multimodal Speaker Verification as a Threat to Speaker Anonymization

Ashi Garg, Cristina Aggazzotti, Leibny Paola García-Perera, Nicholas Andrews

huggingface Score 5.5

Published 2026-07-22 · First seen 2026-07-27

General AI

Abstract

Most automatic speaker verification (ASV) systems operate on individual utterances, despite real-world interactions typically consisting of multiple utterances. As speech accumulates, increasingly rich speaker information becomes available through acoustic, prosodic, and linguistic cues, potentially challenging speaker anonymization methods that primarily target vocal characteristics. We investigate ASV in a multi-utterance, multimodal setting and examine whether aggregating information across anonymized speech impacts privacy. We first study audio-only aggregation across multiple anonymized utterances and observe consistent performance improvements as more speech becomes available. We then incorporate prosodic and linguistic information, showing that multimodal systems outperform unimodal approaches. Finally, we compare aggregation strategies and find that frame-level aggregation yields the lowest EERs. Even with only five anonymized utterances, combining audio and text reduces EER by over 15% relative to audio-only aggregation, demonstrating that substantial speaker-discriminative information remains accessible despite anonymization.

Workflow Status

Review status
pending
Role
unreviewed
Read priority
later
Vote
Not set.
Saved
no
Collections
Not filed yet.
Next action
Not filled yet.

Reading Brief

No structured notes yet. Add `summary_sections`, `why_relevant`, `claim_impact`, or `next_action` in `papers.jsonl` to enrich this view.

Why It Surfaced

No ranking explanation is available yet.

Tags

No tags.

BibTeX

@misc{garg2026multimodal,
  title = {Multimodal Speaker Verification as a Threat to Speaker Anonymization},
  author = {Ashi Garg and Cristina Aggazzotti and Leibny Paola García-Perera and Nicholas Andrews},
  year = {2026},
  abstract = {Most automatic speaker verification (ASV) systems operate on individual utterances, despite real-world interactions typically consisting of multiple utterances. As speech accumulates, increasingly rich speaker information becomes available through acoustic, prosodic, and linguistic cues, potentially challenging speaker anonymization methods that primarily target vocal characteristics. We investigate ASV in a multi-utterance, multimodal setting and examine whether aggregating information across a},
  url = {https://huggingface.co/papers/2607.19636},
  keywords = {code available, huggingface daily},
  eprint = {2607.19636},
  archiveprefix = {arXiv},
}

Metadata

{}