Skip to main navigation Skip to search Skip to main content

SV-Mixer: Replacing the Transformer Encoder with Lightweight MLPs for Self-Supervised Model Compresison in Speaker Verification

  • Jungwoo Heo
  • , Hyun Seo Shin
  • , Chan Yeong Lim
  • , Kyo Won Koo
  • , Seung Bin Kim
  • , Jisoo Son
  • , Ha Jin Yu
  • University of Seoul

Research output: Chapter in Book/Report/Conference proceedingConference contributionpeer-review

1 Scopus citations

Abstract

Self-supervised learning (SSL) has pushed speaker verification accuracy close to state-of-the-art levels, but the Transformer backbones used in most SSL encoders hinder on-device and real-time deployment. Prior compression work trims layer depth or width yet still inherits the quadratic cost of self-attention. We propose SV-Mixer, the first fully MLPbased student encoder for SSL distillation. SV-Mixer replaces Transformer with three lightweight modules: Multi-Scale Mixing for multi-resolution temporal features, Local-Global Mixing for frame-to-utterance context, and Group Channel Mixing for spectral subspaces. Distilled from WavLM, SV-Mixer outperforms a Transformer student by 14.6 % while cutting parameters and GMACs by over half, and at 7 5 % compression, it closely matches the teacher's performance. Our results show that attention-free SSL students can deliver teacher-level accuracy with hardwarefriendly footprints, opening the door to robust on-device speaker verification.

Original languageEnglish
Title of host publicationASRU 2025 - 2025 IEEE Automatic Speech Recognition and Understanding Workshop
PublisherInstitute of Electrical and Electronics Engineers Inc.
ISBN (Electronic)9798331544263
DOIs
StatePublished - 2025
Event2025 IEEE Automatic Speech Recognition and Understanding Workshop, ASRU 2025 - Honolulu, United States
Duration: 6 Dec 202510 Dec 2025

Publication series

NameASRU 2025 - 2025 IEEE Automatic Speech Recognition and Understanding Workshop

Conference

Conference2025 IEEE Automatic Speech Recognition and Understanding Workshop, ASRU 2025
Country/TerritoryUnited States
CityHonolulu
Period6/12/2510/12/25

Keywords

  • knowledge distillation
  • mlp-mixer
  • model compression
  • speaker verification
  • transformer-free architecture

Fingerprint

Dive into the research topics of 'SV-Mixer: Replacing the Transformer Encoder with Lightweight MLPs for Self-Supervised Model Compresison in Speaker Verification'. Together they form a unique fingerprint.

Cite this