Speaker-Invariant Emotion Representations with Gradient Reversal

Authors: Jayathunge, K., Yang, X.

Journal: Lecture Notes in Computer Science

Publication Date: 01/01/2027

Volume: 16818 LNCS

Pages: 466-480

eISSN: 1611-3349

ISSN: 0302-9743

DOI: 10.1007/978-3-032-31397-3_31

Abstract:

Emotion recognition seeks to automatically identify and classify human affective states from various multimodal signals. Traditional supervised learning approaches that rely on categorical emotion labels require large amounts of annotated data, and the labels often reflect cultural assumptions that do not generalise across populations. Contrastive learning-based models offer an alternative by learning embeddings without the need for explicit labels, but often capture speaker-specific features more strongly than emotion-related ones, limiting their effectiveness in cross-sociolinguistic emotion recognition. This paper introduces a novel framework for training unsupervised emotion detection in which a feature encoder is adversarially coupled with a speaker discriminator, discouraging the encoding of speaker-specific information. Simultaneously, audio and visual features are aligned to produce modality-invariant embeddings. The model receives no explicit emotion labels during training, relying solely on an unsupervised contrastive objective. Cross-dataset experiments show significant improvements in target dataset accuracy, particularly as demonstrated in cross-linguistic evaluations transferring from English to Cantonese and vice versa. Code available at https://github.com/jayathungek/gr_suppl.

Source: Scopus