Publication Details

Self-supervised speaker embeddings

STAFYLAKIS Themos, ROHDIN Johan A., PLCHOT Oldřich, MIZERA Petr and BURGET Lukáš. Self-supervised speaker embeddings. In: Proceedings of Interspeech. Graz: International Speech Communication Association, 2019, pp. 2863-2867. ISSN 1990-9772. Available from: https://www.isca-speech.org/archive/Interspeech_2019/pdfs/2842.pdf

Czech title

Embeddingy charakterizující mluvčího se samoučením

Type

conference paper

Language

english

Authors

Stafylakis Themos (OMILIA)
Rohdin Johan A., Dr. (DCGM FIT BUT)
Plchot Oldřich, Ing., Ph.D. (DCGM FIT BUT)
Mizera Petr (OMILIA)
Burget Lukáš, doc. Ing., Ph.D. (DCGM FIT BUT)

URL

Keywords

speaker recognition, self-supervised learning, deep learning

Abstract

Contrary to i-vectors, speaker embeddings such as x-vectors are incapable of leveraging unlabelled utterances, due to the classification loss over training speakers. In this paper, we explore an alternative training strategy to enable the use of unlabelled utterances in training. We propose to train speaker embedding extractors via reconstructing the frames of a target speech segment, given the inferred embedding of another speech segment of the same utterance. We do this by attaching to the standard speaker embedding extractor a decoder network, which we feed not merely with the speaker embedding, but also with the estimated phone sequence of the target frame sequence. The reconstruction loss can be used either as a single objective, or be combined with the standard speaker classification loss. In the latter case, it acts as a regularizer, encouraging generalizability to speakers unseen during training. In all cases, the proposed architectures are trained from scratch and in an endto- end fashion. We demonstrate the benefits from the proposed approach on the VoxCeleb and Speakers in the Wild Databases, and we report notable improvements over the baseline.

Published

2019

Pages

2863-2867

Journal

Proceedings of Interspeech - on-line, vol. 2019, no. 9, ISSN 1990-9772

Proceedings

Proceedings of Interspeech

Conference

Interspeech Conference, Graz, AT

Publisher

International Speech Communication Association

Place

Graz, AT

DOI

10.21437/Interspeech.2019-2842

UT WoS

000831796403001

EID Scopus

2-s2.0-85074683253

BibTeX

@INPROCEEDINGS{FITPUB12092,
   author = "Themos Stafylakis and A. Johan Rohdin and Old\v{r}ich Plchot and Petr Mizera and Luk\'{a}\v{s} Burget",
   title = "Self-supervised speaker embeddings",
   pages = "2863--2867",
   booktitle = "Proceedings of Interspeech",
   journal = "Proceedings of Interspeech - on-line",
   volume = 2019,
   number = 9,
   year = 2019,
   location = "Graz, AT",
   publisher = "International Speech Communication Association",
   ISSN = "1990-9772",
   doi = "10.21437/Interspeech.2019-2842",
   language = "english",
   url = "https://www.fit.vut.cz/research/publication/12092"
}