Publication Details

Improving Noise Robustness of Automatic Speech Recognition via Parallel Data and Teacher-student Learning

MOŠNER Ladislav, WU Minhua, RAJU Anirudh, PARTHASARATHI Sree Hari Krishnan, KUMATANI Kenichi, SUNDARAM Shiva, MAAS Roland and HOFFMEISTER Björn. Improving Noise Robustness of Automatic Speech Recognition via Parallel Data and Teacher-student Learning. In: Proceedings of ICASSP. Brighton: IEEE Signal Processing Society, 2019, pp. 6475-6479. ISBN 978-1-5386-4658-8. Available from: https://ieeexplore.ieee.org/document/8683422

Czech title

Zlepšování odolnosti vůči šumu automatického rozpoznávání řeči pomocí paralelních dat a učení typu učitel-žák

Type

conference paper

Language

english

Authors

Mošner Ladislav, Ing. (DCGM FIT BUT)
Wu Minhua (AmazonCom)
Raju Anirudh (AmazonCom)
Parthasarathi Sree Hari Krishnan (AmazonCom)
Kumatani Kenichi (AmazonCom)
Sundaram Shiva (AmazonCom)
Maas Roland (AmazonCom)
Hoffmeister Björn (AmazonCom)

URL

Keywords

automatic speech recognition, noise robustness, teacher-student training, domain adaptation

Abstract

For real-world speech recognition applications, noise robustness is still a challenge. In this work, we adopt the teacherstudent (T/S) learning technique using a parallel clean and noisy corpus for improving automatic speech recognition (ASR) performance under multimedia noise. On top of that, we apply a logits selection method which only preserves the k highest values to prevent wrong emphasis of knowledge from the teacher and to reduce bandwidth needed for transferring data. We incorporate up to 8000 hours of untranscribed data for training and present our results on sequence trained models apart from cross entropy trained ones. The best sequence trained student model yields relative word error rate (WER) reductions of approximately 10.1%, 28.7% and 19.6% on our clean, simulated noisy and real test sets respectively comparing to a sequence trained teacher.

Published

2019

Pages

6475-6479

Proceedings

Proceedings of ICASSP

Conference

2019 IEEE INTERNATIONAL CONFERENCE ON ACOUSTICS, SPEECH AND SIGNAL PROCESSING (ICASSP), Brighton, GB

ISBN

978-1-5386-4658-8

Publisher

IEEE Signal Processing Society

Place

Brighton, GB

DOI

10.1109/ICASSP.2019.8683422

UT WoS

000482554006141

EID Scopus

2-s2.0-85068975951

BibTeX

@INPROCEEDINGS{FITPUB12098,
   author = "Ladislav Mo\v{s}ner and Minhua Wu and Anirudh Raju and Krishnan Hari Sree Parthasarathi and Kenichi Kumatani and Shiva Sundaram and Roland Maas and Bj{\"{o}}rn Hoffmeister",
   title = "Improving Noise Robustness of Automatic Speech Recognition via Parallel Data and Teacher-student Learning",
   pages = "6475--6479",
   booktitle = "Proceedings of ICASSP",
   year = 2019,
   location = "Brighton, GB",
   publisher = "IEEE Signal Processing Society",
   ISBN = "978-1-5386-4658-8",
   doi = "10.1109/ICASSP.2019.8683422",
   language = "english",
   url = "https://www.fit.vut.cz/research/publication/12098"
}