Modern Methods of Speech Processing

Guarantor

Černocký Jan, prof. Dr. Ing. (DCGM)

Language of instruction

Czech, English

Completion

Examination

Time span

Department

Department of Computer Graphics and Multimedia (UPGM)

Study literature

Moore, B.C.J., : An introduction to the psychology of hearing, Academic Press, 1989
Jelinek, F.: Statistical Methods for Speech Recognition, MIT Press, 1998
Fukunaga, K.: Introduction to Statistical Pattern Recognition, Academic Press, 1990
Vapnik, V. N.: Statistical Learning Theory, Wiley-Interscience, 1998
Dutoit, T.: An Introduction to Text-To-Speech Synthesis, Kluwer Academic Publishers, 1997

Fundamental literature

Psutka, J.: Komunikace s s počítačem mluvenou řečí. Academia, Praha, 1995
Gold, B., Morgan, N.: Speech and audio signal processing, John Wiley & Sons, 2000
Texty z http://www.fit.vutbr.cz/~cernocky/speech/

Syllabus of lectures

Review of notions: signal vectors and parameter matrices, basic statistics.
Stochastic modeling of parameters, modeling of time by state sequences.
Hidden Markov models: basic structure, training.
Recognition of speech using HMM: Viterbi search, token passing.
Pronunciation dictionaries and language models.
Speech production and derived parameters: LPC, Log area ratios, line spectral pairs.
Speech perception and derived parameters: Mel-frequency cepstral coefficients, Perceptual linear prediction.
Temporal properties of hearing - RASTA filtering.
Training the feature extractor on the data - linear discriminant analysis.
Speech databases: standards, contents, speakers, annotations.
Vocoders and modeling of the excitation: multi-pulse and stochastic excitations (GSM coding).
CELP coding: long-term predictor, codebooks. Very low bit-rate coders.
Current methods of speaker identification and verification.

Course inclusion in study plans