Audio-Visual Automatic Speech Recognition Using PZM, MFCC and Statistical Analysis

Saswati Debnath; Pinki Roy

Author	Saswati Debnath Pinki Roy
Keywords	Audio-visual Speech Recognition Lip Tracking Pseudo Zernike Moment Mel Frequency Cepstral Coefficients (MFCC) Incremental Feature Selection (IFS) Statistical Analysis
Abstract	Audio-Visual Automatic Speech Recognition (AV-ASR) has become the most promising research area when the audio signal gets corrupted by noise. The main objective of this paper is to select the important and discriminative audio and visual speech features to recognize audio-visual speech. This paper proposes Pseudo Zernike Moment (PZM) and feature selection method for audio-visual speech recognition. Visual information is captured from the lip contour and computes the moments for lip reading. We have extracted 19th order of Mel Frequency Cepstral Coefficients (MFCC) as speech features from audio. Since all the 19 speech features are not equally important, therefore, feature selection algorithms are used to select the most efficient features. The various statistical algorithm such as Analysis of Variance (ANOVA), Kruskal-wallis, and Friedman test are employed to analyze the significance of features along with Incremental Feature Selection (IFS) technique. Statistical analysis is used to analyze the statistical significance of the speech features and after that IFS is used to select the speech feature subset. Furthermore, multiclass Support Vector Machine (SVM), Artificial Neural Network (ANN) and Naive Bayes (NB) machine learning techniques are used to recognize the speech for both the audio and visual modalities. Based on the recognition rate combined decision is taken from the two individual recognition systems. This paper compares the result achieved by the proposed model and the existing model for both audio and visual speech recognition. Zernike Moment (ZM) is compared with PZM and shows that our proposed model using PZM extracts better discriminative features for visual speech recognition. This study also proves that audio feature selection using statistical analysis outperforms methods without any feature selection technique.
Year of Publication	2021
Journal	International Journal of Interactive Multimedia and Artificial Intelligence
Volume	7
Issue	Regular Issue
Number	2
Number of Pages	121-133
Date Published	12/2021
ISSN Number	1989-1660
URL	https://www.ijimai.org/journal/sites/default/files/2021-11/ijimai7_2_11_0.pdf
DOI	10.9781/ijimai.2021.09.001
	DOI Google Scholar BibTeX EndNote X3 XML EndNote 7 XML Endnote tagged Marc RIS
Attachment	ijimai7_2_11_0.pdf761.73 KB