Feature extraction module
In order to make the computation easy, the incoming audio signal is run through a voice-activity detection system and the signal is divided into speech and non-speech parts of the signals. The voice activity detector uses a 30-component diagonal covariance GMM model. The input to this model is 13-dimensional PLP features, their deltas, and double deltas. The output of GMM is passed to a State Machine that does temporal smoothing.
The output of this GMM-SM module is speech and non-speech parts of the signal.
The speech parts of the signal are further processed to generate the features. The acoustic features are generated based on 40-dimensional log-filterbank energies computed every 10 ms over a window of 25 ms. 10 ...
Become an O’Reilly member and get unlimited access to this title plus top books and audiobooks from O’Reilly and nearly 200 top publishers, thousands of courses curated by job role, 150+ live events each month,
and much more.
Read now
Unlock full access