The network used for this example has three modules:
- A feature extraction module that processes the audio clips into feature vectors
- A deep neural network module that produces softmax probabilities for each word in the input frame of feature vectors
- A posterior handling module that combines the frame-level posterior scores into a single score for each keyword