Chapter 8. Wake-Word Detection: Training a Model
In Chapter 7, we built an application around a model trained to recognize “yes” and “no.” In this chapter, we will train a new model that can recognize different words.
Our application code is fairly general. All it does is capture and process audio, feed it into a TensorFlow Lite model, and do something based on the output. It mostly doesn’t care which words the model is looking for. This means that if we train a new model, we can just drop it into our application and it should work right away.
Here are the things we need to consider when training a new model:
- Input
-
The new model must be trained on input data that is the same shape and format, with the same preprocessing as our application code.
- Output
-
The output of the new model must be in the same format: a tensor of probabilities, one for each class.
- Training data
-
Whichever new words we pick, we’ll need many recordings of people saying them so that we can train our new model.
- Optimization
-
The model must be optimized to run efficiently on a microcontroller with limited memory.
Fortunately for us, our existing model was trained using a publicly available script that was published by the TensorFlow team, and we can use this script to train a new model. We also have access to a free dataset of spoken audio that we can use as training data.
In the next section, we’ll walk through the process of training a model with this script. Then, in “Using the Model in Our Project” ...
Become an O’Reilly member and get unlimited access to this title plus top books and audiobooks from O’Reilly and nearly 200 top publishers, thousands of courses curated by job role, 150+ live events each month,
and much more.
Read now
Unlock full access