Improving LSTMs – beam search
As we saw earlier, the generated text can be improved. Now let's see if beam search, which we discussed in Chapter 7, Long Short-Term Memory Networks, might help to improve the performance. In beam search, we will look ahead a number of steps (called a beam) and get the beam (that is, a sequence of bigrams) that has the highest joint probability calculated separately for each beam. The joint probability is calculated by multiplying the prediction probabilities of each predicted bigram in a beam. Note that this is a greedy search, meaning that we will calculate the best candidates at each depth of the tree iteratively, as the tree grows. It should be noted that this search will not result in the globally best beam. ...
Become an O’Reilly member and get unlimited access to this title plus top books and audiobooks from O’Reilly and nearly 200 top publishers, thousands of courses curated by job role, 150+ live events each month,
and much more.
Read now
Unlock full access