Here, we will look at the steps required to build your own parser with the help of the CKY algorithm. Let's begin summarizing:
- You should have tagged the corpus that has a human-annotated parse tree: if it is tagged as per the Penn Treebank annotation format, then you are good to go.
- With this tagged parse corpus, you can derive the grammar rules and generate the probability for each of the grammar rules.
- You should apply CNF for grammar transformation.
- Use the grammar rules with probability and apply them to the large corpus; use the CKY algorithm with the Viterbi max score to get the most likely parse structure. If you are providing a large amount of data, then you can use the ML learning technique ...