Advanced Techniques: Tokenization

CRM114 sported the interesting combine-successive-tokens scheme. Other schemes include

  • Dealing with separators (e.g., dots or “at” signs in an address) in special ways

  • Recognizing host names, IP addresses, and/or other header information

  • Ignoring certain header information such as dates and message-IDs

  • Choosing only the first n bytes of a message

  • Handling headers specially (e.g., combining the name of the header keyword successively with each alphabetic token on that header’s line)

