December 2015
Intermediate to advanced
160 pages
2h 50m
English
In this chapter, we saw how we can index files such as PDF, Word documents, and spreadsheets in Solr using the powerful features of Apache Tika. There are many more features available for use, but they are beyond the scope of this book. However, you can get a clear picture here on how easy it is to set up Apache Tika with Solr in order to retrieve information from a document. In the next chapter, we'll see how we can use Apache Nutch to crawl web pages and index the information received by the crawler in Solr.
Read now
Unlock full access