Among the techniques available to extract content from the web, we can highlight the following:
- Screen scraping: A technique that allows you to obtain information by moving around the screen, registering user pulsations.
- Web scraping: The aim is to obtain the information of a resource, such as a web page in HTML, and process that information to extract relevant data.
- Report mining: A technique that also tries to obtain information, but in this case from a file (HTML, RDF, CSV, and so on). So, with this approach, we can create a simple and fast mechanism without the need to write an API. A main characteristic is that we can indicate that the system does not need a connection, since it is possible to extract the information ...