July 2017
Beginner to intermediate
312 pages
7h 27m
English
Once the crawling is finished, we have all the data in the MongoDB database. We can now query the database to put all the posts into a pandas dataframe:
import pandas as pdfrom pymongo import MongoClientclient = MongoClient('HOST:PORT')db = client.teamspeedcollection = db.forum_teamspeeddataset = []for element in collection.find(): dataset.append(element)df = pd.DataFrame(dataset)
At this stage, we will also create a new column called full_verbatim, where we concatenate the subject (thread title) and post content:
df['full_verbatim'] = df.apply(lambda x: x['subject'] + " " + x['post'],axis=1)
There exists a direct link between thread title and post, so the textual data included in both variables might be insightful ...
Read now
Unlock full access