July 2017
Beginner to intermediate
312 pages
7h 27m
English
The next step in our analysis is the comparison of popularity between different programming languages. It will be based on samples of the top 1,000 most popular repositories by year.
Firstly, we get the data for last three years:
queries = ["created:>2017-01-01", "created:2015-01-01..2015-12-31", "created:2016-01-01..2016-12-31"]
We reuse the search_repo_paging function to collect the data from the GitHub API and we concatenate the results to a new dataframe.
df = pd.DataFrame() for query in queries: data = search_repo_paging(query) data = pd.io.json.json_normalize(data) df = pd.concat([df, data])
We convert the dataframe to a time series based on the create_atcolumn.
df['created_at'] = df['created_at'].apply(pd.to_datetime) ...
Read now
Unlock full access