Hands-On AI Trading with Python, QuantConnect, and AWS
by Jiri Pik, Ernest P. Chan, Jared Broad, Philip Sun, Vivek Singh
Chapter 4Step 2: Dataset Preparation
Once the problem is defined, the next task is to gather and preprocess relevant historical data used to train and test the predictive model. The quality and comprehensiveness of the dataset directly impacts the model’s performance and ability to generalize to unseen data.
Data Collection
The first phase of dataset preparation involves data collection, which includes gathering historical price data, trading volumes, and other relevant market data for the assets in question. Additionally, macroeconomic indicators, company financial statements, and even alternative data sources such as sentiment analysis from news articles or social media can provide valuable insights. It is crucial to ensure that the data is sourced from reliable providers to maintain accuracy and integrity.
Exploratory Data Analysis
After the data is collected, we need to understand the nature of the dataset and its features. Exploratory Data Analysis (EDA) analyzes and visualizes the dataset to uncover patterns, detect anomalies, and confirm assumptions using summary statistics and charts. It also assists in determining subsequent actions for data preprocessing and model formulation.
In Python, EDA can be efficiently performed using libraries such as pandas (see qnt.co/book-pandas) for data manipulation and tools like Sweetviz (see qnt.co/book-sweetviz) for automated EDA reporting.
To install Sweetviz, run this command
pip install sweetviz
Let’s illustrate how to ...
Become an O’Reilly member and get unlimited access to this title plus top books and audiobooks from O’Reilly and nearly 200 top publishers, thousands of courses curated by job role, 150+ live events each month,
and much more.
Read now
Unlock full access