October 2019
Beginner to intermediate
674 pages
16h 43m
English
Reading a large file in memory at once may consume the entire RAM of the computer and may cause it to throw an error. In such cases, it becomes pertinent to divide the data into chunks. These chunks can then be read sequentially and processed. This is achieved by using the chunksize parameter in read_csv.
The resulting chunks can be iterated over using a for loop. In the following code, we are printing the shape of the chunks:
for chunks in pd.read_csv('Chunk.txt',chunksize=500): print(chunks.shape)
These chunks can then be concatenated to each other using the concat method:
data=pd.read_csv('Chunk.txt',chunksize=500)data=pd.concat(data,ignore_index=True)print(data.shape)
Read now
Unlock full access