How to create a large pandas dataframe from an sql query without running out of memory?

Question

I have trouble querying a table of > 5 million records from MS SQL Server database. I want to select all of the records, but my code seems to fail when selecting to much data into memory. This works: &#8230;but this does not work: It returns this error: I have read here that a similar problem exists when c…

Accepted Answer

Update: Make sure to check out the answer below, as Pandas now has built-in support for chunked loading.You could simply try to read the input table chunk-wise and assemble your full dataframe from the individual pieces afterwards, like this:import pandas as pdimport pandas.io.sql as psqlchunk_size = 10000offset = 0dfs = []while True:  sql = "SELECT * FROM MyTable limit %d offset %d order by ID" % (chunk_size,offset)   dfs.append(psql.read_frame(sql, cnxn))  offset += chunk_size  if len(dfs[-1]) < chunk_size:    breakfull_df = pd.concat(dfs)It might also be possible that the whole dataframe is simply too large to fit in memory, in that case you will have no other option than to restrict the number of rows or columns you&#8217;re selecting.

Advertisement

Answer