Force dask to_parquet to write single file

Question

When using dask.to_parquet(df, filename) a subfolder filename is created and several files are written to that folder, whereas pandas.to_parquet(df, filename) writes exactly one file. Can I use dask's to_parquet (without using compute() to create a pandas df) to just write a single file?

Answer 1

Writing to a single file is very hard within a parallelism system. Sorry, such an option is not offered by Dask (nor probably any other parallel processing library).

You could in theory perform the operation with a non-trivial amount of work on your part: you would need to iterate through the partitions of your dataframe, write to the target file (which you keep open) and accumulate the output row-groups into the final metadata footer of the file. I would know how to go about this with fastparquet, but that library is not being much developed any more.

Answer 2

There is a reasons to have multiple files (in particular when a single big file doesn't fit in memory) but if you really need 1 only you could try this

import dask.dataframe as dd
import pandas as pd
import numpy as np

df = pd.DataFrame(np.random.randn(1_000,5))

df = dd.from_pandas(df, npartitions=4)
df.repartition(npartitions=1).to_parquet("data")

Force dask to_parquet to write single file

Question

2 answers

solution1
2 ACCPTED 2020-04-08 20:38:59

solution2
1 2020-04-08 20:38:43

Force dask to_parquet to write single file

Question

2 answers

solution1 2 ACCPTED 2020-04-08 20:38:59

solution2 1 2020-04-08 20:38:43

solution1
2 ACCPTED 2020-04-08 20:38:59

solution2
1 2020-04-08 20:38:43