Concatenate a large number of HDF5 files

Asked 17/3, 2011 at 23:39 Answered 17/8, 2016 at 17:14

Solved dataset hdf5 scientific-computing

I have about 500 HDF5 files each of about 1.5 GB.

Each of the files has the same exact structure, which is 7 compound (int,double,double) datasets and variable number of samples.

Now I want to concatenate all this files by concatenating each of the datasets so that at the end I have a single 750 GB file with my 7 datasets.

Currently I am running a h5py script which:

creates a HDF5 file with the right datasets of unlimited max
open in sequence all the files
check what is the number of samples (as it is variable)
resize the global file
append the data

this obviously takes many hours, would you have a suggestion about improving this?

I am working on a cluster, so I could use HDF5 in parallel, but I am not good enough in C programming to implement something myself, I would need a tool already written.

Volta answered 17/3, 2011 at 23:39 Comment(8)

One possibility is merging together pairs of files on your cluster; reduce the problem to 250 3GB files, then 125 6 GB files, and so on. This only helps if partially merged files provides any amount of time saving when merging the results later on. – Alee 18/3, 2011 at 0:2

@Alee I am working on hopper at NERSC, theoretical I/O speed is 25 GB/s, also the filesystem is fully parallel and supports MPI I/O. – Volta 18/3, 2011 at 3:3

I was thinking to read maybe 3 or 4 files at a time and write them back all together, but the best would be a c utility that exploits somehow mpi I/O. – Volta 18/3, 2011 at 3:5

Andrea, I am speechless. I figured an array of excellent drives still wouldn't go past a gigabyte per second... – Alee 18/3, 2011 at 3:6

One feature hdf5 has is that you can "mount" several subfiles in a "folder" of the master file. That way it might not be needed to merge them all together into one file. See here: davis.lbl.gov/Manuals/HDF5-1.4.3/Tutor/mount.html – Blackdamp 19/3, 2011 at 20:8

thanks @Blackdamp but I want to concatenate the datasets in order to have a single huge array – Volta 19/3, 2011 at 20:12

@AndreaZonca Could you please post a copy of your script for this? I am currently trying to do something similar and this sounds like it would be very helpful. – Thankless 2/7, 2014 at 22:25

See this snippet: gist.github.com/zonca/8e0dda9d246297616de9 – Volta 3/7, 2014 at 16:54

I found that most of the time was spent in resizing the file, as I was resizing at each step, so I am now first going trough all my files and get their length (it is variable).

Then I create the global h5file setting the total length to the sum of all the files.

Only after this phase I fill the h5file with the data from all the small files.

now it takes about 10 seconds for each file, so it should take less than 2 hours, while before it was taking much more.

Volta answered 21/3, 2011 at 18:8 Comment(0)

I get that answering this earns me a necro badge - but things have improved for me in this area recently.

In Julia this takes a few seconds.

Create a txt file that lists all the hdf5 file paths (you can use bash to do this in one go if there are lots)
In a loop read each line of txt file and use label$i = h5read(original_filepath$i, "/label")
concat all the labels label = [label label$i]
Then just write: h5write(data_file_path, "/label", label)

Same can be done if you have groups or more complicated hdf5 files.

Shillong answered 11/2, 2016 at 7:34 Comment(0)

Ashley's answer worked well for me. Here is an implementation of her suggestion in Julia:

Make text file listing the files to concatenate in bash:

ls -rt $somedirectory/$somerootfilename-*.hdf5 >> listofHDF5files.txt

Write a julia script to concatenate multiple files into one file:

# concatenate_HDF5.jl
using HDF5

inputfilepath=ARGS[1]
outputfilepath=ARGS[2]

f = open(inputfilepath)
firstit=true
data=[]
for line in eachline(f)
    r = strip(line, ['\n'])
    print(r,"\n")
    datai = h5read(r, "/data")
    if (firstit)
        data=datai
        firstit=false
    else
        data=cat(4,data, datai) #In this case concatenating on 4th dimension
    end
end
h5write(outputfilepath, "/data", data)

Then execute the script file above using:

julia concatenate_HDF5.jl listofHDF5files.txt final_concatenated_HDF5.hdf5

Atabrine answered 17/8, 2016 at 17:14 Comment(0)

Make text file listing the files to concatenate in bash:

Write a julia script to concatenate multiple files into one file:

Then execute the script file above using:

Recommended topics

Hot tags