Hi Peter,
in the introduction of the presentation I'd like to mention other attempts and their speedups in comparison to the not optimized CPU implementation. In our meeting you brought up one or two of them. Could you repeat them here (or per mail) including the used technique and the speedup, please?
Hi, I have a SSE4.2 and AVX2 (only supported on Intel Haswell CPUs) solution. I need to invest some time into the latter. If you have a script that launches consecutive tests on some sample data, I can provide you with the needed numbers. NB: Both are not multi-threaded yet. I'll get back to you on this during the week.
We use the original sequence.txt you sent to us, the indexes are: start = 0, end = sequence_length/2.
We did each test about 1800 times, the time in every run was identical (except the very first run of each test which took a _minor_ bit longer).
Hi to all,
I flew over the code and am a bit irritated. The cuda implementation is quite naive and does not take advantage of full PCI Express bandwidth. Please look into using cudaHostAlloc with the appropriate flags to shift the incoming sequence into pinned memory and stream it from there onto the device.
Also, it'd be nice if my account would receive rights to push changes and become a collaborator to this repo. My tests of the opencl code with a nvidia k20c give the following timings (<mean> +/- rms of 50 runs):
opencl 36.7 +/- 0.2 ms
cuda-psteinb 1.1 +/- 0.2 ms
AVX2, 1core 1.1 +/- 0.05 ms
All of the above sheds a quite different light on the subject that we discussed before. The above means that due to the algorithmic complexity GPU and CPU are heads up with a strong notion that the CPU is ahead of things (because I didn't multi-thread the application).
Further, I also saw that you implemented the plain character count. I should have checked earlier more thoroughly, but it'd nice if you would have tried to implement:
https://github.com/XanClic/transalign-killer/blob/master/cpu/transalign_killer.cpp#L48
so results from the original implementation could be compared 1:1.
Best -
Peter
Hi Peter,
thanks for your review! However it is quite... late. Cheng and me implemented the Cuda approaches. I won't be able to rewrite the whole code till tomorrow since I'm currently engaged into another project.
I could fix it within some days though I guess.
It our fault not using the zero copy functions, but is still a bit too late for changing. And the opencl code running on NV device is always not so good.