Other solutions and their speedups

#6 · closed · 8 comments

View on GitHub ↗

roidanton

Hi Peter, in the introduction of the presentation I'd like to mention other attempts and their speedups in comparison to the not optimized CPU implementation. In our meeting you brought up one or two of them. Could you repeat them here (or per mail) including the used technique and the speedup, please?

Comments

psteinb

Hi, I have a SSE4.2 and AVX2 (only supported on Intel Haswell CPUs) solution. I need to invest some time into the latter. If you have a script that launches consecutive tests on some sample data, I can provide you with the needed numbers. NB: Both are not multi-threaded yet. I'll get back to you on this during the week.

psteinb

I need to know of what input data (size of sequence.txt and start/end index) you need the results for SSE4.2 and alike?

roidanton

We use the original sequence.txt you sent to us, the indexes are: start = 0, end = sequence_length/2. We did each test about 1800 times, the time in every run was identical (except the very first run of each test which took a _minor_ bit longer).

psteinb

Hi to all, I flew over the code and am a bit irritated. The cuda implementation is quite naive and does not take advantage of full PCI Express bandwidth. Please look into using cudaHostAlloc with the appropriate flags to shift the incoming sequence into pinned memory and stream it from there onto the device. Also, it'd be nice if my account would receive rights to push changes and become a collaborator to this repo. My tests of the opencl code with a nvidia k20c give the following timings (<mean> +/- rms of 50 runs): opencl 36.7 +/- 0.2 ms cuda-psteinb 1.1 +/- 0.2 ms AVX2, 1core 1.1 +/- 0.05 ms All of the above sheds a quite different light on the subject that we discussed before. The above means that due to the algorithmic complexity GPU and CPU are heads up with a strong notion that the CPU is ahead of things (because I didn't multi-thread the application). Further, I also saw that you implemented the plain character count. I should have checked earlier more thoroughly, but it'd nice if you would have tried to implement: https://github.com/XanClic/transalign-killer/blob/master/cpu/transalign_killer.cpp#L48 so results from the original implementation could be compared 1:1. Best - Peter

janengelmohr

Hi Peter, thanks for your review! However it is quite... late. Cheng and me implemented the Cuda approaches. I won't be able to rewrite the whole code till tomorrow since I'm currently engaged into another project. I could fix it within some days though I guess.

psteinb

Dear all, I know it is late. Just include these thoughts into your discussion tomorrow. Cheers, Peter

freakrobot

It our fault not using the zero copy functions, but is still a bit too late for changing. And the opencl code running on NV device is always not so good.

freakrobot

Thanks, we will do that.