As we have discussed, the overhead of data transfering will be the bottleneck of the GPU implementation. Thus, the project cannot actully benefit from the discrete GPU implementation.
I came up with an idea, and I already discussed it with Jan at noon.
1. We implement both the CUDA version and OpenCL version(already have) version of the algorithm.
2. We measure the Speedup and Efficiency on discrete GPU like NV card, and also on AMD Fusion APU(using OpenCL) which put the CPU and GPU on the same die, which will of course get better speedup than both CPU version and d-GPU version in our case.
3. Then we can compare the result of different devices to make a conclusion on the presentation, which can make this project still meaningful as it should be.
Great idea! Right, using an APU may be the best option we have. Since any element in the sequence is accessed only once and every element in the result buffer is written once and read once, it's actually always the speed of host memory which is the bottleneck instead of the graphics memory. Therefore, any advantage of using graphics memory which is dedicated to parallel access is probably completely negligible.
On the other hand, on an APU, the GPU and the CPU generally still use separated parts of RAM (as far as I know), therefore, we'd still need to copy data. However, if we force the GPU to use host memory, perhaps it is able to access that memory faster than a dedicated GPU through PCIe, though. The best way to go would probably be a PlayStation 4 with its unified GDDR5 memory model. ;-)
Sounds like very good ideas! Two thoughts come to my mind:
1) implement a CUDA/OpenCL version that gives the same (aka correct) results as the serial version, then port the application to other architectures (and potentially profile your code to optimise it on different architectures). Always remember, that your implementation should not only be fast but also correct.
2) in case you'll ask at some point, I neither can give you access to an AMD APU nor a PlayStation 4 (as much as I wish I could)
Please commit any code you have as soon as you have implemented and validated it. I have recently gotton my hands on an Haswell machine that will be equipped with a AMD ATI GPU at some point. So I can contribute some numbers as well, if needed.
Best -
@freakrobot, do you have access to an AMD APU?
> he GPU and the CPU generally still use separated parts of RAM (as far as I know) [...] PlayStation 4 with its unified GDDR5 memory model
Yep, AMDs PC APUs [starting with Kaveri](http://www.hardwareluxx.de/index.php/news/hardware/prozessoren/29272-ces-2014-amd-enthuellt-weitere-details-zu-kaveri.html) will support unified memory.
I have written a request for support to the academic tier of the HSA foundation. Let's see if they could help us test the algorithm (once working, which I don't see currently).
Thank you very much! :-)
> once working, which I don't see currently
Well, the code in my branch seemed to yield the correct result for every input I tested it with, so we decided to merge it to master (which I did just now).
Code is working and giving reasonable results. Since the functionality does not match the one given in the original implementation, I can only say that it gives reasonable results.
Closing before presentation ...