Purpose of the project

#3 · closed · 9 comments

View on GitHub ↗

freakrobot

As we have discussed, the overhead of data transfering will be the bottleneck of the GPU implementation. Thus, the project cannot actully benefit from the discrete GPU implementation. I came up with an idea, and I already discussed it with Jan at noon. 1. We implement both the CUDA version and OpenCL version(already have) version of the algorithm. 2. We measure the Speedup and Efficiency on discrete GPU like NV card, and also on AMD Fusion APU(using OpenCL) which put the CPU and GPU on the same die, which will of course get better speedup than both CPU version and d-GPU version in our case. 3. Then we can compare the result of different devices to make a conclusion on the presentation, which can make this project still meaningful as it should be.

Comments

XanClic

Great idea! Right, using an APU may be the best option we have. Since any element in the sequence is accessed only once and every element in the result buffer is written once and read once, it's actually always the speed of host memory which is the bottleneck instead of the graphics memory. Therefore, any advantage of using graphics memory which is dedicated to parallel access is probably completely negligible. On the other hand, on an APU, the GPU and the CPU generally still use separated parts of RAM (as far as I know), therefore, we'd still need to copy data. However, if we force the GPU to use host memory, perhaps it is able to access that memory faster than a dedicated GPU through PCIe, though. The best way to go would probably be a PlayStation 4 with its unified GDDR5 memory model. ;-)

psteinb

Sounds like very good ideas! Two thoughts come to my mind: 1) implement a CUDA/OpenCL version that gives the same (aka correct) results as the serial version, then port the application to other architectures (and potentially profile your code to optimise it on different architectures). Always remember, that your implementation should not only be fast but also correct. 2) in case you'll ask at some point, I neither can give you access to an AMD APU nor a PlayStation 4 (as much as I wish I could) Please commit any code you have as soon as you have implemented and validated it. I have recently gotton my hands on an Haswell machine that will be equipped with a AMD ATI GPU at some point. So I can contribute some numbers as well, if needed. Best -

roidanton

@freakrobot, do you have access to an AMD APU? > he GPU and the CPU generally still use separated parts of RAM (as far as I know) [...] PlayStation 4 with its unified GDDR5 memory model Yep, AMDs PC APUs [starting with Kaveri](http://www.hardwareluxx.de/index.php/news/hardware/prozessoren/29272-ces-2014-amd-enthuellt-weitere-details-zu-kaveri.html) will support unified memory.

psteinb

we could ask at global foundaries if they have spare ones! c't reported at the end of 2013 that the Kaveri APUs might be produced in Dresden.

janengelmohr

That'd be awesome! An APU might be one of the fastest solution for this problem as I already discussed with Cheng. Maybe AMD will be generous :-)

psteinb

I have written a request for support to the academic tier of the HSA foundation. Let's see if they could help us test the algorithm (once working, which I don't see currently).

XanClic

Thank you very much! :-) > once working, which I don't see currently Well, the code in my branch seemed to yield the correct result for every input I tested it with, so we decided to merge it to master (which I did just now).

XanClic

Argh, a misclick, didn't want to close, sorry.

psteinb

Code is working and giving reasonable results. Since the functionality does not match the one given in the original implementation, I can only say that it gives reasonable results. Closing before presentation ...