Okay, the implementation with an optimized use of the sequence and result buffers (that is, both are as small as possible) is now as fast as the original CPU implementation and considerably faster for greater sequence lengths (2x to 2.5x for 256 MB). However, there is a problem: Minimizing the result buffer can be either done in an unsafe variant (branch “unsafe”; every “c_add” kernel just reads BASE values, which is not correct for the last one which has a pretty high chance of reading too much data in this case (i.e. result data of the first few kernels in the same level)) or in a safe variant. Doing this safe may either be achieved by limiting the source range of the final buffer (which makes everything incredibly slow) or by simply executing the final kernel of each level on the CPU (branch “shared-only”); this, however, requires the use of host memory for the result buffer which makes this incompatible with using local/shared memory.
I guess we'll just have to compare the speed of this shared-only approach and the one using local/shared memory and then decide on what to use.
Note that the SSE implementation is still five times as fast as “shared-only”, even for a sequence length of 1 GB.
Okay, so I've done a little calculation: My main memory is DDR3-1600 with a peak transfer rate of 12.8 GB/s (according to Wikipedia). Reading 1 GB of sequence data would therefore require at least 78 ms. My SSE implementation takes 35 ms which is due to the fact that it stops before reading everything (still, that is fast than expected; the index returned for 2^29 (sequence length divided by two) is 640M into the sequence, therefore, it should have taken 50 ms).
Our current implementation works on the whole sequence, which makes it take **always** more than 78 ms (my graphics card is connected via PCIe 3.0 x16, whose bandwidth is up to 15.75 GB/s, so this should not be the bottleneck). Currently, it takes about 180 to 200 ms for 1 GB.
The bandwidth of my GPU's memory is 153.6 GB/s which would be great if the data wouldn't have to be read from host memory at least once.
Therefore, if the initial data is in DDR3 host memory, we don't stand a chance. The SSE implemention is as fast as it gets (completely memory bound, there's no way it gets any faster).
There are only two ways of getting any faster:
1. The sequence is not in host memory initially.
2. We find a new algorithm that doesn't involve reading the whole sequence.
(1) could be possible either if the original data is calculated on the GPU anyway, so the program doing that calculation could just leave it there for us to take (that was a joke); or, if we have a system with a faster host memory (**cough cough** the PS4 has a bandwidth of 176 GB/s **cough cough**). But even in that case, the SSE implementation is probably still faster since parallelising an already memory-bound algorithm does not help at all.
As for (2), I don't think that's possible. We could just take a guess, but in the end, we have to verify that our guess is correct: And to do so, we have to read the whole sequence up to the guessed index. And that last part is exactly what the CPU implementations do, therefore, there is simply no way this can be faster.
(This is also the reason why the CPU algorithms are the fastest memory-bound algorithms imaginable: At some point, you **have** to read the sequence until the requested index, so these memory accesses are mandatory; and since these accesses are also all the accesses the CPU implementations actually perform, there is no way to get them faster as long as they are indeed memory bound.)
So, how about we give up? ;-)
PS: Actually, I verified that the SSE implementation is indeed memory bound (and I'm not just making this up from its runtime): I rand cachegrind and callgrind (from valgrind) over it and it said that basically 85 % (estimation based on “memory accesses with a cache miss take 111x as much time as all others” (which I highly doubt, it's probably more)) of the execution time is spent during the SSE memory accesses (which is not surprising considering all these accesses are bound to be cache misses).
Hey,
so I was just going to start implementing the CUDA algorithm but I highly doubt that there won't be a speedup at all now. I will implement it anyways just to validate our guesses ;) Actually I'm really interested in those APU shared memory thing. Does anyone know someone how owns an APU?
Well, a simple APU won't do – the most important thing is that the host memory has a higher bandwidth than usual, i.e., that you don't simply have an APU with DDR3 memory (as is generally the case).
Dear all,
never guess, always measure. How many times did you repeat your measurements to obtain the numbers above? You take your memory bandwidth from Wikipedia??? Use the _stream_ benchmark for this and obtain real-world numbers.
As I said before, there is no guarantee that this application benefits from GPUs. Have you tried asynchronous data transfers? Don't know how this works in OpenCL, but in CUDA this is documented quite well. Have you experimented more with latency hiding on the device (once the data is on the device). I welcome the CUDA implementation for this, since it allows us to somewhat compare the results from different programming paradigms.
Also, I am still asking around for a Kaveri APU. But had no luck so far. The HSA Foundation has not replied to my mail at all. Due to the recent release of the Kaveri APU, this might be due to work(over)load. Maybe there is a German Chair of Computer Science somewhere that has Kaveri APUs under their fingers.
I am still setting up our "OpenCL" machine. I'll probably give your code a try towards the end of the week. In the meantime, I suggest, that you start composing the final presentation, add the numbers and conclusions you have obtained so far. This might save you time and pressure later.
One more thing, I have the feeling so far that workload has been quite asymmetric among the group members. I'll hence suggest, that those individuals that have not coded/committed anything yet prepare and give the presentation (e.g. not max). After the presentation, I'll ask questions to each one of you. I need to make sure that grading your team is done fair as much as possible.
> never guess, always measure
I did not guess, I only measured and calculated whether my results are plausible.
> How many times did you repeat your measurements to obtain the numbers above?
About twenty times with a standard deviation of maybe 10 % each.
> You take your memory bandwidth from Wikipedia???
I believed memory modules had specific timings which made the bandwidth easily calculable, but it appears I was wrong.
> Use the stream benchmark for this and obtain real-world numbers.
stream says 13.4 GB/s for copy and scale and 14.7 GB/s for triad and add, so 12.8 GB/s were at least not too far off.
> Have you tried asynchronous data transfers?
As I said, we have to read the data at some point. Deferring this will (in my very very humble opinion) only make matters worse. We have to use the memory at maximum capacity. Okay, this probably sounds dumb and buzzwordy, so I'll rephrase:
There is no doubt from my side that we absolutely must read the memory from the start until the point in the sequence we're looking at. I can't imagine any other way to verify that point being indeed the solution. There is no way we can read memory faster than with the full memory bandwidth. If the memory bandwidth is 15 GB/s and we're reading through the whole sequence (which is what you proposed to use as a comparison point in the other thread), there is no way we can be any faster than size / 15 GB/s, i.e., 67 ms for 1 GB of sequence data.
It has to be read. The memory initially containing the sequence data limits our speed. We cannot exclude the read from host memory from our calculation; unless, of course, you'd allow us to. In that case we'd just assume the data is already available on the GPU which would then render the host memory speed limitation irrelevant. _Then_ we could go to things such as optimizing memory accesses on the GPU itself.
> In the meantime, I suggest, that you start composing the final presentation, add the numbers and conclusions you have obtained so far. This might save you time and pressure later.
> One more thing, I have the feeling so far that workload has been quite asymmetric among the group members.
I simply did the original OpenCL implementation, now it's up to the others to port it to CUDA and refine it (and make the presentation); the tasks have already been distributed. ;-)
And of course I'm here to answer your questions. :-)