increase the upper boundary

#5 · closed · 4 comments

View on GitHub ↗

psteinb

The upper search index is far too small!! In the cpu code: https://github.com/XanClic/transalign-killer/blob/master/cpu/transalign_killer.cpp#L107 in the opencl code: https://github.com/XanClic/transalign-killer/blob/master/opencl/main.c#L118 in the _manual_ sse code (why were compilers created again?): https://github.com/XanClic/transalign-killer/blob/master/sse/main.c#L64 Especially if you try to blame memory transfers for something and then you only use half of the data, AFAIK.

Comments

XanClic

It's not safe to increase the upper boundary since we don't know it beforehand. We could fix it since we know it for the given data, however. But currently, the logical index points two thirds into the buffer, therefore, if we simply add a 50 % penalty to the time required, we receive the worst case timing. Furthermore, I personally think comparing the average case is probably better here than comparing the worst case, but obviously using the worst case is better for us, so that's for you to decide. ;-) Until now, I was fine with knowing that I had to add about 50 % for the CPU implementations to get the worst-case result. If one is to add that penalty to my results posted in the other thread, the SSE implementation is still twice as fast.

psteinb

My concern with this is simply that you compare the file size (say 10MB) to transfers in memory of your host (which might only use 50% to work upon) and extrapolate linearly. Due to the fact that modern CPUs have a caching hierachy this extrapolation might be off by quite a bit. As I have explained earlier, in the production system biology guarantees that you'll never run over the end of the sequence! That is why I set the upper boundary so close to the end of the sequence. And even if you append-clone data sets, this should be OK. So please do it. Also ... _Please always measure_ and make observations/interpretations on that basis. Not on speculations or extrapolations. In Computer Science you are in the luxurious situation that repeating your experiment is usually quite cheap! So please use this virtue. Run the code on a given data set (say) 50 times. Obtain the mean and sample variance and use this as a basis for interpretation.

XanClic

> Due to the fact that modern CPUs have a caching hierachy In the other thread, I mentioned that the result index is 640 MB into the sequence. Not even Xeon L3 caches are that big. > That is why I set the upper boundary so close to the end of the sequence. To be precise, it wasn't close to the sequence end, but exceeded it – this is the reason why I changed it in the first place. > Please always measure and make observations/interpretations on that basis. We do measure. But we're studying computer science and are not plain hackers, so at some point we have to approach the problem systematically and that naturally involves proving whether we're actually able to make our implementation faster in any way. > Not on speculations or extrapolations. If my assumptions were bare speculations, I ask you very kindly to prove (or even better: make) them wrong. Don't we have to read the data from host memory? If we actually don't, that's great, of course! :-) (and I honestly mean that: It'd be a real shame if we couldn't get this faster on a GPU simply because the host memory is slow) > In Computer Science you are in the luxurious situation that repeating your experiment is usually quite cheap! Yes, but computer science involves approaching a problem systematically in addition to just throwing everything we can think of at it, which is hacking. Hacking is great and mere computer science often lacks a good chunk of hacking, but I always say both fields cooperate quite nicely and therefore we shouldn't just rely on one of them. > Run the code on a given data set (say) 50 times. Obtain the mean and sample variance and use this as a basis for interpretation. Yes, automatizing and “standardizing” the test is a very good idea. Thank you! :-) So, any volunteers? ;-) PS: Oh, and about > in the manual sse code (why were compilers created again?): I'm always astounded by the results GCC is able to produce on -O3. However, I've yet to see it truly use SSE/MMX (and I don't mean the scalar instructions) other than when finding a memset or memcpy implementation. One could use the intrinsics, but the last time I did that, the result was basically hell. Every result was stored to memory and every operand was loaded from memory. Never will I ever use intrinsics again. What did work, however, was GCC's vector extension, i.e. **attribute**((vector_size(16))) or something like that. But I don't think I'll get it to produce exactly that code I have there for horizontal addition of bytes (especially considering there is in fact no instruction for doing so; I use psadbw which does something similar but not exactly that). In the end, I probably would have to write the same amount of C code as I've done now in assembly, so… Compilers, especially GCC on x86, are really good, I know that. In a general case, I'd never write assembly my own (except for fun). But for vectorizing, at least GCC still has a long way to go (that's at least my experience). If your experience differs, just know that there is a reason why I've written assembly and it isn't “omg compilers are so bad”. They're generally great, but for SSE, they're not so much (except for maybe the Intel compiler, but I never tried it, I have to admit).

psteinb

Ok, we'll leave size/2 for now.