Not that the article does this, but since my current bandwagon is that assessing performance by comparing the source of two programs without consideration of the compiler and target processor is silly, I decided to try out the C++ version with several compilers and options. Renaming the file to 'lpath.cpp' and compiling with 'cc lpath.cpp -std=gnu++11 -Wall -Oxxx -march=native -o lpath-cc-Oxxx' here's what I found an an i7 Haswell running at 3.4 GHz:
lpath-clang3.5-O2 8981 LANGUAGE C++ 763
lpath-gcc4.7-O2 8981 LANGUAGE C++ 769
lpath-gcc4.8-O2 8981 LANGUAGE C++ 746
lpath-icpc14-O2 8981 LANGUAGE C++ 750
lpath-icpc15-O2 8981 LANGUAGE C++ 735
lpath-clang3.5-O3 8981 LANGUAGE C++ 734
lpath-gcc4.7-O3 8981 LANGUAGE C++ 943
lpath-gcc4.8-O3 8981 LANGUAGE C++ 946
lpath-icpc14-O3 8981 LANGUAGE C++ 664
lpath-icpc15-O3 8981 LANGUAGE C++ 655
The last column is the time reported by the program in milliseconds. What this shows is that the same source compiled with Intel's icpc 14 or 15 -O3 is about 50% faster than the same source compiled with g++ 4.7 or 4.8 and -O3, and about 20% faster against g++ -O2. The point isn't that Intel's compiler is so much better, but that this degree of variation is normal for a benchmark like this. There are times when Clang or GCC will come out ahead by the same margin. The lesson is that you aren't benchmarking source code, you are benchmarking a particular compiler with particular options running on a particular processor.
The article is wonderfully specific about what was used, but one should be very careful extrapolating to different combinations. In addition to the compiler differences, note for example that although I did my tests on a processor running less than 1.5x faster, I got runtimes that were almost 2-3x faster. Most likely, this is because the Haswell processor I tested on is more efficient than the several generation old Westmere that the author used. There's nothing right or wrong about either choice, but the degree of difference is why it's always important to specify.
I glanced briefly at the code with 'perf record -F10000', and my quick conclusion was the the Intel version was running faster because it was making better use of the branchless cmov's than the other compilers, and thus has 10,000,000 fewer branch prediction errors. At 20 cycles per miss, this accounts for over half the difference between icpc and the others. The use of the new fast variable shift instructions (shlx) and the once-again fast bit test (bt) instruction is probably the rest of the difference.
The difference between g++ -O2 and -O3 seems to be that the -O3 version is doing almost everything off the stack rather than in registers. It's bad enough that this is probably a performance bug rather than intended behavior.
Measurement is highly specific -- the time taken for this benchmark task, by this toy program, with this programming language implementation, with these options, on this computer, with these workloads.
Same toy program, same computer, same workload -- but much slower.
The article is wonderfully specific about what was used, but one should be very careful extrapolating to different combinations. In addition to the compiler differences, note for example that although I did my tests on a processor running less than 1.5x faster, I got runtimes that were almost 2-3x faster. Most likely, this is because the Haswell processor I tested on is more efficient than the several generation old Westmere that the author used. There's nothing right or wrong about either choice, but the degree of difference is why it's always important to specify.
I glanced briefly at the code with 'perf record -F10000', and my quick conclusion was the the Intel version was running faster because it was making better use of the branchless cmov's than the other compilers, and thus has 10,000,000 fewer branch prediction errors. At 20 cycles per miss, this accounts for over half the difference between icpc and the others. The use of the new fast variable shift instructions (shlx) and the once-again fast bit test (bt) instruction is probably the rest of the difference.
The difference between g++ -O2 and -O3 seems to be that the -O3 version is doing almost everything off the stack rather than in registers. It's bad enough that this is probably a performance bug rather than intended behavior.