Thanks for that... didn't notice that we had a repeated section (I just updated it with the proper text for that section)
As for comparison to the Phi... it is the fact that our actual core size (and thus the whole chip) is MUCH smaller. The Phi's cores are actually based on the original Pentium architecture (they are just downsized P54Cs) with added AVX instructions. Contrary to popular belief, they are not full x86-64 cores.
Intel has not released the official die size for the Phi, but has said it is about ~5 billion transistors (at a 22nm process), and independent "guesstimates" have pegged the die at 600-700mm2 (there is one place that says 350mm2, but that is false). For the top of the line 61 core Xeon Phi, it uses 300W, with a theoretical peak performance of 1.2TFLOP of double precision performance. That gets you to about 4GFLOP/Watt.
In comparison, our entire will be under 100 million gates, with each core (excluding memory) being around 100k gates. At a 28nm process, our core size (without memory) is a little under 0.1mm2. Our theoretical peak performance per compute chip is 256 GFLOPs double precision, while it should be using around 3 Watts, giving us a performance per watt ratio of ~85GFLOPs/Watt.
Intel has even said that their next generation Xeon Phi, made at their 14nm process, will be at 14 to 16GFLOPs/Watt. At SC14 last week, they made a soft announcement for the following generation at 10nm process will be in the ~2018 timeframe, and that is estimated at only being around 25GFLOPs/Watt.
The bottom line is Intel is just following Moore's law, and is sticking to big and complex systems, which while retaining legacy compatibility, kill you when it comes to efficiency.