P4 was such a curious thing in it's own right; a lot of ambition that was perhaps too forceful.

Hell, if you -could- keep the pipeline from mispredicting and fed with data, one or two of it's internal ALUs actually ran at 2x the main CPU clock. Alas, that's an even bigger ask than adding SSE2 branching, and they decided to do RDRAM (Which, AFAIR was worse for overall latency than SDR or DDR)