Earlier quoted context omitted.
I worked on a large scale project that was itanium based and let me tell you, the compilers / software toolchains were an absolute dumpster fire. We were planning on being between 500 and 2000 processors in our cluster. We used HP boxes and I want to say it was around 2002/2003 when this was going on. We were supposed to be a huge public showpiece client for both intel and HP. It… did not go very well. I remember the…
That's the tragedy behind it, since the whole premise on how Itanium was done was that compilers will be able to, eventually, any day now, optimize for it and then you'll see, just wait and then you'll see.. any day now!
All that Itanium was, at the end of the day, was an in-order PA-RISC/SPARC hybrid that exposed a lot of the superscalar innards to the programmer (compiler)
VLIW scheduling even back then wasn't as much of a mystery as the usenet and register flamewars implied. It really is pretty straightforward for a compiler to behave like an in-order superscalar scheduler. And since the compiler has a much global information about the instruction stream it can do much more optimizations and static ordering than a normal in-order HW superscalar scheduler alone.
Plus itanium had plenty of dynamic microarchitectural components to complement the static ordering done by the compiler and increase FU utilization.
If anything, things like predication and its adverse effect in power consumption had a much worse impact on itanium than the compilers.
What killed the itanium was just simple economics; its performance was fine (for the time, at least for itanium2). It's performance/price ratio, however, was not.