Every time I see people trying to eliminate branches, I wonder, do we realize that having long pipelines where a branch misprediction stalls the pipeline is not actually a necessary part of architecture? The pipeline is long because we do lots of analysis and translation on the fly, just in time, which could easily be done in most cases ahead of time, as it's not a very stateful algorithm. This is how Transmeta Cruso…
Transmeta's translation did not somehow eliminate branch costs.
In fact I distinctly remember a comp.arch thread and post from Linus himself while he worked at Transmeta that paraphrased was "the cpu's job is to generate cache misses as fast as possible."
Compulsory misses are a thing. No form of JIT is capable of eliminating them. Capacity misses are inevitable in the real world, even with the monster caches we have now.
Itanium thought it could eliminate branch costs with static analysis. How did that work out?
I really wish programmers would actually read a book about computer architecture before they so confidently conclude that it's simple to build things better than state of the art processors. You are underestimating the scale of smart that has gone into current processors by at least 7 orders of magnitude in my opinion.