Earlier quoted context omitted.
> but that's only because no one actually cares about these architectures. This. Optimizing performance on BE is a waste of everyone's time and resources.
Except, of course, for those small number of us who primarily do run on a BE platform. I don't mind doing this work, just let us do it.
(But the status quo on BE is that it does a load followed by a byte swap, which is probably pretty cheap anyway. The compiler might even already know how to optimize that into the appropriate LE-load instruction.)