| | nJ/insn |
| MSP430 | 0.9 |
| PIC24 | 2? |
| 1990s StrongARM | 1 |
| LPC1110 | 0.3 |
| Pentium | 10 |
| STM32L0 | 0.23 |
| Ickes DSP 2008 | 0.01 |
| Subliminal 2006 | 0.0026 |
MSP430, PIC24: http://www.ti.com/general/docs/lit/getliterature.tsp?baseLit...StrongARM: http://www.researchgate.net/profile/Kristofer_Pister/publica...
LPC1110: http://www.nxp.com/documents/data_sheet/LPC111X.pdf
Pentium: http://www.newscientist.com/blog/technology/2006/08/explodin...
Ickes DSP 2008: http://www-mtl.mit.edu/researchgroups/icsystems/pubs/confere...
Subliminal 2006: http://web.eecs.umich.edu/~taustin/papers/VLSI06-sublim.pdf 2009: https://web.eecs.umich.edu/~taustin/papers/TVLSI09-sublimina...
This is sort of comparing apples to oranges. The Pentium (all of them) uses wildly varying amounts of power for different instructions, and has 32-bit or 64-bit instructions, with hardware floating point. The STM32L0, LPC1110, and StrongARM are 32-bit processors with no hardware floating point. The MSP430 is a 16-bit CPU, while the PIC24 is an 8-bit CPU. The Ickes et al. device includes a 16-bit FFT accelerator and only runs at 4MHz, but it was only fabricated as a prototype; you can't buy it. The Zhai et al. Subliminal device, also only fabricated as a prototype, only runs at 833kHz, and it doesn't even include an integer multiply instruction, but its somewhat limited ALU is 32 bits.
However, _all_ of these numbers are far from the Landauer bound (kT ln 2). Suppose that we don't make any concessions to reversibility in our CPU design, like Metronome and its successors. A 32-bit instruction, then, erases the 32-bit register where its result is stored, costing 32 kT ln 2, or in some situations, an average of 16 kT ln 2. Supposing T = 300 K, 32 kT ln(2) ≈ 0.092 attojoules. That's _over seven orders of magnitude_ better than the prototype Subliminal processor mentioned above, and nine orders of magnitude better than the Cortex-M0-based commercial processors mentioned above.
Wolpert has published https://arxiv.org/pdf/1806.04103.pdf (mentioned in the article, but not linked AFAICT) which gives better expressions for the cost of computation than the Landauer bound. I haven't finished reading it, but it looks really interesting.