Very interesting. I wonder how it would cope with some noise -- for example playing the resultant file and recording it with a microphone for example. I suspect there will be content lost at the bottom and the top of the image depending on the frequency response of the microphone/speaker.
This uses a linear frequency scale (which is just the nature of Fourier transforms), whereas our ears are sensitive on a log frequency scale. In other words, the information that's most important to our hearing, which is what a mic & speaker will preserve the best, is in the bottom 10% of the image.
A cheap speaker & mic will probably lose a lot of content above about 10 kHz - which is the entire top half of the image. Even though this wouldn't be that huge a difference to our ears, it would sure look bad in the image.
As for background noise, the difference would probably look like the difference here: http://www.sweetwater.com/insync/media/2010/09/RXAdv-e-xlarg... (that's a screenshot of audio restoration software that removes noise, so it's technically doing the opposite process as best it can, but the difference would be similar).