Purely out of curiosity, has anyone used this in conjunction with a linguistics study to visualise sounds made during speech? I did a quick literature search but couldn't find anything. Not sure whether if features would be too subtle to see anything interesting... Edit: Spoke too soon. Did find one paper[1] (though I can't read it) which uses it to compare the production of 's' and 'z' sounds. Would love to know if…
We tried to image sound waves, but the density gradient for sound is much less than that produced by changes in temperature. We attempted to make a resonant chamber and use a high-intensity ultrasonic source, and make a standing wave that we might capture photographically, and while we saw something once we could not produce it. It's likely that the ultrasonic source we were using was gradually degrading.
The research paper you link to almost surely uses the heat differences to see the jets of air coming out of the mouth, and not actual sound waves.