Their solution is to use a very long array (i.e., a towed array) of microphones (i.e., transducers), and perform algorithmic beam-forming to narrow down the possible points of origin for the sound.
This seems like a pretty obvious approach, but I haven't noticed anyone proposing it.