To me "tap to pay" encodes immediate-term incidentality and proximity.
It's three syllables, it rolls off the tongue when you say it, it takes maybe 450ms to speak, and the brain can encode it using similar (or maybe the same?) mechanisms to how brand names and word combinations are captured without parsing/questioning when very young. This is one of many factors that contributes to a sort of "flow" that combats the "oooooooooh that sounds complicated"-of-death that stands to kill new complex digital products that need to be adopted en mass to function.
"Tap" also encodes "move near reader" but does so by suggesting that you move too close. Thus you'll either have people moving well within the active area (at which point the transaction may even be able to start and complete by the time it's been tapped). It's also physically easier for me to physically execute "tap object against other object" than "hover object 1cm in front of other object", especially when moving.
If you want something to scale, you need fail-safe design. My local bus transit system has a giant (but featureless) NFC pad with a screen saying "Tap here " above it. The number of people I see tapping the screen is... the screen should obviously say "below", the active area should have a ring of LEDs around it, etc etc etc; it's broken design. However, I've also seen people doing the same thing (tapping the screen) with payment terminals. In situations like this, you're designing the system to be viable for the dumbest user.