Earlier quoted context omitted.
Single-mic source separation is possible in an unsupervised manner today that could probably work better than beamforming both compute-wise and with regard to implementation difficulty (you'd just need a lot of recordings to represent the space of sounds you want to separate).
Do you have any references for this or a link to a commercial service? I'm currently in the process or trying to extract some background voices in a video (an interview where the faint background conversation is in English and the loud overdub is in Bulgarian). I tried Melodyne but it seems to only separate on pitch, not volume, and the pitchs are too similar (mono, three voices, all female) and words are made of lot…
I haven't seen it provided as a commercial service or free model yet, but there is open source code for Mixit that lets you train using the open source / canned FUSS dataset.