Earlier quoted context omitted.
This is an interesting take, and I'd guess that the training data for this probably did use podcasts as a source. Getting very realistic / real world conversational training data for an ai would be hard. Only a subset of us appear on podcasts, radio or tv and probably all speak in a slightly artificial manner when we do.
I agree, I thinks it's probably very easy to find billions of hours of conversation on YouTube, but non of it is set to training data with a good transcript.
AI agents like this are trying to recreate personal intimacy I guess, which does feel like it might be different somehow.