Earlier quoted context omitted.
Yeah, Eleven Labs must be raking it in. You can get hours of audio out of it for free with Eleven Reader, which suggests that their inference costs aren't that high. Meanwhile, those same few hours of audio, at the exact same quality, would cost something like $100 when generated through their website or API, a lot more than any other provider out there. Their pricing (and especially API pricing) makes no sense, not…
Kokoro gives great results especially when speaking english. Model is small enough to run even on smartphone ~3x faster than realtime.
With the budget one tenth that of Stable Diffusion and less ethical qualms, you could easily 10x or 100x this.