Earlier quoted context omitted.
>People don't need self-help advise, they need a fair redistribution of increased productivity. The increased productivity is pretty much entirely coming from AI researchers and the companies investing in huge amounts of GPUs, and they are the ones receiving most of the windfalls, how's that not fair?
The biggest contribution is still from the training set, whose original authors get 0 because of "fair use" in the copyright.
My thinking here is coming from the paper "From Entropy to Epiplexity"[1] which partly discussed why you can train on synthetic data: it's the structure of the data that enables learning, not just the amount of "information". Authors of images and videos may have worked just as hard as authors of text training sets, but they didn't contribute to AI as much because there just wasn't the same kind of structure to discover there. It's the people who found the usable structure, not the people who accidentally generated it, who created the value.