Earlier quoted context omitted.
at this point we're _good data_ limited, which has little to do with scraping.
Why kind of data that isn’t public would be so valuable for AI training? Seems like there’s a fuck ton. All of Wikipedia, GitHub for code, etc. I can understand targeting certain sites like Reddit, etc. but not random websites
If you look closely even Google does this. This is probably why many popular sites started getting down ranked in the last 2 years. Now they're below the fold and Google can present their content as their own through the AI box.