Is this such a big problem? You could still scrape all the data, or not?
For example, my hobby search engine got started because I found out about these dumps and decided it would be an interesting challenge to try to work with them[1]. If I’d needed to build a scraper first the project would never have gotten off the ground.
[1]: https://search.feep.dev/blog/post/2021-09-04-stackexchange