Earlier quoted context omitted.
Bah. If they average thousands of downloads per file, there's plenty of room in there for some crawlers. And the total data set is less than a terabyte; seed a torrent somewhere for $20. The user-pays S3 bucket also exists as a good thing but S3 is much more expensive than data needs to be.
They started in 1991, when a terabyte was an unimaginably huge quantity of data and it was common for anonymous FTP servers like xxx.lanl.gov to request that you not connect until after business hours to avoid interfering with the main purpose of the machines. When I joined the internet in 1992, our 7.5-MHz VAX had a 56-kbps frame-relay link to New Mexico Technet (TECNET on our DECNET), which I think may also have pr…
The collection was a lot smaller then. The whole thing has fit on one hard drive for a long time. And as far as interpreting my post as criticism goes, apply it to the last ten years only.
> This is the context in which the arXiv's hostile stance toward spidering was established.
They should have reconsidered it at some point.
> series of torrents
Sure. A torrent for each 500MB chunk they already collate, or yearly, or both.