Live data from Hacker News

An update on Wayback Machine access

blog.archive.org

301–310 of 374 posts

Re: An update on Wayback Machine access

#301

Earlier quoted context omitted.

I’ve read a few different experiences with hosts having had success with Anubis to cut down on excessive scraping. One that comes to mind is the Dolphin project.[1] I’ve read a few of those threads, but often it’s people at cross-purposes with the goals of Anubis. Is there a chance you could clarify the “not working” bit? [1] https://dolphin-emu.org/blog/2025/06/04/dolphin-progress-rep...

From a few weeks ago: https://news.ycombinator.com/item?id=49500040 Basically, the cost of an optimized solution is orders of magnitude lower than the cost of an in-browser solution. Anyone dedicated can easily afford to solve workloads higher than your users will tolerate. You might stop casual scrapers, but you're not going to stop someone who cares. AI scraping companies care.

You're assuming that the part of Anubis that stops bots is the PoW. It's not.

Re: An update on Wayback Machine access

#302

There was a feature on Amazon Web Services for a while, and I wish it was still there... Downloader pays. I make some content and upload it. When you want to download it, you pay Amazon the egress fees. And maybe I get to charge just a bit more, to help me with the Ingress, storage, content creation, etc. I mean, I know that there's going to be problems with rate limiting, etc. And yes, we have those problems with LL…

Note the Amazon egress fee is one hundred times anywhere sane's egress fee.

Re: An update on Wayback Machine access

#304
post #52
post #45

Earlier quoted context omitted.

I think that'd raise serious copyright concerns, if the Wayback machine started selling other people's intellectual property.

Shouldn’t it follow that it’s illegal for the AI labs to profit off of all of that stolen copyrighted data too?

It should, but it doesn't.

Re: An update on Wayback Machine access

#305

Earlier quoted context omitted.

Micropayments would solve all the problems except for the problem that people absolutely loathe micropayments. Like, vein-popping furiously hate them. Whenever the topic of micropayments for internet content comes up, a bunch of people start talking about payment processors and their floor on prices, and so on. That's not wrong, but it can be designed around and I think it's a scapegoat to avoid confronting the fact…

Micropayments would solve so many problems for the internet. And, cryptocurrencies would solve so many problems for micropayments. But, it's a non-starter because any proposal gets flooded with people popping veins about how crypto can't solve anything.

If it's so easy to solve, go solve it. Set up a test site with micropayments. Maybe scrape CNN and see how many people will pay you micro for a copy of CNN.

Re: An update on Wayback Machine access

#306

Earlier quoted context omitted.

It doesn't need to be crypto, or payment processor based. My ideal experience would be I load $20 into the browser somewhere like a wallet in one block (that could be a payment processor step). If I visit a participating page, it decrements my wallet $.01 or whatever. The downside is the possibility of abuse and tracking by governments, which would have to be handled at the source, not the symptom.

I worked on this before. You can solve the privacy issue using "statistical payments", by lack of a better word. You visit website A,A,A,B,C,D,A,A At the end of the month, you send your entire 20$ randomly to one of the websites you visited. This will level out everyone's contribution and reward websites with lots of traffic. It eliminates the need for micropayments.

So I set up my own website and put a book on the F5 key to nearly guarantee I get my own $20 back. I also hotlink images from it to Reddit.

You can look into international call termination fee fraud in the public telephone network, for more on this.

Re: An update on Wayback Machine access

#307

Earlier quoted context omitted.

Once upon a time some people explored backing up the Internet Archive. However, that experiment ended. They mention there were some learnings and they then say: > The Internet Archive continues to explore methods and code to decentralize the collection, to have a mirror running in various ways - these include IPFS, FileCoin, and others. The INTERNETARCHIVE.BAK project also added general mirroring and tracking code to…

The Internet Archive's torrents are a sick joke. I've yet to find one that actually manages to complete. They always get stuck at 90-something percent but that final blocks always fail verification and get retried, fail, and the process repeats forever. Because they're web seeds they're hitting IA infrastructure and not offloading to a real swarm. So their broken torrents are just screwing themselves.

Interesting, I haven’t encountered this problem yet.

Re: An update on Wayback Machine access

#308
post #22

Earlier quoted context omitted.

My guess? Because even with a paid endpoint, the type of unscrupulous yahoo that is DDOSing IA today would probably still abuse the free endpoints because they can. The revenue that might come from a paid endpoint could help to scale up, but with how slow IA usually seems, I suspect there is an upper limit to how much traffic they can serve without a LOT more revenue. This is a major "this is why we can't have nice t…

IA might be large enough to earn consideration, but generally scrapers just don't care about being good citizens. I work in the GLAM space and we offer OAI-PMH interfaces for the harvesting of our collections data - which doesn't stop companies from preferring to scrape our website for worse (less complete, less structured, less standardized) data instead.

Do they know it exists? On my site, some types of blocked bots are getting plain-text instructions saying why I'm blocking them and what they can do instead - and it seems to have worked in some cases.

Re: An update on Wayback Machine access

#309
post #66
post #64

Earlier quoted context omitted.

What if you're not charging for the content, but as compensation for the network bandwidth / server resources consumed by serving that content? The idea isn't to profit from content (the IA is a nonprofit anyway), just to allow the IA to continue to serve its purpose as an archive of public data without being overwhelmed by bots.

Exactly, it's a question for the lawyers to sort out.

And they're likely on risky enough ground after the book lending thing

Re: An update on Wayback Machine access

#310
post #53

Earlier quoted context omitted.

No one is upset that the AI companies are scraping the web, they’re upset how poorly implemented the scrapers, but the scraping itself is fine. Lots of people and businesses scrape the web.

There are people commenting parallel to you saying they are upset about AI companies scraping the web.

Because of the request load though. The ethical thing is separate.
Post reply on HN