Live data from Hacker News

An update on Wayback Machine access

blog.archive.org

321–330 of 374 posts

Re: An update on Wayback Machine access

#321
post #17

I really appreciate the Archive team's efforts to make the Wayback Machine more responsive, and have donated a few times to support them. Unfortunately, the restrictions have been way too strict for the last few months: from my residential IP, simply moving the mouse on the calendar for a specific URL is enough to get stuck on 429 error messages for a while; and from corporate ISPs (for example, on airport WiFi), you…

Generally speaking I feel like detecting the bots might be a lost cause. For someone like the Internet Archive I don't know how to deal with it, for smaller sites, cache everything, static pages whenever possible.

Sadly I see rate-limiting usage in general becoming a thing. With residential proxies and more sophisticated bots either pretending to be Chrome or directly piloting Chrome, it's going to become impossible to tell a real user from a bot. Only solution is to pretend that everyone is a bot and design for it.

Re: An update on Wayback Machine access

#322
post #17

I really appreciate the Archive team's efforts to make the Wayback Machine more responsive, and have donated a few times to support them. Unfortunately, the restrictions have been way too strict for the last few months: from my residential IP, simply moving the mouse on the calendar for a specific URL is enough to get stuck on 429 error messages for a while; and from corporate ISPs (for example, on airport WiFi), you…

I wonder if your browser is prefetching every link you move the mouse over.

I don't think so, but moving the mouse over any date in the calendar makes a request for the list of snapshots taken on that day.

Re: An update on Wayback Machine access

#323
post #118

The anti virus industry needs to crack down on crawler and proxy malware, plus ISPs FINALLY need to replace CGNATs with iov6 to stop crawlers banning everyone behind a NAT.

Anti-virus, maybe, the ISP definitively needs to step up and just shut off people internet when large amounts of bot traffic is detected. IPv6 is going to do nothing, because right now you're getting scrapped/attacked/DDoS/whatever with a single request from millions of IPs at once.

Re: An update on Wayback Machine access

#324
post #17

I really appreciate the Archive team's efforts to make the Wayback Machine more responsive, and have donated a few times to support them. Unfortunately, the restrictions have been way too strict for the last few months: from my residential IP, simply moving the mouse on the calendar for a specific URL is enough to get stuck on 429 error messages for a while; and from corporate ISPs (for example, on airport WiFi), you…

I wonder if your browser is prefetching every link you move the mouse over.

I have had similar experiences just opening archived pages with a couple of embedded images. It's at a level where just using the site normally is painful.

Re: An update on Wayback Machine access

#325

Earlier quoted context omitted.

It’s the only “genocide” in history where the population being genocided grew during their genocide.

this is not correct

Are you disputing that Gaza had more births than reported deaths during this period, or do you have other examples of “genocides” where the population grew during their genocide?

Re: An update on Wayback Machine access

#326

Shouldn't the solution be to gate bulk access for automated services for a price? Not just the wayback machine, any personal or corporate website? My blogs are getting slammed and there are issues with cloudflare or captchas.

> Shouldn't the solution be to gate bulk access for automated services for a price?

The problem is that many of the people who are scraping this data doesn't want to pay. These are organisations who would rather not clone your git repo, and instead scrape every single page on your Forgejo installation. These are NOT nice people.

Re: An update on Wayback Machine access

#328

Earlier quoted context omitted.

From a few weeks ago: https://news.ycombinator.com/item?id=49500040 Basically, the cost of an optimized solution is orders of magnitude lower than the cost of an in-browser solution. Anyone dedicated can easily afford to solve workloads higher than your users will tolerate. You might stop casual scrapers, but you're not going to stop someone who cares. AI scraping companies care.

You're assuming that the part of Anubis that stops bots is the PoW. It's not.

The other parts are even more trivial for someone who cares to work around.

Re: An update on Wayback Machine access

#329
post #97

I've been getting this error a lot. Asking users to email them with details of their OS, browser, IP address is just crazy. Their support is supposedly already swamped and they are asking for more!? Changes made by IA shouldn't become my responsibility.

> Asking users to email them with details of their OS, browser, IP address is just crazy. It's surely to serve as data to help tell humans apart from bots. > Changes made by IA shouldn't become my responsibility. They're a free service. It's ultimately not their responsibility to service you either.

> They're a free service. It's ultimately not their responsibility to service you either.

They are a nonprofit with a mission and continue to solicitate donations based on that mission.

Re: An update on Wayback Machine access

#330

Earlier quoted context omitted.

You're assuming that the part of Anubis that stops bots is the PoW. It's not.

The other parts are even more trivial for someone who cares to work around.

"someone who cares" is the part that creates the firewall.
Post reply on HN