Live data from Hacker News

An update on Wayback Machine access

blog.archive.org

311–320 of 374 posts

Re: An update on Wayback Machine access

#311
post #26

Earlier quoted context omitted.

Yeah cuz X is fucking shitty, why would I want to give them any traffic?

“It’s ok to do bad things to people/things I don’t like” Feels good when you get to dish it out doesn’t it?

if kidnapping is so bad why do we put the Unabomber in prison

Re: An update on Wayback Machine access

#312
post #111

Earlier quoted context omitted.

Yes, arguably, and for reasons I already gave. I’d genuinely spend a bit more time reading and thinking rather than replying. Your replies are pithy but you’re missing details and frankly making cognitive mistakes. (Apologies if this seems harsh, I don’t mean it as an insult, but this thread has blown up entirely unnecessarily- and yes, I know I’m not helping either!)

You seem to think that scraping websites "for the public good" is somehow different than scraping websites for any other reason . The end result is exactly the same.

I'm guessing you use search engines, right? Those use scrapers and have to use scrapers. It's how they work.

A "scraper" is simply an automated process that fetches URLs intended for display to a human, and processes it. The act of scraping doesn't imply anything about:

1. The frequency of the fetches,

2. The way that the resulting page is processed.

Search engines scrape. Again, they have to. Same goes for archive.org.

Thing is, there aren't tens of thousands of search engines/archive.orgs that can overload a site at once.

Re: An update on Wayback Machine access

#313

Earlier quoted context omitted.

What do you expect them to do though? You have to be a reasonable person.

Data analysis would be a good start, better blocking heuristics, an off-the-shelf solution used by other organisations that don't have this problem, etc.

Which other organisation is as valuable to scrape as IA?

Re: An update on Wayback Machine access

#314

Earlier quoted context omitted.

archive.today uses clients to perform DDoS, I would not recommend using their site.

[flagged]

This is well-known. If you are very sure it's not happening anymore, please provide a reputable source.

Re: An update on Wayback Machine access

#315
post #285

Earlier quoted context omitted.

Correct:Also, archive.* has actively edited archived sites to promote their agenda. Why folks continue to use them confuses me. One would think the big wikipedia purge would curb such behavior.

> Why folks continue to use them Because there's no working alternative.

unwall.app works for at least some sites.

Re: An update on Wayback Machine access

#316

Earlier quoted context omitted.

Wikipedia is not a challenge to the system. If it were, it would be excluded from search results, much like blogs are now.

This became obvious to me when from 2023-2025 they refused to call it anything other than "Israel-Hamas war" It's since been renamed to "Gaza genocide", as it should have been all along - but it took forever to get there. This is because of their policy of only mirroring what mainstream media outlets are saying.

It’s the only “genocide” in history where the population being genocided grew during their genocide.

Re: An update on Wayback Machine access

#317

Earlier quoted context omitted.

This became obvious to me when from 2023-2025 they refused to call it anything other than "Israel-Hamas war" It's since been renamed to "Gaza genocide", as it should have been all along - but it took forever to get there. This is because of their policy of only mirroring what mainstream media outlets are saying.

It’s the only “genocide” in history where the population being genocided grew during their genocide.

this is not correct

Re: An update on Wayback Machine access

#318

Earlier quoted context omitted.

People doing this say it makes things worse because then the bots download both.

Because not enough people do this earnestly, and many more do it maliciously (bot endpoints that lie, or provide significantly less information than people endpoints) or put it behind a business contract (yes, APIs), so the bots or agents can't trust it in general. Also let's not forget that innocent sites suffering from floods of scrapers are actually the minority here - this is just a special case; the main reason…

It's defensively, not maliciously. Malice would imply that the the site owner is morally obliged to serve the bots.

Re: An update on Wayback Machine access

#319
post #224
post #3

> Here’s what’s going on. The Internet Archive’s Wayback Machine has been hit by waves of high-volume automated traffic, and we’ve put protections in place to keep the service running. I'm pretty certain this is scrapers that are trying to workaround blocks on accessing original sites by hitting the Wayback Machine copy instead. Appalling behavior. In addition to the load it puts on this vital non-profit piece of Int…

My use of wayback has skyrocketed recently due to anti-bot measures. I often cannot get past captchas, and archive.org is one of the fallbacks I try. However, archive.is, etc are more reliable. I wish the internet archive acted more like a library system, where multiple organizations could mirror the content. They are a big single point of failure, and I’m shocked Trump/SCOTUS haven’t intentionally burnt the archives…

Ironically, archive.is itself has a captcha that doesn't like my home FF install.
Post reply on HN