Earlier quoted context omitted.
Yeah cuz X is fucking shitty, why would I want to give them any traffic?
“It’s ok to do bad things to people/things I don’t like” Feels good when you get to dish it out doesn’t it?
An update on Wayback Machine access
311–320 of 374 posts
Re: An update on Wayback Machine access
#312Earlier quoted context omitted.
Yes, arguably, and for reasons I already gave. I’d genuinely spend a bit more time reading and thinking rather than replying. Your replies are pithy but you’re missing details and frankly making cognitive mistakes. (Apologies if this seems harsh, I don’t mean it as an insult, but this thread has blown up entirely unnecessarily- and yes, I know I’m not helping either!)
You seem to think that scraping websites "for the public good" is somehow different than scraping websites for any other reason . The end result is exactly the same.
A "scraper" is simply an automated process that fetches URLs intended for display to a human, and processes it. The act of scraping doesn't imply anything about:
1. The frequency of the fetches,
2. The way that the resulting page is processed.
Search engines scrape. Again, they have to. Same goes for archive.org.
Thing is, there aren't tens of thousands of search engines/archive.orgs that can overload a site at once.
Re: An update on Wayback Machine access
#313Earlier quoted context omitted.
What do you expect them to do though? You have to be a reasonable person.
Data analysis would be a good start, better blocking heuristics, an off-the-shelf solution used by other organisations that don't have this problem, etc.
Re: An update on Wayback Machine access
#314Re: An update on Wayback Machine access
#315Earlier quoted context omitted.
Correct:Also, archive.* has actively edited archived sites to promote their agenda. Why folks continue to use them confuses me. One would think the big wikipedia purge would curb such behavior.
> Why folks continue to use them Because there's no working alternative.
Re: An update on Wayback Machine access
#316Earlier quoted context omitted.
Wikipedia is not a challenge to the system. If it were, it would be excluded from search results, much like blogs are now.
This became obvious to me when from 2023-2025 they refused to call it anything other than "Israel-Hamas war" It's since been renamed to "Gaza genocide", as it should have been all along - but it took forever to get there. This is because of their policy of only mirroring what mainstream media outlets are saying.
Re: An update on Wayback Machine access
#317Earlier quoted context omitted.
This became obvious to me when from 2023-2025 they refused to call it anything other than "Israel-Hamas war" It's since been renamed to "Gaza genocide", as it should have been all along - but it took forever to get there. This is because of their policy of only mirroring what mainstream media outlets are saying.
It’s the only “genocide” in history where the population being genocided grew during their genocide.
Re: An update on Wayback Machine access
#318Earlier quoted context omitted.
People doing this say it makes things worse because then the bots download both.
Because not enough people do this earnestly, and many more do it maliciously (bot endpoints that lie, or provide significantly less information than people endpoints) or put it behind a business contract (yes, APIs), so the bots or agents can't trust it in general. Also let's not forget that innocent sites suffering from floods of scrapers are actually the minority here - this is just a special case; the main reason…
Re: An update on Wayback Machine access
#319> Here’s what’s going on. The Internet Archive’s Wayback Machine has been hit by waves of high-volume automated traffic, and we’ve put protections in place to keep the service running. I'm pretty certain this is scrapers that are trying to workaround blocks on accessing original sites by hitting the Wayback Machine copy instead. Appalling behavior. In addition to the load it puts on this vital non-profit piece of Int…
My use of wayback has skyrocketed recently due to anti-bot measures. I often cannot get past captchas, and archive.org is one of the fallbacks I try. However, archive.is, etc are more reliable. I wish the internet archive acted more like a library system, where multiple organizations could mirror the content. They are a big single point of failure, and I’m shocked Trump/SCOTUS haven’t intentionally burnt the archives…