Live data from Hacker News

An update on Wayback Machine access

blog.archive.org

211–220 of 374 posts

Re: An update on Wayback Machine access

#211
post #201

Earlier quoted context omitted.

> Why folks continue to use them confuses me. I use them. I haven't ever heard mention that the content is edited. Do you have a source?

See the "Background" section of the Wikipedia RFC on banning archive.today links: https://en.wikipedia.org/wiki/Wikipedia:Requests_for_comment... They bulk replaced one string (a name) with another one across many archived pages, and added malicious code to all archive pages that would rapidly send requests to gyrovague.com in an attempt to DDOS them.

Just out of curiosity, how do we know that what is in the Wikipedia comments is accurate? I have no skin in the game. I was just wondering. Anybody can post anything on Wikipedia comments. I find it odd that Ars Technica would use that as a source. Maybe it's fine for gossip and speculation but it shouldn't be in Ars Technica then.

Re: An update on Wayback Machine access

#212
post #45
post #35

AI companies should pay billions to wayback machine for access

I think that'd raise serious copyright concerns, if the Wayback machine started selling other people's intellectual property.

Isn't that already a big part of reddit's business model?

Re: An update on Wayback Machine access

#213

Earlier quoted context omitted.

The links are usually to Archive.today (aka archive.ph and a bunch of other domains), not Wayback Machine (which is run by Internet Archive).

Correct:Also, archive.* has actively edited archived sites to promote their agenda. Why folks continue to use them confuses me. One would think the big wikipedia purge would curb such behavior.

> Why folks continue to use them confuses me.

Seems like the “confused” is disingenuous when not paying for content you read is a clear motivation.

Re: An update on Wayback Machine access

#214
post #31

Earlier quoted context omitted.

Then... don't? If it sucks so much why do you need to read the tweets? Great take: "This private website is owned by a man I don't like, so I refuse to pay for it - or even give it the possibility to monetize my traffic with ads!" Still quite mainstream take: "... so I'll use an adblocker on it" Immature take: "This private website that I hate and boycott is also an important part of our culture, but the posts on it…

You're confusing the site for the content on it. Some of the content is good, the site sucks and is run by a guy who seig-heils crowds. Even if the content sucked, your post has big "you want to improve `X`, yet you participate in `X`"[0] energy. 0 – https://kitzy.com/content/assets/images/we-should-improve-so...

Twitter is not society. It's a privately-owned web site, nothing more. Always has been. Participating in it is giving your bogeyman power.

Re: An update on Wayback Machine access

#215

I've been getting this error a lot. Asking users to email them with details of their OS, browser, IP address is just crazy. Their support is supposedly already swamped and they are asking for more!? Changes made by IA shouldn't become my responsibility.

OTOH I do appreciate their openness. It beats those times when I try visiting a site, only to get a cryptic 403 error or similar and no suggestion the site would like to hear from me.

Re: An update on Wayback Machine access

#217
post #3

> Here’s what’s going on. The Internet Archive’s Wayback Machine has been hit by waves of high-volume automated traffic, and we’ve put protections in place to keep the service running. I'm pretty certain this is scrapers that are trying to workaround blocks on accessing original sites by hitting the Wayback Machine copy instead. Appalling behavior. In addition to the load it puts on this vital non-profit piece of Int…

I run some small websites, including a tiny forum that’s been a goldmine for scrapers. I had to significantly tweak some firewall rules and configuration after scrapers behind residential proxies suddenly accounted for over 99% of requests. However, I also relaxed rules for automated traffic that was well-behaved, and I went out of my way to ensure that the Wayback Machine was able to hit everything. I should kick a…

This is the way to go! Cut the cat and mouse, win win ish.

Re: An update on Wayback Machine access

#218
post #3

> Here’s what’s going on. The Internet Archive’s Wayback Machine has been hit by waves of high-volume automated traffic, and we’ve put protections in place to keep the service running. I'm pretty certain this is scrapers that are trying to workaround blocks on accessing original sites by hitting the Wayback Machine copy instead. Appalling behavior. In addition to the load it puts on this vital non-profit piece of Int…

I run some small websites, including a tiny forum that’s been a goldmine for scrapers. I had to significantly tweak some firewall rules and configuration after scrapers behind residential proxies suddenly accounted for over 99% of requests. However, I also relaxed rules for automated traffic that was well-behaved, and I went out of my way to ensure that the Wayback Machine was able to hit everything. I should kick a…

How do you separate the Wayback Machine from malicious bots that pretend to be the Wayback Machine? Are you whitelisting their IP blocks?

Because I get a ton of scraper requests that forge Googlebot, Bing, and Yandex user-agents that are totally not coming from their IP ranges. In fact, sometimes they all come from the same IP...

Re: An update on Wayback Machine access

#219

Earlier quoted context omitted.

I run some small websites, including a tiny forum that’s been a goldmine for scrapers. I had to significantly tweak some firewall rules and configuration after scrapers behind residential proxies suddenly accounted for over 99% of requests. However, I also relaxed rules for automated traffic that was well-behaved, and I went out of my way to ensure that the Wayback Machine was able to hit everything. I should kick a…

How do you separate the Wayback Machine from malicious bots that pretend to be the Wayback Machine? Are you whitelisting their IP blocks? Because I get a ton of scraper requests that forge Googlebot, Bing, and Yandex user-agents that are totally not coming from their IP ranges. In fact, sometimes they all come from the same IP...

I thought at least google (and possibly others) provided a way to verify the user agent?

Re: An update on Wayback Machine access

#220
post #211
post #201

Earlier quoted context omitted.

See the "Background" section of the Wikipedia RFC on banning archive.today links: https://en.wikipedia.org/wiki/Wikipedia:Requests_for_comment... They bulk replaced one string (a name) with another one across many archived pages, and added malicious code to all archive pages that would rapidly send requests to gyrovague.com in an attempt to DDOS them.

Just out of curiosity, how do we know that what is in the Wikipedia comments is accurate? I have no skin in the game. I was just wondering. Anybody can post anything on Wikipedia comments. I find it odd that Ars Technica would use that as a source. Maybe it's fine for gossip and speculation but it shouldn't be in Ars Technica then.

Because a lot of us watched the drama unfold in real time.
Post reply on HN