Live data from Hacker News

An update on Wayback Machine access

blog.archive.org

271–280 of 374 posts

Re: An update on Wayback Machine access

#272

Earlier quoted context omitted.

How do you separate the Wayback Machine from malicious bots that pretend to be the Wayback Machine? Are you whitelisting their IP blocks? Because I get a ton of scraper requests that forge Googlebot, Bing, and Yandex user-agents that are totally not coming from their IP ranges. In fact, sometimes they all come from the same IP...

They have their own ASN, which I’ve explicitly allowed requests from. https://www.peeringdb.com/asn/7941

TIL they have their own ASN. This is helpful, thanks.

Re: An update on Wayback Machine access

#273
post #3

> Here’s what’s going on. The Internet Archive’s Wayback Machine has been hit by waves of high-volume automated traffic, and we’ve put protections in place to keep the service running. I'm pretty certain this is scrapers that are trying to workaround blocks on accessing original sites by hitting the Wayback Machine copy instead. Appalling behavior. In addition to the load it puts on this vital non-profit piece of Int…

I run some small websites, including a tiny forum that’s been a goldmine for scrapers. I had to significantly tweak some firewall rules and configuration after scrapers behind residential proxies suddenly accounted for over 99% of requests. However, I also relaxed rules for automated traffic that was well-behaved, and I went out of my way to ensure that the Wayback Machine was able to hit everything. I should kick a…

Kudos on that. I don't know what site you're hosting, but I appreciate knowing that it is there

Re: An update on Wayback Machine access

#276

Earlier quoted context omitted.

archive.today uses clients to perform DDoS, I would not recommend using their site.

[flagged]

It's been posted multiple times of the last few months as they were removed from wikipedia and cloudflare. [1] [2] [3]

1. https://news.ycombinator.com/item?id=46843805

2. https://news.ycombinator.com/item?id=47092006

3. https://news.ycombinator.com/item?id=47474255

Re: An update on Wayback Machine access

#277

Earlier quoted context omitted.

I’ve read a few different experiences with hosts having had success with Anubis to cut down on excessive scraping. One that comes to mind is the Dolphin project.[1] I’ve read a few of those threads, but often it’s people at cross-purposes with the goals of Anubis. Is there a chance you could clarify the “not working” bit? [1] https://dolphin-emu.org/blog/2025/06/04/dolphin-progress-rep...

> Is there a chance you could clarify the “not working” bit? I see anubis, 90% of the time I close the tab before it finishes.

But why?

Re: An update on Wayback Machine access

#278
post #211

Earlier quoted context omitted.

Just out of curiosity, how do we know that what is in the Wikipedia comments is accurate? I have no skin in the game. I was just wondering. Anybody can post anything on Wikipedia comments. I find it odd that Ars Technica would use that as a source. Maybe it's fine for gossip and speculation but it shouldn't be in Ars Technica then.

Because a lot of us watched the drama unfold in real time.

We have always been at war with Eastasia.

Re: An update on Wayback Machine access

#279
post #259

Earlier quoted context omitted.

Is scale what we're discussing though? e.g. a prompt of "fetch and summarise it for me" is very close to what a human would be doing with a web browser, and doesn't seem to involve any kind of scaling issue.

Sure, but all the time I'll ask Claude a question, and then I'll see it fetch 5-10 different URLs to come up with answer. I certainly would not be fetching those URLs at that rate if I were doing it myself. I would probably be visiting those pages, one by one, over the span of 10-20 minutes. That's the scale argument.

As would I when researching anything myself. I'll do a web search, and if I see some highly relevant results, I'll middle-click them so they open in a new tab, and I'll easily do 5+ at a time, before then going to read the first one.

Same with browsing HN, btw. I have a row of 9 HN tabs open, all of them opened at the same time, as I scrolled the front page and middle-clicked on thread link to anything interesting.

Post reply on HN