Live data from Hacker News

An update on Wayback Machine access

blog.archive.org

331–340 of 374 posts

Re: An update on Wayback Machine access

#331

Earlier quoted context omitted.

this is not correct

Are you disputing that Gaza had more births than reported deaths during this period, or do you have other examples of “genocides” where the population grew during their genocide?

https://en.wikipedia.org/wiki/Brandolini%27s_law

Account created 1 hour ago and only made these two comments, no others. Dang should investigate where these bots are coming from.

Re: An update on Wayback Machine access

#332

There was a feature on Amazon Web Services for a while, and I wish it was still there... Downloader pays. I make some content and upload it. When you want to download it, you pay Amazon the egress fees. And maybe I get to charge just a bit more, to help me with the Ingress, storage, content creation, etc. I mean, I know that there's going to be problems with rate limiting, etc. And yes, we have those problems with LL…

Note the Amazon egress fee is one hundred times anywhere sane's egress fee.

My desired usage pattern stands... Someone who publishes content shouldn't be punished for everyone else wanting to access it, and shouldn't have to resort to product placement, advertising, sponsorship, or begging to fund it.

I don't know, maybe WebTorrent should have been the answer? For upcoming, viral content?

But for the deep archives, like the Wayback Machine? I feel like I'd happily pay for egress, and a bit to support them. If it was automatic and built in...

I wish Flattr or something like it had thrived...

Re: An update on Wayback Machine access

#333

Earlier quoted context omitted.

If you're donating, then you are evidently willing to pay without any "buts". So aren't you arguing against your own actions?

Not at all. I donate to Internet Archive, but we're talking here about paying for unobstructed access to but one part of their service: Wayback Machine. Two different things.

100% of the people who write "I would pay, if..." or "I would pay, but..." are people who are never going to pay even a dime. You might be the exception, and sorry for bunching you up with them. You have paid already by donation.

I think that we should all pay when asked for things which we find useful, even if they aren't perfect. If nobody else is offering anything, then we have to take what's being offered. When there's a market, more providers will begin offering their versions.

Re: An update on Wayback Machine access

#334
post #17

I really appreciate the Archive team's efforts to make the Wayback Machine more responsive, and have donated a few times to support them. Unfortunately, the restrictions have been way too strict for the last few months: from my residential IP, simply moving the mouse on the calendar for a specific URL is enough to get stuck on 429 error messages for a while; and from corporate ISPs (for example, on airport WiFi), you…

Generally speaking I feel like detecting the bots might be a lost cause. For someone like the Internet Archive I don't know how to deal with it, for smaller sites, cache everything, static pages whenever possible. Sadly I see rate-limiting usage in general becoming a thing. With residential proxies and more sophisticated bots either pretending to be Chrome or directly piloting Chrome, it's going to become impossible…

we’ve been using Datadome at work for this because it’s nice to get someone else to think about the constant bot cat and mouse, and we all benefit from rules and fixes created from other client data. Not an ad- that service is eye watering expensive but I think it makes sense as something to offload.

Re: An update on Wayback Machine access

#335

Earlier quoted context omitted.

Generally speaking I feel like detecting the bots might be a lost cause. For someone like the Internet Archive I don't know how to deal with it, for smaller sites, cache everything, static pages whenever possible. Sadly I see rate-limiting usage in general becoming a thing. With residential proxies and more sophisticated bots either pretending to be Chrome or directly piloting Chrome, it's going to become impossible…

we’ve been using Datadome at work for this because it’s nice to get someone else to think about the constant bot cat and mouse, and we all benefit from rules and fixes created from other client data. Not an ad- that service is eye watering expensive but I think it makes sense as something to offload.

We use spur.us which provides us with information on IPs they believe is running proxies. It's a nice service, in terms of pricing it's completely reasonable, for our use case. I can't imagine how much work goes into compiling their data.

Re: An update on Wayback Machine access

#336

Mad props to the people at the Archive. You are the heros we need in a formerly open internet that is surrendering to evil big corps and closing down free access. The Internet Archive is in a really bad spot being attacked from multiple sides at once. But — while service has not been consistent — they have maintained open access. I can still access anonymously from Tor without Cloudflare or some other centralized gat…

I donate to them every year b/c I fully agree they’re doing a thankless critical job very well.

I have a recurring donation, there is so much cool stuff in the archive. It really is the library of the internet, and I love it.

Re: An update on Wayback Machine access

#337

Earlier quoted context omitted.

Gatekeeping information is not the solution

Do you have a library card?

I can walk into a library, pull a book from a shelf, sit down, and read it front to back. No library card needed. Only need one to take a book home. Completely different scenario.

Edit:clarification

Re: An update on Wayback Machine access

#338
post #287

Earlier quoted context omitted.

Why would any chatbot provider attack Wikipedia *legally*? Captcha is fully solved, and agents are fully capable of acting as editors, pushing any agenda desired by the user.

Strangers can't really edit Wikipedia any more, especially if their edit is suspicious. It's a closed system despite the appearance. An anti-vandal bot or human will quickly revert your edit.

This is not really true in my experience aside from protected articles; however it is true that sometimes another human will make an edit which degrade the article quality or insist on keeping something that shouldn't be kept. Also, Articles on contentious subjects tend to be problematic.

Re: An update on Wayback Machine access

#339
post #236

Earlier quoted context omitted.

> Why folks continue to use them confuses me. Seems like the “confused” is disingenuous when not paying for content you read is a clear motivation.

Convenience as a higher order motivator than disgust at the bad behavior of archive.{today,ph,...} mentioned elsewhere, I think is the point of the comment to which you replied.

Convenience is by definition one step.

Disgust requires research, analysis, and decision. The research alone is beyond most users' general practice.

Why would anyone walk in the front door when they could hop a fence, pry open a window, and crawl in?

Re: An update on Wayback Machine access

#340

Earlier quoted context omitted.

Are you disputing that Gaza had more births than reported deaths during this period, or do you have other examples of “genocides” where the population grew during their genocide?

https://en.wikipedia.org/wiki/Brandolini%27s_law Account created 1 hour ago and only made these two comments, no others. Dang should investigate where these bots are coming from.

Funny how you led with a desire to get Wikipedia to align with one side in a conflict, and an opposition to the “mainstream” having no diversity of opinion in your view, yet the moment you are confronted with alternative viewpoints and facts you refuse to engage with the content, and instead try to silence views that you don’t like.
Post reply on HN