Live data from Hacker News

An update on Wayback Machine access

blog.archive.org

231–240 of 374 posts

Re: An update on Wayback Machine access

#232

Earlier quoted context omitted.

I’ve read a few different experiences with hosts having had success with Anubis to cut down on excessive scraping. One that comes to mind is the Dolphin project.[1] I’ve read a few of those threads, but often it’s people at cross-purposes with the goals of Anubis. Is there a chance you could clarify the “not working” bit? [1] https://dolphin-emu.org/blog/2025/06/04/dolphin-progress-rep...

From a few weeks ago: https://news.ycombinator.com/item?id=49500040 Basically, the cost of an optimized solution is orders of magnitude lower than the cost of an in-browser solution. Anyone dedicated can easily afford to solve workloads higher than your users will tolerate. You might stop casual scrapers, but you're not going to stop someone who cares. AI scraping companies care.

That's not even the fundamental problem. Even if the payload runs optimally in the browser, the cost of CPU is so small that it's basically irrelevant.

If you waste your user's time with something that would take a full minute to run on a datacenter core, you're costing the scraper something like $0.000005: 360 W TDP on a 128-core EPYC 9754 * $0.10/kWh. In reality, it'll be substantially less than that, because CPUs don't use 0W at idle.

The only way this would make any sense is if there were many more scrapers than users and scrapers cared more about latency than real users, but that's the exact opposite of reality. The entire endeavor is so fundamentally misguided that it almost seems like a psyop.

Re: An update on Wayback Machine access

#234
post #228

Earlier quoted context omitted.

It doesn't matter who or what the service is, how much they have, or whatever else. They created a problem and now users have to pay for the inconvenience by emailing(!) specific details that could be captured automatically through web logs: OS, browser, IP address. It's ridiculous.

> details that could be captured automatically through web logs You can't be serious. Are you ok? The entire point is that they're trying to tell bots and humans apart. They're trusting email (and how you write your email) as a good signal that you're human. What are you talking about getting it from the log? The point is to correlate. How do you expect them to know who you are in the log unless you give them that in…

Lollipops?

Re: An update on Wayback Machine access

#235

Earlier quoted context omitted.

I do donate. But that doesn't mean I have to thank them before every meal or think that the service is perfect.

What do you expect them to do though? You have to be a reasonable person.

Data analysis would be a good start, better blocking heuristics, an off-the-shelf solution used by other organisations that don't have this problem, etc.

Re: An update on Wayback Machine access

#236

Earlier quoted context omitted.

Correct:Also, archive.* has actively edited archived sites to promote their agenda. Why folks continue to use them confuses me. One would think the big wikipedia purge would curb such behavior.

> Why folks continue to use them confuses me. Seems like the “confused” is disingenuous when not paying for content you read is a clear motivation.

Convenience as a higher order motivator than disgust at the bad behavior of archive.{today,ph,...} mentioned elsewhere, I think is the point of the comment to which you replied.

Re: An update on Wayback Machine access

#237

Earlier quoted context omitted.

This is absolutely something that's happening. There are even paid scraper API's that offer "Wayback Machine fallback" as a feature.

It's an open secret that you can often circumvent paywalls by searching Wayback.

One site was doing that which archive in their name but wasn't apart of archive.org

Re: An update on Wayback Machine access

#239

Earlier quoted context omitted.

It's an open secret that you can often circumvent paywalls by searching Wayback.

One site was doing that which archive in their name but wasn't apart of archive.org

archive.is or archive.today?

Why being shy in the era of stealing AI?

Re: An update on Wayback Machine access

#240
post #199

Earlier quoted context omitted.

> Why folks continue to use them confuses me. I use them. I haven't ever heard mention that the content is edited. Do you have a source?

See https://arstechnica.com/tech-policy/2026/02/wikipedia-might-... https://en.wikipedia.org/wiki/Wikipedia:Archive.today_guidan... ? Besides tampering with content, the site was also using visitors to DDOS a blog that mentioned the owner of archive.today.

[dead]
Post reply on HN