Live data from Hacker News

An update on Wayback Machine access

blog.archive.org

241–250 of 374 posts

Re: An update on Wayback Machine access

#241

Earlier quoted context omitted.

Gatekeeping information is not the solution

Why not? If its the difference between the information being available at all, then I choose login any day of the week.

This really only inconveniences people, not bots.

Re: An update on Wayback Machine access

#242

Earlier quoted context omitted.

How do you separate the Wayback Machine from malicious bots that pretend to be the Wayback Machine? Are you whitelisting their IP blocks? Because I get a ton of scraper requests that forge Googlebot, Bing, and Yandex user-agents that are totally not coming from their IP ranges. In fact, sometimes they all come from the same IP...

I thought at least google (and possibly others) provided a way to verify the user agent?

Sorry, not following. I thought the user agent is a string that the caller can set to anything. There is no immediate, reliable way to tell whether the string is correct.

Yeah, a real browser would produce certain patterns and never certain others. So in some cases one could clearly say it's not a human using a browser. But a scraper could also make efforts to mimic human browsing. Mostly the frequency of requests can tell with high probability tell it's not human. But such algorithms might occasionally give false positives for real users, exactly like it has obviously happened for the archive.

What google service are you referring to? Not sure whether the archove uses any of Google's tracking. I have pretty strong blocking of trackers and ads. But the archive works for me.

Re: An update on Wayback Machine access

#243

Earlier quoted context omitted.

I was thinking the same thing... paid access for high volume users or scrapers could actually help fund the non-profit. Maybe let website owners decide which scrapers are allowed to use their content, or allow them to get paid for use of it. If news and other sites were getting paid, maybe they could go back to optimizing for good content instead of clicks.

I think that would get into murky water really quickly with the rights holders (/content creators) not exactly being thrilled the Wayback Machine is essentially monetizing their IP behind their back.

Interesting, in all the years I have never noticed that IP has 2 meanings (well probably more...) Yeah, I am an engineer and usually try to avoid the legal BS. Although I hate that AI has made stealing legal if you are big enough.

Re: An update on Wayback Machine access

#244
post #187

Earlier quoted context omitted.

I think that would get into murky water really quickly with the rights holders (/content creators) not exactly being thrilled the Wayback Machine is essentially monetizing their IP behind their back.

If the money was guaranteed to only be used to pay the costs, there probably wouldn't be any problems

No. What happened to their remote library scheme? They did not make if for profit, but still...

Re: An update on Wayback Machine access

#245
post #53

Earlier quoted context omitted.

No one is upset that the AI companies are scraping the web, they’re upset how poorly implemented the scrapers, but the scraping itself is fine. Lots of people and businesses scrape the web.

There are people commenting parallel to you saying they are upset about AI companies scraping the web.

Yeah I dont think they know why they think that.

Re: An update on Wayback Machine access

#247
I recently remembered a wonderful comic blog from 2010’s that is not online anymore. It was a sonderful Finnish LGTG-thened comic blog that I use to read, then forgot completely until few weeks ago. WM had it stored of course, so I could read through this amazing piece of internet art again.

I really so through some money their way, they do wonderful work.

Re: An update on Wayback Machine access

#248

Earlier quoted context omitted.

We used to always "scrape" the wayback machine for any sort of news article we actually paid to consume. I was absolutely shocked by major news sites making very important edits to an article without any sort of editorial notice! Sadly this sort of thing is probably not really possible anymore, but I can't really blame anyone for making this sort of decision. I can't imagine how much more traffic they get now vs 2021…

What if they charged money? Is it something you'd pay for?

I mean we already paid for the article from the source itself. I guess I'd expect a better "diff" source from them, but if they dont even update the article itself, i guess i wouldn't expect a paid service to have those updates either?

Re: An update on Wayback Machine access

#250

Earlier quoted context omitted.

What if they charged money? Is it something you'd pay for?

I mean we already paid for the article from the source itself. I guess I'd expect a better "diff" source from them, but if they dont even update the article itself, i guess i wouldn't expect a paid service to have those updates either?

ah, i think i misunderstood your original post. if you mean the wayback-machine/arkive, then I suspect it'd be hard to justify? You are essentially paying a third-party source to validate that diffs didn't go through on the source material.

with llms, at some point it probably becomes easier to use your paid api connection to manage your own cached version yourself?

Post reply on HN