Live data from Hacker News

An update on Wayback Machine access

blog.archive.org

141–150 of 374 posts

Re: An update on Wayback Machine access

#141

Earlier quoted context omitted.

fully expect them and wikipedia to get assaulted by AI companies to monopolize data source access in the near future. The future is bleak :\

I'm expecting a full court legal attack on Wikipedia at some point. It's just too good for information, and companies would rather you use their chatbot to regurgitate that info now that they have their own copies. Google was already built to pretend as if they had some magic system giving you "Answers" that were 95% just the text of the infobox they used to have for wikipedia on the side. It's going to suddenly be e…

on what basis?

Re: An update on Wayback Machine access

#142

Earlier quoted context omitted.

fully expect them and wikipedia to get assaulted by AI companies to monopolize data source access in the near future. The future is bleak :\

I'm expecting a full court legal attack on Wikipedia at some point. It's just too good for information, and companies would rather you use their chatbot to regurgitate that info now that they have their own copies. Google was already built to pretend as if they had some magic system giving you "Answers" that were 95% just the text of the infobox they used to have for wikipedia on the side. It's going to suddenly be e…

Wikipedia is not a challenge to the system. If it were, it would be excluded from search results, much like blogs are now.

Re: An update on Wayback Machine access

#143

Earlier quoted context omitted.

Then make agent friendly content. Take the text and make a markdown version.

People doing this say it makes things worse because then the bots download both.

Not to mention that it solves none of the rate issues. If the scrapers are hitting your site 10,000 times a day, adding markdown isn’t going to change that at all.

Re: An update on Wayback Machine access

#144

Mad props to the people at the Archive. You are the heros we need in a formerly open internet that is surrendering to evil big corps and closing down free access. The Internet Archive is in a really bad spot being attacked from multiple sides at once. But — while service has not been consistent — they have maintained open access. I can still access anonymously from Tor without Cloudflare or some other centralized gat…

I donate to them every year b/c I fully agree they’re doing a thankless critical job very well.

Thank you for the inspiration! I just made my first donation.

Re: An update on Wayback Machine access

#145
post #2

Why not just offer a paid endpoint for the crawlers? It's not like the demand is going to go away anytime soon. It serves nobody except CloudFlare and hardware companies when one side set up blockers and the other side spend money putting VPN SDKs in consumer TVs. I am also curious how the (Russian?) paywall bypass mirror archive.is is doing given that they are probably subject to similar amounts of traffic.

> It serves nobody except CloudFlare and hardware companies when one side set up blockers and the other side spend money putting VPN SDKs in consumer TVs

The problem with this perspective is that it ignores the victimization which is happening to all sorts of sites right now.

On one hand, you have content owners/suppliers which are trying to place restrictions on how much free bulk use is allowed.

When scrapers go to exotic lengths to evade the blocks, eg by using thousands of ephemeral IP addresses to collect an entire corpus, saying stuff like that makes it sound like it's all a wash.

"Oh, what a silly situation... How did we ever end up like this? It's not good for anyone ..."

No, there is a victim trying to defend themselves from rampant theft of resources, and a corporate asshole which doesn't care about the effects of their actions.

Re: An update on Wayback Machine access

#146

Earlier quoted context omitted.

I assume that's because the IP range of your company network overlaps with a range used by some scrapers, and if it doesn't happen on your phone even in the corporate network, then IA probably checks some extra signals like the user agent in addition to the IP

No - my phone is not connected to work's WiFi. Wonder who the bad actors in my company are...

It might just be your corporate VPN and whatever ASN it’s being routed through.

Re: An update on Wayback Machine access

#147

Shouldn't the solution be to gate bulk access for automated services for a price? Not just the wayback machine, any personal or corporate website? My blogs are getting slammed and there are issues with cloudflare or captchas.

> Shouldn't the solution be to gate bulk access for automated services for a price?

Fine in theory but determined scrapers will use residential proxies in bulk.

Re: An update on Wayback Machine access

#148
post #95

I wonder why doesn’t the Internet Archive require logging-in prior to accessing the Wayback Machine. It would probably help them distinguish humans from bots, at a very little cost to humans.

Gatekeeping information is not the solution

Why not? If its the difference between the information being available at all, then I choose login any day of the week.

Re: An update on Wayback Machine access

#149
post #115
post #77

Earlier quoted context omitted.

Has it been settled whether robots.txt applies to user-driven chat sessions and if things like the crawl delay should be applied to say an end-user, an ip address, a harness provider, etc? My understanding is robots.txt is more for training exclusions, but less so for agent work.

AI bros think they should be exempt from robots.txt. Administrators of big services beg to differ. No solid consensus has arisen. I bet it's gonna take a lawsuit or two to see how it shakes out.

From the start, robots.txt has always been an indicator of a site's preference with no actual legal significance.

Re: An update on Wayback Machine access

#150

Earlier quoted context omitted.

I want agents to be able to act on my behalf, that’s the entire point. An agent should be able to do anything I can do sitting at my browser.

I suspect most people would be ok with this if they could only do it at the rate and frequency you yourself can do it. The problem is largely one of scale.

Is scale what we're discussing though?

e.g. a prompt of "fetch and summarise it for me" is very close to what a human would be doing with a web browser, and doesn't seem to involve any kind of scaling issue.

Post reply on HN