Earlier quoted context omitted.
Gatekeeping information is not the solution
Why not? If its the difference between the information being available at all, then I choose login any day of the week.
An update on Wayback Machine access
241–250 of 374 posts
Re: An update on Wayback Machine access
#242Earlier quoted context omitted.
How do you separate the Wayback Machine from malicious bots that pretend to be the Wayback Machine? Are you whitelisting their IP blocks? Because I get a ton of scraper requests that forge Googlebot, Bing, and Yandex user-agents that are totally not coming from their IP ranges. In fact, sometimes they all come from the same IP...
I thought at least google (and possibly others) provided a way to verify the user agent?
Yeah, a real browser would produce certain patterns and never certain others. So in some cases one could clearly say it's not a human using a browser. But a scraper could also make efforts to mimic human browsing. Mostly the frequency of requests can tell with high probability tell it's not human. But such algorithms might occasionally give false positives for real users, exactly like it has obviously happened for the archive.
What google service are you referring to? Not sure whether the archove uses any of Google's tracking. I have pretty strong blocking of trackers and ads. But the archive works for me.
Re: An update on Wayback Machine access
#243Earlier quoted context omitted.
I was thinking the same thing... paid access for high volume users or scrapers could actually help fund the non-profit. Maybe let website owners decide which scrapers are allowed to use their content, or allow them to get paid for use of it. If news and other sites were getting paid, maybe they could go back to optimizing for good content instead of clicks.
I think that would get into murky water really quickly with the rights holders (/content creators) not exactly being thrilled the Wayback Machine is essentially monetizing their IP behind their back.
Re: An update on Wayback Machine access
#244Earlier quoted context omitted.
I think that would get into murky water really quickly with the rights holders (/content creators) not exactly being thrilled the Wayback Machine is essentially monetizing their IP behind their back.
If the money was guaranteed to only be used to pay the costs, there probably wouldn't be any problems
Re: An update on Wayback Machine access
#245Earlier quoted context omitted.
No one is upset that the AI companies are scraping the web, they’re upset how poorly implemented the scrapers, but the scraping itself is fine. Lots of people and businesses scrape the web.
There are people commenting parallel to you saying they are upset about AI companies scraping the web.
Re: An update on Wayback Machine access
#246Re: An update on Wayback Machine access
#247I really so through some money their way, they do wonderful work.
Re: An update on Wayback Machine access
#248Earlier quoted context omitted.
We used to always "scrape" the wayback machine for any sort of news article we actually paid to consume. I was absolutely shocked by major news sites making very important edits to an article without any sort of editorial notice! Sadly this sort of thing is probably not really possible anymore, but I can't really blame anyone for making this sort of decision. I can't imagine how much more traffic they get now vs 2021…
What if they charged money? Is it something you'd pay for?
Re: An update on Wayback Machine access
#249Re: An update on Wayback Machine access
#250Earlier quoted context omitted.
What if they charged money? Is it something you'd pay for?
I mean we already paid for the article from the source itself. I guess I'd expect a better "diff" source from them, but if they dont even update the article itself, i guess i wouldn't expect a paid service to have those updates either?
with llms, at some point it probably becomes easier to use your paid api connection to manage your own cached version yourself?