Earlier quoted context omitted.
> Google shouldn't be the center of the Web. I agree, but are you suggesting it's going to be better if WayBackMachine is?
Yes. At least Archive.org isn't an evil mega corporation destroying the internet. Yet.
Why I link to Wayback Machine instead of original web content
51–60 of 262 posts
Re: Why I link to Wayback Machine instead of original web content
#52Re: Why I link to Wayback Machine instead of original web content
#53Re: Why I link to Wayback Machine instead of original web content
#54Earlier quoted context omitted.
Would be nice if there's an automatic way to have a link revert to the Wayback Machine once the original link stops working. I can't think of an easy way to do that, though.
Either a browser extension, or an 'active' system where your site checks the health of the pages it links to.
E.g. https://addons.mozilla.org/firefox/addon/wayback-machine_new...
Re: Why I link to Wayback Machine instead of original web content
#55In practice however, archive.org did censor content based on political preference.
Sounds plausible, but I sure would like a citation for that claim.
Re: Why I link to Wayback Machine instead of original web content
#56https://web.archive.org/web/20200908090515/https://hawaiigen...
Re: Why I link to Wayback Machine instead of original web content
#57Earlier quoted context omitted.
INTERNETARCHIVE.BAK: The INTERNETARCHIVE.BAK project (also known as IA.BAK or IABAK) is a combined experiment and research project to back up the Internet Archive's data stores, utilizing zero infrastructure of the Archive itself (save for bandwidth used in download) and, along the way, gain real-world knowledge of what issues and considerations are involved with such a project. Started in April 2015, the project alr…
I wish there were a way to get a low-rez copy of their entire archive. So, only text, no images, binaries, PDFs (other than PDFs converted to text which they seem to do). As it stands the archive is so huge, the barrier to mirroring is high.
When scoping out the size of Google+, one of ArchiveTeam's recent projects, it emerged that the typical size of a post was roughly 120 bytes, but total page weight a minimum of 1 MB, for a 1% payload to throw-weight ratio. This seems typical of much the modern Web. And that excludes external assets: images, JS, CSS, etc.
If just the source text and sufficient metadata were preserved, all of G+ would be startlingly small -- on the order of 100 GB I believe. Yes, posts could be longer (I wrote some large ones), and images (associated with about 30% of posts by my estimate) blew things up a lot. But the scary thing is actually how little content there really was. And while G+ certainly had a "ghost town" image (which I somewhat helped define), it wasn't tiny --- there were plausibly 100 - 300 million users with substantial activity.
But IA's WBM has a goal and policy of preserving the Web as it manifests, which means one hell of a lot of cruft and bloat. As you note, increasingly a liability.
Re: Why I link to Wayback Machine instead of original web content
#58Earlier quoted context omitted.
Google shouldn't be the center of the Web. They could also easily determine where the archive link is pointing to and not penalize. But I guess making sure we align with Google's incentives is more important than just using the Web.
> Google shouldn't be the center of the Web. I agree, but are you suggesting it's going to be better if WayBackMachine is?
We as a community need to think bigger rather than resigning ourselves to our fate.
Re: Why I link to Wayback Machine instead of original web content
#59Earlier quoted context omitted.
WayBackMachine alternative, archive.is, has an option to download zip archive of HTML with images and CSS (but no JS) - this way you can preserve and host a copy of original webpage on your own website
Or just wget -rk... Mirroring a website isn't so hard that you need a service to do it for you. Your browser even has such a function; try ctrl-s.
Re: Why I link to Wayback Machine instead of original web content
#60Earlier quoted context omitted.
There's no reason that pagerank couldn't be adapted to take into account wayback machine urls, there is a link with a url pointing at https://web.archive.org/web/*/https://news.ycombinator.com/ google could easily register that as a link to both resources - one to web.archive, the other to the site. there is also no reason why that has to become a slippery slope, if anyone is going to ask "but where do you stop!!"
After all, they did change their search to accommodate AMP. Changing it to take WebArchive into account is a) peanuts and b) is actually better for the web
Some kind of CDN-edge-archive hybrid.