Live data from Hacker News

Internet Archive Scholar

scholar.archive.org

31–40 of 59 posts

Re: Internet Archive Scholar

#32

Earlier quoted context omitted.

> I've started saving I do similar. I've had https://github.com/ArchiveBox/ArchiveBox bookmarked for a while as something to try better organise all that, but like a great many things I haven't go around to it yet.

I use Raindrop for this. It’s a pretty great bookmark manager made by an indie dev, but it also can create archives of bookmarked pages. https://help.raindrop.io/backups#permanent-library

“Only available in Pro plan”. No information on whether that is a general limitation or if these features are present if self-hosted, so I assume the former. And there seems to be little obvious information for self-hosting.

So probably not one for my use case.

Re: Internet Archive Scholar

#33

Earlier quoted context omitted.

Some people use this tool (of mine) for saving web content from either bookmarks or just everything you browse: https://github.com/crisdosyago/Diskernet There's also plent¥ of other similar tools: - https://github.com/ArchiveBox/ArchiveBox - https://github.com/gildas-lormeau/SingleFile

I'd never heard of SingleFile before but it looks excellent. It would be great if Firefox could incorporate it into its save function too. Firefox save page works but as shown on the SingleFile demo video, it's not really what a user would expect, it's often not complete and splitting it across multiple files/directories isn't ideal either.

It's unfortunate as Firefox used to have excellent MHTML support (which similarly achieves an all-in-one file) via addons, particularly the feature-rich UnMHT. While Chromium and its derivatives support MHTML saving natively (and in the past Opera Presto and IE).

If they brought back MHTML saving support it'd be a great win.

Re: Internet Archive Scholar

#34

Earlier quoted context omitted.

While the author(s) are still alive, they are often a productive contact. (In one case I was able to give back: bundling up the several scans an author had of a half-century old paper from their student days into a single, hopefully cromulent, PDF) Edit: recall also that accepting that links are one-way and might be dead was the key simplification that allowed HTTP to take off after prior attempts at hypermedia had f…

Good (Edit) point! It's good that the web accepts dead links by design, we can't expect perfection from our distributed information, but it seems that the bitrot of information is too high compared to the information storage technologies available. 2 spinning rust drive can store the library of congress. ~ 2,000 drives would store the web (1). How many millions of these drives get manufactured per year? Our technolog…

2000 drives x $100 = $200,000. Double that for backup, $400,000. Admin, maintenance, let's say total, $1M/year. So, Wikipedia could end its own deadlink problem (IF the reference sources would agree.

But stuff goes missing at Wayback because people don't agree to their pages being backed-up. Copyright, whatever. So it's like Global Heating, the tech is there, but people just can't agree. So 'pirate' backer-uppers go to jail. And island-nations and expensive ocean-side properties are being submerged. So it goes.

Re: Internet Archive Scholar

#35
post #5

Some days, nearly half the links I click are dead so I've found myself relying on the waybackmachine more and more over the past few months. It's really shocking just how fast digital obsolescence reared its ugly head. Of course angelcites etc. were a clear early blow, but nowadays... I've started saving the html (including the css seems like too much overhead, and often it's incomplete or relies on downloads still -…

Some people use this tool (of mine) for saving web content from either bookmarks or just everything you browse: https://github.com/crisdosyago/Diskernet There's also plent¥ of other similar tools: - https://github.com/ArchiveBox/ArchiveBox - https://github.com/gildas-lormeau/SingleFile

> Coming to a future release, soon!: The ability to publish your own search engine that you curated with the best resources based on your expert knowledge and experience.

This would be fantastic, being able to browse a curated internet made of accumulated lists from other trusted users on the net, similar to how ad blocking lists work today.

You are genuinely trying to steer the internet into what it used to be: a museum of knowledge and expert discussion.

Edit: Ah but wow, Polyform license. Huh.

Re: Internet Archive Scholar

#36

Earlier quoted context omitted.

While the author(s) are still alive, they are often a productive contact. (In one case I was able to give back: bundling up the several scans an author had of a half-century old paper from their student days into a single, hopefully cromulent, PDF) Edit: recall also that accepting that links are one-way and might be dead was the key simplification that allowed HTTP to take off after prior attempts at hypermedia had f…

Good (Edit) point! It's good that the web accepts dead links by design, we can't expect perfection from our distributed information, but it seems that the bitrot of information is too high compared to the information storage technologies available. 2 spinning rust drive can store the library of congress. ~ 2,000 drives would store the web (1). How many millions of these drives get manufactured per year? Our technolog…

Even if that estimate is off by an order of magnitude, which given the weight of modern web pages it easily could be, 20,000 drives to store the entire web seems way more doable than I ever would have imagined.

Re: Internet Archive Scholar

#37
post #5

Some days, nearly half the links I click are dead so I've found myself relying on the waybackmachine more and more over the past few months. It's really shocking just how fast digital obsolescence reared its ugly head. Of course angelcites etc. were a clear early blow, but nowadays... I've started saving the html (including the css seems like too much overhead, and often it's incomplete or relies on downloads still -…

Some people use this tool (of mine) for saving web content from either bookmarks or just everything you browse: https://github.com/crisdosyago/Diskernet There's also plent¥ of other similar tools: - https://github.com/ArchiveBox/ArchiveBox - https://github.com/gildas-lormeau/SingleFile

Beware the very strange and bad license for Diskernet, which is "Polyform Strict License 1.0.0"

Re: Internet Archive Scholar

#38

Earlier quoted context omitted.

Some people use this tool (of mine) for saving web content from either bookmarks or just everything you browse: https://github.com/crisdosyago/Diskernet There's also plent¥ of other similar tools: - https://github.com/ArchiveBox/ArchiveBox - https://github.com/gildas-lormeau/SingleFile

Beware the very strange and bad license for Diskernet, which is "Polyform Strict License 1.0.0"

For people looking for more info on these strange licenses:

https://www.reddit.com/r/linux/comments/coazye/what_does_rli...

Re: Internet Archive Scholar

#39
post #5

Some days, nearly half the links I click are dead so I've found myself relying on the waybackmachine more and more over the past few months. It's really shocking just how fast digital obsolescence reared its ugly head. Of course angelcites etc. were a clear early blow, but nowadays... I've started saving the html (including the css seems like too much overhead, and often it's incomplete or relies on downloads still -…

[deleted]

Re: Internet Archive Scholar

#40
After a little testing, this looks like a good information source, although the combination of Google Scholar and sci-hub is probably still the best option, i.e. I couldn't find anything with Internet Scholar that wasn't available with the other options, and the quality of results on searchs is somewhat higher with Google Scholar (this may be because Google Scholar utilizes citation count as a search parameter, which Internet Archive Scholar doesn't seem to do).

Internet Archive is a great resource, however, it should get state funding as it provides a fundamentally important archival service. It's too bad it has to rely so much on private philanthropic donations (although state support comes with possible political interference, i.e. censorship of material that some politician doesn't like, maybe that's less of a problem with private donations, although then you could have some billionaire doing the same thing).

Post reply on HN