Live data from Hacker News

WorldBrain's Memex: Bookmarking for the power users of the web

getmemex.com

151–160 of 215 posts

Re: WorldBrain's Memex: Bookmarking for the power users of the web

#151
post #149

Earlier quoted context omitted.

Cool. Let me know when I can extract the full text contents from such files using common, built-in tools on your average fresh install of MacOS/Linux. PDF works just fine. It presents a feasible view of the original data, and allows for data harvesting with ease.

Let me know when I can extract the full text You can extract the full text from these (with whatever tools you like) with better fidelity than you can from a pdf, which is a lossy conversion from the same source. This seems to barely merit debating, unless I'm missing something. PDF works just fine. I'm sure it works for you and I'm not harbouring any delusions I'm going to talk you out of your decades-established wo…

>PDF is a lossy conversion

I am not finding this to be true. Pretty much every PDF I have has been usable for extracting the text content - unless the Web authors intentionally work to obfuscate/disable this functionality, i.e. using images to display text content.

>PDF is not a good way to archive web pages, either manually or programmatically.

I disagree, entirely, with your conclusion - you haven't made a strong argument. 20,000+ fully-searchable, indexable, accessible-in-offline PDF files vs. your opinion so far. I don't see any of the issues you've stated are insurmountable - in fact, I find the reality to be completely the opposite to your stated opinion. Please expand on this if you have the energy.

Re: WorldBrain's Memex: Bookmarking for the power users of the web

#152
post #149

Earlier quoted context omitted.

Let me know when I can extract the full text You can extract the full text from these (with whatever tools you like) with better fidelity than you can from a pdf, which is a lossy conversion from the same source. This seems to barely merit debating, unless I'm missing something. PDF works just fine. I'm sure it works for you and I'm not harbouring any delusions I'm going to talk you out of your decades-established wo…

>PDF is a lossy conversion I am not finding this to be true. Pretty much every PDF I have has been usable for extracting the text content - unless the Web authors intentionally work to obfuscate/disable this functionality, i.e. using images to display text content. >PDF is not a good way to archive web pages, either manually or programmatically. I disagree, entirely, with your conclusion - you haven't made a strong a…

fully-searchable, indexable, accessible-in-offline

Probably the shortest version of the point I'm trying to make is that every current browser does a much better job of providing you this than printing to PDF. If you rely on this as a personal web archiving system, you're going to lose data in the most irritating way - data you thought you collected but actually didn't.

Re: WorldBrain's Memex: Bookmarking for the power users of the web

#153

This is an interesting perhaps meta-relevant topic for HN. How many of us bookmark or otherwise record interesting posts from here and elsewhere? How many of us ever refer that accumulated digital memory? I have about 7,000 links with notes accumulated over the last few decades. I’ve read a lot of them, but the hard to acknowledge reality is that even with a refined workflow, recording my links in a near perfect taxo…

I just... copy the links and paste them into Google Keep :D Really fast and searchable, I usually find myself searching for the same shit after a while, so Keep is another destination.

But now I realized I may want to back them up somewhere... technically these notes aren't important and can be lost, but I'd like to keep them

Re: WorldBrain's Memex: Bookmarking for the power users of the web

#154

This is still not as powerful as my one, simple trick to handle all bookmarks, ever: Print to PDF. I've been doing it since last century, and I have 10's of thousands of PDF's of every single web page I've ever found interesting, sitting right there in a directory on my computer. Its indexable, searchable, grok'able, available off-line, allows me to harvest data without fuss, and gives me access to anything I can rem…

It's too bad browsers don't have an easy way to print to browser-page-sized PDF. Standard 8.5x11/A4 paper sized PDFs of webpages tend to look pretty terrible. I used to use the Scrapbook plugin for Firefox but I realized for the most part just plaintext might be best. So I'm in the process of setting up a workflow that will save article in markdown in one click and sync between my phone and my computer.

> So I'm in the process of setting up a workflow

There shouldn't be a need to setup a process. This functionality exists in many places.

For example, I use Joplin to save articles in Markdown format. It's the best web to markdown conversion tool I have found. Then at some point later, I'll pick what's still interesting to me and export from Joplin to PDF if I like.

Insapaper is $3 / month and is a save for later tool. You can then export all your articles in Epub and other formats.

I'm sure there are loads others.

Re: WorldBrain's Memex: Bookmarking for the power users of the web

#155
post #152

Earlier quoted context omitted.

>PDF is a lossy conversion I am not finding this to be true. Pretty much every PDF I have has been usable for extracting the text content - unless the Web authors intentionally work to obfuscate/disable this functionality, i.e. using images to display text content. >PDF is not a good way to archive web pages, either manually or programmatically. I disagree, entirely, with your conclusion - you haven't made a strong a…

fully-searchable, indexable, accessible-in-offline Probably the shortest version of the point I'm trying to make is that every current browser does a much better job of providing you this than printing to PDF. If you rely on this as a personal web archiving system, you're going to lose data in the most irritating way - data you thought you collected but actually didn't.

I don't understand your point at all. In what way is a browser going to give me information that is not available to me unless I'm online? PDF's of sites I've visited have all the data I need - the stuff I read that then prompted me to print to PDF. I've searched and I'm yet to find a single PDF in my collection that doesn't have the info that prompted me to save it in the first place. I understand you believe your point is strong - its still not being made in a way that I can relate.

Re: WorldBrain's Memex: Bookmarking for the power users of the web

#156
post #61

Earlier quoted context omitted.

You could try the awesome SingleFile extension: https://github.com/gildas-lormeau/SingleFile It might be a good compromise between PDF and plain text. It's pretty nice because it essentially serialises a snapshot of the current DOM tree, so it works with all kinds of JS-generated pages. The files should be relatively grep-able, because it's normal HTML. Of course, you might want to strip HTML tags for more sophistica…

SingleFile is a really great extension, but I wanted something a bit more pared down that I could easily use on both mobile and desktop and sync between them using Syncthing. So I'm trying to copy some of SingleFile's UI and graft it on to Markdown-Clipper.[1] And also add the ability to save the images that get picked up by Readability (which Markdown-Clipper uses). [1] https://github.com/enrico-kaack/markdown-clipp…

Joplin already has this feature via browser extension. It has a mobile app, but never tested it myself.

Re: WorldBrain's Memex: Bookmarking for the power users of the web

#157

Something I noticed when I use "Read Later" style applications to save pages is that I will, most of the time, forget about how I arrived at a certain page. This is important to me because it gives me the context to decide a perspective on the page. If I was able to save pages while also knowing where I found them and maybe make a comment about why I found it interesting, then I would be able to organize my knowledge…

Using Pinboard [0] I currently solve this by using tags like "via-twitter", "via-hackernews", or even people like "via-john". I also occasionally add a note to my pin (bookmark) to remind me why I bookmarked it.

[0] https://pinboard.in/

Re: WorldBrain's Memex: Bookmarking for the power users of the web

#158
post #132

Earlier quoted context omitted.

To be a little pedantic, while this is a fantastic idea, is it really bookmarking? What you’ve done instead is compiled a personal digital library; akin to a Kindle

Inasmuch as every PDF still has the URL to the original page, for my uses - I'd say yes, it is bookmarking. Its not like I'd cut the spine off every book in my library and create an index out of the covers ..

Apparently only Firefox saves the source URL. In PDFs saved from Safari or Chrome on macOS I can't find the URL anywhere. Maybe I'm missing something.

Re: WorldBrain's Memex: Bookmarking for the power users of the web

#159

Earlier quoted context omitted.

Hmm, how could the results have been disappointing? It just searches for text, how bad can it be?

I would search for terms that I knew were on pages that should be indexed but they wouldn't be in the results list.

I installed it yesterday and noticed that it doesn't actually index much. It should be, but it's not, the pages aren't added. If they are in the index, it finds them in a search, but very few pages are.

Re: WorldBrain's Memex: Bookmarking for the power users of the web

#160
I have been using Pocket [0], Instapaper [1] and Pinboard [2] over the years.

I am currently using Pocket and Pinboard in parallel: articles / websites that I want to read later are sent to Pocket (untagged), websites that I might want to get to back later are tagged and sent to Pinboard.

While my archive on Pinboard works quite well I am very disappointed by the support. Either the developer does not answer at all or months later. Not acceptable for a paid service.

While Memex looks interesting having no API makes it a pass for me (for now).

[0] https://getpocket.com/ [1] https://www.instapaper.com/ [2] https://pinboard.in/

Post reply on HN