Live data from Hacker News

WorldBrain's Memex: Bookmarking for the power users of the web

getmemex.com

61–70 of 215 posts

Re: WorldBrain's Memex: Bookmarking for the power users of the web

#61

This is still not as powerful as my one, simple trick to handle all bookmarks, ever: Print to PDF. I've been doing it since last century, and I have 10's of thousands of PDF's of every single web page I've ever found interesting, sitting right there in a directory on my computer. Its indexable, searchable, grok'able, available off-line, allows me to harvest data without fuss, and gives me access to anything I can rem…

It's too bad browsers don't have an easy way to print to browser-page-sized PDF. Standard 8.5x11/A4 paper sized PDFs of webpages tend to look pretty terrible. I used to use the Scrapbook plugin for Firefox but I realized for the most part just plaintext might be best. So I'm in the process of setting up a workflow that will save article in markdown in one click and sync between my phone and my computer.

You could try the awesome SingleFile extension: https://github.com/gildas-lormeau/SingleFile

It might be a good compromise between PDF and plain text. It's pretty nice because it essentially serialises a snapshot of the current DOM tree, so it works with all kinds of JS-generated pages.

The files should be relatively grep-able, because it's normal HTML. Of course, you might want to strip HTML tags for more sophisticated searching.

Re: WorldBrain's Memex: Bookmarking for the power users of the web

#62

Earlier quoted context omitted.

What sorts of web pages do you do this for? What if the pdf version is difficult to read?

Every web page I'll ever want to refer to, ever again. There are no good reasons for exceptions to this technique, imho. If the PDF version is difficult to read - which it rarely is, by the way - all I need to do is open the PDF and use the links in the page header to go visit the site again - all the details about the page are still there in the PDF, links are still clickable, etc. And if its really important, and I…

> There are no good reasons for exceptions to this technique, imho

Dynamic web sites? A PDF of my bank website isn't going to help me much.

Re: WorldBrain's Memex: Bookmarking for the power users of the web

#63
post #49

Earlier quoted context omitted.

nothing beats print-to-PDF What's the advantage over browser's built-in 'save entire page' option? Print to PDF loses formatting and obscures the URL you got the thing from.

$ pdftotext | grep sometext ^^ first advantage Also: a) Formatting is not lost, it just changes to fit the default paper size I've got selected (A4) but doesn't really make much difference, since its a snapshot, and b) URL is right there in the Header of the PDF, and is clickable, so no - not really an issue. This archive also functions as a bookmark collection as well as an offline copy for future reference .. (Disc…

You can grep the saved archives and they often save working copies of local interactive content in a way PDF doesn't. Internal structure and annotation is also preserved. I'm not sure I understand the formatting comment, you seem to be saying formatting is not lost and supporting that with an example of how formatting is lost. Don't get me wrong, it should definitely be easier to save, index and otherwise manipulate web pages. But out of the the trivial methods, 'print to PDF' is one of the poorer methods.

Re: WorldBrain's Memex: Bookmarking for the power users of the web

#64

Earlier quoted context omitted.

Every web page I'll ever want to refer to, ever again. There are no good reasons for exceptions to this technique, imho. If the PDF version is difficult to read - which it rarely is, by the way - all I need to do is open the PDF and use the links in the page header to go visit the site again - all the details about the page are still there in the PDF, links are still clickable, etc. And if its really important, and I…

> There are no good reasons for exceptions to this technique, imho Dynamic web sites? A PDF of my bank website isn't going to help me much.

If the intention is to save data from your bank website, you're probably going to have to jump through hoops anyway, assuming your bank is doing its job. (Or just remember to use Reader mode first..)

However, if the intention is to just save a link to the bank website for future reference, my technique still works since every page in the PDF produced contains a header with the URL - just like a normal bookmark.

Re: WorldBrain's Memex: Bookmarking for the power users of the web

#65
post #63

Earlier quoted context omitted.

$ pdftotext | grep sometext ^^ first advantage Also: a) Formatting is not lost, it just changes to fit the default paper size I've got selected (A4) but doesn't really make much difference, since its a snapshot, and b) URL is right there in the Header of the PDF, and is clickable, so no - not really an issue. This archive also functions as a bookmark collection as well as an offline copy for future reference .. (Disc…

You can grep the saved archives and they often save working copies of local interactive content in a way PDF doesn't. Internal structure and annotation is also preserved. I'm not sure I understand the formatting comment, you seem to be saying formatting is not lost and supporting that with an example of how formatting is lost. Don't get me wrong, it should definitely be easier to save, index and otherwise manipulate…

It depends on the site - but I haven't found 'lost formatting' to be an issue at all - since, when I want to do a granular search I'm using 'pdftotext' to search on plaintext, and when I find a PDF of interest, I open it and can go directly back to the web page from which it was printed by way of the footer/header which contains a clickable URL.

Most of the time though, the formatting isn't an issue. It depends on the site though - some authors produce stuff that doesn't look good as PDF, even if the content is still there. That doesn't bug me much.

Re: WorldBrain's Memex: Bookmarking for the power users of the web

#66
post #49

Earlier quoted context omitted.

nothing beats print-to-PDF What's the advantage over browser's built-in 'save entire page' option? Print to PDF loses formatting and obscures the URL you got the thing from.

$ pdftotext | grep sometext ^^ first advantage Also: a) Formatting is not lost, it just changes to fit the default paper size I've got selected (A4) but doesn't really make much difference, since its a snapshot, and b) URL is right there in the Header of the PDF, and is clickable, so no - not really an issue. This archive also functions as a bookmark collection as well as an offline copy for future reference .. (Disc…

Grep is not much of an advantage when Mac, Windows and Linux all have as-you-type full-text search of common formats like HTML and PDF.

Re: WorldBrain's Memex: Bookmarking for the power users of the web

#67
Does WorldBrain Memex save any data about the sites I bookmark?

I‘ve been using Onenote for the past 10 years to bookmark or save websites.

It had worked OK to share from mobile but my Onenote notebook is now approaching 10 GB in size.

And I have a pretty bad experience with syncing as it doesn‘t reliably sync in the background if I don‘t regularly open the app on mobile (especially on iOS).

Re: WorldBrain's Memex: Bookmarking for the power users of the web

#68
This is an interesting perhaps meta-relevant topic for HN.

How many of us bookmark or otherwise record interesting posts from here and elsewhere?

How many of us ever refer that accumulated digital memory?

I have about 7,000 links with notes accumulated over the last few decades.

I’ve read a lot of them, but the hard to acknowledge reality is that even with a refined workflow, recording my links in a near perfect taxonomy, to a repository with full text search and spaced repetition reminder cards, the things I remember are those that I took the time to read.

I suspect most people here has a comparable metric to share.

Maybe the best bookmark repository is nul:

Re: WorldBrain's Memex: Bookmarking for the power users of the web

#69

Something I noticed when I use "Read Later" style applications to save pages is that I will, most of the time, forget about how I arrived at a certain page. This is important to me because it gives me the context to decide a perspective on the page. If I was able to save pages while also knowing where I found them and maybe make a comment about why I found it interesting, then I would be able to organize my knowledge…

https://histre.com/ has tree-style web history, taking notes on those web pages, and more. Disclaimer: I'm the founder. It automatically creates a knowledge base for you. The paths you took to arrive at a piece of information is just one part of the puzzle that it puts together for you. The main idea is that we throw away a lot of the signal we generate while doing things online and this can be put to good use for ou…

Hmm, I thought I was the only one who thought like that. I've just been exporting entire browser trees from Tree Style Tabs (with hierarchy) at once and attaching them to a page in my Zettelkasten or another part of my knowledge base.

It is great to have the entire context of my browsing session to go back to.

Re: WorldBrain's Memex: Bookmarking for the power users of the web

#70
My tools of choice for advanced bookmarking and offline read:

* org-mode [1].

* org-board [2] for offline archiving.

* Org Capture [3] for getting links or text chunks from browser.

* git repo for tracking history.

With org-mode I can create really complex connections between articles and citations, add tags, have TODO lists and many more. To visualize things and connections, org-mind-map [4] can be useful. Because everything is text, grep, ripgrep, ag, xapian and other similar tools works without problems.

I'm aware this setup isn't for everyone (you need to be Emacs user), but I still need to find proper alternative with this amount of flexibility, keeping everything in plain text format.

[1] https://orgmode.org/

[2] https://github.com/scallywag/org-board

[3] https://chrome.google.com/webstore/detail/org-capture/kkkjlf...

[4] https://github.com/the-humanities/org-mind-map

Post reply on HN