Live data from Hacker News

WorldBrain's Memex: Bookmarking for the power users of the web

getmemex.com

141–150 of 215 posts

Re: WorldBrain's Memex: Bookmarking for the power users of the web

#141
post #72

Earlier quoted context omitted.

Ok, so we seem to agree print-to-pdf loses formatting. I share your interest in and fascination with this (weirdly irksome and edgecasey) problem but just about any modern browser provides better facilities for saving web pages with higher fidelity than 'print to pdf'. Print to pdf is so easy to beat, you'd have to go out of your way to find a way to not-beat it - say, saving just 'page source'.

PDF gives me an off-line readable version of the website, and is pretty compatible with pdftotext as a pipeline tool... I'm not sure that formatting is such a huge issue - I'm yet to find a web page I can't extract some meaningful info from, later on ..

an off-line readable version of the website

As does 'Save as: Web Archive' in Safari or, if you tell it to store offline, 'Add to Reading List'. In Chrome you can save pages to .mht. All of these are single-keystroke, better ways to locally archive a web page.

Re: WorldBrain's Memex: Bookmarking for the power users of the web

#143

This is still not as powerful as my one, simple trick to handle all bookmarks, ever: Print to PDF. I've been doing it since last century, and I have 10's of thousands of PDF's of every single web page I've ever found interesting, sitting right there in a directory on my computer. Its indexable, searchable, grok'able, available off-line, allows me to harvest data without fuss, and gives me access to anything I can rem…

Wouldn't being able to print to epub be better? epubs are just zipped html files after all so the conversion is more direct and you don't lose any info.

Re: WorldBrain's Memex: Bookmarking for the power users of the web

#144
post #90

This is an interesting perhaps meta-relevant topic for HN. How many of us bookmark or otherwise record interesting posts from here and elsewhere? How many of us ever refer that accumulated digital memory? I have about 7,000 links with notes accumulated over the last few decades. I’ve read a lot of them, but the hard to acknowledge reality is that even with a refined workflow, recording my links in a near perfect taxo…

I've been fighting my ~3000 (~70% untagged) bookmarks for a while. Right now, I've gave up on silly tags like "Postgres" or "Python". Currently, I'm trying to adapt the bookmark concept into different uses cases. The main one is sessions, but I have a few others niche ones, like "read later" and "a tool a day". Honestly, my takeaway from managing my bookmarks, is that, snapshotting a session, is the closest thing I h…

I've had a similar experience.

I used to meticulously sort and tag individual bookmarks but rarely review them. Storing sessions and other "playlists" of bookmarks puts them in a form that I actually return to.

Plus this method takes far less time and effort than tagging and bagging pages according to an ever-expanding set of custom taxonomies.

I'm sure others have been using bookmarks this way for a while but it felt like a revelation to me :)

Re: WorldBrain's Memex: Bookmarking for the power users of the web

#145

I tried this for a couple of months but the search results were disappointing. I think something like "Stealth" ( https://github.com/cookiengineer/stealth ) will prove to be a better strategy.

Hmm, how could the results have been disappointing? It just searches for text, how bad can it be?

I would search for terms that I knew were on pages that should be indexed but they wouldn't be in the results list.

Re: WorldBrain's Memex: Bookmarking for the power users of the web

#146
post #132

This is still not as powerful as my one, simple trick to handle all bookmarks, ever: Print to PDF. I've been doing it since last century, and I have 10's of thousands of PDF's of every single web page I've ever found interesting, sitting right there in a directory on my computer. Its indexable, searchable, grok'able, available off-line, allows me to harvest data without fuss, and gives me access to anything I can rem…

To be a little pedantic, while this is a fantastic idea, is it really bookmarking? What you’ve done instead is compiled a personal digital library; akin to a Kindle

Inasmuch as every PDF still has the URL to the original page, for my uses - I'd say yes, it is bookmarking.

Its not like I'd cut the spine off every book in my library and create an index out of the covers ..

Re: WorldBrain's Memex: Bookmarking for the power users of the web

#147

This is still not as powerful as my one, simple trick to handle all bookmarks, ever: Print to PDF. I've been doing it since last century, and I have 10's of thousands of PDF's of every single web page I've ever found interesting, sitting right there in a directory on my computer. Its indexable, searchable, grok'able, available off-line, allows me to harvest data without fuss, and gives me access to anything I can rem…

Wouldn't being able to print to epub be better? epubs are just zipped html files after all so the conversion is more direct and you don't lose any info.

PDF just has better tooling for search, which supplants my need for proper formatting since by the point I'm searching for things in my history its the content that matters to me anyway, exclusive of the original formatting.

Re: WorldBrain's Memex: Bookmarking for the power users of the web

#148
post #141

Earlier quoted context omitted.

PDF gives me an off-line readable version of the website, and is pretty compatible with pdftotext as a pipeline tool... I'm not sure that formatting is such a huge issue - I'm yet to find a web page I can't extract some meaningful info from, later on ..

an off-line readable version of the website As does 'Save as: Web Archive' in Safari or, if you tell it to store offline, 'Add to Reading List'. In Chrome you can save pages to .mht. All of these are single-keystroke, better ways to locally archive a web page.

Cool. Let me know when I can extract the full text contents from such files using common, built-in tools on your average fresh install of MacOS/Linux.

PDF works just fine. It presents a feasible view of the original data, and allows for data harvesting with ease.

Re: WorldBrain's Memex: Bookmarking for the power users of the web

#149
post #141

Earlier quoted context omitted.

an off-line readable version of the website As does 'Save as: Web Archive' in Safari or, if you tell it to store offline, 'Add to Reading List'. In Chrome you can save pages to .mht. All of these are single-keystroke, better ways to locally archive a web page.

Cool. Let me know when I can extract the full text contents from such files using common, built-in tools on your average fresh install of MacOS/Linux. PDF works just fine. It presents a feasible view of the original data, and allows for data harvesting with ease.

Let me know when I can extract the full text

You can extract the full text from these (with whatever tools you like) with better fidelity than you can from a pdf, which is a lossy conversion from the same source. This seems to barely merit debating, unless I'm missing something.

PDF works just fine.

I'm sure it works for you and I'm not harbouring any delusions I'm going to talk you out of your decades-established workflow. But for anyone looking for ways to keep track of web pages, thinking about building tools in this space, etc - no, PDF is not a good way to archive web pages, either manually or programmatically.

Re: WorldBrain's Memex: Bookmarking for the power users of the web

#150
post #55

Earlier quoted context omitted.

I would be careful with using this method and check the generated PDF versions with your eyes before writing them off as "all is good, it is archived now". I recently got bitten by that, when I was trying to print out some page in Chrome, and it was rendering as a bunch of white space surrounded by some elements from the page, but without any actual content I cared about. Turns out, my situation isn't that uncommon f…

I've since learned in this thread that Chrome and Firefox are not as good as Safari for this technique - it hasn't impacted me much since I only use Chrome/Firefox for development, mostly. And although I do occasionally check the produced PDF's, the layout doesn't matter to me at all since I use a cmd-line grep or combination of 'pdftotext' to find the page, open the PDF, and click the link to go to the original web…

[deleted]
Post reply on HN