Live data from Hacker News

WorldBrain's Memex: Bookmarking for the power users of the web

getmemex.com

161–170 of 215 posts

Re: WorldBrain's Memex: Bookmarking for the power users of the web

#161

This is still not as powerful as my one, simple trick to handle all bookmarks, ever: Print to PDF. I've been doing it since last century, and I have 10's of thousands of PDF's of every single web page I've ever found interesting, sitting right there in a directory on my computer. Its indexable, searchable, grok'able, available off-line, allows me to harvest data without fuss, and gives me access to anything I can rem…

> Yeah, my computer is fast enough that I can just do "find . -name '*.pdf' -exec pdftotext {} \; | grep -i someSearchTerm" and come back later.

You may be interested in my ripgrep-all [1] tool, it should allow you to search those tens of thousands of PDFs in under a second (with hot cache)

[1] https://phiresky.github.io/blog/2019/rga--ripgrep-for-zip-ta...

Re: WorldBrain's Memex: Bookmarking for the power users of the web

#162
post #152

Earlier quoted context omitted.

fully-searchable, indexable, accessible-in-offline Probably the shortest version of the point I'm trying to make is that every current browser does a much better job of providing you this than printing to PDF. If you rely on this as a personal web archiving system, you're going to lose data in the most irritating way - data you thought you collected but actually didn't.

I don't understand your point at all. In what way is a browser going to give me information that is not available to me unless I'm online? PDF's of sites I've visited have all the data I need - the stuff I read that then prompted me to print to PDF. I've searched and I'm yet to find a single PDF in my collection that doesn't have the info that prompted me to save it in the first place. I understand you believe your p…

In what way is a browser going to give me information that is not available to me unless I'm online?

The alternatives I'm talking about are absolutely available offline. I don't understand why you keep arguing against a position I've never taken.

Re: WorldBrain's Memex: Bookmarking for the power users of the web

#163
post #160

I have been using Pocket [0], Instapaper [1] and Pinboard [2] over the years. I am currently using Pocket and Pinboard in parallel: articles / websites that I want to read later are sent to Pocket (untagged), websites that I might want to get to back later are tagged and sent to Pinboard. While my archive on Pinboard works quite well I am very disappointed by the support. Either the developer does not answer at all o…

Oli here from WorldBrain.io

A bit of a timing issue here :) We are about to have a big release with lots of improvements, including an API, performance, UX and bug fixes.

Our API will be served via Storex https://medium.com/@WorldBrain/storexhub-an-offline-first-op...

Re: WorldBrain's Memex: Bookmarking for the power users of the web

#164

I'm surprised nobody has mentioned https://web.hypothes.is/ --- it's a non-profit trying to solve the same idea. They are actually trying to advance on the ideas of the w3c annotation's working group and do everything open source.

It really frustrates me that its 2020 and they still don't have a real extension for Firefox, the only thing preventing me from regularly using it.

We are about to develop a bi-directional integration with Hypothes.is and Memex together with the Hypothes.is team.

(Oli here from WorldBrain.io)

Re: WorldBrain's Memex: Bookmarking for the power users of the web

#165
post #123

I remember using this software last time, it is wayyyy~ too buggy, it stalls, crashes, and slows down the browser. Also that import feature is actually crawling the site, beware if you are using a proxy or something with rate limit.

Is it still buggy? I had a similar experience, but it looks and feels a lot smoother and faster now.

It's way better now but still some way to go!

Kind of a bummer that this trended one week too early - next week we'll publish a big release with performance, UX and stability improvements.

Re: WorldBrain's Memex: Bookmarking for the power users of the web

#166
post #60

This is still not as powerful as my one, simple trick to handle all bookmarks, ever: Print to PDF. I've been doing it since last century, and I have 10's of thousands of PDF's of every single web page I've ever found interesting, sitting right there in a directory on my computer. Its indexable, searchable, grok'able, available off-line, allows me to harvest data without fuss, and gives me access to anything I can rem…

I discovered Zotero for this use. I don't have any use of its bibliographical abilities, but it stores web pages and PDF articles fine, and is searchable, etc.

Yes, because of all the metadata you can preserve with the save.

Also, frequently what yo want to save is a link to a book or an article (or a Wikipedia article, say) and Zotero recognizes many of these formats and saves them correctly, it's toolbar icon even changes to let you know.

It's a fantastic tool. It's got cloud storage, and it has an API (which I've used -- it's super easy).

Re: WorldBrain's Memex: Bookmarking for the power users of the web

#167

I was one of the first backers who paid for their lifetime subscription. Except it was nowhere to be found and my account was essentially "free". Nice way to treat your early adopters, guys!

Oli here from WorldBrain.io

I am confused. We never offered lifetime subscriptions. However what we did is give people who supported us between 4 and 5 times the supporter amount in credits they can use to upgrade. We sent an email around to everyone at the end of last year.

(You're the only "Nikolay" in our customer DB, so I gave it a check and you have tons of credits still left)

The reason it was "free" for you at checkout is because the credits were applied.

Hope that clarifies things.

Re: WorldBrain's Memex: Bookmarking for the power users of the web

#168

Earlier quoted context omitted.

I would search for terms that I knew were on pages that should be indexed but they wouldn't be in the results list.

I installed it yesterday and noticed that it doesn't actually index much. It should be, but it's not, the pages aren't added. If they are in the index, it finds them in a search, but very few pages are.

Oli here from Memex.

It may be that you have not touched the indexing preferences (which only index pages that are visited for more than 5 seconds)

Is that the reason, or does it still not work?

Re: WorldBrain's Memex: Bookmarking for the power users of the web

#169

I had used this for some time in the past (on and off), periodically. One caveat I found was it was taking a huge toll on my browser (often, I felt the lags). Not sure if that's the problem now or not. Eventually, I ended up not using it and started using other tools (specific tools for specific tasks).

Kind of a bummer that this trended one week too early - next week we'll publish a big release with performance, UX and stability improvements.

Especially we focused on indexing and page load performance.

Re: WorldBrain's Memex: Bookmarking for the power users of the web

#170
post #152

Earlier quoted context omitted.

fully-searchable, indexable, accessible-in-offline Probably the shortest version of the point I'm trying to make is that every current browser does a much better job of providing you this than printing to PDF. If you rely on this as a personal web archiving system, you're going to lose data in the most irritating way - data you thought you collected but actually didn't.

I don't understand your point at all. In what way is a browser going to give me information that is not available to me unless I'm online? PDF's of sites I've visited have all the data I need - the stuff I read that then prompted me to print to PDF. I've searched and I'm yet to find a single PDF in my collection that doesn't have the info that prompted me to save it in the first place. I understand you believe your p…

PDF conversion throws away almost all of the structure contained in HTML. Tools like pdf2text then try to reconstruct some of that structure (such as the correct sequence of letters and words) using complex heuristics that don't always work.

They often do that successfully enough, especially if all you want is to grep for words. pdf2text also has a table mode that attempts to reconstruct table structured content. This is far less successful.

So depending on how you want to process your stored data, saving as PDF may or may not preserve sufficient information.

If you were storing pages for the purpose of extracting specific properties of things you are researching (say product information or the tree structure of HN threads), then throwing away all that structure makes it a lot more difficult or even impossible to reconstruct the information you need.

If I were storing pages for unknown future purposes, I wouldn't want to throw away any information I might need, and therefore I would never use PDF as an archival format.

But I understand that you store PDF files for a very specific purpose for which lossy PDF conversion happens to be good enough. So that's fine of course.

The only question I have is where I can find the source URL of the stored pages.

Post reply on HN