Live data from Hacker News

Show HN: Open Paperless – Scan, index, and archive paper documents

github.com

71–80 of 104 posts

Re: Show HN: Open Paperless – Scan, index, and archive paper documents

#72

Any good open source desktop software with linux support to do this? I don't see why I would personally want a web app for this.

Well, if you have a home server, having a web app works quite well. But if you don't, then a desktop app would probably be better.

Re: Show HN: Open Paperless – Scan, index, and archive paper documents

#73

Earlier quoted context omitted.

You'd be surprised how useful having all of your documents digitally available can be. I've been doing this for years already (albeit using Google Drive - yes, some people dont trust Google - I understand that, but doesn't bother me) and these are some of the common use-cases where I find it really useful: * Tax returns. This alone makes it worthwhile. * Call-centers. You'll often have the reference/account number/et…

> Now I top-up with a home scanner attached to my LAN - if you're looking to buy a scanner for home, make sure you get one with an automatic document feeder that can do both sides at once so you can just chuck the papers in and hit go, then collect the PDFs from your network drive. Any suggestions for one? I have a Canon P-208, but it's close to useless for batch scanning.

I'm partial to Fujitsu ScanSnap document scanners. It's worth paying the extra money for a Fujitsu, imo.

The only catch is that you are supposed to replace the pick roller and pad assembly every year or so.

I would recommend buying a couple of extras in advance, as they will get harder to find and more expensive as time goes on.

Having said that, I didn't replace my pick roller/pad assembly until after 9 years of operation. After a few years it would have trouble feeding multipage documents, but most of the stuff I was scanning was only one or two pages. When I did eventually replace those parts, the scanner was basically as good as new.

Re: Show HN: Open Paperless – Scan, index, and archive paper documents

#74

Earlier quoted context omitted.

I like this idea. I only need to keep 10 years and almost never need to retrieve anything. So if I use 10 boxes I can just throw away the oldest.

Would it cause any real issues if you were unable to retrieve something? If the answer is no, toss the other 9 boxes

This. I dispose of almost all the paper I ever receive. Anything really important can usually be reissued by whoever issued it.

Re: Show HN: Open Paperless – Scan, index, and archive paper documents

#75

This is arguably a lot more than I need. I'm a hoarder in that I have every email I've ever sent or received (bar junkmail), and every piece of paper I've ever received. Most of my paper is now scanned - I think I have two boxes left in my garden shed. I don't bother with OCR because search doesn't help me when I don't know what to search for (e.g. invoice for a jumper I bought in 2010 - fashion labels rarely call th…

A few years ago I was involved with a startup that built a document management system for consumers, and we actually got pretty good results with OCR + automatic tagging based on a very simple database that maps keywords to tags.

Let's say you want to auto-tag bills and other documents from your ISP. So you add the ISP's name, phone number, website address etc. into the database - any uniquely-identifying keywords that typically appear on the documents that they send. Now any document that contains these keywords will get tagged as "ISP", making it very easy to find in the future.

Even if the OCR quality isn't perfect, at least one of these keywords will most likely get matched.

Another example - you could add the names of your family members as keywords, making it easy to find all documents related to Jenny or Susan.

You could argue that full-text search would achieve the same result, but uploading documents into the system and having them auto-tagged as "ISP", "car-payments", "Walmart", "Susan" and so on feels a little bit like magic, as if the system is actively helping you organize your papers.

The keyword approach is also very easy to understand and tweak, unlike more rigorous but opaque methods of document clustering (such as tf-idf).

Re: Show HN: Open Paperless – Scan, index, and archive paper documents

#76
post #31

a nontrivial name conflict with Paperless ( https://github.com/danielquinn/paperless ) ...

This had to be on purpose, which is quite disappointing.

If you had been following the story of these projects you would know that it is paperless (and not open paperles or mayan) that has been leaching off other projects and people.

Re: Show HN: Open Paperless – Scan, index, and archive paper documents

#77
post #43
post #31

a nontrivial name conflict with Paperless ( https://github.com/danielquinn/paperless ) ...

They even share the exact same tagline: "Scan, index, and archive all of your paper documents". If Open Paperless some kind of fork of Paperless?

Don't go there. Leave sleeping dogs lie as they say.

Re: Show HN: Open Paperless – Scan, index, and archive paper documents

#78
post #75

This is arguably a lot more than I need. I'm a hoarder in that I have every email I've ever sent or received (bar junkmail), and every piece of paper I've ever received. Most of my paper is now scanned - I think I have two boxes left in my garden shed. I don't bother with OCR because search doesn't help me when I don't know what to search for (e.g. invoice for a jumper I bought in 2010 - fashion labels rarely call th…

A few years ago I was involved with a startup that built a document management system for consumers, and we actually got pretty good results with OCR + automatic tagging based on a very simple database that maps keywords to tags. Let's say you want to auto-tag bills and other documents from your ISP. So you add the ISP's name, phone number, website address etc. into the database - any uniquely-identifying keywords th…

Everything you say is true, and the value, I think, is clear. The part I don't like is that I have to create a database manually. Granted, the results will save me time as I don't have to manually tag the routine.

Food for thought.

Re: Show HN: Open Paperless – Scan, index, and archive paper documents

#79
post #34

Earlier quoted context omitted.

you can just ask your recipeint to send you some uuid on the letter after they receive it. i have no idea what problem you are trying to solve though.

i once heard of a case were the employer send an unrelated document to an employee. later on, the payment stopped and the employer claimed that they fired him at that time. I don't recall how that case ultimately turned out, but maybe something like that? would be incredibly rare though and of dubious worth for mostly anyone

I believe there are courier services that help guarantee document delivery to the correct person, i.e. for sending/serving court summons.

Re: Show HN: Open Paperless – Scan, index, and archive paper documents

#80
post #3

Will this automatically center and apply perspective transforms to pictures taken with phone cameras?

The best version of this that I've seen is Scanner Pro by Readdle. I had to scan three months worth of food receipts for an insurance claim and this feature was a lifesaver.
Post reply on HN