Live data from Hacker News

Show HN: Open Paperless – Scan, index, and archive paper documents

github.com

81–90 of 104 posts

Re: Show HN: Open Paperless – Scan, index, and archive paper documents

#81

Earlier quoted context omitted.

This had to be on purpose, which is quite disappointing.

If you had been following the story of these projects you would know that it is paperless (and not open paperles or mayan) that has been leaching off other projects and people.

how did you get that idea? and why would DanielQuinn even bother? He doesnt accept any donations and instead encourages people to donate to charity...

First commits:

* open Paperless Dec 18, 2017

* Paperless Dec 20, 2015

* mayan edms Feb 3, 2011

Re: Show HN: Open Paperless – Scan, index, and archive paper documents

#82
For those wondering about the relationship between Mayan EDMS, Paperless and Open Paperless here is a story line summary of the saga.

Roberto Rosario (the creator of Mayan) is a very well known name in the Django, Python, document management, maker, hacking, open health and open source in the goverment circles.

- https://speakerdeck.com/siloraptor - https://en.wikipedia.org/wiki/Roberto_Rosario - https://www.pycon.it/conference/p/roberto-rosario - http://pyvideo.org/djangocon-us-2014/liberation-and-moderniz... - https://cpucadviceletters.org/login/?next=/ - https://twit.tv/shows/floss-weekly/episodes/253 - https://en.wikipedia.org/wiki/Mayan_(software) - https://www.youtube.com/watch?v=rubzEAojf-k

Mayan EDMS was initially released in February 3, 2011 (Wikipedia and git log). In June 2015, Roberto gave a workshop in DjangoCon named From zero to paperless with Mayan EDMS (https://archive.is/FDpYS). Daniel Quinn (the creator of Paperless) also attended and presented at the same DjangoCon event (https://vimeo.com/135907408) and 6 months later after working on it for several months (Daniel's own words), he released Paperless on December 20, 2015 (https://github.com/danielquinn/paperless/commits/master?afte...). By January 24, 2016, Paperless had "exploded in popularity" (https://twitter.com/danielagquinn/status/691242822431830016).

Both projects used Python, Django, same Django 3rd party apps like DjangoSuit, same document consumer model, same OCR engine, REST API, among other things. On the surface it appeared that Paperless was a copy of Mayan EDMS concepts and implementations without giving credit or mention. Many additions were planned for Paperless that were features and implementations already in Mayan (https://www.reddit.com/r/selfhosted/comments/44mh88/scan_ind...).

A separate point of contention was that the name "Paperless" had been in use by other projects much earlier that Daniel's Paperless (https://github.com/search?utf8=%E2%9C%93&q=paperless&type=). Since there is no trademark on the name or description, other projects appeared with the same name and description (https://github.com/lrnt/paperless).

On March 15, 2016, Daniel presented Paperless at CodeNode (https://skillsmatter.com/skillscasts/7843-intro-to-paperless).

It was Daniel's February 27, 2016 tweet suggesting to be paid to work on Paperless that sparked the animosity between the users of the two projects (https://twitter.com/danielagquinn/status/703629488932970500).

Many heated debates ensued. Even then, the main critique of Paperless remained technical, but lack of maturity and implemenation was described by one Reddit users as: "I've looked into paperless and it currently lacks a lot of...nearly well everything. Maybe in a year or two" (https://www.reddit.com/r/linux/comments/6m9evn/want_to_go_pa...)

On April 9, 2016, Daniel added a reference to Mayan to the documentation of Paperless (https://github.com/danielquinn/paperless/commit/674d54ec3878...).

On April 17, 2016, Daniel posted on his old twitter account: "It looks like my idea for Paperless wasn't all that unique. This other project uses a lot of the same tools: http://www.mayan-edms.com" (https://twitter.com/danielagquinn/status/721726208606646272).

On April 14, 2017, Daniel Quinn posted in his blog a summary of his experiences at DjangoCon Europe 2017 where he mentions meeting Roberto in person. He describes Roberto as a "rival geek" in what appears to be jest and uses positive adjectives to describe Roberto in the rest of the post. (https://danielquinn.org/blog/djangocon-2017/)

On April 16, 2017 Daniel posted a tweet mentioning the popularity Paperless (https://twitter.com/danielagquinn/status/853701257051205632).

The last release of Paperless is made on Sep 9, 2017.

On Oct 18, 2017 Daniel posted: "I changed my Twitter name! This isn't me any more, so if you're looking for me, you should keep head over to @danielagquinn." (https://twitter.com/searchingfortao/status/92077862371561062...). Only 7 commits have been made to Paperless since with the last commit happening on Novermber 5, 2017.

On December 18, 2017 a user named "zhoubear" anounced on Reddit's selfhoted "Open Paperless: Scan, index, and archive all of your paper documents" (https://www.reddit.com/r/selfhosted/comments/7kjocg/scan_ind...). It turned out that Open Paperless was a forked Mayan EDMS with cosmetic changes but with copyrights changed and no attribution to Mayan EDMS. After a much heated debate, copyrights and attributions were restored and the project's description has been updated to show that it is a new front end for Mayan among other usability changes meant for home users.

In 4 days, Open Paperless has surpassed Mayan EDMS in popularity on Github.

No posts or comments from Roberto can be found in reference of Paperless or Open Paperless.

https://twitter.com/search?q=paperless%20from%3Asearchingfor...

Re: Show HN: Open Paperless – Scan, index, and archive paper documents

#83

Earlier quoted context omitted.

You'd be surprised how useful having all of your documents digitally available can be. I've been doing this for years already (albeit using Google Drive - yes, some people dont trust Google - I understand that, but doesn't bother me) and these are some of the common use-cases where I find it really useful: * Tax returns. This alone makes it worthwhile. * Call-centers. You'll often have the reference/account number/et…

> Now I top-up with a home scanner attached to my LAN - if you're looking to buy a scanner for home, make sure you get one with an automatic document feeder that can do both sides at once so you can just chuck the papers in and hit go, then collect the PDFs from your network drive. Any suggestions for one? I have a Canon P-208, but it's close to useless for batch scanning.

HP OfficeJet Pro 8620. It's an all-in-one, but interestingly the best scanner for such purposes I've found so far (I'd love to find a better one).

Scans both sides, up to 50 pages in a batch. Open-source drivers, works well with Linux. Ethernet. Can drop it on on your SMB share.

Re: Show HN: Open Paperless – Scan, index, and archive paper documents

#84
post #81

Earlier quoted context omitted.

If you had been following the story of these projects you would know that it is paperless (and not open paperles or mayan) that has been leaching off other projects and people.

how did you get that idea? and why would DanielQuinn even bother? He doesnt accept any donations and instead encourages people to donate to charity... First commits: * open Paperless Dec 18, 2017 * Paperless Dec 20, 2015 * mayan edms Feb 3, 2011

- Feb 27, 2016, "I wish I could be paid to work on my side projects like @TheTweetPile and Paperless". (https://twitter.com/danielagquinn/status/703629488932970500)

- Mar 13, 2016, "Added Gratipay" (https://github.com/danielquinn/paperless/commit/500cee1d9d06...)

- Mar 13, 2016, "Gratipay sucks because it depends on PayPal. Long live bitcoin" (https://github.com/danielquinn/paperless/commit/705a5b8a8468...)

- Apr 9, 2016, "Added reference to Mayan EDMS" (https://github.com/danielquinn/paperless/commit/674d54ec3878...)

- Finally on Dec 26, 2016, "Give to the UNHCR" (https://github.com/danielquinn/paperless/commit/9ea39aeecb4f...)

Re: Show HN: Open Paperless – Scan, index, and archive paper documents

#85
post #81

Earlier quoted context omitted.

If you had been following the story of these projects you would know that it is paperless (and not open paperles or mayan) that has been leaching off other projects and people.

how did you get that idea? and why would DanielQuinn even bother? He doesnt accept any donations and instead encourages people to donate to charity... First commits: * open Paperless Dec 18, 2017 * Paperless Dec 20, 2015 * mayan edms Feb 3, 2011

In contrast, donations by Roberto Rosario:

- Gold sponsor of Django schema migration work (http://www.aeracode.org/2013/04/11/whole-load-kickstarting/)

- Gold sponsor of the Django multi template engine work (https://myks.org/en/multiple-template-engines-for-django/fun...)

- Gold sponsor for the Django REST framework work (http://www.django-rest-framework.org/topics/kickstarter-anno...)

- Gives money and time to groups looking to increase participation of women to STEM fields like #IncludeGirls (https://medium.com/ladies-storm-hackathons/include-girls-137...) (https://aldia.microjuris.com/2013/10/07/ciclo-de-conferencia...) (https://www.eventbrite.com/e/python-django-workshop-tickets-...).

- Has started initiatives like PythonLatino and DjangoLatino to increase Latinamerican participation in Django and Python.

- Was chairman of the group that got the PyCon started in Cuba after travel restrictions were relaxed.

- Author or many Python and Django popular projects that help a lot of people like Awesome Django (https://github.com/rosarior/awesome-django).

- Sponsored STEM books.

- PSF candidate with endorsements by Python leaders like Anna Ravencroft, co-author of the Python Cookbook along with her husband Alex Martelli, fellow of the Python Software Foundation (https://en.wikipedia.org/wiki/Alex_Martelli) (https://wiki.python.org/moin/PythonSoftwareFoundation/BoardC...).

Many more but links are dead and don't want the spend the rest of the day on the internet archive.

Re: Show HN: Open Paperless – Scan, index, and archive paper documents

#86
post #30

Looks great. Love the idea behind it, but... There is at least one country (mine - Switzerland) which is not able to use software like yours. The problems are the current laws that force people and organizations to store physical copies of the documents (for several years). Electronic documents have no value in front of the law, which is why we have no choice but to do all of that offline, manually. I've tried many a…

You could store them chronically on paper and do the actual sorting on your computer or server.

Re: Show HN: Open Paperless – Scan, index, and archive paper documents

#87
post #81

Earlier quoted context omitted.

how did you get that idea? and why would DanielQuinn even bother? He doesnt accept any donations and instead encourages people to donate to charity... First commits: * open Paperless Dec 18, 2017 * Paperless Dec 20, 2015 * mayan edms Feb 3, 2011

In contrast, donations by Roberto Rosario: - Gold sponsor of Django schema migration work ( http://www.aeracode.org/2013/04/11/whole-load-kickstarting/ ) - Gold sponsor of the Django multi template engine work ( https://myks.org/en/multiple-template-engines-for-django/fun... ) - Gold sponsor for the Django REST framework work ( http://www.django-rest-framework.org/topics/kickstarter-anno... ) - Gives money and time t…

There is no name conflict with mayan edms though...?

i mean, yeah, Roberto Rosario is a great guy and mayan edms is objectively better than Paperless, i've no idea how that has any kind of relevance to the name conflict...?

Re: Show HN: Open Paperless – Scan, index, and archive paper documents

#88
post #46
post #36

Earlier quoted context omitted.

Thank you. This brings up another problem we have with our laws: Corporates are not allowed to store such documents on servers which are outside of our country. And that's usually the case with clouds services, because of obvious reasons.

That’s not a problem with your laws, but more a problem with SaaS. Data should always stay as local as possible, ideally even within of your own organization.

How is keeping data local a good idea? Given how often there is a major breach reported, everything points to this being the opposite -- there aren't enough security specialists to go around for every company. Why not take advantage of the work done by the major cloud providers or SaaS companies instead of redoing everything in house all the time. It's like people want to actually solve the same problems over and over for eternity.

Re: Show HN: Open Paperless – Scan, index, and archive paper documents

#89
post #75

This is arguably a lot more than I need. I'm a hoarder in that I have every email I've ever sent or received (bar junkmail), and every piece of paper I've ever received. Most of my paper is now scanned - I think I have two boxes left in my garden shed. I don't bother with OCR because search doesn't help me when I don't know what to search for (e.g. invoice for a jumper I bought in 2010 - fashion labels rarely call th…

A few years ago I was involved with a startup that built a document management system for consumers, and we actually got pretty good results with OCR + automatic tagging based on a very simple database that maps keywords to tags. Let's say you want to auto-tag bills and other documents from your ISP. So you add the ISP's name, phone number, website address etc. into the database - any uniquely-identifying keywords th…

Out of curiosity what is the state of the art today for extracting text or other data from scanned documents (forms, legal docs, receipts, etc) ?

Re: Show HN: Open Paperless – Scan, index, and archive paper documents

#90
post #88
post #46

Earlier quoted context omitted.

That’s not a problem with your laws, but more a problem with SaaS. Data should always stay as local as possible, ideally even within of your own organization.

How is keeping data local a good idea? Given how often there is a major breach reported, everything points to this being the opposite -- there aren't enough security specialists to go around for every company. Why not take advantage of the work done by the major cloud providers or SaaS companies instead of redoing everything in house all the time. It's like people want to actually solve the same problems over and ove…

> How is keeping data local a good idea?

Sometimes you want to avoid competitors, or foreign governments, or intelligence agencies from accessing your data. Sometimes you want to avoid any kind of tracking or metadata analysis. Sometimes you want to avoid external points of failure.

Often it is cost - even including wages, AWS or GCP are a factor 10 to 100 more expensive than what you can run on rented dedicated systems from a local hoster, or by colocating. Even your own datacenter can be cheaper, and end up with identical uptime.

And then there's a question of trust. A simple question: Would you trust a Chinese SaaS company to securely store all your customer's data, and run all your services? Would you trust a Romanian one? Why would I trust an American one?

Post reply on HN