I've been doing the data storage tango for thirty years now, and the one constant is that everything requires maintenance or has to be moved to a different platform and software eventually. Today, everything is in Google Drive and Google Photos; about 600 GB worth of data and media. Maintenance consists of an annual Takeout ordeal with copies going on two different removable drives, one of which is stored in a so-cal…
Where do you guys live to need to store birth certificate in a fireproof bag? I can just login into my government website with a digital identity provider and print one legally valid copy any time. Same for any identity related certificate.
Ask HN: How do you manage your “family data warehouse”?
81–90 of 141 posts
Re: Ask HN: How do you manage your “family data warehouse”?
#82I have a fairly deep paper bin that I lay papers flat in. Any marginal paper documents or receipts that aren't obviously garbage and aren't obviously valuable go into this bin, always placed on the top.
Every 10-12 months I pick up the stack and remove, then shred, the bottom 1/4 of it - without review or remorse. If I didn't need them by then I never would.
I find that I actually use this buffer:
Every month or two there actually is some marginal paper document that I need to shuffle through and find. It's in reverse chronological order so that is very simple.
Re: Ask HN: How do you manage your “family data warehouse”?
#83Re: Ask HN: How do you manage your “family data warehouse”?
#84Earlier quoted context omitted.
You get access to a shell via ssh, https://www.rsync.net/products/platform.html Looking at, https://duckduckgo.com/?q=site%3Arsync.net%2Fproducts&t=ipho... You have rclone, sftp, rsync, restic, borg, synology and other. My understanding is there’s no bandwidth pricing. (I’m sure they’ll show up and clarify it. Amazing they always seem to do so)
Thanks - this is awesome. I will investigate. Might solve the problem I’ve had forever!
We would be very happy to have you.
Re: Ask HN: How do you manage your “family data warehouse”?
#85That’s it.
Re: Ask HN: How do you manage your “family data warehouse”?
#86A Synology NAS running Portainer ( https://www.portainer.io/ ) running Paperless NGX ( https://github.com/paperless-ngx/paperless-ngx ) This works better than I can possibly tell you. I have an Epson WorkForce ES-580W that I bought when my mother passed away to bulk scan documents and it scans everything, double-sided if required, multi-page PDFs if required, at very high speed and uploads everything to OneDrive, at…
I haven't looked at Portainer before, what does it give you that just running a Docker instance on the NAS doesn't do?
For home server usage, I just use it as a frontend to docker-compose files that gives me a pretty UI I can access from a web browser, instead of needing to SSH in to the home server.
Re: Ask HN: How do you manage your “family data warehouse”?
#87Re: Ask HN: How do you manage your “family data warehouse”?
#881. Beancount via org-mode notes, org-babel for transactions, bean queries from the org mode file for reporting, non perfect, a new thing, in the past I've used GNU Cash but I prefer to put anything in org so...
2-5. equally in org-mode org roam managed notes, non textual stuff as org-attachments linked directly in notes or indirectly "elisp:(start-process-shell-command" alike to run different stuff than mere opening.
I have to retain something for fiscal reasons, something for personal reasons and something are extras. It's not that perfectly crafted, but accessing anything by titles (org-roam-node-find) is like having anything in a graph, accessed via comfy search&narrow. When titles fails consult-rg on the org-mode root (~/notes) does the rest.
I've tries org-ql, to query notes for certain infos like "active contracts I have", but even with templating (yasnippet) it's not that immediate, I have some queries and templates but I do not really use if not very rarely. I've try LLama (Khoj) on my notes but well, results are very bad compared to direct textual mach access. So, yes, I have some clutter, I do not see it much, it does not consume significant amount of storage nor pollute my searches so... It's manageable, rarely I found something is the clutter, so it's not totally useless anyway. I am able to find anything noted so far, of course I do not log my life completely so sometimes I'm looking for things I do not have, that's rare but happen. In this cases the answer is "it depend".
I choose to live digitally in Emacs mostly because of integration: I have mails, todos, notes, anything in the same place/tool/model/paradigm like my mind is one for all.
Downsides: it's a personal thing, not really usable by other family members, not designed for collaborative works, it's possible works together but only as ISOLATED individual, all on their own desktop, possibly Emacs but not necessarily and definitively not "the same Emacs" sharing steps/results by various means like patches via mail, direct textual mails "I found this and that, I conclude that" etc.
Re: Ask HN: How do you manage your “family data warehouse”?
#89Safe deposit box at the bank for the physical records. A labeled manila envelope stored locally for each of us with the typical "identity" documents, SSN card, birth certificate, etc. Everything is stored locally, unencrypted, and backed up on a local removable HDD, and on BackBlaze. BackBlaze has saved my bacon more than once. As for physical objects, they're all temporarily mine, until I pass on. They're either too…
Re: Ask HN: How do you manage your “family data warehouse”?
#90A Synology NAS running Portainer ( https://www.portainer.io/ ) running Paperless NGX ( https://github.com/paperless-ngx/paperless-ngx ) This works better than I can possibly tell you. I have an Epson WorkForce ES-580W that I bought when my mother passed away to bulk scan documents and it scans everything, double-sided if required, multi-page PDFs if required, at very high speed and uploads everything to OneDrive, at…
I have Paperless watching a network folder from my NAS. The scanner scans directly into that folder and Paperless picks it up and processes it within a couple seconds. Even if Paperless went down for a bit (which happened before I upgraded my server), the files remain on the shared drive until it comes back online. Without that, I had paper piling up waiting to be scanned.
Beside that is a nice idea, but too focused on scanned stuff, today many docs are not scanned nor pdfs.