Live data from Hacker News

Ask HN: How do you organize your files

news.ycombinator.com

31–40 of 44 posts

Re: Ask HN: How do you organize your files

#31
I made an automatic document tagger and categorizer. It collects any docs or HTML pages saved to Dropbox, dropped into a Telegram channel, saved with zotero, Slack, Mattermost, private webdav, etc, cleans the docs, pulls the text, performs topic modeling, along with a bunch of other NLP stuff, then renames all docs into something meaningful, sorts docs into a custom directory structure where folder names match the topics discovered, tags docs with relevant keywords, and visually maps the documents as an interactive graph. Full text search for each doc via solr. HTML docs are converted to clean text PDFs after ads are removed. This 'knowledge base' is contained in a single ECMS, external accounts for data input are configured from a single yaml file. There's also a web scraper that takes crawl templates as json files and uploads data into the CMS as files to be parsed with the rest of the docs. The idea is to be able to save whatever you are reading right now with one click whether you are on your mobile or desktop, or if you are collaborating in a group, and have a single repository where all the organizing is done actively 24/7 with ML.

Currently reconstructing the entire thing to production spec, as an AWS AMI, perhaps later polished into a personal knowledge base saas where the cleaned and sorted content is public accessible with REST/cmis api.

This project has single handedly eaten almost a third of my life.

Re: Ask HN: How do you organize your files

#32
My home directory:

  - bin :: quick place to put simple scripts and have available everywhere
  - build :: download projects for inspection and building, not for actively
       working on them
  - work-for :: where to put all projects; all project folders are available to
       me in zsh like ~proj-1/ so getting to them is quick despite depth.
    - me :: private projects for my use only
      - proj-1
    - all :: open source
      - proj-2
    - client :: for clients
      - client-1
        - proj-3
  - org :: org mode files
    - diary :: notes relating to the day
      - 2017-06-21.org :: navigated with binding `C-c d` defaulting to today
    - work-for :: notes for project with directory structure reflecting that of
         ~/work-for
      - client
        - client-1
          - proj-3.org
  - know :: things to learn from: txt's, books, papers, and other interesting
       documents
  - mail :: maildirs for each account
    - addr-1
  - downloads :: random downloads from the internet
  - media :: entertainment
    - music
    - vids
    - pics
      - wallpaper
  - t :: for random ad-hoc tests requiring directories/files; e.g. trying things
       with git
  - repo :: where to put bare git repositories for private projects (i.e. ~work-for/me/)
  - .password-store :: (for `pass` password manager)
    - type-1 :: ssh, web, mail (for smtp and imap), etc.
      - host-1 :: news.ycombinator.com, etc.
        - account-1 :: jol, jolmg, etc.
Not all folders are available on all machines, like ~/repo is on a private server, but they follow the same structure.

Re: Ask HN: How do you organize your files

#33
For ebooks I created folders for main-categories and some sub-categories (inspired by Amazon.com or some other ebook shop structure).

For photos folders per device/year/month.

For Office documents pre-pending date using the ISO date format (2017-06-21 or 170621) works great. (for sharing with others over various channels like mail/chat/fileserver/cloud/etc)

Re: Ask HN: How do you organize your files

#34
From a recent backup, there are

417,361 files

in my main collection of files for my startup, computing, applied math, etc.

All those files are well enough organized.

Here's how I do it and how I do related work more generally (I've used the techniques for years, and they are all well tested).

(1) Principle 1: For the relevant file names, information, indices, pointers, abstracts, keywords, etc., to the greatest extent possible, stay with the old 8 bit ASCII character set in simple text files easy to read by both humans and simple software.

(2) Principle 2: Generally use the hierarchy of the hierarchical file system, e.g., Microsoft's Windows HPFS (high performance file system), as the basis (framework) for a taxonomic hierarchy of the topics, subjects, etc. of the contents of the files.

(3) To the greatest extent possible, I do all reading and writing of the files using just my favorite programmable text editor KEdit, a PC version of the editor XEDIT written by an IBM guy in Paris for the IBM VM/CMS system. The macro language is Rexx from Mike Cowlishaw from IBM in England. Rexx is an especially well designed language for string manipulation as needed in scripting and editing.

(4) For more, at times make crucial use of Open Object Rexx, especially its function to generate a list of directory names, with standard details on each directory, of all the names in one directory subtree.

(5) For each directory x, have in that directory a file x.DOC that has whatever notes are appropriate for good descriptions of the files, e.g., abstracts and keywords of the content, the source of the file, e.g., a URL, etc. Here the file type of an x.DOC file is just simple ASCII text and is not a Microsoft Word document.

There are some obvious, minor exceptions, that is, directories with no file named x.DOC from me. E.g., directories created just for the files used by a Web page when downloading a Web page are exceptions and have no x.DOC file.

(6) Use Open Object Rexx for scripts for more on the contents of the file system. E.g., I have a script that for a current directory x displays a list of the (immediate) subdirectories of x and the size of all the files in the subtree rooted at that subdirectory. So, for all the space used by the subtree rooted at x, I get a list of where that space is used by the immediate subdirectories of x.

(7) For file copying, I use Rexx scripts that call the Windows commands COPY or XCOPY, called with carefully selected options. E.g., I do full and incremental backups of my work using scripts based on XCOPY.

For backup or restore of the files on a bootable partition, I use the Windows program NTBACKUP which can backup a bootable partition while it is running.

(8) When looking at or manipulating the files in a directory, I make heavy use of the DIR (directory) command of KEdit. The resulting list is terrific, and common operations on such files can be done with commands to KEdit (e.g., sort the list), select lines from the list (say, all files x.HTM), delete lines from the list, copy lines from the list to another file, use short macros written in Kexx (the KEdit version of Rexx), often from just a single keystroke to KEdit, to do other common tasks, e.g., run Adobe's Acrobat on an x.PDF file, have Firefox display an x.HTM file.

More generally, with one keystroke, have Firefox display a Web page where the URL is the current line in KEdit, etc.

I wrote my own e-mail client software. Then given the date header line of an e-mail message, one keystroke displays the e-mail message (or warns that the date line is not unique, but it always has been).

So, I get to use e-mail message date lines as 'links' in other files. So, if some file T1 has some notes about some subject and some e-mail message is relevant, then, sure, in file T1 just have the date line as a link.

This little system worked great until I converted to Microsoft's Outlook 2003. If I could find the format of the files Outlook writes, I'd implement the feature again.

(9) For writing software, I type only into KEdit.

Once I tried Microsoft's Visual Studio and for a first project, before I'd typed anything particular to the project, I got 50 MB or so of files nearly none of which I understood. That meant that whenever anything went wrong, for a solution I'd have to do mud wrestling with at least 50 MB of files I didn't understand; moreover, understanding the files would likely have been a long side project. No thanks.

E.g., my startup needs some software, and I designed and wrote that software. Since I wrote the software in Microsoft's Visual Basic .NET, the software is in just simple ASCII files with file type VB.

There are 24,000 programming language statements.

So, there are about 76,000 lines of comments for documentation which is IMPORTANT.

So, all the typing was done into KEdit, and there are several KEdit macros that help with the typing.

In particular, for documentation of the software I'm using -- VB.NET, ASP.NET, ADO.NET, SQL Server, IIS, etc. -- I have 5000+ Web pages of documentation, from Microsoft's MSDN, my own notes, and elsewhere.

So, at some point in the code where some documentation is needed for clarity for the code, I have links to my documentation collection, each link with the title of the documentation. Then one keystroke in KEdit will display the link, typically have Firefox open the file of the MSDN HTML documentation.

Works great.

The documentation is in four directories, one for each of VB, ASP, SQL, and Windows. Each directory has a file that describes each of the files of documentation in that directory. Each description has the title of the documentation, the URL of the source (if from the Internet which is the usual case), the tree name of the documentation in my file system, an abstract of the documentation, relevant keywords, and sometimes some notes of mine. KEdit keyword searches on this file (one for each of the four directories) are quite effective.

(10) Environment Variables

I use Windows environment variables and the Windows system clipboard to make a lot of common tasks easier.

E.g., the collection of my files of documentation of Visual Basic is in my directory

H:\data05\projects\software\vb\

Okay, on the command line of a console window, I can type

G VB

and then have that directory current.

Here 'G' abbreviates 'go to'!

So, to command G, argument 'VB' acts like a short nickname for directory

H:\data05\projects\software\vb\

Actually that means that I have -- established when the system boots -- a Windows environment variable MARK.VB with value

H:\data05\projects\software\vb\

I have about 40 such MARK.x environment variables.

So, sure, I could use the usual Windows tree walking commands to navigate to directory

H:\data05\projects\software\vb\

but typing

G VB

is a lot faster. So, such nicknames are justified for frequently used directories fairly deep in the directory tree.

Environment variables

MARK.TO

MARK.FROM

are used by some other programs, especially my scripts that call COPY and XCOPY.

So, to copy from directory A to directory B, I navigate to directory A and type

MARK FROM

which sets environment variable

MARK.FROM

to the directory tree name of directory A. Similarly for directory B.

Then my script

COPYFT1.RXS

takes as argument the file name and does the copy.

My script

COPYFT2.RXS

takes two arguments, the file name of the source and the file name to be used for the copy.

I have about 200 KEdit macros and about 200 Rexx scripts. They are crucial tools for me.

(11) FACTS

About 12 years ago I started a file FACTS.DAT. The file now has 74,317 lines, is

2,268,607

bytes long, and has 4,017 facts.

Each such fact is just a short note, sure, on average

2,268,607 / 4,017 = 565

bytes long and

74,317 / 4,017 = 18.5

lines long.

And that is about

12 * 365 / 4,017 = 1.09

that is, an average of right at one new fact a day.

Each new fact has its time and date, a list of keywords, and is entered at the end of the file.

The file is easily used via KEdit and a few simple macros.

I have a little Rexx script to run KEdit on the file FACTS.DAT. If KEdit is already running on that file, then the script notices that and just brings to the top of the Z-order that existing instance of KEdit editing the file -- this way I get single threaded access to the file.

So, such facts include phone numbers, mailing addresses, e-mail addresses, user IDs, passwords, details for multi-factor authentication, TODO list items, and other little facts about whatever I want help remembering.

No, I don't need special software to help me manage user IDs and passwords.

Well, there is a problem with the taxonomic hierarchy: For some files, it might be ambiguous which directory they should be in. Yes, some hierarchical file systems permitted to be listed in more than one directory, but AFAIK the Microsoft HPFS file system does not.

So, when it appears that there is some ambiguity in what directory a new file should go, I use the x.DOC files for those directories to enter relevant notes.

Also my file FACTS.DAT may have such notes.

Well, (1)-(11) is how I do it!

Re: Ask HN: How do you organize your files

#35
My file layout is quite uninteresting. The most noteworthy thing is that I have an additional toplevel directory /x/ where I keep all the stuff that would otherwise be in $HOME, but which I don't want to put in $HOME because it doesn't need to be backed up.

- /x/src contains all Git repos that are pushed somewhere. Structure is the same as wanted by Go (i.e., GOPATH=/x/). I have a helper script and accompanying shell function `cg` (cd to git repo) where I give a Git repo URL and it puts me in the repo directory below /x/src, possibly cloning the repo from that URL if I don't have it locally yet.

  $ pwd
  /home/username
  $ cg gh:foo/bar # understands Git URL aliases, too
  $ pwd
  /x/src/github.com/foo/bar
As I said, that's not in the backup, but my helper script maintains an index of checked-out repos in my home directory, so that I can quickly restore all checkouts if I ever have to reinstall.

- /x/bin is $GOBIN, i.e. where `go install` puts things, and thus also in my PATH. Similar role to /usr/local/bin, but user-writable.

- /x/steam has my Steam library.

- /x/build is a location where CMake can put build artifacts when it does an out-of-source build. It mimics the structure of the filesystem, but with /x/build prefixed. For example, if I have a source tree that uses CMake checked out at /home/username/foo/bar, then the build directory will be at /x/build/home/username/foo/bar. I have a `cd` hook that sets $B to the build directory for $PWD, and $S to the source directory for $PWD whenever I change directories, so I can flip between source and build directory with `cd $B` and `cd $S`.

- /x/scratch contains random junk that programs expect to be in my $HOME, but which I don't want to backup. For example, many programs use ~/.cache, but I don't want to backup that, so ~/.cache is a symlink to the directory /x/scratch/.cache here.

Re: Ask HN: How do you organize your files

#36
post #24
post #18

I try not to over think it, just: ~/$MAJOR_TOPIC | |--- ./$MORE_SPECIFIC | |--- ./$MORE_SPECIFIC | |--- ./general-file.type | | ./general-file.type | |--- ./$MORE_SPECIFIC | |--- ./general-file.type etc As you find yourself collecting more general files under a directory that can be logically grouped, create a new directory and move them to it. Also keep all your directories in the same naming convention (idk maybe I…

That's pretty much how I did it on my NAS. My top level is basically "fiction", "non-fiction", "music", "pictures", "software" - then it just goes from there. But it has problems. For instance, I like to collect information about robotics and artificial intelligence. In many cases I have papers with titles like "Using Computer Vision to Control a Robot Arm via a CNN". Do I put it under "robotics/sensors/vision" or "a…

If you're running Windows, you can use Everything [1] to instantly find files on your computer just by knowing their name.

1. https://www.voidtools.com/downloads/

Re: Ask HN: How do you organize your files

#37
I use `mess` [1]. Short descrption: New stuff that is not filed away instantly goes into a folder "current" linked to the youngest folder in a tree (mess_root > year > week). If needed at a later time: file it accordingly, otherwise old folders are purged if disk space is low. Taking it a step further: synching everything across work and personal machines using `syncthing`.

[1] http://chneukirchen.org/blog/archive/2006/01/keeping-your-ho...

Re: Ask HN: How do you organize your files

#38
post #31

I made an automatic document tagger and categorizer. It collects any docs or HTML pages saved to Dropbox, dropped into a Telegram channel, saved with zotero, Slack, Mattermost, private webdav, etc, cleans the docs, pulls the text, performs topic modeling, along with a bunch of other NLP stuff, then renames all docs into something meaningful, sorts docs into a custom directory structure where folder names match the to…

This sounds really interesting -- can you share anything else, or pieces of the pipeline...especially topic modeling?

Re: Ask HN: How do you organize your files

#39
Downloads

└─Filename preserved, ordered by date or grouped in arbitrary functional folders

Drivers

├─Video

├─Sound

└─MB

Music

└─Primary Artist

  └─YYYY.AlbumName (Keeps albums in date order)

    └─AlbumName Track# Title.mp3 (truncates sensibly on a car stereo)
Pictures

└─YYYY-MM-DD.Event Description (DD is optional)

Projects

├─scripts - reusable across clients

│ └─language

│ └─purpose

└─clientname

  ├─source code

  └─documents
Utils (single-executable files that don't require an install)

I use Beyond Compare as my primary file manager at home and work. Folder comparison is the easiest way to know if a file copy fully completed. Multi-threaded move/copy is nice too.

Post reply on HN