Live data from Hacker News

Designing better file organization around tags, not hierarchies (2017)

nayuki.io

91–100 of 169 posts

Re: Designing better file organization around tags, not hierarchies (2017)

#91
Hello everyone, thank you for all the comments. Seeing this on the HN front page caught me by surprise. In the past year I shared this article publicly (Reddit) and privately (with tech-savvy acquaintances) for comment, and the general sentiment I received was that these ideas were not ready to be read by a mass audience. The article is way too long and pulls in many disparate ideas; it explains both why traditional features are problematic and how new features would work better. In the end, it is unclear what a real implementation would look like, and what concrete benefits and annoyances would come out of real-world usage. I was hoping to build an ugly prototype before asking for feedback.

Regarding the comments on this HN thread, it seems the general discussion is around tagging. This is indeed the title of the article and the main idea that motivated my exploration, but I believe the other ideas are just as important. I explored notions like no-filenames, strong preference for hash addressing and references, location independence, immutability, backups and deduplication, preference for external (non-embedded) file metadata, first-class media libraries, and more.

I think the debate about tagging is quite adequate, and would be happy to hear comments about the other features/non-features, and whether all the ideas fit or don't fit cohesively as a system.

Re: Designing better file organization around tags, not hierarchies (2017)

#92
Every time I read an article about attempts at non-hierarchical filesystems, I try to figure out how I'd take the huge piles of stuff I generate when I'm drawing (and publishing) a graphic novel and reorganize it under tags. It's never pretty.

Like, okay, sure, I tag everything with the name of the project, that's a no-brainer. But if I just do that then I get the hundreds of files I generate (one per page) mixed up with everything else - web-res renderings of each page, model sheets (and their source files), promotional material, stuff sent to publishers to try and convince them to deal with that part of the process, and the huge mass of files I generate for each book I print (which can be more than one for a single multi-year project): source files tweaked for print, print-res renderings thereof, files for the kickstarter for each book... So I tag all of these attributes too, and imagining putting all these tags on a file as I save it sure is a lot of fun, even if I imagine some sort of save requestor that keeps a list of all my previously-used tags, including ways to filter those - I don't care about any of the tags attached to my music collection or my collection of cartoon porn or my programming projects when I'm working on my comics projects, for instance, so I'd want to quickly narrow it down to just tags found in my art projects, and...

Ultimately it just starts to look like a hierarchical structure in my mind, except for the fact that I'm interacting with it by some kind of tag-filtering file browser on top of a huge filesystem that mixes everything together in a non-human-browseable structure.

Re: Designing better file organization around tags, not hierarchies (2017)

#93
post #88

Every time I use a tagging-based system, I become more convinced that tags are what I want for almost all things, not just files.

How do you find an untagged file? Are wrongly tagged files undiscoverable?

This is more of a UI/UX question, but I thought about it on and off for months and have a partial answer. Look at how Gmail works - you have All Mail, but you also have Inbox, Sent, and your own custom labels (tags). Every message can always be found in All Mail, and you can restrict your search by date range. Similarly on an image board like Danbooru, even if you don't tag an image, it will appear on the chronological stream of every image ever uploaded to the system. So, I'm hoping the design will end up something like these two examples. You should be able to list every file, preferably in chronological order; you should also be able to exclude files that have at least one tag; and the software might have a special nag section showing you all the files that you left uncategorized.

Re: Designing better file organization around tags, not hierarchies (2017)

#95
https://web.archive.org/web/20070927003401/http://www.namesy...

Hans Reiser had some good ideas.

And by the way there is a good analogy with www - originally we had just the addresses (somehow hierarchical), then we had hierarchical catalogue of Yahoo, and then quickly it became too much for that and we now rely on search.

Re: Designing better file organization around tags, not hierarchies (2017)

#96
post #58

Earlier quoted context omitted.

Once upon a time there was Google Desktop, and it was incredibly useful. I wouldn't trust Google on my machine anymore, but an open source replace ment would be great. https://en.m.wikipedia.org/wiki/Google_Desktop

OSX has Spotlight. It searches pretty much all the same stuff Google Desktop is listed as searching. If you want something open source instead then check out Quicksilver [ https://qsapp.com ]. Dunno about Windows or Your Favorite Unix. alternativeto.net may be of help: https://alternativeto.net/software/google-desktop/

Thanks. My subtle point was why use tags when you could use search? Search engines index the whole internet - millions of hierarchical file systems of all kinds. Why reinvent file systems?

I don't use Mac. Windows 10 "Cortana" is no match for where Google Desktop was 10 years ago. I rarely use it unless I have to. Google searched the contents. Cortana just searches filenames and tries to route searches to the web. Linux...I don't see anything - Ubuntu has Cortana-like feature, but it's not Google Desktop. It's very odd to go backwards technology-wise.

Re: Designing better file organization around tags, not hierarchies (2017)

#97
post #89
post #67

Earlier quoted context omitted.

Tags are fun when you have a few thousand items to test your MVP with. It gets much less fun when you have millions of items with thousands of tags, all on a flat hierarchy. On the other hand, when you're stuck with a flat hierarchy anyway (e.g. thousands of pictures, all named DCIMxxxx.jpg), tags can be more useful. But only if they're automatic. I want the best of both worlds. I want to organize my stuff into folde…

I absolutely agree with this, except to add that it seems, in theory, in the case of a 100% tag-based file-system, there would never (*very rarely) be a flat list where you have to scroll through millions of files. The UX of a single flat list with millions of files named DCIMxxxx.jpg is a limitation of the current format and doesn't make sense when so much information can be generated about our files upon creation.…

Part of the adoption problem here would be trust. Hierarchical filesystems are something we're used to, and we can trust they're implemented correctly. That means, if I visit a folder and see some files, I know what I see is all of the files there (+/- hidden file settings); if something is missing, it's not there, period.

Tag search is a search. Can be broken. Can be optimized in a way that causes it to lie. I look at the results, and I'm not sure if they're complete. Maybe the file I'm looking for is really not there, or maybe the search gave up too early. Or the tag was slightly misformatted?

Maybe I'm too used to the old thing, but I like the notion that there's one canonical tree structure that makes all my data reachable. In case I've misplaced something, the search space of all paths through the filesystem tree (or a subtree of interest) is vastly smaller than the search space of all possible values of all relevant tags.

Re: Designing better file organization around tags, not hierarchies (2017)

#98
post #95

https://web.archive.org/web/20070927003401/http://www.namesy... Hans Reiser had some good ideas. And by the way there is a good analogy with www - originally we had just the addresses (somehow hierarchical), then we had hierarchical catalogue of Yahoo, and then quickly it became too much for that and we now rely on search.

There's a difference in searching the Internet and searching your own data. On the Internet, there's much more of everything than you'd ever need, and close to none of it is something you've created, or even seen before. So you use search to get some reasonably relevant results. The search doesn't have to be - and isn't - complete nor correct.

On the filesystem, I'd like to know the data browser isn't hiding files from me by reporting only "top 100 relevant results", or not indexing half of it because $reasons, or not showing them because of faulty query. Being able to iterate through all files on your disk in a tree-like fashion seems like a feature.

Re: Designing better file organization around tags, not hierarchies (2017)

#99

Have you considered a file system organized as a timeline that _also_ supports tagging? I find one of the key concepts that's not a first-class concept is _when_ the file was modified. Rather than a file-and-folder physical analogy for the file system UI, I think a timeline-oriented UI could present some advantages for the way that humans actually think and work. Tags would be a helpful orthogonal organization scheme…

Not sure if it generalizes for the entire filesystem - not all files are modified due to explicit user action. Software keeps logging all the time. Many applications update their config files when closing, or during runtime. Saving a document in a program may trigger saving another 3 files elsewhere. All in all, it seems like a recipe for seeing the least relevant data first.

Re: Designing better file organization around tags, not hierarchies (2017)

#100

Earlier quoted context omitted.

I eventually had to use a series of command-line tools to shepherd my badly-organized photo collection into something like a usable state, and it basically involved finding all the JPEGs on my system, de-duplicating them, and dropping them into folders based on their date-taken metadata. It was a huge pain, and entirely an artifact of the tyranny of the folder-based filesystem.

for i in *.jpg; do date=$(exif $i --machine-readable --tag=0x9003 | cut -d ' ' -f 1 | tr : -) mkdir -p $date mv $i $date done But don't many photo-browsing programs allow browsing by date, or even GPS-tagged location? e.g. on KDE's DigiKam I choose "timeline" view (or map view).

Oh, damn. The formatting was lost, I had line breaks in there!
Post reply on HN