Earlier quoted context omitted.
Please write that big post! Sounds interesting
I second this! Sounds like an interesting read! :)
I am endlessly fascinated with content tagging systems
241–250 of 279 posts
Re: I am endlessly fascinated with content tagging systems
#242Earlier quoted context omitted.
Isolated hashtags on IG were often not useful for finding specific content. To find what you want you'd need an interface to a page that showed posts across multiple hashtags which they didn't provide. For example, posts that contain both #jenniferlawrence and #hoodedeyes would be what you are looking for, and if you interacted with enough posts containing both of those tags you would end up seeing the content you wa…
> For example, posts that contain both #jenniferlawrence and #hoodedeyes would be what you are looking for I mean, not to overwork it, but in this example, this isn't accurate: I just want pictures of Jennifer Lawrence, not the much smaller slice of content where the poster had her particular eye shape in mind. Also, I don't want to only "end up seeing" pictures of Jennifer Lawrence in my main feed, I want to be able…
Excellently put.
To me, it seems that the problem is that tagging is adversarial in any system where spamming can be rewarded.
Re: I am endlessly fascinated with content tagging systems
#243I worked with the Wikipedia category system a few years ago, and you could see the problems with hierarchical tagging systems right in action back then. (Though it may have gotten better in the meantime) The system appeared simple: There were just two relations, "Article A is a member of category B" and "Category X is a subcategory of category Y". However, in practice, the community was using this system to represent…
There are only two kinds of relation here, “subset of” and “instance of” (aka “element of”, type-token). The category-category relations are intended to always be a subset relation. The article-category relations are intended to always be an instance-of relation. - "19th century American writers" is a subset of "American writers“. => Both are a category, so no problem. - “Novelists” is a subset of “Writers”. => Both…
The problem with using such a strict type system as a tagging-system is exacerbated by the cases where:
1. Someone adds an article to a category (they tag the article), but then want to add a subcategory to that category. Now the category contains both articles and subcategories (violating the constraint). So the user would have to move all its articles into its subcategories for the constraint to be satisfied. This can be an enormous amount of work (needing to invent new subcategories for all the articles in the category not fitting into the specific subcategory they had in mind).
2. Someone wants to add an article, but only has a vague idea of a super-category in which it would fit. Now they have to exhaustively crawl/navigate the tree of sub-categories, until they find only the leaf sub-categories which only contains articles, which is a place they could put it. The input barrier thus becomes high (which is antithetical to how people expect to use tags).
Re: I am endlessly fascinated with content tagging systems
#244Earlier quoted context omitted.
I don't think you necessarily need multiple graphs; just labeled edges.
You just need some way to interact with it as multiple graphs. Some variation of labeled edges is probably the best.
tagged_with_italian_tag vs. tagged_with_spanish_tag ?
tagged_with_genus_tag vs. tagged_with_geo_tag ?
Would that afford such multiple graphs?
Re: I am endlessly fascinated with content tagging systems
#245Earlier quoted context omitted.
There are only two kinds of relation here, “subset of” and “instance of” (aka “element of”, type-token). The category-category relations are intended to always be a subset relation. The article-category relations are intended to always be an instance-of relation. - "19th century American writers" is a subset of "American writers“. => Both are a category, so no problem. - “Novelists” is a subset of “Writers”. => Both…
> Another way to put this is that categories have to be typed: a category contains either (just) articles, or it contains (just) categories. The problem with using such a strict type system as a tagging-system is exacerbated by the cases where: 1. Someone adds an article to a category (they tag the article), but then want to add a subcategory to that category. Now the category contains both articles and subcategories…
Analogy: The set of real numbers has the set of natural numbers as a subset, but it doesn’t have the set of natural numbers as an element, because the set of natural numbers is not a real number — the individual natural numbers are.
Likewise, the category “Occupations” may contain articles describing occupations, and it may have a subcategory “Clerical occupations” (subset-of relation), but it cannot contain the category “Writers” as an element (in an instance-of relation as with the articles), because writers are not a subset of occupations.
Furthermore, as an example of a meta-category, the category “Categories with more than 100 entries” may contain the category “Occupations” as an element (but not as a subcategory!), and hence cannot contain any articles as elements.
The element type of the category “Categories with more than 100 entries” is categories, and the element type of “Occupations” is articles. The point is that you can’t mix both types of elements within the same category. This is independent from subcategories. Any category can have subcategories, the only condition being that the subcategories must have the same element type as the suoercategory.
The idea is that a category can have both subset-of and instance-of relations at the same time (and each relation needs to be marked as such in the system), but the instance-of relation is restricted to be either articles or categories, but not both.
Re 2: I believe that problem goes away, given the above.
Re: I am endlessly fascinated with content tagging systems
#246I worked with the Wikipedia category system a few years ago, and you could see the problems with hierarchical tagging systems right in action back then. (Though it may have gotten better in the meantime) The system appeared simple: There were just two relations, "Article A is a member of category B" and "Category X is a subcategory of category Y". However, in practice, the community was using this system to represent…
Many comments below are hinting at - but not naming - triplestores. "A has relationship X with B". This is how wikidata works. Learning about those and learning how to query wikidata just blew my mind.
Re: I am endlessly fascinated with content tagging systems
#247A lot of the items described are problems in ontologies
Yeah. A tag is a predicate. Sub-tags are implication (male author => author). Tag aliases are equivalence (implication in both directions).
Re: I am endlessly fascinated with content tagging systems
#248Earlier quoted context omitted.
You just need some way to interact with it as multiple graphs. Some variation of labeled edges is probably the best.
In your examples, would the edges be like: tagged_with_italian_tag vs. tagged_with_spanish_tag ? tagged_with_genus_tag vs. tagged_with_geo_tag ? Would that afford such multiple graphs?
https://en.wikipedia.org/wiki/Entity%E2%80%93attribute%E2%80...
Re: I am endlessly fascinated with content tagging systems
#249Instagram's tagging system was actually really effective at categorizing content and discovery because each hashtag was treated as a node in a (giant) graph, where each node has multiple properties, including post count (number of posts using a tag), 'velocity' (number of posts using a particular tag per unit time), etc. I could write up a big post about it as I made a study of it in when I created a web app for find…
I went and got my laptop to type up a reply to this: Instagram's tagging system was and is atrocious in combination with their discovery mechanisms and the incentives they create. A real example, this has been true for years: I want to look at pictures of Jennifer Lawrence's makeup because she, like me, has hooded eyes and that makes useful reference. I go to instagram imagining that I will find fan accounts posting…
I suppose it probably still works better than Instagram.
Re: I am endlessly fascinated with content tagging systems
#250Earlier quoted context omitted.
I've been a librarian for more than 15 years and I can only speak from personal experience when I say that I am the apex predator of nothing. Every once and a while I will get it in my head to systematize my personal knowledge base with a controlled vocabulary and ontology and I just fall on my face. I really want it for some twisted reason, though. Turns out LC subject headings -- for all their failures -- are prett…
> controlled vocabulary Are you using English? English words can almost mean whatever you want them to. Perhaps design your own language that removes ambiguity. Probably requires a knowledge of philosophy to distinguish between say concrete and abstract, good luck. Maybe start with correcting the ontology of: https://cuberule.com/ (which takes a geometric approach to defining food types). Also perhaps decide whether…
That's what a controlled vocabulary is. It's essentially a set of tags which are clearly defined. So instead of #horses being defined purely by the word "horses," it has an attached definition along the lines of, "The category 'horses' includes equine biology, sports relating to horses, the cultural history of horses, and all other topics involving real horses. Metaphorical horses such as saw horses are not included." Tags like #horse would be redirected to #horses, since there is only one canonical horse tag in the vocabulary.