This is interesting, but raises the metaquestion of when compiled listings or search are appropriate.
DDG at this writing advertises 13,564 bangs:
https://duckduckgo.com/bang?q=
This page, as with the Show HN, groups bangs by category (Entertainment
Multimedia
News
Online Services
Research
Shopping
Tech
Translation).
There is a search feature (match on bang or URL) as well, with, of course, its own bang: !bang
In my experience, it's possible to, very roughly, define bounds to search results and usefulness or approach:
- 0 results is the null set. Obvious but worth noting.
- 1 result is an identity. That is, a search returning precisely one entity identifies that entity. (This relationship may not be persistent over time).
- To 10 results is a near universally useful set. SERPs, news sites, and many other interfaces focus on this scale. Ontologies typically branch to about 2-30 items. HN's front page lists 30 items. The Presidential Daily Brief runs about ten items. This apparently taxes the attention of some incumbents.
- To 100 results remains tractable to the determined user. This means a capacity to identify specific item(s) of interest without necessarily resorting to recordkeeping. More detailed news sites might include as many articles. A print newspaper typically runs 100-250 individual stories per day.
- To 1,000 results is methodologically tractable, though some recording and filtering system is all but certainly required. These need not be programmatic or algorithmic though often are. News wires (AP, AFP, Reuters, UPI) typically run 1,000--5,000 items/day.
- To 1,000,000 results, jumping a few orders of magnitude, typically requires some technical searching or sorting capacity. Again, not necessarily algorithmic (a large print library collection of a million or more volumes can be managed through paper based systems, though seldom is now, and requires extensive structure of stacks, index, circulation, and reshelving).
- To ~1 billion items is all but certainly programmatic.
- Greater thaan a trilliion items enters the realm of probabalistic matches, relevance, and AI.
I'm not aware of specific research in this area but would be quite interested in references.
The classification would suggest bang search rather than comprehensive listing as being of interest to most peoople though.
Sources: Seveeral of the values above are discussed here:
https://old.reddit.com/r/dredmorbius/comments/6c220n/media_a...
https://old.reddit.com/r/dredmorbius/comments/7qya12/informa...