How is search so bad?My broad take is that previously search worked (Altavista era through early-mid Google) because it referenced organic links put in place by real humans and keywords, plus basic metadata like physical location of servers, freshness of content, frequency of update, metadata behind domains, etc.
Since the mid 1990s that has increasingly been gamed heavily, PageRank style approaches have come and sort of gone, and a vast majority of content accessed by consumers has moved to one of a small number of platforms or walled gardens, often mobile applications. I don't know for sure, but I'd assume with confidence that the majority of result inclusion decisions made by Google are now based on rejection blacklists, 'known good' safe hits and effectively minimizing anomalous results above the fold. Simultaneously, the internet has become an international place and the bar has been raised for new entrants such that an incapacity to return meaningful results in multiple languages bars a search engine from any significant market position. A huge percentage of results are either Wikipedia/reference pages, local news or Q&A sites. Further, huge amounts of what is out there is behind Cloudflare or similar firewalls which will probably frustrate new and emerging spiders.
The existing monopolies, having some established capacity and reputation in this regard, may have become somewhat entrenched and lazy, and do not care enough about improvement. They are literally able to sail happily on market inertia, while generating ridiculous advertising revenues. In China we have Baidu, and in most rest of the world, Google.
Who will bring about a new search engine? Greg Lindahl https://news.ycombinator.com/user?id=greglindahl who formerly made Blekko is apparently working on another one.
I once wrote a small one (~2001) which was based upon the concept of multilingual semantic indices (a sort of non-rigorously obtained language-neutral epistemology was the core database). I still think this would be a meaningful approach to follow, since so much is lost in translation, particularly around current events. One problem with evolving public utilities in this area are that such approaches border on open source intelligence (OSI) and most people with linguistic or computational chops in that area leave academia and get eaten up by the military industrial complex or Google.
Now we have https://commoncrawl.org/the-data/get-started/ which makes reasonable quality sample crawl data super-available. Now "all" we need is people to hack on algorithms and a means to commercialize them as alternatives to the status quo.