Live data from Hacker News

Ask HN: Can we create a new internet where search engines are irrelevant?

news.ycombinator.com

141–150 of 395 posts

Re: Ask HN: Can we create a new internet where search engines are irrelevant?

#141
I'm too late, but yes, it is not easy but it definitely seems doable: http://comboy.pl/wot.html

I'm sorry it's a bit long, TL;DR you need to be explicit about people you trust. Those people do the same an then thanks to the small world effect you can establish your trust to any entity that is already trusted by some people.

No global ranking is the key. How good some information is, is relative and depends on who do you trust (which is basically form of encoding your beliefs). And yes, you can avoid information bubble much better than now but writing more when I'm so late to the thread seems a bit pointless.

Re: Ask HN: Can we create a new internet where search engines are irrelevant?

#142
post #128
post #110

Everyone has missed the most important aspect of search engines, from the point of view of their core function of information retrieval: they're the internet equivalent of a library index. Either you find a way to make information findable in a library without an index (how?!?) or you find a novel way to make a neutral search engine - one that provides as much value as Google but whose costs are paid in a different w…

The problem is that current search engines are indexing what is essentially a stack of random books thrown together by anonymous library goers. Before being able to guide readers to books, librarians have to the following non-trivial tasks over the entire collection: - identify the book's theme - measure the quality of the information - determine authenticity / malicious content - remember the position of the book in…

>I welcome any attempts to build a more organized internet. I don't think the communal book pile approach is scaling very well.

Let me know if I misunderstand your comment but to me, this has already been tried.

Yahoo's founders originally tried to "organize" the internet like a good librarian. Yahoo in 1994 was originally called, "Jerry and David's Guide to the World Wide Web"[0] with hierarchical directories to curated links.

However, Jerry & David noticed that Google's search results were more useful to web surfers and Yahoo was losing traffic. Therefore, in 2000 they licensed Google's search engine. Google's approach was more scaleable than Yahoo's.

I often see several suggestions that the alternative to Google is curated directories but I can't tell if people are unaware of the early internet's history and don't know that such an idea was already tried and how it ultimately failed.

[0] http://static3.businessinsider.com/image/57977a3188e4a714088...

Re: Ask HN: Can we create a new internet where search engines are irrelevant?

#143

That was what the early internet was like (I was there). People built indexes by hand, lists of pages on certain topics. There was the Gopher protocol that was supposed to help with finding things. But this was all top-down stuff, the first indexing/crawling search engines were bottom-up and it worked so much better. And for a while we had an ecosystem of different search engines until Google came along, was genuinel…

In the very early days, you didn't need a search engine because there weren't that many web sites and you knew most of the main ones anyway (or later on had them in your own hotlists in Mosaic). Nowadays you need a search because there is so much content.

The problem is that the amount of content and the size of the potential user base are so large that is is impossible to offer search as a free service, i.e. it has to be funded in some way. Perhaps instead of having a free advertising-driven search, there would be space for a subscription-based model? Subscription based (and advert free) models seem to be working in other areas, e.g. TV/films and music.

Another problem though is that more and more content seems to be becoming unsearchable, e.g. behind walled gardens or inside apps.

Re: Ask HN: Can we create a new internet where search engines are irrelevant?

#145
post #128
post #110

Everyone has missed the most important aspect of search engines, from the point of view of their core function of information retrieval: they're the internet equivalent of a library index. Either you find a way to make information findable in a library without an index (how?!?) or you find a novel way to make a neutral search engine - one that provides as much value as Google but whose costs are paid in a different w…

The problem is that current search engines are indexing what is essentially a stack of random books thrown together by anonymous library goers. Before being able to guide readers to books, librarians have to the following non-trivial tasks over the entire collection: - identify the book's theme - measure the quality of the information - determine authenticity / malicious content - remember the position of the book in…

The current search engines are also indexing books maliciously inserted in the library in a way to maximize their exposure e.g. a million "different" pamphlets advertising Bob's Bible Auto Repair Service inserted in the Bible category.

A "better library" can't be permissionless and unfiltered; Dewey Decimal System relies on the metadata being truthful, and the internet is anything but.

You can't rely on information provided by content creators; Manual curation is an option but doesn't scale (see the other answer re: early Yahoo and Google).

Re: Ask HN: Can we create a new internet where search engines are irrelevant?

#146

I see a lot of good comments here, I got inspired to write this: What if this new Internet instead of using URI based on ownership (domains that belong to someone), would rely on topic? In examples: netv2://speakers/reviews/BW netv2://news/anti-trump netv2://news/pro-trump netv2://computer/engineering/react/i-like-it netv2://computer/engineering/electron/i-dont-like-it A publisher of webpage (same html/http) would pu…

Inventing/extending a new NNTP is nice idea too.

The Internet has become synonymous with the web/http protocol. The web alternatives to NNTP won instead of newer versions of Usenet. New versions of IRC, UUCP, S/FTP, SMTP, etc., instead of webifying everything would be nice. But those services are still there and fill an important niche for those not interested in seeing everything eternal septembered.

Re: Ask HN: Can we create a new internet where search engines are irrelevant?

#147
post #142
post #128

Earlier quoted context omitted.

The problem is that current search engines are indexing what is essentially a stack of random books thrown together by anonymous library goers. Before being able to guide readers to books, librarians have to the following non-trivial tasks over the entire collection: - identify the book's theme - measure the quality of the information - determine authenticity / malicious content - remember the position of the book in…

>I welcome any attempts to build a more organized internet. I don't think the communal book pile approach is scaling very well. Let me know if I misunderstand your comment but to me, this has already been tried. Yahoo's founders originally tried to "organize" the internet like a good librarian. Yahoo in 1994 was originally called, "Jerry and David's Guide to the World Wide Web" [0] with hierarchical directories to cu…

I remember trying to get one of my company's sites listed on Yahoo! back in the late 1990s. Despite us being an established company (founded in 1985) with a good domain name (cardgames.com) and a bunch of good, free content (rules for various card games, links to various places to play those games online, etc.), it took months.

Re: Ask HN: Can we create a new internet where search engines are irrelevant?

#148

I see a lot of good comments here, I got inspired to write this: What if this new Internet instead of using URI based on ownership (domains that belong to someone), would rely on topic? In examples: netv2://speakers/reviews/BW netv2://news/anti-trump netv2://news/pro-trump netv2://computer/engineering/react/i-like-it netv2://computer/engineering/electron/i-dont-like-it A publisher of webpage (same html/http) would pu…

How would topic validity get enforced?

For example, if a publisher has a particular pro-Trump article, they would likely want (for obvious financial reasons) to push it to both etv2://news/anti-trump and netv2://news/pro-trump . What would prevent them from doing that?

Also, a publisher of "GET RICH QUICK NOW!!!" article would want to push it to both netv2://news/anti-trump and netv2://computer/engineering/electron/i-dont-like-it topics.

You can't simply have topics, you can have communities like news/pro-trump that are willing to spend the labor required for moderation i.e. something like reddit. But not all content has such communities willing and able to do so well.

Re: Ask HN: Can we create a new internet where search engines are irrelevant?

#149
Sorry, but this was too tempting: https://imgur.com/a/6UcAOnF

But seriously, I'm not sure it is feasible, I wish the internet could auto-index itself and still be decentralized, where any type of content can be "discovered" as soon as it is connected to the "grid".

The advantage would be that users could search any content without filters, without AI tempering with the order based on some rules ... BUT on the other hand, people use search engines because their results are relevant (what ever that means these days), so having an internet that is searchable by default would probably never be a good UX and hence not replace existing search engines. It not just about the internet being searchable, it would have to solve all the problems search engines have solved in the last ten years too

Re: Ask HN: Can we create a new internet where search engines are irrelevant?

#150
post #82

Earlier quoted context omitted.

That's very workable.Any agent should have a private key with which it signs it's pushes. Age of an agent and score of feedback for that agent determine its ranking.Though that still leaves gaming possible with the feedback. But heavy feeback like "this is malicious content" could be moderated. (So that people cant just report stuff they don't like).

>But heavy feeback like "this is malicious content" could be moderated. // You're just shifting around your trust problem. You need to handle 4chan level manipulation (million of users coordinating to manipulate polls), or Scientology depth (getting thousands of people in to USA government jobs in order to get recognised as a religion). If it's "we'll catch it in moderation" then whoever wants to manipulate it just g…

You can't get around the problem of manipulation if your trustworthiness metric for content will be the same for all people, as it is on reddit, hacker news, or Amazon for example. Having moderators just concentrates the issue into a smaller number of people and you haven't solved the central problem--manipulation is profitable.

But think of how we solve this problem in our personal interactions with other people, and this should be a clue for how to solve it with computational help. We have a pretty good idea of which people are trustworthy (or capable, or dependable, or any other characteristic) in our daily lives, and based on our interactions with them we update these internal measures of trustworthiness. If we need to get information from someone we don't know, we form a judgement of their trustworthiness based off of input from people we trust--e.g. giving a reference. This is really just Bayesian inference at its core.

We should be able to come up with a computational model for how this personal measure of trustworthiness works. It would act as a filter over content that we obtain. Throw a search engine on top of this, sure, but in the end you'd still need to get trustworthiness weights onto information if you want it to be manipulation-resistant. This labeling is what I mean by manual curation. You can't leave that up to the search engine or the aggregator because those can be gamed, like the examples you gave for aggregators and SEO for search engines have shown.

Post reply on HN