You're effectively crawling portions of the web based on your query, at runtime! It's a pretty neat technique. But you obviously have to trust the sources and the links to provide you with relevant data.
Ask HN: Can we create a new internet where search engines are irrelevant?
181–190 of 395 posts
Re: Ask HN: Can we create a new internet where search engines are irrelevant?
#182Earlier quoted context omitted.
>But heavy feeback like "this is malicious content" could be moderated. // You're just shifting around your trust problem. You need to handle 4chan level manipulation (million of users coordinating to manipulate polls), or Scientology depth (getting thousands of people in to USA government jobs in order to get recognised as a religion). If it's "we'll catch it in moderation" then whoever wants to manipulate it just g…
You can't get around the problem of manipulation if your trustworthiness metric for content will be the same for all people, as it is on reddit, hacker news, or Amazon for example. Having moderators just concentrates the issue into a smaller number of people and you haven't solved the central problem--manipulation is profitable. But think of how we solve this problem in our personal interactions with other people, an…
We really don't. People get surprised all the time that someone had an affair, or cheated, or ripped someone off, or whatever. "But I trusted you" ...
It's actually relatively easy to fool people in to trusting you, as many red team members will probably confirm.
Look at someone like Boris Johnson, people are trusting him to lead the country knowing that he's well known to betray people's trust and that he even had a court case lodged against him based on his very blatant lying to the entire country. You can even watch the video of him being interviewed where the interviewers says (paraphrasing) "but we all know that's a half truth" and BoJo just pushes it and pushes it and refuses to accept that it's anything other than absolute truth.
>If we need to get information from someone we don't know, we form a judgement of their trustworthiness based off of input from people we trust--e.g. giving a reference. //
This is domain authority again - trust some domains manually, let it flow from there. If that domain trusts another domain then they link to it, trust flows to the other domain, and so on. Maintaining such trust for a long time adds to a particular domains trust factor, linking to domains not trusted by others detracts from it.
Re: Ask HN: Can we create a new internet where search engines are irrelevant?
#183It was thought that one way of finding information is to ask your network (Facebook and Twitter would be examples), and then they would pass on the message and a chain of trusted sources would get the information back to you. I am being purposefully vague because I don't think people know what an effective version of that would look like, but its worth exploring. If you have some data you might ask questions like: 1.…
Not convinced any kind of formalised 'question answering network' could replace search. It would be both slow, and require an enormous asymmetric investment of time, for a diffuse and unspecified reward.
Re: Ask HN: Can we create a new internet where search engines are irrelevant?
#184Another original intent: that URLs would not need to be user-visible, and you wouldn't need to type them in.
Re: Ask HN: Can we create a new internet where search engines are irrelevant?
#185Earlier quoted context omitted.
> I often see several suggestions that the alternative to Google is curated directories but I can't tell if people are unaware of the early internet's history and don't know that such an idea was already tried and how it ultimately failed. ¿Por qué no los dos? 1) The idea is that a more organized structure is easier for a librarian to index. Today, libraries still have librarians. The book pile just wouldn't take dec…
There are plenty of these, wikipedia has a list [1]. I think these efforts get bogged down in the huge amount of content out there, the impermanence of that content and also the difficulty in placing sites into ontologies. And at the end of the day, there's not a large enough value proposition to balance the immense effort. I think, if you were to do it today, you would want to work on / with the internet archive, so…
What would make the approach viable is if there were a nice way to automate and crowd source most/all of the effort. Maybe that means changing the idea of what makes a website. Maybe there could just be little grass roots reddit-esque communities that are indexed/verified (google already favors reddit/hn links). Who knows, but it's an interesting problem to kick around.
Re: Ask HN: Can we create a new internet where search engines are irrelevant?
#186Earlier quoted context omitted.
That's very workable.Any agent should have a private key with which it signs it's pushes. Age of an agent and score of feedback for that agent determine its ranking.Though that still leaves gaming possible with the feedback. But heavy feeback like "this is malicious content" could be moderated. (So that people cant just report stuff they don't like).
The reason I mentioned that the trust metric should be transitive and distributed is so that it prevents gaming as much as possible. You wouldn't want to have a trusted central authority (for everyone) because that could always be corrupted or gamed if it's profitable enough. Rather every individual would have a set of trusted peers with different "trust" weights for each based on the individual's perception of their…
Re: Ask HN: Can we create a new internet where search engines are irrelevant?
#187Earlier quoted context omitted.
There are plenty of these, wikipedia has a list [1]. I think these efforts get bogged down in the huge amount of content out there, the impermanence of that content and also the difficulty in placing sites into ontologies. And at the end of the day, there's not a large enough value proposition to balance the immense effort. I think, if you were to do it today, you would want to work on / with the internet archive, so…
Obviously a naïve web directory isn't going to cut it. What would make the approach viable is if there were a nice way to automate and crowd source most/all of the effort. Maybe that means changing the idea of what makes a website. Maybe there could just be little grass roots reddit-esque communities that are indexed/verified (google already favors reddit/hn links). Who knows, but it's an interesting problem to kick…
But to me, crowdsourcing is also what Jerry & David did. The users submitted links to Yahoo. AltaVista also had a form for users to submit new links.
Also, Wikipedia's list of links are also crowdsourced in the sense that many outside websurfers (not just staff editors) make suggested edits to the wiki pages. Looking at a "revision history" of a particular wiki page makes the crowdsourced edits more visible: https://en.wikipedia.org/w/index.php?title=List_of_web_direc...
Re: Ask HN: Can we create a new internet where search engines are irrelevant?
#188Re: Ask HN: Can we create a new internet where search engines are irrelevant?
#189Everyone has missed the most important aspect of search engines, from the point of view of their core function of information retrieval: they're the internet equivalent of a library index. Either you find a way to make information findable in a library without an index (how?!?) or you find a novel way to make a neutral search engine - one that provides as much value as Google but whose costs are paid in a different w…
Maybe personal whitelist/blacklist for domains and authors could improve things. Sort of "Web of trust" but done properly.
Not completely without search engines, but for example, if every website was responsible for maintaining it's own index, we could effectively run our own search engines after initialising "base" trusted website lists. Let's say I'm new to this "new internet", I ask around what are some good websites for information I'm interested in. My friend tells me wikipedia is good for general information, webmd for health queries, stackoverflow for programming questions, and so on. I add wikipedia.org/searchindex, webdm.com/searchindex and stackoverflow.com/searchindex to my personal search engine instance, and every time I search something, these three are queried. This could be improved with local cache, synonyms, etc. As you carry on using it, you expand your "library". Of course it would increase workload of individual resources, but has potential to give feel of that web 1.0 once again.
Re: Ask HN: Can we create a new internet where search engines are irrelevant?
#190Earlier quoted context omitted.
>, where all the smarts [...] reside on your device, not in the cloud, is the most promising. [...] An on-device search agent could potentially be the best solution [...] Maybe I misunderstand your proposal but to me, this is not technically possible. We can think of a modern search engine as a process that reduces a raw dataset of exabytes [0] into a comprehensible result of ~5000 bytes (i.e. ~5k being the 1st page…
The smarts living on-device is not necessarily the same as the smarts executing on-device. We already have the means to execute arbitrary code (JS) or specific database queries (SQL) on remote hosts. It's not inconceivable, to me, that my device "knowing me" could consist of building up a local database of the types of things that I want to see, and when I ask it to do a new search, it can assemble a small program wh…
Though to your point, google probably ends up storing this information in the cloud