Live data from Hacker News

Ask HN: Why doesn't anyone create a search engine comparable to 2005 Google?

news.ycombinator.com

211–220 of 492 posts

Re: Ask HN: Why doesn't anyone create a search engine comparable to 2005 Google?

#211
post #175

I'd like to see a "just search" engine, all it does it search for a specific string, case insensitively, across the entire web. No curation or anything, just sorted in lexicological order closest match first maybe falling back to page age if it has more then one exact match. Perhaps give me some regular expressions as well.

Maybe a “stability factor” could be calculated. Whereas earlier new content was king, I now value a stable long term source of information. So domain age + page age + content variability + dependency on ads. That might give more honest sources a go.

That's a good idea, I'd make it a option. Do you want newest first, oldest first or by stability?

Re: Ask HN: Why doesn't anyone create a search engine comparable to 2005 Google?

#212
post #189
post #171

Earlier quoted context omitted.

This is not correct. Brave Search owns its own (growing) index and relies on third-parties like Bing for some fraction of the requests. Which is not the same thing as relying fully on Bing or third-parties for results like so many meta-search engines. More detailed answer here: https://search.brave.com/help/independence Edit: Forgot to say that I work on Brave Search.

brave 'falls back' to bing. which in my experience is most of the time. in fact, out of all the queries i did a while back, they all seemed to come directly from bing. is there a way to disable the reliance on bing and get pure 'brave only' results? and can you be more specific as to what this fraction is? do you blend at all?

You can check exactly which fraction of the results were fetched from Brave's index vs. third-parties using the "independence score" found in the setting drawers (opening can be done with the cog icon at the top right of any page on search.brave.com). There is there a global and personalized score of independence (respectively aggregated on all user's and for your queries only).

Explanation is also found here with screenshots: https://search.brave.com/help/independence

Re: Ask HN: Why doesn't anyone create a search engine comparable to 2005 Google?

#213
post #195

Earlier quoted context omitted.

SEO often seems to be a compensation for the fact that a site doesn't have particularly worthwhile content. So punishing SEO surprisingly does promote higher quality of search results.

Yes and no. A lot of those sites are small local businesses trying to get found. A front page listing can be the difference between surviving and going under. Much of the time the blog spam is what floats hours, contact info, and services provided to the first page.

Be that as it may, search ranking is a zero sum game. The unfair advantage SEO gives this particular struggling business means another goes under. I'd rather punish the guy trying to game the system than the one with enough principles not to.

Re: Ask HN: Why doesn't anyone create a search engine comparable to 2005 Google?

#214
post #2

This is how Private Search [1] works since it decouples the search from the user. This means nobody knows both who searched and what they searched for. This is a huge leap for privacy in search. [1] https://private.sh

Just tried it and it worked for me.

Re: Ask HN: Why doesn't anyone create a search engine comparable to 2005 Google?

#215

I'm probably the only person who doesn't think Google search has deteriorated. I play security CTFs, so a lot of times I have to search for peculiar technical details on various software. Also, like any other human being, I also make generic queries. In both cases, I feel like I almost always get to the desired webpage within the top few results.

It honestly depends on what you are searching for.

Case 1: You just want the name of the website, or an article, example "Facebook" -> fb.com, "Gordan Ramsay" -> Wiki/official website/Celeb gossip website you are good. Not much competition here.

Case 2: You are looking for something technical like "GNU rnano CVE-abcde"/"OpenBSD ARM64 Qualcomm Wifi driver not working", you are again in the fine territory, not much if any money to be made here so very less competition. There will be the official forums, websites, maybe some conference websites in this category.

Case 3: "Chicken potpie recipe", "How to be more organised": This is the category where people are trying to game the SEO algos. How the hell do recipe websites with 27 popups, 12000 word essay on the secret family history ends up on top ? There are a huge number of passionately made simple recipe websites but they have to be "found" by us. For the second query I mention about being more organized I think most people are looking for some sort of a review article which looks at some various schools of thoughts regarding discipline, cleanliness pointing to further resources and exploring the why and what to do for this. Here the search engine needs to determine the context of the query which is fairly abstract and then the internal heuristics it uses are supposed to drive it to a meaningful list of websites. Maybe the average joe would like to click cosmopolitan's article but I would never do that. Based on my previous click history maybe google should determine what I kind of links am I looking for. But when they figure that out they'd much faster use this behavioral insight for advertisers. A great search engine is basically a primitive personal librarian, I'd pay a yearly subscription for one.

The internet is vast and it has stuff that I don't know about. How my 7 word abstract query is gonna get me there is the question mark. Also, for a lot of queries the top results can be plagued by spammy/fraud results which are on top because they managed to trick the SEO algos. These bad actors were not as prevalent for 2005 google.

Re: Ask HN: Why doesn't anyone create a search engine comparable to 2005 Google?

#216
post #212
post #189

Earlier quoted context omitted.

brave 'falls back' to bing. which in my experience is most of the time. in fact, out of all the queries i did a while back, they all seemed to come directly from bing. is there a way to disable the reliance on bing and get pure 'brave only' results? and can you be more specific as to what this fraction is? do you blend at all?

You can check exactly which fraction of the results were fetched from Brave's index vs. third-parties using the "independence score" found in the setting drawers (opening can be done with the cog icon at the top right of any page on search.brave.com). There is there a global and personalized score of independence (respectively aggregated on all user's and for your queries only). Explanation is also found here with sc…

So Brave is still dependent on Google and Bing it seems. Also is this Brave's CEO: https://www.bbc.com/news/technology-26868536 https://www.nytimes.com/2020/12/22/business/brave-brendan-ei... ? "Brendan Eich's opposition to same-sex marriage cost him his job at Mozilla." "Covid comments get a tech C.E.O. in hot water, again."

Re: Ask HN: Why doesn't anyone create a search engine comparable to 2005 Google?

#218
post #91

1) Google is better at AI, for example let's take this sloppy search: "some joke where you can't tell if it is serious or joke" It is called Poe's law, and Google returned it at #4. Bing or Duckduckgo don't have a clue... 2) They have a years of user's data, like for specific term, they see what users clicked most, so they see which results were perceived as most relevant. It is hard to catch up if you dont have such…

> 1) Google is better at AI, for example let's take this sloppy search: "some joke where you can't tell if it is serious or joke" > It is called Poe's law, and Google returned it at #4. Bing or Duckduckgo don't have a clue... Interesting, I was looking for a good benchmark like this. For me Google returned it at #5 with an image/related terms carousel before it which places it physically more around #7 on the page. B…

The Brave results though seem to contain “good sites” whereas the Google results are content mill blogspam. The exact placement of Poe’s Law is somewhat less important.

Re: Ask HN: Why doesn't anyone create a search engine comparable to 2005 Google?

#219
post #185

I would use a search engine that only indexed Reddit, Stack Exchange, Wikipedia, and a small number of other sites. And that specifically blocked Pinterest, Quora, most non-personal “blogs”, etc. People suggest DDG ! operators, but I don’t want to use a site’s (bad, single-site) search box. I want a multi-site SERP that only displays results from known good sites, which are customizable.

Even rules such as “if there is a Wikipedia result in the top 10, display it first”.

Re: Ask HN: Why doesn't anyone create a search engine comparable to 2005 Google?

#220

I've been using kagi.com for a month or so now, and it consistently beats DDG and Ecosia for result quality. I'd guess it beats Google too, since last time I used Google it was nothing but ads and spam which is why I stopped.

Thank you for the vote of confidence! Better than Google is our goal, glad you perceive it that way.

You're welcome. I'm really impressed with it most of the time. Still not made it on to the Orion beta though ;)
Post reply on HN