Earlier quoted context omitted.
Perhaps vote on results like on Reddit posts? Gets the junk sites down (and out of the index eventually).
Reddit is a heavily gatekeeped community by the mods in regards to specific topics
Ask HN: Why doesn't anyone create a search engine comparable to 2005 Google?
251–260 of 492 posts
Re: Ask HN: Why doesn't anyone create a search engine comparable to 2005 Google?
#252Earlier quoted context omitted.
I tried out four search words with your search engine, and I am not convinced that it is mainly the index size and not the algorithm that is to blame for bad search results. There are way too much high ranking false positives. Here is what I tried: a) "Berlin": 1. The movie festival "Berlinale" 2. The Wikipedia entry about Berlin 3. Something about a venue "Little Berlin", but the link resolves to an online gaming si…
I don't know about others, but when I think of the "good old google days" I'm _not_ expecting the results for your example queries to be any good. In those days querying took some effort but the effort paid off. The results for "history" just couldn't matter less in this mindset. You search for "USA history" or "house commons history" or "lake whatever history" instead. If the results come up with unexpected things m…
Bingo! If you cede control to Google, it _will_ do what it's optimized to do, and not what _you_ are looking for.
Re: Ask HN: Why doesn't anyone create a search engine comparable to 2005 Google?
#253Ha, yes, I've done that at https://gigablast.com/ . The biggest problems now are the following: 1) Too hard to spider the web. Gatekeeper companies like Cloudflare (owned in part by Google) and Cloudfront make it really difficult for upstart search engines to download web pages. 2) Hardware costs are too high. It's much more expensive now to build a large index (50B+ pages) to be competitive. I believe my algorithms…
Re: Ask HN: Why doesn't anyone create a search engine comparable to 2005 Google?
#254Some people try: https://www.mojeek.com/ https://fireball.com/ https://search.brave.com/
Mojeek founder story here: https://blog.mojeek.com/2021/03/to-track-or-not-to-track.htm... No-tracking and independent from the start. Now at 4.6 billion pages with own infrastructure and IP. Went to market in 2020 with contextual ads and API. Self-disclosure: CEO
Re: Ask HN: Why doesn't anyone create a search engine comparable to 2005 Google?
#255Ha, yes, I've done that at https://gigablast.com/ . The biggest problems now are the following: 1) Too hard to spider the web. Gatekeeper companies like Cloudflare (owned in part by Google) and Cloudfront make it really difficult for upstart search engines to download web pages. 2) Hardware costs are too high. It's much more expensive now to build a large index (50B+ pages) to be competitive. I believe my algorithms…
I tried out four search words with your search engine, and I am not convinced that it is mainly the index size and not the algorithm that is to blame for bad search results. There are way too much high ranking false positives. Here is what I tried: a) "Berlin": 1. The movie festival "Berlinale" 2. The Wikipedia entry about Berlin 3. Something about a venue "Little Berlin", but the link resolves to an online gaming si…
Re: Ask HN: Why doesn't anyone create a search engine comparable to 2005 Google?
#256Ha, yes, I've done that at https://gigablast.com/ . The biggest problems now are the following: 1) Too hard to spider the web. Gatekeeper companies like Cloudflare (owned in part by Google) and Cloudfront make it really difficult for upstart search engines to download web pages. 2) Hardware costs are too high. It's much more expensive now to build a large index (50B+ pages) to be competitive. I believe my algorithms…
I wionder how much this is true, and how much (despite all our rhetoric to the contrary) it's because we have actually come to expect Google's modern proprietary page ranking, which counts more than just inbound links but all sorts of other signals (freshness, relevance to our previous queries, etc.).
We dislike the additional signals when it feels like Google is trying to second-guess our intentions, but we probably don't notice how well they work when they give us the result we expect in the first three links.
Re: Ask HN: Why doesn't anyone create a search engine comparable to 2005 Google?
#257Re: Ask HN: Why doesn't anyone create a search engine comparable to 2005 Google?
#258I think DuckDuckGo is closer to what you want. Same results for everyone, better privacy, and they're proactive about improving their results. https://duckduckgo.com/ Part of the problem is that there's a lot more low-quality content to wade through now than there was in 2005. I think the Google of 2005 would have trouble delivering quality results today also.
Re: Ask HN: Why doesn't anyone create a search engine comparable to 2005 Google?
#259Re: Ask HN: Why doesn't anyone create a search engine comparable to 2005 Google?
#260Earlier quoted context omitted.
Have you ever looked at the Amazon file? I'll see if I can track down the link but I remember somebody sharing a dump with me from Amazon that apparently was a recent scrape. Edit: https://registry.opendata.aws/commoncrawl/
That's Common Crawl, they do the spidering of some billions of webpages but that's still a tiny percentage of the web versus Google or Bing.