Live data from Hacker News

Ask HN: Why doesn't anyone create a search engine comparable to 2005 Google?

news.ycombinator.com

251–260 of 492 posts

Re: Ask HN: Why doesn't anyone create a search engine comparable to 2005 Google?

#251
post #113
post #78

Earlier quoted context omitted.

Perhaps vote on results like on Reddit posts? Gets the junk sites down (and out of the index eventually).

Reddit is a heavily gatekeeped community by the mods in regards to specific topics

Reddit is an extreme example of group think. Try posting something pro-Trump (I mean, surely even that guy has a positive thing or two to be said about him) and you'll get banned in some subs. Or you may get banned simply because the mod doesn't like the fact that you don't toe the party line.

Re: Ask HN: Why doesn't anyone create a search engine comparable to 2005 Google?

#252

Earlier quoted context omitted.

I tried out four search words with your search engine, and I am not convinced that it is mainly the index size and not the algorithm that is to blame for bad search results. There are way too much high ranking false positives. Here is what I tried: a) "Berlin": 1. The movie festival "Berlinale" 2. The Wikipedia entry about Berlin 3. Something about a venue "Little Berlin", but the link resolves to an online gaming si…

I don't know about others, but when I think of the "good old google days" I'm _not_ expecting the results for your example queries to be any good. In those days querying took some effort but the effort paid off. The results for "history" just couldn't matter less in this mindset. You search for "USA history" or "house commons history" or "lake whatever history" instead. If the results come up with unexpected things m…

> The idea that the engine should "understand what you mean" is what took us to the current state. Now it feels like queries don't matter anymore. Google thinks it knows the semantics better than you, and steering it off its chosen path is sometimes obnoxiously hard.

Bingo! If you cede control to Google, it _will_ do what it's optimized to do, and not what _you_ are looking for.

Re: Ask HN: Why doesn't anyone create a search engine comparable to 2005 Google?

#253
post #12

Ha, yes, I've done that at https://gigablast.com/ . The biggest problems now are the following: 1) Too hard to spider the web. Gatekeeper companies like Cloudflare (owned in part by Google) and Cloudfront make it really difficult for upstart search engines to download web pages. 2) Hardware costs are too high. It's much more expensive now to build a large index (50B+ pages) to be competitive. I believe my algorithms…

Do you have some sort of PageRank?

Re: Ask HN: Why doesn't anyone create a search engine comparable to 2005 Google?

#254

Some people try: https://www.mojeek.com/ https://fireball.com/ https://search.brave.com/

Mojeek founder story here: https://blog.mojeek.com/2021/03/to-track-or-not-to-track.htm... No-tracking and independent from the start. Now at 4.6 billion pages with own infrastructure and IP. Went to market in 2020 with contextual ads and API. Self-disclosure: CEO

Do you use some sort of PageRank?

Re: Ask HN: Why doesn't anyone create a search engine comparable to 2005 Google?

#255
post #12

Ha, yes, I've done that at https://gigablast.com/ . The biggest problems now are the following: 1) Too hard to spider the web. Gatekeeper companies like Cloudflare (owned in part by Google) and Cloudfront make it really difficult for upstart search engines to download web pages. 2) Hardware costs are too high. It's much more expensive now to build a large index (50B+ pages) to be competitive. I believe my algorithms…

I tried out four search words with your search engine, and I am not convinced that it is mainly the index size and not the algorithm that is to blame for bad search results. There are way too much high ranking false positives. Here is what I tried: a) "Berlin": 1. The movie festival "Berlinale" 2. The Wikipedia entry about Berlin 3. Something about a venue "Little Berlin", but the link resolves to an online gaming si…

OK, I'll bite. How would _you_ rank the results for each of those queries?

Re: Ask HN: Why doesn't anyone create a search engine comparable to 2005 Google?

#256
post #12

Ha, yes, I've done that at https://gigablast.com/ . The biggest problems now are the following: 1) Too hard to spider the web. Gatekeeper companies like Cloudflare (owned in part by Google) and Cloudfront make it really difficult for upstart search engines to download web pages. 2) Hardware costs are too high. It's much more expensive now to build a large index (50B+ pages) to be competitive. I believe my algorithms…

> You do a search on Gigablast and say, well, why didn't it get this result that Google got. And that's because the index isn't big enough

I wionder how much this is true, and how much (despite all our rhetoric to the contrary) it's because we have actually come to expect Google's modern proprietary page ranking, which counts more than just inbound links but all sorts of other signals (freshness, relevance to our previous queries, etc.).

We dislike the additional signals when it feels like Google is trying to second-guess our intentions, but we probably don't notice how well they work when they give us the result we expect in the first three links.

Re: Ask HN: Why doesn't anyone create a search engine comparable to 2005 Google?

#258

I think DuckDuckGo is closer to what you want. Same results for everyone, better privacy, and they're proactive about improving their results. https://duckduckgo.com/ Part of the problem is that there's a lot more low-quality content to wade through now than there was in 2005. I think the Google of 2005 would have trouble delivering quality results today also.

DuckDuckGo isn’t really a search engine, it’s a website that uses bing’s api.

Re: Ask HN: Why doesn't anyone create a search engine comparable to 2005 Google?

#260
post #167

Earlier quoted context omitted.

Have you ever looked at the Amazon file? I'll see if I can track down the link but I remember somebody sharing a dump with me from Amazon that apparently was a recent scrape. Edit: https://registry.opendata.aws/commoncrawl/

That's Common Crawl, they do the spidering of some billions of webpages but that's still a tiny percentage of the web versus Google or Bing.

Do you have any stats on that? I've always wondered about the coverage of Common Crawl, if you include all the historical crawl files too.
Post reply on HN