Live data from Hacker News

Ask HN: Why doesn't anyone create a search engine comparable to 2005 Google?

news.ycombinator.com

241–250 of 492 posts

Re: Ask HN: Why doesn't anyone create a search engine comparable to 2005 Google?

#241
post #12

Ha, yes, I've done that at https://gigablast.com/ . The biggest problems now are the following: 1) Too hard to spider the web. Gatekeeper companies like Cloudflare (owned in part by Google) and Cloudfront make it really difficult for upstart search engines to download web pages. 2) Hardware costs are too high. It's much more expensive now to build a large index (50B+ pages) to be competitive. I believe my algorithms…

It's much more expensive now to build a large index (50B+ pages)

Do you have a cost estimate? Also could you be more selective in indexing, e.g. by having users requests sites to be crawled.

Re: Ask HN: Why doesn't anyone create a search engine comparable to 2005 Google?

#242
post #12

Ha, yes, I've done that at https://gigablast.com/ . The biggest problems now are the following: 1) Too hard to spider the web. Gatekeeper companies like Cloudflare (owned in part by Google) and Cloudfront make it really difficult for upstart search engines to download web pages. 2) Hardware costs are too high. It's much more expensive now to build a large index (50B+ pages) to be competitive. I believe my algorithms…

I tried out four search words with your search engine, and I am not convinced that it is mainly the index size and not the algorithm that is to blame for bad search results. There are way too much high ranking false positives. Here is what I tried: a) "Berlin": 1. The movie festival "Berlinale" 2. The Wikipedia entry about Berlin 3. Something about a venue "Little Berlin", but the link resolves to an online gaming si…

I think you have some great feedback here but for me it also highlights how subjective search results can be for individuals - for example, these false positives that you mention (b2, b3) appear as the top result on Google for me for that query.

It makes me think there must be some fairly large segment of the population that want that domain returned as a result for their query, no?

Re: Ask HN: Why doesn't anyone create a search engine comparable to 2005 Google?

#243
post #12

Ha, yes, I've done that at https://gigablast.com/ . The biggest problems now are the following: 1) Too hard to spider the web. Gatekeeper companies like Cloudflare (owned in part by Google) and Cloudfront make it really difficult for upstart search engines to download web pages. 2) Hardware costs are too high. It's much more expensive now to build a large index (50B+ pages) to be competitive. I believe my algorithms…

It's much more expensive now to build a large index (50B+ pages) Do you have a cost estimate? Also could you be more selective in indexing, e.g. by having users requests sites to be crawled.

Requiring users to know what sites they want in advance somewhat defeats the purpose of a search engine, no?

Re: Ask HN: Why doesn't anyone create a search engine comparable to 2005 Google?

#244
post #12

Ha, yes, I've done that at https://gigablast.com/ . The biggest problems now are the following: 1) Too hard to spider the web. Gatekeeper companies like Cloudflare (owned in part by Google) and Cloudfront make it really difficult for upstart search engines to download web pages. 2) Hardware costs are too high. It's much more expensive now to build a large index (50B+ pages) to be competitive. I believe my algorithms…

This is great! I found something other engines do not pick up! apparently I signed an agile manifesto in 2010 https://agilemanifesto.org/display/000000190.html

Re: Ask HN: Why doesn't anyone create a search engine comparable to 2005 Google?

#245
post #26

They do[0] but nobody cares anymore. Google controls web distribution through Google Chrome. I think we are at the point of no return. There won't be any competition anytime soon no matter what US government does. Only innovation can displace Google. [0] https://search.marginalia.nu/

Marginalia is great to find blog posts, personal sites and other long form content, but it's not a replacement for Google nor intends to.

Funny. Marginalia has an option for No JavaScript but I cannot even do an HTTP “POST” with JavaScript disable at my web browser.

Disclaimer: I study for malicious JS stuff.

Re: Ask HN: Why doesn't anyone create a search engine comparable to 2005 Google?

#246

Earlier quoted context omitted.

It's much more expensive now to build a large index (50B+ pages) Do you have a cost estimate? Also could you be more selective in indexing, e.g. by having users requests sites to be crawled.

Requiring users to know what sites they want in advance somewhat defeats the purpose of a search engine, no?

Not at all. You only have to fail the first request. It is an approach I took with my own attempt at a search engine way back! In fact I know personally that there is at least one patent out there that suggests initial 1st time request users being asked to provide the appropriate response as an efficient way to teach systems for future users.

Obviously failing first requests isn't ideal but for popular requests it quickly becomes insignificant. Wikipedia might (if they don't already) want to make a similar suggestion for users to contribute when finding a low content/missing page.

Re: Ask HN: Why doesn't anyone create a search engine comparable to 2005 Google?

#248
post #82

Earlier quoted context omitted.

Heh, I guess you mean "trawling" - trolling the entire web is something very different :)

"Trolling" is fine, see e.g. https://grammarist.com/usage/trawl-troll/#:~:text=Troll%20fo... .

Not in this context - "trolling" as described there would apply to targeted indexing of a specific site; while "trawling" would refer to a wide net that attempts to catch all the sites.

Re: Ask HN: Why doesn't anyone create a search engine comparable to 2005 Google?

#249
post #12

Ha, yes, I've done that at https://gigablast.com/ . The biggest problems now are the following: 1) Too hard to spider the web. Gatekeeper companies like Cloudflare (owned in part by Google) and Cloudfront make it really difficult for upstart search engines to download web pages. 2) Hardware costs are too high. It's much more expensive now to build a large index (50B+ pages) to be competitive. I believe my algorithms…

I tried out four search words with your search engine, and I am not convinced that it is mainly the index size and not the algorithm that is to blame for bad search results. There are way too much high ranking false positives. Here is what I tried: a) "Berlin": 1. The movie festival "Berlinale" 2. The Wikipedia entry about Berlin 3. Something about a venue "Little Berlin", but the link resolves to an online gaming si…

I don't know about others, but when I think of the "good old google days" I'm _not_ expecting the results for your example queries to be any good.

In those days querying took some effort but the effort paid off. The results for "history" just couldn't matter less in this mindset. You search for "USA history" or "house commons history" or "lake whatever history" instead. If the results come up with unexpected things mixed in, you refine the query.

It was almost like a dialog. As a user, you brought in some context. The engine showed you its results, with a healthy mix of examples of everything it thought was in scope. Then you narrowed the scope by adding keywords (or forcing keywords out). Rinse and repeat. As a user, you were in command and the results reflected that.

The idea that the engine should "understand what you mean" is what took us to the current state. Now it feels like queries don't matter anymore. Google thinks it knows the semantics better than you, and steering it off its chosen path is sometimes obnoxiously hard.

Re: Ask HN: Why doesn't anyone create a search engine comparable to 2005 Google?

#250
post #91

1) Google is better at AI, for example let's take this sloppy search: "some joke where you can't tell if it is serious or joke" It is called Poe's law, and Google returned it at #4. Bing or Duckduckgo don't have a clue... 2) They have a years of user's data, like for specific term, they see what users clicked most, so they see which results were perceived as most relevant. It is hard to catch up if you dont have such…

> Google is better at AI, for example let's take this sloppy search: "some joke where you can't tell if it is serious or joke"

My problem there is that I don't expect or want my search engine to do that. The counter case is where I remember a quote from and article and want to find the article. Old Google would help me find matching text and I could quickly find the original article. Current Google will try to interpret the text and give me some nonsense based on that.

AI has ruined other Google features... the "search by image" feature now analyzes the image, returns a generic tag like "woman", and shows me the wikipedia article on women as the first result.

Old search by image had tineye like functionality and you could find the source of images.

Post reply on HN