Live data from Hacker News

Ask HN: Why doesn't anyone create a search engine comparable to 2005 Google?

news.ycombinator.com

271–280 of 492 posts

Re: Ask HN: Why doesn't anyone create a search engine comparable to 2005 Google?

#271

Earlier quoted context omitted.

I tried out four search words with your search engine, and I am not convinced that it is mainly the index size and not the algorithm that is to blame for bad search results. There are way too much high ranking false positives. Here is what I tried: a) "Berlin": 1. The movie festival "Berlinale" 2. The Wikipedia entry about Berlin 3. Something about a venue "Little Berlin", but the link resolves to an online gaming si…

I don't know about others, but when I think of the "good old google days" I'm _not_ expecting the results for your example queries to be any good. In those days querying took some effort but the effort paid off. The results for "history" just couldn't matter less in this mindset. You search for "USA history" or "house commons history" or "lake whatever history" instead. If the results come up with unexpected things m…

I get what you mean, but part of the whole initial appeal of Google was that it gave much more relevant results initially than Altavista or the other options. That was why Google put in the audacious "I'm feeling lucky" button.

Re: Ask HN: Why doesn't anyone create a search engine comparable to 2005 Google?

#272
post #12

Ha, yes, I've done that at https://gigablast.com/ . The biggest problems now are the following: 1) Too hard to spider the web. Gatekeeper companies like Cloudflare (owned in part by Google) and Cloudfront make it really difficult for upstart search engines to download web pages. 2) Hardware costs are too high. It's much more expensive now to build a large index (50B+ pages) to be competitive. I believe my algorithms…

I really love how the results organize multiple matching pages from the same domain. This is really cool.

Re: Ask HN: Why doesn't anyone create a search engine comparable to 2005 Google?

#273
post #262
post #12

Ha, yes, I've done that at https://gigablast.com/ . The biggest problems now are the following: 1) Too hard to spider the web. Gatekeeper companies like Cloudflare (owned in part by Google) and Cloudfront make it really difficult for upstart search engines to download web pages. 2) Hardware costs are too high. It's much more expensive now to build a large index (50B+ pages) to be competitive. I believe my algorithms…

> Cloudflare (owned in part by Google) Please elaborate. Is there a special relationship between Cloudflare and Google?

Google Capital is an investor: https://www.forbes.com/sites/katevinton/2015/09/22/google-mi...

Re: Ask HN: Why doesn't anyone create a search engine comparable to 2005 Google?

#274

Earlier quoted context omitted.

I don't know about others, but when I think of the "good old google days" I'm _not_ expecting the results for your example queries to be any good. In those days querying took some effort but the effort paid off. The results for "history" just couldn't matter less in this mindset. You search for "USA history" or "house commons history" or "lake whatever history" instead. If the results come up with unexpected things m…

I get what you mean, but part of the whole initial appeal of Google was that it gave much more relevant results initially than Altavista or the other options. That was why Google put in the audacious "I'm feeling lucky" button.

Yeah but it's from that same philosophy that Google Search is useless as it optimises for the first result.

There is no search engine that searches literally for what you asked and nothing else. Search is shit in 2021 because it tries to be too clever. I'm more clever than it, let me do the refining.

Re: Ask HN: Why doesn't anyone create a search engine comparable to 2005 Google?

#276
post #12

Ha, yes, I've done that at https://gigablast.com/ . The biggest problems now are the following: 1) Too hard to spider the web. Gatekeeper companies like Cloudflare (owned in part by Google) and Cloudfront make it really difficult for upstart search engines to download web pages. 2) Hardware costs are too high. It's much more expensive now to build a large index (50B+ pages) to be competitive. I believe my algorithms…

curious how you implemented the index, memory based or disk based? Either way you are right, HW costs are extremely expensive and you would need a lot of high RAM/high core count machines to return such a large index to the endusers in a low latency fashion.

Re: Ask HN: Why doesn't anyone create a search engine comparable to 2005 Google?

#277
post #12

Ha, yes, I've done that at https://gigablast.com/ . The biggest problems now are the following: 1) Too hard to spider the web. Gatekeeper companies like Cloudflare (owned in part by Google) and Cloudfront make it really difficult for upstart search engines to download web pages. 2) Hardware costs are too high. It's much more expensive now to build a large index (50B+ pages) to be competitive. I believe my algorithms…

If you're serious about this, add a paid tier. Until it's free, I don't trust you will not ever sell my data to make bank.

Re: Ask HN: Why doesn't anyone create a search engine comparable to 2005 Google?

#279

Earlier quoted context omitted.

It's much more expensive now to build a large index (50B+ pages) Do you have a cost estimate? Also could you be more selective in indexing, e.g. by having users requests sites to be crawled.

Requiring users to know what sites they want in advance somewhat defeats the purpose of a search engine, no?

since sites are so desperate to be indexed, doesn't it seem better to put the onus on them to announce themselves? it would be great if dns registries publshed public keys .. maybe they do in newer schemes?

Re: Ask HN: Why doesn't anyone create a search engine comparable to 2005 Google?

#280
post #231

Earlier quoted context omitted.

Nice. I'd pay 5-10$/mo for a search engine that didn't just funnel me into the revenue-extracting regions of the web like Google does.

A subscriber-supported search engine sounds cool to me. Any precedent?

Kagi.com does this. In closed beta at the moment, but you can email and request access.
Post reply on HN