Live data from Hacker News

Ask HN: Why doesn't anyone create a search engine comparable to 2005 Google?

news.ycombinator.com

31–40 of 492 posts

Re: Ask HN: Why doesn't anyone create a search engine comparable to 2005 Google?

#31
post #12

Ha, yes, I've done that at https://gigablast.com/ . The biggest problems now are the following: 1) Too hard to spider the web. Gatekeeper companies like Cloudflare (owned in part by Google) and Cloudfront make it really difficult for upstart search engines to download web pages. 2) Hardware costs are too high. It's much more expensive now to build a large index (50B+ pages) to be competitive. I believe my algorithms…

Perhaps trolling the entire web is not useful today? I’d love a search engine where I can whitelist sites or take an existing whitelist from trusted curators.

Re: Ask HN: Why doesn't anyone create a search engine comparable to 2005 Google?

#32
post #26

They do[0] but nobody cares anymore. Google controls web distribution through Google Chrome. I think we are at the point of no return. There won't be any competition anytime soon no matter what US government does. Only innovation can displace Google. [0] https://search.marginalia.nu/

Marginalia is great to find blog posts, personal sites and other long form content, but it's not a replacement for Google nor intends to.

It does operate on a scale and principle fairly similar to early 2000s google, so the comparison isn't that far off, but yeah, it's quite some way before it's viable for general search. Dunno if I'll ever get there, but it does consistently seem to get better so who knows.

Re: Ask HN: Why doesn't anyone create a search engine comparable to 2005 Google?

#33
post #26

They do[0] but nobody cares anymore. Google controls web distribution through Google Chrome. I think we are at the point of no return. There won't be any competition anytime soon no matter what US government does. Only innovation can displace Google. [0] https://search.marginalia.nu/

Marginalia is great to find blog posts, personal sites and other long form content, but it's not a replacement for Google nor intends to.

But it is a good start and foundation for something bigger and better.

Re: Ask HN: Why doesn't anyone create a search engine comparable to 2005 Google?

#34
post #22

No mention of DDG in the comments? Is there a reason I'm not seeing or it's just not the preferred alt-search on HN? Seems to have been working fine for me when I struggle to get past the funnels and content mills on Google.

I dont find search results to be too relevant (at least for me, also Spaniard here). It is my default search engine only for the bang commands.

Re: Ask HN: Why doesn't anyone create a search engine comparable to 2005 Google?

#35
post #6

While I feel that Google has become worse in last couple of years, I'm pretty sure it is still better now when 15 years ago. Maybe it is just some kind of nostalgia?

the internet has changed, partially due to google's influence

instead of discussion forums and Q&A sites, everyone's on facebook/twitter/discord/slack/snapchat/tiktok/etc... none of that is really very google friendly

online marketing and SEO is a much larger industry now, so with less (by % of total) searchable content generated by people (which is on social media) a lot of the high-ranking content that appears in search is highly optimized marketing

then you have other kind of weird things like... half of all internet traffic being bots

Re: Ask HN: Why doesn't anyone create a search engine comparable to 2005 Google?

#36
post #12

Ha, yes, I've done that at https://gigablast.com/ . The biggest problems now are the following: 1) Too hard to spider the web. Gatekeeper companies like Cloudflare (owned in part by Google) and Cloudfront make it really difficult for upstart search engines to download web pages. 2) Hardware costs are too high. It's much more expensive now to build a large index (50B+ pages) to be competitive. I believe my algorithms…

I've used Gigablast off and on for a long time (I think I first discovered Gigablast in 2006 or so). Would be cool to have a registration service for legitimate spiders. I used to run a team that scraped jobs and delivered them (by fax, email, us mail as require by law) to local veteran's employment staffers for compliance. We were contracted by huge companies (at one point about 700 of the fortune 1000) to do so, and often our spiders would be blocked by the employer's IT department even though the HR team was paying us big bucks to do so.

Re: Ask HN: Why doesn't anyone create a search engine comparable to 2005 Google?

#37
Cliqz wanted to build new search engine but failed. It's just too difficult to operate at that scale and break the existing monopoly of big G.

https://www.burda.com/en/news/cliqz-closes-areas-browser-and...

https://news.ycombinator.com/item?id=23031520

https://0x65.dev/blog/2019-12-06/building-a-search-engine-fr...

Re: Ask HN: Why doesn't anyone create a search engine comparable to 2005 Google?

#38
post #12

Ha, yes, I've done that at https://gigablast.com/ . The biggest problems now are the following: 1) Too hard to spider the web. Gatekeeper companies like Cloudflare (owned in part by Google) and Cloudfront make it really difficult for upstart search engines to download web pages. 2) Hardware costs are too high. It's much more expensive now to build a large index (50B+ pages) to be competitive. I believe my algorithms…

Perhaps trolling the entire web is not useful today? I’d love a search engine where I can whitelist sites or take an existing whitelist from trusted curators.

Trusted curators is a dangerous dependency

Re: Ask HN: Why doesn't anyone create a search engine comparable to 2005 Google?

#39
post #19
post #12

Ha, yes, I've done that at https://gigablast.com/ . The biggest problems now are the following: 1) Too hard to spider the web. Gatekeeper companies like Cloudflare (owned in part by Google) and Cloudfront make it really difficult for upstart search engines to download web pages. 2) Hardware costs are too high. It's much more expensive now to build a large index (50B+ pages) to be competitive. I believe my algorithms…

Regarding the Gatekeeper companies like Cloudflare, it sounds like anti-competitive behavior that could potentially be targeted with anti-trust legislation, correct?

i would assume its mostly anti scraping protection which is mostly for privacy. you don't want to allow everyone scrap your website, pull and use your info. for example from fb, ig, LinkedIn, github, .... you can build a really big profiling db on people that way. so websites need to know you are a legit search engine first

Re: Ask HN: Why doesn't anyone create a search engine comparable to 2005 Google?

#40

Natural Language Processing is a pox on modern search engines. I suspect that Google et. al. wanted to transform their product into an answer engine that powers voice assistants like Siri and just assumed everyone would naturally like the new way better. I can't stand how Google is always trying to guess what I want, rather than simply returning non-personalized results solely based on exactly what I typed in the tex…

What’s a pox?
Post reply on HN