Ha, yes, I've done that at https://gigablast.com/ . The biggest problems now are the following: 1) Too hard to spider the web. Gatekeeper companies like Cloudflare (owned in part by Google) and Cloudfront make it really difficult for upstart search engines to download web pages. 2) Hardware costs are too high. It's much more expensive now to build a large index (50B+ pages) to be competitive. I believe my algorithms…
Nice. I'd pay 5-10$/mo for a search engine that didn't just funnel me into the revenue-extracting regions of the web like Google does.
Ask HN: Why doesn't anyone create a search engine comparable to 2005 Google?
231–240 of 492 posts
Re: Ask HN: Why doesn't anyone create a search engine comparable to 2005 Google?
#232Earlier quoted context omitted.
Nice. I'd pay 5-10$/mo for a search engine that didn't just funnel me into the revenue-extracting regions of the web like Google does.
A subscriber-supported search engine sounds cool to me. Any precedent?
Re: Ask HN: Why doesn't anyone create a search engine comparable to 2005 Google?
#233Earlier quoted context omitted.
Plus, it scales less well than pure algorithmic search. This fight already happened, with a much smaller internet.
It works really, really well for libraries. Research libraries (and research librarians) are phenomenally valuable. I've missed them any time I'm not at a university. Both curators and algorithms are valuable. This goes for finding books, for finding facts and figures, for finding clothes, for finding dishwashers, and for pretty much everything else. I love the fact that I have search engines and online shopping, but…
It scales extremely poorly. It works very well for situations where there are customers/sponsors are willing to spend lots of money for quality, because then the cost scaling doesn't matter as much; research libraries, Lexis/Nexus Westlaw, etc. all do this, but it's not cheap, and the cost scaling with the size of the corpus sucks compared to algorithmic search.
It is among the approaches to internet search that lost to more purely algorithmic search, because it scales poorly in cost.
Re: Ask HN: Why doesn't anyone create a search engine comparable to 2005 Google?
#234Ha, yes, I've done that at https://gigablast.com/ . The biggest problems now are the following: 1) Too hard to spider the web. Gatekeeper companies like Cloudflare (owned in part by Google) and Cloudfront make it really difficult for upstart search engines to download web pages. 2) Hardware costs are too high. It's much more expensive now to build a large index (50B+ pages) to be competitive. I believe my algorithms…
Regarding the Gatekeeper companies like Cloudflare, it sounds like anti-competitive behavior that could potentially be targeted with anti-trust legislation, correct?
Um, this is America. Every market is basically a trust, cartel, or monopoly.
And I don't know if you can hear that, but there is literally laughter in the halls of power. All the show hearings by congress on social media and tech companies only has to do with two things:
1) one political party thinking the other is getting an advantage by them
2) shaking them down for more lobbying and campaign donations
No one in the halls of power give two shits about competition. Larger companies mean larger campaign donations, and more powerful people to hobnob with if/when you leave or lose your political office.
Of course I think that breaking up the cartels in every major sector would lead to massive improvements: more companies is more employment, more domestic employment, more people trying to innovate in management and product development, more product choice, lower prices, more competition, more redundancy/tolerance to supply chain disruption, less corruption in government and possibly better regulation.
Every large company brazenly does market abuse up and to the point of one and only one limiter: the "bad PR" line. So I guess we have that.
Re: Ask HN: Why doesn't anyone create a search engine comparable to 2005 Google?
#235Ha, yes, I've done that at https://gigablast.com/ . The biggest problems now are the following: 1) Too hard to spider the web. Gatekeeper companies like Cloudflare (owned in part by Google) and Cloudfront make it really difficult for upstart search engines to download web pages. 2) Hardware costs are too high. It's much more expensive now to build a large index (50B+ pages) to be competitive. I believe my algorithms…
Re: Ask HN: Why doesn't anyone create a search engine comparable to 2005 Google?
#236Re: Ask HN: Why doesn't anyone create a search engine comparable to 2005 Google?
#237Because we're not having a 2005-Web anymore. More to the point, SEO & Google have evolved together. To have barely relevant results today you need to be good . That takes stellar talent which costs huge amounts of money. Thus, the Google of today, which is optimized to extract that money from us.
An easy way to become way better than google — detect google ads on pages, and penalize these pages in the index. For obvious reason, google search is incapable of doing so.
Re: Ask HN: Why doesn't anyone create a search engine comparable to 2005 Google?
#238Ha, yes, I've done that at https://gigablast.com/ . The biggest problems now are the following: 1) Too hard to spider the web. Gatekeeper companies like Cloudflare (owned in part by Google) and Cloudfront make it really difficult for upstart search engines to download web pages. 2) Hardware costs are too high. It's much more expensive now to build a large index (50B+ pages) to be competitive. I believe my algorithms…
Which is why the next big search engine should be distributed: https://yacy.net.
Re: Ask HN: Why doesn't anyone create a search engine comparable to 2005 Google?
#239Earlier quoted context omitted.
Nice. I'd pay 5-10$/mo for a search engine that didn't just funnel me into the revenue-extracting regions of the web like Google does.
A subscriber-supported search engine sounds cool to me. Any precedent?
Re: Ask HN: Why doesn't anyone create a search engine comparable to 2005 Google?
#240Ha, yes, I've done that at https://gigablast.com/ . The biggest problems now are the following: 1) Too hard to spider the web. Gatekeeper companies like Cloudflare (owned in part by Google) and Cloudfront make it really difficult for upstart search engines to download web pages. 2) Hardware costs are too high. It's much more expensive now to build a large index (50B+ pages) to be competitive. I believe my algorithms…
> 2) Hardware costs are too high. Which is why the next big search engine should be distributed: https://yacy.net .