Ha, yes, I've done that at https://gigablast.com/ . The biggest problems now are the following: 1) Too hard to spider the web. Gatekeeper companies like Cloudflare (owned in part by Google) and Cloudfront make it really difficult for upstart search engines to download web pages. 2) Hardware costs are too high. It's much more expensive now to build a large index (50B+ pages) to be competitive. I believe my algorithms…
> 2) Hardware costs are too high. Which is why the next big search engine should be distributed: https://yacy.net .
Ask HN: Why doesn't anyone create a search engine comparable to 2005 Google?
341–350 of 492 posts
Re: Ask HN: Why doesn't anyone create a search engine comparable to 2005 Google?
#342Re: Ask HN: Why doesn't anyone create a search engine comparable to 2005 Google?
#343Re: Ask HN: Why doesn't anyone create a search engine comparable to 2005 Google?
#344I would use a search engine that only indexed Reddit, Stack Exchange, Wikipedia, and a small number of other sites. And that specifically blocked Pinterest, Quora, most non-personal “blogs”, etc. People suggest DDG ! operators, but I don’t want to use a site’s (bad, single-site) search box. I want a multi-site SERP that only displays results from known good sites, which are customizable.
I've done similar optimizations elsewhere to counter Google's trash results, e.g., I've been beefing up my personal recipe database, with the goal being that I can avoid a google search altogether whenever possible, only hitting google as a last resort.
More and more I wonder, with the modern internet, is it even a feature that the whole web is indexed? Might be a bug.
Re: Ask HN: Why doesn't anyone create a search engine comparable to 2005 Google?
#345Earlier quoted context omitted.
Google Capital is an investor: https://www.forbes.com/sites/katevinton/2015/09/22/google-mi...
That is not the same as being owned by Google.
- Sincerely, a Google employee who has nothing to do with the investment branch of the company
Re: Ask HN: Why doesn't anyone create a search engine comparable to 2005 Google?
#346Earlier quoted context omitted.
This might work well in some situations (e.g. research, development), however it would also increase the effect of echo chambers I think.
Part of the thing with echo chambers is that the search terms themselves can be indicative of a particular bubble. For example, there's a difference in the people that refer to the Bureau of Alcohol, Tobacco, and Firearms by the official initialism, "ATF", and those that use "BATF". There's a strong antigun control bent to the `"BATF" guns` query, compared to the `"ATF" guns` query. If you're indexing forums or socia…
Interestingly, back then, Google was big on neutrality and refused to do anything, stating that it reflected the way people used the word. It was finally addressed using "Google bombing" techniques. Something that Google didn't care much about back them because of its low impact.
Re: Ask HN: Why doesn't anyone create a search engine comparable to 2005 Google?
#347Re: Ask HN: Why doesn't anyone create a search engine comparable to 2005 Google?
#348Ha, yes, I've done that at https://gigablast.com/ . The biggest problems now are the following: 1) Too hard to spider the web. Gatekeeper companies like Cloudflare (owned in part by Google) and Cloudfront make it really difficult for upstart search engines to download web pages. 2) Hardware costs are too high. It's much more expensive now to build a large index (50B+ pages) to be competitive. I believe my algorithms…
I tried out four search words with your search engine, and I am not convinced that it is mainly the index size and not the algorithm that is to blame for bad search results. There are way too much high ranking false positives. Here is what I tried: a) "Berlin": 1. The movie festival "Berlinale" 2. The Wikipedia entry about Berlin 3. Something about a venue "Little Berlin", but the link resolves to an online gaming si…
- Berlin:
Wiki
Berlin travel site (visit Berlin)
website for Berlin
Youtube videos
Britannica for Berlin
Bunch of US town sites named Berlin
- Philosophy:
Same skincare website is first result
Wiki is second
Britannica is third
Stanford
News stories
Other dictionaries and encyclopedias
- History"
history.com is first result
Then is the "my activity" google site, maybe this is actually relevant
Youtube, lots of history channel stuff
Twitter history tag
Wikipedia for "History"
How to delete your Chrome browser history
Dictionary definitions
- Caesar:
Wiki for Julius Caesar
Britannica
BBC for JC
Google maps telling me how to get to Little Caesar's Pizza
Dictionary
Apparently some uni has a system called CAESAR
biography.com
Caesar salad recipe
history.com
images for Caesar
Re: Ask HN: Why doesn't anyone create a search engine comparable to 2005 Google?
#349Earlier quoted context omitted.
That is not the same as being owned by Google.
Especially since Cloudflare went public back in 2019, at which point any investors cashed out. - Sincerely, a Google employee who has nothing to do with the investment branch of the company
Well, actually that is also not true. At IPO preferred stocks convert to common but the investors can keep their ownership, they can but don't have to cash out or can only partially cash out.
Investors can also keep board seats in many (or most?) cases.
Re: Ask HN: Why doesn't anyone create a search engine comparable to 2005 Google?
#350The consistent theme every time this comes up is that dealing with the sheer weight of the internet is almost impossible today. SEO spam is hard to fight and the index gets too heavy. However, I wonder if this is a sign that we're looking at the problem wrong. What if instead of even trying to index the entire web, we moved one step back towards the curated directories of the early web? Give users a search engine and…
I like the idea of a subset of the web, and for a niche purpose. Not sure about user-hosting. Capital is the huge barrier to entry today: Larry Page's genius was to extend google's tech, consumer-habit and PR barriers-to-entry into a capital-based advantage: massive geo server farms, giving faster responses. Consumers have a demonstrated huge preference for faster response.