Live data from Hacker News

Ask HN: Why doesn't anyone create a search engine comparable to 2005 Google?

news.ycombinator.com

341–350 of 492 posts

Re: Ask HN: Why doesn't anyone create a search engine comparable to 2005 Google?

#341
post #12

Ha, yes, I've done that at https://gigablast.com/ . The biggest problems now are the following: 1) Too hard to spider the web. Gatekeeper companies like Cloudflare (owned in part by Google) and Cloudfront make it really difficult for upstart search engines to download web pages. 2) Hardware costs are too high. It's much more expensive now to build a large index (50B+ pages) to be competitive. I believe my algorithms…

> 2) Hardware costs are too high. Which is why the next big search engine should be distributed: https://yacy.net .

"distributed" doesn't make things more hardware efficient... It literally always make them less efficient. If e.g : mastodon had the same number of users as Twitter it would use 10x the ressources for the same traffic.

Re: Ask HN: Why doesn't anyone create a search engine comparable to 2005 Google?

#344
post #185

I would use a search engine that only indexed Reddit, Stack Exchange, Wikipedia, and a small number of other sites. And that specifically blocked Pinterest, Quora, most non-personal “blogs”, etc. People suggest DDG ! operators, but I don’t want to use a site’s (bad, single-site) search box. I want a multi-site SERP that only displays results from known good sites, which are customizable.

I've been thinking about this as well. As Google search results get increasingly worse, I find myself subconsciously filtering out all the garbage and gravitating towards a small number of known sites; and, as many other HNers do, I frequently mitigate this filtering step altogether by adding add "reddit" to any search in which I'm seeking out real human sentiment.

I've done similar optimizations elsewhere to counter Google's trash results, e.g., I've been beefing up my personal recipe database, with the goal being that I can avoid a google search altogether whenever possible, only hitting google as a last resort.

More and more I wonder, with the modern internet, is it even a feature that the whole web is indexed? Might be a bug.

Re: Ask HN: Why doesn't anyone create a search engine comparable to 2005 Google?

#345
post #330

Earlier quoted context omitted.

Google Capital is an investor: https://www.forbes.com/sites/katevinton/2015/09/22/google-mi...

That is not the same as being owned by Google.

Especially since Cloudflare went public back in 2019, at which point any investors cashed out.

- Sincerely, a Google employee who has nothing to do with the investment branch of the company

Re: Ask HN: Why doesn't anyone create a search engine comparable to 2005 Google?

#346

Earlier quoted context omitted.

This might work well in some situations (e.g. research, development), however it would also increase the effect of echo chambers I think.

Part of the thing with echo chambers is that the search terms themselves can be indicative of a particular bubble. For example, there's a difference in the people that refer to the Bureau of Alcohol, Tobacco, and Firearms by the official initialism, "ATF", and those that use "BATF". There's a strong antigun control bent to the `"BATF" guns` query, compared to the `"ATF" guns` query. If you're indexing forums or socia…

Kind of like when searching for "jew" on Google led to antisemitic websites, that's because jews usually prefer the term "jewish".

Interestingly, back then, Google was big on neutrality and refused to do anything, stating that it reflected the way people used the word. It was finally addressed using "Google bombing" techniques. Something that Google didn't care much about back them because of its low impact.

Re: Ask HN: Why doesn't anyone create a search engine comparable to 2005 Google?

#348
post #12

Ha, yes, I've done that at https://gigablast.com/ . The biggest problems now are the following: 1) Too hard to spider the web. Gatekeeper companies like Cloudflare (owned in part by Google) and Cloudfront make it really difficult for upstart search engines to download web pages. 2) Hardware costs are too high. It's much more expensive now to build a large index (50B+ pages) to be competitive. I believe my algorithms…

I tried out four search words with your search engine, and I am not convinced that it is mainly the index size and not the algorithm that is to blame for bad search results. There are way too much high ranking false positives. Here is what I tried: a) "Berlin": 1. The movie festival "Berlinale" 2. The Wikipedia entry about Berlin 3. Something about a venue "Little Berlin", but the link resolves to an online gaming si…

Let's compare with google:

- Berlin:

Wiki

Berlin travel site (visit Berlin)

website for Berlin

Youtube videos

Britannica for Berlin

Bunch of US town sites named Berlin

- Philosophy:

Same skincare website is first result

Wiki is second

Britannica is third

Stanford

News stories

Other dictionaries and encyclopedias

- History"

history.com is first result

Then is the "my activity" google site, maybe this is actually relevant

Youtube, lots of history channel stuff

Twitter history tag

Wikipedia for "History"

How to delete your Chrome browser history

Dictionary definitions

- Caesar:

Wiki for Julius Caesar

Britannica

BBC for JC

Google maps telling me how to get to Little Caesar's Pizza

Dictionary

Apparently some uni has a system called CAESAR

biography.com

Caesar salad recipe

history.com

images for Caesar

Re: Ask HN: Why doesn't anyone create a search engine comparable to 2005 Google?

#349
post #345
post #330

Earlier quoted context omitted.

That is not the same as being owned by Google.

Especially since Cloudflare went public back in 2019, at which point any investors cashed out. - Sincerely, a Google employee who has nothing to do with the investment branch of the company

> at which point any investors cashed out.

Well, actually that is also not true. At IPO preferred stocks convert to common but the investors can keep their ownership, they can but don't have to cash out or can only partially cash out.

Investors can also keep board seats in many (or most?) cases.

Re: Ask HN: Why doesn't anyone create a search engine comparable to 2005 Google?

#350

The consistent theme every time this comes up is that dealing with the sheer weight of the internet is almost impossible today. SEO spam is hard to fight and the index gets too heavy. However, I wonder if this is a sign that we're looking at the problem wrong. What if instead of even trying to index the entire web, we moved one step back towards the curated directories of the early web? Give users a search engine and…

I like the idea of a subset of the web, and for a niche purpose. Not sure about user-hosting. Capital is the huge barrier to entry today: Larry Page's genius was to extend google's tech, consumer-habit and PR barriers-to-entry into a capital-based advantage: massive geo server farms, giving faster responses. Consumers have a demonstrated huge preference for faster response.

I’ve often thought Alexa.com top n sites Would be a good starting point.
Post reply on HN