Live data from Hacker News

Ask HN: Why doesn't anyone create a search engine comparable to 2005 Google?

news.ycombinator.com

471–480 of 492 posts

Re: Ask HN: Why doesn't anyone create a search engine comparable to 2005 Google?

#471
post #180

Earlier quoted context omitted.

I disagree. A lot of people I know already switched to Duckduckgo. Google’s ability to get relevant results is dropping like a brick, while the quality of DDG has been improving slowly but steadily.

I wish I could agree but from my experience, DDG's search results aren't really that great. Often even worse than Google's. And another private company is not the answer I believe. We need something more drastic, an open-source search engine organized as a genuine non-profit organization. Something like that. Otherwise, whatever replaces Google will just turn into another Google as soon as it gets any momentum.

I think open source will be tough because you're going to need a lot of saints to work on a search engine of Google's caliber.

Maybe an alternative revenue model instead of ads.

Re: Ask HN: Why doesn't anyone create a search engine comparable to 2005 Google?

#472
post #12

Ha, yes, I've done that at https://gigablast.com/ . The biggest problems now are the following: 1) Too hard to spider the web. Gatekeeper companies like Cloudflare (owned in part by Google) and Cloudfront make it really difficult for upstart search engines to download web pages. 2) Hardware costs are too high. It's much more expensive now to build a large index (50B+ pages) to be competitive. I believe my algorithms…

Oh my god! This works so much better than every Internet search engine I have tried.

Re: Ask HN: Why doesn't anyone create a search engine comparable to 2005 Google?

#473

Earlier quoted context omitted.

Oh of course you can criticise the result, I more found it interesting that a billion dollar, optimized search experience thought your false positive was actually a top result. A huge variance in the subjectivity between your experience and their invested reasoning. But while we're speculating on how the domain the appears at the top of the list, let me hazard a guess... Philosophy.com was registered in 1999 and acco…

Okay, you convinced me that it should (inter-subjectivly) count not as a real false positive as I first thought. Nevertheless, when I try to analyze what is going on here, I would rather use the word "context" instead of "subjectivity", since I think (or at least hope) that my surprise to find this brand on place #2 in my Google results for "philosophy" is shared by quite a lot of people who lack the context to give…

We can agree on that, yes =)

I was thinking about this and when you look at the top keyword searches on Google, it's dominated by people searching brands each year, so I think Google is just naturally optimised for this. I think any Search Engine designed for the masses would probably have to behave like this too. https://www.siegemedia.com/seo/most-popular-keywords

I agree, I think the early web was used more for general information rather than specific brand information (and was more useful for people like myself). I'm not sure what is needed to get more results such as university papers or personal web-sites as I think that people use the internet differently now and that the link structure reflects that.

It's interesting that Google isn't used to search for people anymore (I couldn't see any people in the recent top 100 keyword search data).

Re: Ask HN: Why doesn't anyone create a search engine comparable to 2005 Google?

#474
post #12

Ha, yes, I've done that at https://gigablast.com/ . The biggest problems now are the following: 1) Too hard to spider the web. Gatekeeper companies like Cloudflare (owned in part by Google) and Cloudfront make it really difficult for upstart search engines to download web pages. 2) Hardware costs are too high. It's much more expensive now to build a large index (50B+ pages) to be competitive. I believe my algorithms…

did you ever try to raise funds? why/not? not accusing, just curious.

did you ever think, let me just focus on Italy-relevant results? or job search only? or some slice of search.

Re: Ask HN: Why doesn't anyone create a search engine comparable to 2005 Google?

#475
post #256

Earlier quoted context omitted.

> You do a search on Gigablast and say, well, why didn't it get this result that Google got. And that's because the index isn't big enough I wionder how much this is true, and how much (despite all our rhetoric to the contrary) it's because we have actually come to expect Google's modern proprietary page ranking, which counts more than just inbound links but all sorts of other signals (freshness, relevance to our pre…

I think people also have an inflated recollection of how good Google actually was back in 2005. Back then Google was only going up against indexes and link-rings, not 2021 Google/Bing/DDG/etc.

I hate google now. Every time I use it by accident I’m reminded how infuriating it is. I know DuckDuckGo is just bing in a Halloween mask, but I’ll gladly use something that’s not awesome as long as it’s also not infuriating. I’d take 2005 google any day.

Re: Ask HN: Why doesn't anyone create a search engine comparable to 2005 Google?

#476

Earlier quoted context omitted.

I think people also have an inflated recollection of how good Google actually was back in 2005. Back then Google was only going up against indexes and link-rings, not 2021 Google/Bing/DDG/etc.

Google was good, actually very good back in 2000s. Their PageRank algorithm practically eliminated spam pages that were simply a list of keywords. Before Google, those pages came up on the first page of Altavista. I don't specifically remember 2005, but the quality went down with more modern but still shady SEO practices.

No, quality went down because google shat the bed. All the changes have been deliberate.

Re: Ask HN: Why doesn't anyone create a search engine comparable to 2005 Google?

#478

Earlier quoted context omitted.

Okay, you convinced me that it should (inter-subjectivly) count not as a real false positive as I first thought. Nevertheless, when I try to analyze what is going on here, I would rather use the word "context" instead of "subjectivity", since I think (or at least hope) that my surprise to find this brand on place #2 in my Google results for "philosophy" is shared by quite a lot of people who lack the context to give…

We can agree on that, yes =) I was thinking about this and when you look at the top keyword searches on Google, it's dominated by people searching brands each year, so I think Google is just naturally optimised for this. I think any Search Engine designed for the masses would probably have to behave like this too. https://www.siegemedia.com/seo/most-popular-keywords I agree, I think the early web was used more for ge…

Some observations:

Most of the "brands" in the top 100, especially at the beginning, are rather Internet services. These search terms seem to have been entered not with the intention to "search" in the sense to find some new information, but as a substitute for a bookmark to the respectice service. Who searches for #1 "youtube" does not want information about youtube, but wants to use the youtube Web-site as a portal to find videos there.

I would also guess that most of these searches haven't been initiated through the Google Web-site, but directly from the browser's adress/search bar or a smartphone app. They exhibit a specific usage pattern, but do not show what the people, that entered them, were really searching for, if they were searching at all. What are those people who search for "youtube" doing next: either search again on youtube or log into their youtube account and browse their youtube bookmarks.

The early Internet did not have so many different service people used at a daily basis, and those that existed were more diverse (think of the many differen online email providers in those days) so that the search terms spread out more. Also browsers had no direct integration with a search engine. The incentive was higher to use bookmarks for your favourite service, since otherwise you had to use a boomark to a search engine anyway.

Perhaps it would be more approbriate to compare the use of the early Google not with the current Google, but the current Google Scholar?

Re: Ask HN: Why doesn't anyone create a search engine comparable to 2005 Google?

#479
post #330

Earlier quoted context omitted.

Google Capital is an investor: https://www.forbes.com/sites/katevinton/2015/09/22/google-mi...

That is not the same as being owned by Google.

Actually, being an investor in a company is the same as owning that company in part.

Re: Ask HN: Why doesn't anyone create a search engine comparable to 2005 Google?

#480
post #12

Ha, yes, I've done that at https://gigablast.com/ . The biggest problems now are the following: 1) Too hard to spider the web. Gatekeeper companies like Cloudflare (owned in part by Google) and Cloudfront make it really difficult for upstart search engines to download web pages. 2) Hardware costs are too high. It's much more expensive now to build a large index (50B+ pages) to be competitive. I believe my algorithms…

Perhaps trolling the entire web is not useful today? I’d love a search engine where I can whitelist sites or take an existing whitelist from trusted curators.

If the user requests a website, you could at least crawl on request, which would be an excuse to bypass the rules in robots.txt. It would be a loophole, let’s say.
Post reply on HN