Live data from Hacker News

Ask HN: Why doesn't anyone create a search engine comparable to 2005 Google?

news.ycombinator.com

451–460 of 492 posts

Re: Ask HN: Why doesn't anyone create a search engine comparable to 2005 Google?

#451

Earlier quoted context omitted.

I think people also have an inflated recollection of how good Google actually was back in 2005. Back then Google was only going up against indexes and link-rings, not 2021 Google/Bing/DDG/etc.

> I think people also have an inflated recollection of how good Google actually was back in 2005. I've been pointing this out for at least close to a decade. I know since I bothered to screenshot and blog about it in 2012. I'll admit mistakes happened back then too, but they were more forgivable like keyword stuffing on unrelated pages. Back then Google were on our side and removed those as fast as possible. Today ho…

> Thinking about it it seems logical that for a search engine that practically speaking has monopoly both on users and as mattgb points out - tonsome degree also on indexing - serving the correct answer first is just dumb: if they can keep me going between their search results and tech blogs with their ads embedded one, two or five times extra that means one, two or five times more ad impressions.

This would mean that google were measuring the quality of their search results by the number of ad impressions which seems unlikely to me. Maybe in some big, wooly sense this is sort of true but it seems pretty unlikely that anyone interested in search quality (i.e. the search team at google) is looking at ad impressions.

Re: Ask HN: Why doesn't anyone create a search engine comparable to 2005 Google?

#452
post #12

Ha, yes, I've done that at https://gigablast.com/ . The biggest problems now are the following: 1) Too hard to spider the web. Gatekeeper companies like Cloudflare (owned in part by Google) and Cloudfront make it really difficult for upstart search engines to download web pages. 2) Hardware costs are too high. It's much more expensive now to build a large index (50B+ pages) to be competitive. I believe my algorithms…

Make sure to file complaints to any competition market authority you have in your country.

Re: Ask HN: Why doesn't anyone create a search engine comparable to 2005 Google?

#453
post #16

Early 2000s google index ran in a garage. The current google index has dedicated power stations. It's a bit like the car industry - you could run a startup from your garage in the early days but you need titanic amounts of capital to compete now thanks to vertical integration. Major governments and billionaires can compete but everybody else is locked out of the market (most "startups" use bings index).

Google's datacenters are huge because they save user behavior data, not because their web search index is particularly big. Also, Google Search wastes a lot of resources on the "search as you type" feature. Running a search engine in your garage is feasible today because hardware and connectivity have improved much faster than the size of the WWW.

It's the frequency of updates that chews power.

Also, that user data is used to improve search results and mitigate webspam that didnt exist in 2005.

Re: Ask HN: Why doesn't anyone create a search engine comparable to 2005 Google?

#454

The consistent theme every time this comes up is that dealing with the sheer weight of the internet is almost impossible today. SEO spam is hard to fight and the index gets too heavy. However, I wonder if this is a sign that we're looking at the problem wrong. What if instead of even trying to index the entire web, we moved one step back towards the curated directories of the early web? Give users a search engine and…

I wonder if we could use some kind of federation (ActivityPub?) to build an aggregate of the search indexes of a curated community. Something like a giant federated whitelist of domains to index.

Re: Ask HN: Why doesn't anyone create a search engine comparable to 2005 Google?

#455

Earlier quoted context omitted.

I would not deny that a large part of subjectivity is involved. This is why I used several markers of subjectivity in my evaluation ("what I can see", "that leaves me", "they seem to me", "I would say", etc.). And related to that: I also agree with other responses that a search often needs to be refined. So my four examples where in no way an exhaustive evaluation, but an explorative experiment, where I just used two…

Oh of course you can criticise the result, I more found it interesting that a billion dollar, optimized search experience thought your false positive was actually a top result. A huge variance in the subjectivity between your experience and their invested reasoning. But while we're speculating on how the domain the appears at the top of the list, let me hazard a guess... Philosophy.com was registered in 1999 and acco…

Okay, you convinced me that it should (inter-subjectivly) count not as a real false positive as I first thought.

Nevertheless, when I try to analyze what is going on here, I would rather use the word "context" instead of "subjectivity", since I think (or at least hope) that my surprise to find this brand on place #2 in my Google results for "philosophy" is shared by quite a lot of people who lack the context to give it meaning, because this brand is unknown to them. I have the excuse that it is a North American brand irrelevant in my German context. Interestingly, when I search for "philosophy" on amazon.com (without refining the search), I get almost exclusively beauty products and related items as a result, but when I search for "philosophy" on amazon.de it is only books. Google nevertheless has the beauty brand as #2 in Germany. Can we agree that Amazon is better at considering the context of the search for "philosophy" than Google?

As an aside: Your "amazon" example reminds me when I was searching for "Davidson" expecting to find information about Donald Davidson, but received a lot of results about Harley-Davidson. (But since I was aware of the importance of this brand, it was understandable to me.)

Re: Ask HN: Why doesn't anyone create a search engine comparable to 2005 Google?

#456

Earlier quoted context omitted.

You are going to pay for a generalized web search when DDG/Google/Bing/etc are free?

Yes. I use Brave Search and I hope they add a paid tier, which I think they have confirmed they'll add at a later date. If you don't pay, you are the product. Simple as that.

Telegram, Signal, Mozilla are counterexamples... Have a large charitable donated cash balance sitting in your account, and your organisational motivation is all different

Re: Ask HN: Why doesn't anyone create a search engine comparable to 2005 Google?

#457
post #399

Earlier quoted context omitted.

I tried out four search words with your search engine, and I am not convinced that it is mainly the index size and not the algorithm that is to blame for bad search results. There are way too much high ranking false positives. Here is what I tried: a) "Berlin": 1. The movie festival "Berlinale" 2. The Wikipedia entry about Berlin 3. Something about a venue "Little Berlin", but the link resolves to an online gaming si…

I'll admit I had not been working on the quality of single term queries as much as I should have lately. However, especially for such simple queries, having a database of link text (inbound hyperlinks and the associated hypertest) is very, very important. And you don't get the necessary corpus of link text if you have a small index. So in this particular case the index size is, indeed, quite likely a factor. And than…

Let me add just one thought on the single term searches: I do not think that a good search result for such terms as "philosophy" should focus on the primary meaning of the term alone. As someone else had pointed out, the beauty brand can be quite important for some people. If we look at a search engine as a tool that needn't present me with near perfect results from the outset, but something I can have a dialogue with to find something, than it is best that results for single terms presents me with a variety of different special meanings (and probably some useful suggestions how to refine my search). Perhaps you can scrap the Wikipedia disambiguation pages and use it somehow to refine your search results.

Re: Ask HN: Why doesn't anyone create a search engine comparable to 2005 Google?

#458
post #406

Earlier quoted context omitted.

No bad idea. At the risk of being sidelined: "philosophy" was not so a bad term either. Start with an arbitrary Wikipedia link and click on the first keyword of the summary after the linguistic annotations (or other annotations in brackets) and repeat the process until you reach a loop. You will almost always end with "philosophy" -> "metaphysics" -> "philosophy" -> ... This works for "Berlin", "history" and "Caesar"…

that's tripped out. where did you hear about that?

I can't remember. Probably on Hacker News.

Re: Ask HN: Why doesn't anyone create a search engine comparable to 2005 Google?

#459
post #12

Ha, yes, I've done that at https://gigablast.com/ . The biggest problems now are the following: 1) Too hard to spider the web. Gatekeeper companies like Cloudflare (owned in part by Google) and Cloudfront make it really difficult for upstart search engines to download web pages. 2) Hardware costs are too high. It's much more expensive now to build a large index (50B+ pages) to be competitive. I believe my algorithms…

Perhaps trolling the entire web is not useful today? I’d love a search engine where I can whitelist sites or take an existing whitelist from trusted curators.

ha nice to hear this idea. I'm planning to work on this as a side project, just started recently

Re: Ask HN: Why doesn't anyone create a search engine comparable to 2005 Google?

#460
post #402

Earlier quoted context omitted.

lmao, hopefully the C code isn't nearly as bad as your html

Please make your substantive points without snark or swipes. We ban accounts that do the latter, because it's poisonous to the culture we're trying to develop here. If you wouldn't mind reviewing https://news.ycombinator.com/newsguidelines.html and taking the intended spirit of the site more to heart, we'd be grateful.

Probably been here longer than you, so really irrelevant. Anyway every single page of his site has html errors, not pointing it out is more poisonous than doing so.
Post reply on HN