Live data from Hacker News

Ask HN: Why doesn't anyone create a search engine comparable to 2005 Google?

news.ycombinator.com

411–420 of 492 posts

Re: Ask HN: Why doesn't anyone create a search engine comparable to 2005 Google?

#411

Earlier quoted context omitted.

2005? There were loads of other search engines (SE), and many meta-SE: hotbot, dogpile, metacrawler, ... (IIRC), plenty more. There was also indexes, which Yahoo, AOL (remember them!) had but there was, what was it called, dmoz?, the open web directory. When Google started, being in the right web directory gave you a boost in SERPs as it was used as a domain trust indicator, and the categories were used for keywords.…

I've tried but can't remember what SE that was, Omni-something?? Google replaced Altavista in my usage, who in turn were usually better than their predecessors.

I used them all and kept using the ones that gave me unique results. Google was hands down better because of pagerank and boosts to dmoz listed sites and because they scanned the whole page ignoring keywords.

Re: Ask HN: Why doesn't anyone create a search engine comparable to 2005 Google?

#412
post #80

Earlier quoted context omitted.

I've had extensively dealing with Cloudflare. They have a complex whitelisting system that is difficult to get on, and they also have an 'AI' system that determines if you should be kicked off that whitelist for whatever reason. Furthermore, they give Google preferred treatment in their UIs and backend algos because it is the incumbent and nobody cares about other smaller search engines. So there's a lot of detail to…

It isn't up to them to give everyone a fair shot. That isn't what their customers actually want. Cloudflare aren't in the "fair shots for all search engines" business. They are in the "stop requests you don't want hitting your servers" business.

I'd argue that a level playing field and more competition in the search space is a good thing.

Re: Ask HN: Why doesn't anyone create a search engine comparable to 2005 Google?

#414

Earlier quoted context omitted.

I think you have some great feedback here but for me it also highlights how subjective search results can be for individuals - for example, these false positives that you mention (b2, b3) appear as the top result on Google for me for that query. It makes me think there must be some fairly large segment of the population that want that domain returned as a result for their query, no?

I would not deny that a large part of subjectivity is involved. This is why I used several markers of subjectivity in my evaluation ("what I can see", "that leaves me", "they seem to me", "I would say", etc.). And related to that: I also agree with other responses that a search often needs to be refined. So my four examples where in no way an exhaustive evaluation, but an explorative experiment, where I just used two…

Oh of course you can criticise the result, I more found it interesting that a billion dollar, optimized search experience thought your false positive was actually a top result. A huge variance in the subjectivity between your experience and their invested reasoning.

But while we're speculating on how the domain the appears at the top of the list, let me hazard a guess...

Philosophy.com was registered in 1999 and according to waybackmachine, has been selling cosmetics on the site since 2000 (20+ years). The company sold in 2010 for ~$1B to a holding company with revenues of $10B+ today (Unfortunately I couldn't find how much it contributes to that revenue). According to Wikipedia, the Philosophy brand has been endorsed by celebrities, including "long-time endorser" Oprah Winfrey, possibly the biggest endorsement you could get for their industry/demographic.

I think it is a long established business, with strong revenues and there's more people online searching for cosmetic brands than for philosophers.

In the same way (admittedly in the extreme) when I'm researching deforestation and I query to see how things are going for the 'amazon', the top result is another successful company registered pre 2000, with strong revenues that most likely attracts more visitors..

Re: Ask HN: Why doesn't anyone create a search engine comparable to 2005 Google?

#415
post #256

Earlier quoted context omitted.

> You do a search on Gigablast and say, well, why didn't it get this result that Google got. And that's because the index isn't big enough I wionder how much this is true, and how much (despite all our rhetoric to the contrary) it's because we have actually come to expect Google's modern proprietary page ranking, which counts more than just inbound links but all sorts of other signals (freshness, relevance to our pre…

I think people also have an inflated recollection of how good Google actually was back in 2005. Back then Google was only going up against indexes and link-rings, not 2021 Google/Bing/DDG/etc.

Google was good, actually very good back in 2000s. Their PageRank algorithm practically eliminated spam pages that were simply a list of keywords. Before Google, those pages came up on the first page of Altavista.

I don't specifically remember 2005, but the quality went down with more modern but still shady SEO practices.

Re: Ask HN: Why doesn't anyone create a search engine comparable to 2005 Google?

#416
post #12

Ha, yes, I've done that at https://gigablast.com/ . The biggest problems now are the following: 1) Too hard to spider the web. Gatekeeper companies like Cloudflare (owned in part by Google) and Cloudfront make it really difficult for upstart search engines to download web pages. 2) Hardware costs are too high. It's much more expensive now to build a large index (50B+ pages) to be competitive. I believe my algorithms…

I just tried it and the UI is kinda old and not mobile friendly but the English results I got were satisfying. Not the case for French though. I'll try again in the future, diversity in this landscape is important.

Re: Ask HN: Why doesn't anyone create a search engine comparable to 2005 Google?

#417

Earlier quoted context omitted.

Mojeek founder story here: https://blog.mojeek.com/2021/03/to-track-or-not-to-track.htm... No-tracking and independent from the start. Now at 4.6 billion pages with own infrastructure and IP. Went to market in 2020 with contextual ads and API. Self-disclosure: CEO

HN is wild: 30m after something is mentiond, the CEO chimes in.

Now, if we could just get that on the Facebook thread... ;)

Re: Ask HN: Why doesn't anyone create a search engine comparable to 2005 Google?

#418
post #256

Earlier quoted context omitted.

> You do a search on Gigablast and say, well, why didn't it get this result that Google got. And that's because the index isn't big enough I wionder how much this is true, and how much (despite all our rhetoric to the contrary) it's because we have actually come to expect Google's modern proprietary page ranking, which counts more than just inbound links but all sorts of other signals (freshness, relevance to our pre…

Well if the result didn't appear in the first 5-10 pages, it's probably not in the index. You can see it with other search engines. I challenge you to come up with a Google query for which a first-page result won't be seen within the first 10 pages of Bing results for the same query. (Bonus points if that result is relevant). There's only so much tweaking that personalization and other heuristic can do. But if someth…

I would like to see the least relevant search result Google comes up with. :)

Yes, I realize this is probably trivial with an API call, but I always found it interesting there isn't a way to see what the site with the lowest pagerank in the index is.

Re: Ask HN: Why doesn't anyone create a search engine comparable to 2005 Google?

#419

Earlier quoted context omitted.

Requiring users to know what sites they want in advance somewhat defeats the purpose of a search engine, no?

since sites are so desperate to be indexed, doesn't it seem better to put the onus on them to announce themselves? it would be great if dns registries publshed public keys .. maybe they do in newer schemes?

That works once your search engine is more widely used, but not a lot of sites are going to register with a niche search engines. Many users on the other hand really want a search engine like this and would be willing to invest some time.

Re: Ask HN: Why doesn't anyone create a search engine comparable to 2005 Google?

#420
post #77

Earlier quoted context omitted.

Trusted consumers are better. The original page-rank algo was organic and bottom-up. But now it's the person not the page. Businesses compete for interaction not inbound links. So if you can make a modern page-rank that follows interaction instead of links and isn't a walled garden then I'd invest.

I could make that work, but what do you mean by "walled garden" in this context?

the business and allies of google - those entrenched interests that limit the current visibility of the web to themselves
Post reply on HN